Text-Conditioned Resampler For Long Form Video Understanding

University of Oxford

Google

Google DeepMind

TU Munich


Abstract
Videos are highly redundant data source and it is often enough to identify a few key moments to solve any given task. In this paper, we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequences for a task. TCR localises relevant visual features from the video given a text condition and provides them to a LLM to generate a text response. Due to its lightweight design and use of cross-attention, TCR can process more than 100 frames at a time allowing the model to use much longer chunks of video than earlier works. We make the following contributions: (i) we design a transformer-based sampling architecture that can process long videos conditioned on a task, together with a training method that enables it to bridge pre-trained visual and language models; (ii) we empirically validate its efficacy on a wide variety of evaluation tasks, and set a new state-of-the-art on NextQA, EgoSchema, and the EGO4D-LTA challenge; and (iii) we determine tasks which require longer video contexts and that can thus be used effectively for further evaluation of long-range video models.


๋น„๋””์˜ค๋Š” ๊ณ ๋„๋กœ ์ค‘๋ณต๋œ ๋ฐ์ดํ„ฐ ์†Œ์Šค์ด๋ฉฐ, ์ฃผ์–ด์ง„ ์ž‘์—…์„ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด ๋ช‡ ๊ฐ€์ง€ ํ•ต์‹ฌ ์ˆœ๊ฐ„์„ ์‹๋ณ„ํ•˜๋Š” ๊ฒƒ๋งŒ์œผ๋กœ ์ข…์ข… ์ถฉ๋ถ„ํ•ฉ๋‹ˆ๋‹ค. ์ด ๋…ผ๋ฌธ์—์„œ, ์šฐ๋ฆฌ๋Š” ๊ธด ๋น„๋””์˜ค ์‹œํ€€์Šค๋ฅผ ์ž‘์—…์— ๋งž๊ฒŒ ์ฒ˜๋ฆฌํ•˜๊ธฐ ์œ„ํ•ด ์‚ฌ์ „ ํ›ˆ๋ จ๋œ ๋น„์ฃผ์–ผ ์ธ์ฝ”๋”์™€ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ์‚ฌ์šฉํ•˜๋Š” ํ…์ŠคํŠธ ์กฐ๊ฑด๋ถ€ ๋น„๋””์˜ค ๋ฆฌ์ƒ˜ํ”Œ๋Ÿฌ(TCR) ๋ชจ๋“ˆ์„ ์ œ์‹œํ•ฉ๋‹ˆ๋‹ค. TCR์€ ํ…์ŠคํŠธ ์กฐ๊ฑด์ด ์ฃผ์–ด์ง„ ๋น„๋””์˜ค์—์„œ ๊ด€๋ จ ๋น„์ฃผ์–ผ ํŠน์ง•์„ ์ฐพ์•„๋‚ด๊ณ , ์ด๋ฅผ LLM์— ์ œ๊ณตํ•˜์—ฌ ํ…์ŠคํŠธ ์‘๋‹ต์„ ์ƒ์„ฑํ•ฉ๋‹ˆ๋‹ค. ๊ฐ€๋ฒผ์šด ์„ค๊ณ„์™€ ํฌ๋กœ์Šค ์–ดํ…์…˜์˜ ์‚ฌ์šฉ์œผ๋กœ ์ธํ•ด, TCR์€ ํ•œ ๋ฒˆ์— 100๊ฐœ ์ด์ƒ์˜ ํ”„๋ ˆ์ž„์„ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ์œผ๋ฉฐ, ์ด๋Š” ์ด์ „ ์ž‘์—…๋“ค๋ณด๋‹ค ํ›จ์”ฌ ๊ธด ๋น„๋””์˜ค ์กฐ๊ฐ์„ ์‚ฌ์šฉํ•  ์ˆ˜ ์žˆ๊ฒŒ ํ•ฉ๋‹ˆ๋‹ค. ์šฐ๋ฆฌ๋Š” ๋‹ค์Œ๊ณผ ๊ฐ™์€ ๊ธฐ์—ฌ๋ฅผ ํ•ฉ๋‹ˆ๋‹ค: (i) ์ž‘์—…์— ๋”ฐ๋ผ ๊ธด ๋น„๋””์˜ค๋ฅผ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ๋Š” ํŠธ๋žœ์Šคํฌ๋จธ ๊ธฐ๋ฐ˜ ์ƒ˜ํ”Œ๋ง ์•„ํ‚คํ…์ฒ˜๋ฅผ ์„ค๊ณ„ํ•˜๊ณ , ์‚ฌ์ „ ํ›ˆ๋ จ๋œ ๋น„์ฃผ์–ผ ๋ฐ ์–ธ์–ด ๋ชจ๋ธ ๊ฐ„์˜ ์—ฐ๊ฒฐ์„ ๊ฐ€๋Šฅํ•˜๊ฒŒ ํ•˜๋Š” ํ›ˆ๋ จ ๋ฐฉ๋ฒ•์„ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค; (ii) ๋‹ค์–‘ํ•œ ํ‰๊ฐ€ ์ž‘์—…์—์„œ ๊ทธ ํšจ๊ณผ๋ฅผ ์‹ค์ฆ์ ์œผ๋กœ ๊ฒ€์ฆํ•˜๊ณ , NextQA, EgoSchema, EGO4D-LTA ์ฑŒ๋ฆฐ์ง€์—์„œ ์ƒˆ๋กœ์šด ์ตœ๊ณ  ๊ธฐ๋ก์„ ์„ธ์›๋‹ˆ๋‹ค; ๊ทธ๋ฆฌ๊ณ  (iii) ๋” ๊ธด ๋น„๋””์˜ค ์ปจํ…์ŠคํŠธ๋ฅผ ํ•„์š”๋กœ ํ•˜๊ณ  ๋”ฐ๋ผ์„œ ์žฅ๊ฑฐ๋ฆฌ ๋น„๋””์˜ค ๋ชจ๋ธ์˜ ์ถ”๊ฐ€ ํ‰๊ฐ€์— ํšจ๊ณผ์ ์œผ๋กœ ์‚ฌ์šฉ๋  ์ˆ˜ ์žˆ๋Š” ์ž‘์—…์„ ๊ฒฐ์ •ํ•ฉ๋‹ˆ๋‹ค.






https://arxiv.org/abs/2312.11897