University of Oxford
Google
Google DeepMind
TU Munich
Abstract
Videos are highly redundant data source and it is often enough to identify a few key moments to solve any given task. In this paper, we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequences for a task. TCR localises relevant visual features from the video given a text condition and provides them to a LLM to generate a text response. Due to its lightweight design and use of cross-attention, TCR can process more than 100 frames at a time allowing the model to use much longer chunks of video than earlier works. We make the following contributions: (i) we design a transformer-based sampling architecture that can process long videos conditioned on a task, together with a training method that enables it to bridge pre-trained visual and language models; (ii) we empirically validate its efficacy on a wide variety of evaluation tasks, and set a new state-of-the-art on NextQA, EgoSchema, and the EGO4D-LTA challenge; and (iii) we determine tasks which require longer video contexts and that can thus be used effectively for further evaluation of long-range video models.
๋น๋์ค๋ ๊ณ ๋๋ก ์ค๋ณต๋ ๋ฐ์ดํฐ ์์ค์ด๋ฉฐ, ์ฃผ์ด์ง ์์
์ ํด๊ฒฐํ๊ธฐ ์ํด ๋ช ๊ฐ์ง ํต์ฌ ์๊ฐ์ ์๋ณํ๋ ๊ฒ๋ง์ผ๋ก ์ข
์ข
์ถฉ๋ถํฉ๋๋ค. ์ด ๋
ผ๋ฌธ์์, ์ฐ๋ฆฌ๋ ๊ธด ๋น๋์ค ์ํ์ค๋ฅผ ์์
์ ๋ง๊ฒ ์ฒ๋ฆฌํ๊ธฐ ์ํด ์ฌ์ ํ๋ จ๋ ๋น์ฃผ์ผ ์ธ์ฝ๋์ ๋๊ท๋ชจ ์ธ์ด ๋ชจ๋ธ(LLM)์ ์ฌ์ฉํ๋ ํ
์คํธ ์กฐ๊ฑด๋ถ ๋น๋์ค ๋ฆฌ์ํ๋ฌ(TCR) ๋ชจ๋์ ์ ์ํฉ๋๋ค. TCR์ ํ
์คํธ ์กฐ๊ฑด์ด ์ฃผ์ด์ง ๋น๋์ค์์ ๊ด๋ จ ๋น์ฃผ์ผ ํน์ง์ ์ฐพ์๋ด๊ณ , ์ด๋ฅผ LLM์ ์ ๊ณตํ์ฌ ํ
์คํธ ์๋ต์ ์์ฑํฉ๋๋ค. ๊ฐ๋ฒผ์ด ์ค๊ณ์ ํฌ๋ก์ค ์ดํ
์
์ ์ฌ์ฉ์ผ๋ก ์ธํด, TCR์ ํ ๋ฒ์ 100๊ฐ ์ด์์ ํ๋ ์์ ์ฒ๋ฆฌํ ์ ์์ผ๋ฉฐ, ์ด๋ ์ด์ ์์
๋ค๋ณด๋ค ํจ์ฌ ๊ธด ๋น๋์ค ์กฐ๊ฐ์ ์ฌ์ฉํ ์ ์๊ฒ ํฉ๋๋ค. ์ฐ๋ฆฌ๋ ๋ค์๊ณผ ๊ฐ์ ๊ธฐ์ฌ๋ฅผ ํฉ๋๋ค: (i) ์์
์ ๋ฐ๋ผ ๊ธด ๋น๋์ค๋ฅผ ์ฒ๋ฆฌํ ์ ์๋ ํธ๋์คํฌ๋จธ ๊ธฐ๋ฐ ์ํ๋ง ์ํคํ
์ฒ๋ฅผ ์ค๊ณํ๊ณ , ์ฌ์ ํ๋ จ๋ ๋น์ฃผ์ผ ๋ฐ ์ธ์ด ๋ชจ๋ธ ๊ฐ์ ์ฐ๊ฒฐ์ ๊ฐ๋ฅํ๊ฒ ํ๋ ํ๋ จ ๋ฐฉ๋ฒ์ ์ ๊ณตํฉ๋๋ค; (ii) ๋ค์ํ ํ๊ฐ ์์
์์ ๊ทธ ํจ๊ณผ๋ฅผ ์ค์ฆ์ ์ผ๋ก ๊ฒ์ฆํ๊ณ , NextQA, EgoSchema, EGO4D-LTA ์ฑ๋ฆฐ์ง์์ ์๋ก์ด ์ต๊ณ ๊ธฐ๋ก์ ์ธ์๋๋ค; ๊ทธ๋ฆฌ๊ณ (iii) ๋ ๊ธด ๋น๋์ค ์ปจํ
์คํธ๋ฅผ ํ์๋ก ํ๊ณ ๋ฐ๋ผ์ ์ฅ๊ฑฐ๋ฆฌ ๋น๋์ค ๋ชจ๋ธ์ ์ถ๊ฐ ํ๊ฐ์ ํจ๊ณผ์ ์ผ๋ก ์ฌ์ฉ๋ ์ ์๋ ์์
์ ๊ฒฐ์ ํฉ๋๋ค.
๋๊ธ 2