Google DeepMind
Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents. We show how rewards for visual achievement of a variety of language goals can be derived from the CLIP family of models, and used to train RL agents that can achieve a variety of language goals. We showcase this approach in two distinct visual domains and present a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.
ํ๋ถํ๊ณ ๊ฐ๋ฐฉ๋ ํ๊ฒฝ์์ ๋ง์ ๋ชฉํ๋ฅผ ๋ฌ์ฑํ ์ ์๋ ๋ฒ์ฉ ์์ด์ ํธ๋ฅผ ๊ตฌ์ถํ๋ ๊ฒ์ ๊ฐํ ํ์ต ์ฐ๊ตฌ์ ์ต์ ์ ์ค ํ๋์
๋๋ค. RL์ ์ฌ์ฉํ์ฌ ๋ฒ์ฉ ์์ด์ ํธ๋ฅผ ๊ตฌ์ถํ๋ ๋ฐ ์์ด ์ฃผ์ ์ ํ ์์ธ์ ๋ค์ํ ๋ชฉํ๋ฅผ ๋ฌ์ฑํ๊ธฐ ์ํ ๋ค์์ ๋ณด์ ํจ์๊ฐ ํ์ํ๋ค๋ ๊ฒ์ด์์ต๋๋ค. ์ฐ๋ฆฌ๋ ๊ฐํ ํ์ต ์์ด์ ํธ์ ๋ํ ๋ณด์์ ์์ฒ์ผ๋ก ํ์กดํ๋ ์๊ฐ-์ธ์ด ๋ชจ๋ธ, ๋๋ VLM๋ค์ ์ฌ์ฉํ๋ ๊ฐ๋ฅ์ฑ์ ์กฐ์ฌํฉ๋๋ค. ์ฐ๋ฆฌ๋ CLIP ๊ณ์ด ๋ชจ๋ธ์์ ํ์๋ ๋ค์ํ ์ธ์ด ๋ชฉํ์ ์๊ฐ์ ์ฑ์ทจ๋ฅผ ์ํ ๋ณด์์ด ์ด๋ป๊ฒ ๋์ถ๋ ์ ์๋์ง ๋ณด์ฌ์ฃผ๋ฉฐ, ์ด๋ฅผ ์ฌ์ฉํ์ฌ ๋ค์ํ ์ธ์ด ๋ชฉํ๋ฅผ ๋ฌ์ฑํ ์ ์๋ RL ์์ด์ ํธ๋ฅผ ํ๋ จ์ํต๋๋ค. ์ฐ๋ฆฌ๋ ์ด ์ ๊ทผ ๋ฐฉ์์ ๋ ๊ฐ์ง ๋ค๋ฅธ ์๊ฐ์ ์์ญ์์ ๋ณด์ฌ์ฃผ๊ณ , ๋ ํฐ VLM๋ค์ด ์๊ฐ์ ๋ชฉํ ๋ฌ์ฑ์ ์ํ ๋ณด๋ค ์ ํํ ๋ณด์์ผ๋ก ์ด์ด์ง๋ ๊ฒฝํฅ์ ๋ณด์ฌ์ฃผ๋ฉฐ, ์ด๋ ๋ค์ ๋ ์ ๋ฅํ RL ์์ด์ ํธ๋ฅผ ๋ง๋ค์ด๋
๋๋ค.
https://arxiv.org/abs/2312.09187
Vision-Language Models as a Source of RewardsBuilding generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents. We show how rewards for visual achievement of a variety of language goals can be derived from the CLIP family of models, and used to train RL agents that can achieve a variety of language goals. We showcase this approach in two distinct visual domains and present a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.arxiv.org
๋๊ธ 1