Vision-Language Models as a Source of Rewards

Google DeepMind


Abstract

Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents. We show how rewards for visual achievement of a variety of language goals can be derived from the CLIP family of models, and used to train RL agents that can achieve a variety of language goals. We showcase this approach in two distinct visual domains and present a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.


ํ’๋ถ€ํ•˜๊ณ  ๊ฐœ๋ฐฉ๋œ ํ™˜๊ฒฝ์—์„œ ๋งŽ์€ ๋ชฉํ‘œ๋ฅผ ๋‹ฌ์„ฑํ•  ์ˆ˜ ์žˆ๋Š” ๋ฒ”์šฉ ์—์ด์ „ํŠธ๋ฅผ ๊ตฌ์ถ•ํ•˜๋Š” ๊ฒƒ์€ ๊ฐ•ํ™” ํ•™์Šต ์—ฐ๊ตฌ์˜ ์ตœ์ „์„  ์ค‘ ํ•˜๋‚˜์ž…๋‹ˆ๋‹ค. RL์„ ์‚ฌ์šฉํ•˜์—ฌ ๋ฒ”์šฉ ์—์ด์ „ํŠธ๋ฅผ ๊ตฌ์ถ•ํ•˜๋Š” ๋ฐ ์žˆ์–ด ์ฃผ์š” ์ œํ•œ ์š”์ธ์€ ๋‹ค์–‘ํ•œ ๋ชฉํ‘œ๋ฅผ ๋‹ฌ์„ฑํ•˜๊ธฐ ์œ„ํ•œ ๋‹ค์ˆ˜์˜ ๋ณด์ƒ ํ•จ์ˆ˜๊ฐ€ ํ•„์š”ํ•˜๋‹ค๋Š” ๊ฒƒ์ด์—ˆ์Šต๋‹ˆ๋‹ค. ์šฐ๋ฆฌ๋Š” ๊ฐ•ํ™” ํ•™์Šต ์—์ด์ „ํŠธ์— ๋Œ€ํ•œ ๋ณด์ƒ์˜ ์›์ฒœ์œผ๋กœ ํ˜„์กดํ•˜๋Š” ์‹œ๊ฐ-์–ธ์–ด ๋ชจ๋ธ, ๋˜๋Š” VLM๋“ค์„ ์‚ฌ์šฉํ•˜๋Š” ๊ฐ€๋Šฅ์„ฑ์„ ์กฐ์‚ฌํ•ฉ๋‹ˆ๋‹ค. ์šฐ๋ฆฌ๋Š” CLIP ๊ณ„์—ด ๋ชจ๋ธ์—์„œ ํŒŒ์ƒ๋œ ๋‹ค์–‘ํ•œ ์–ธ์–ด ๋ชฉํ‘œ์˜ ์‹œ๊ฐ์  ์„ฑ์ทจ๋ฅผ ์œ„ํ•œ ๋ณด์ƒ์ด ์–ด๋–ป๊ฒŒ ๋„์ถœ๋  ์ˆ˜ ์žˆ๋Š”์ง€ ๋ณด์—ฌ์ฃผ๋ฉฐ, ์ด๋ฅผ ์‚ฌ์šฉํ•˜์—ฌ ๋‹ค์–‘ํ•œ ์–ธ์–ด ๋ชฉํ‘œ๋ฅผ ๋‹ฌ์„ฑํ•  ์ˆ˜ ์žˆ๋Š” RL ์—์ด์ „ํŠธ๋ฅผ ํ›ˆ๋ จ์‹œํ‚ต๋‹ˆ๋‹ค. ์šฐ๋ฆฌ๋Š” ์ด ์ ‘๊ทผ ๋ฐฉ์‹์„ ๋‘ ๊ฐ€์ง€ ๋‹ค๋ฅธ ์‹œ๊ฐ์  ์˜์—ญ์—์„œ ๋ณด์—ฌ์ฃผ๊ณ , ๋” ํฐ VLM๋“ค์ด ์‹œ๊ฐ์  ๋ชฉํ‘œ ๋‹ฌ์„ฑ์„ ์œ„ํ•œ ๋ณด๋‹ค ์ •ํ™•ํ•œ ๋ณด์ƒ์œผ๋กœ ์ด์–ด์ง€๋Š” ๊ฒฝํ–ฅ์„ ๋ณด์—ฌ์ฃผ๋ฉฐ, ์ด๋Š” ๋‹ค์‹œ ๋” ์œ ๋Šฅํ•œ RL ์—์ด์ „ํŠธ๋ฅผ ๋งŒ๋“ค์–ด๋ƒ…๋‹ˆ๋‹ค.








https://arxiv.org/abs/2312.09187

Vision-Language Models as a Source of RewardsBuilding generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for building generalist agents with RL has been the need for a large number of reward functions for achieving different goals. We investigate the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents. We show how rewards for visual achievement of a variety of language goals can be derived from the CLIP family of models, and used to train RL agents that can achieve a variety of language goals. We showcase this approach in two distinct visual domains and present a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.arxiv.org