Google Research
Answering complex natural language questions often necessitates multi-step reasoning and integrating external information. Several systems have combined knowledge retrieval with a large language model (LLM) to answer such questions. These systems, however, suffer from various failure cases, and we cannot directly train them end-to-end to fix such failures, as interaction with external knowledge is non-differentiable. To address these deficiencies, we define a ReAct-style LLM agent with the ability to reason and act upon external knowledge. We further refine the agent through a ReST-like method that iteratively trains on previous trajectories, employing growing-batch reinforcement learning with AI feedback for continuous self-improvement and self-distillation. Starting from a prompted large model and after just two iterations of the algorithm, we can produce a fine-tuned small model that achieves comparable performance on challenging compositional question-answering benchmarks with two orders of magnitude fewer parameters.
๋ณต์กํ ์์ฐ์ด ์ง๋ฌธ์ ๋๋ตํ๋ ๊ฒ์ ์ข
์ข
๋ค๋จ๊ณ ์ถ๋ก ๊ณผ ์ธ๋ถ ์ ๋ณด์ ํตํฉ์ ํ์๋ก ํฉ๋๋ค. ์ฌ๋ฌ ์์คํ
๋ค์ด ๋๊ท๋ชจ ์ธ์ด ๋ชจ๋ธ(Large Language Model, LLM)๊ณผ ์ง์ ๊ฒ์์ ๊ฒฐํฉํ์ฌ ์ด๋ฌํ ์ง๋ฌธ์ ๋ตํด์์ต๋๋ค. ๊ทธ๋ฌ๋ ์ด๋ฌํ ์์คํ
๋ค์ ๋ค์ํ ์คํจ ์ฌ๋ก๋ฅผ ๊ฒช๊ณ ์์ผ๋ฉฐ, ์ธ๋ถ ์ง์๊ณผ์ ์ํธ์์ฉ์ด ๋ฏธ๋ถ ๋ถ๊ฐ๋ฅํ๊ธฐ ๋๋ฌธ์ ์ด๋ฌํ ์คํจ๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด ์ง์ ์ ์ผ๋ก ์ข
๋จ ๊ฐ(end-to-end)์ผ๋ก ํ๋ จํ ์ ์์ต๋๋ค. ์ด๋ฌํ ๊ฒฐ์ ๋ค์ ํด๊ฒฐํ๊ธฐ ์ํด, ์ฐ๋ฆฌ๋ ์ธ๋ถ ์ง์์ ์ดํดํ๊ณ ์ด์ ๋ํด ํ๋ํ ์ ์๋ ๋ฅ๋ ฅ์ ๊ฐ์ง ReAct ์คํ์ผ์ LLM ์์ด์ ํธ๋ฅผ ์ ์ํฉ๋๋ค. ์ฐ๋ฆฌ๋ ์ด ์์ด์ ํธ๋ฅผ ReST์ ๊ฐ์ ๋ฐฉ๋ฒ์ผ๋ก ๋์ฑ ์ ์ ํ์ฌ, ์ด์ ๊ถค์ ๋ค์ ๋ํด ๋ฐ๋ณต์ ์ผ๋ก ํ๋ จํ๊ณ , ์ง์์ ์ธ ์๊ธฐ ๊ฐ์ ๊ณผ ์๊ฐ ์ฆ๋ฅ๋ฅผ ์ํด ์ฑ์ฅ ๋ฐฐ์น(growing-batch) ๊ฐํ ํ์ต๊ณผ ์ธ๊ณต์ง๋ฅ ํผ๋๋ฐฑ์ ์ฌ์ฉํฉ๋๋ค. ์ฒ์์๋ ํฐ ๋ชจ๋ธ๋ก ์์ํ์ฌ ์๊ณ ๋ฆฌ์ฆ์ ๋จ ๋ ๋ฒ์ ๋ฐ๋ณต ํ์, ์ฐ๋ฆฌ๋ ๋งค๊ฐ๋ณ์๊ฐ ๋ ๊ฐ์ ํฌ๊ธฐ ์์๋ก ์ ์ ์ธ๋ฐํ๊ฒ ์กฐ์ ๋ ์์ ๋ชจ๋ธ์ ์์ฑํ ์ ์๋๋ฐ, ์ด๋ ๋์ ์ ์ธ ๊ตฌ์ฑ์ ์ง๋ฌธ-์๋ต ๋ฒค์น๋งํฌ์์ ๋น๊ต ๊ฐ๋ฅํ ์ฑ๋ฅ์ ๋ฌ์ฑํฉ๋๋ค.
key contributions
โข We build a flavor of ReAct agent with self-critique for the task of long-form question answering.
โข We define a proxy evaluation metric for the agent based on Bamboogle and BamTwoogle datasets, with a strong emphasis on auto-eval.
โข We demonstrate that the performance of the agent could be effectively improved through Rest-style iterative fine-tuning on its reasoning traces.
โข Furthermore, we do it purely from stepwise AI feedback without using human-labeled training data.
โข Finally, we show that the synthetic data produced as part of this iterative process could
be used for distilling the agent into one or two orders of magnitude smaller models with
performance comparable to the pre-trained teacher agent.
์ฃผ์ ๊ธฐ์ฌ ์ฌํญ:
โข ์ฐ๋ฆฌ๋ ์ฅ๋ฌธํ ์ง๋ฌธ ์๋ต ์์ ์ ์ํ ์๊ฐ ๋นํ ๊ธฐ๋ฅ์ ๊ฐ์ถ ReAct ์์ด์ ํธ์ ํ ํํ๋ฅผ ๊ตฌ์ถํฉ๋๋ค.
โข ์ฐ๋ฆฌ๋ Bamboogle ๋ฐ BamTwoogle ๋ฐ์ดํฐ์ ์ ๊ธฐ๋ฐ์ผ๋ก ํ ์์ด์ ํธ์ ๋๋ฆฌ ํ๊ฐ ๋ฉํธ๋ฆญ์ ์ ์ํฉ๋๋ค. ์ด ํ๊ฐ๋ ์๋ ํ๊ฐ(auto-eval)์ ํฐ ์ค์ ์ ๋๊ณ ์์ต๋๋ค.
โข ์ฐ๋ฆฌ๋ ์์ด์ ํธ์ ์ถ๋ก ๊ถค์ ์ ๋ํ ReST ์คํ์ผ์ ๋ฐ๋ณต์ ์ธ ๋ฏธ์ธ ์กฐ์ ์ ํตํด ์์ด์ ํธ์ ์ฑ๋ฅ์ ํจ๊ณผ์ ์ผ๋ก ํฅ์์ํฌ ์ ์์์ ๋ณด์ฌ์ค๋๋ค.
โข ๋์ฑ์ด, ์ฐ๋ฆฌ๋ ์ธ๊ฐ ๋ ์ด๋ธ์ด ๋ถ์ ํ๋ จ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ์ง ์๊ณ ์์ํ๊ฒ ๋จ๊ณ๋ณ AI ํผ๋๋ฐฑ์ ํตํด ์ด๋ฅผ ์ํํฉ๋๋ค.
โข ๋ง์ง๋ง์ผ๋ก, ์ฐ๋ฆฌ๋ ์ด ๋ฐ๋ณต์ ์ธ ๊ณผ์ ์ ์ผ๋ถ๋ก ์์ฑ๋ ํฉ์ฑ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ์ฌ ์์ด์ ํธ๋ฅผ ์ฌ์ ํ๋ จ๋ ๊ต์ฌ ์์ด์ ํธ์ ๋น๊ตํ ์ ์๋ ์ฑ๋ฅ์ ๊ฐ์ง ํจ์ฌ ๋ ์์ ๋ชจ๋ธ๋ก ์ฆ๋ฅํ ์ ์์์ ๋ณด์ฌ์ค๋๋ค.
https://huggingface.co/papers/2312.10003
Paper page - ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM AgentPaper page - ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agenthuggingface.co
LLM์ ์๊ฐ ๋นํ์ด ๋ถ๊ฐ๋ฅ ํ๋ค๋ ๋ง๋ ์์๋๋ฐ ํด๊ฒฐ ๋๊ฑด๊ฐ
์ด ๊ธ ํ์ ์ง๊ฐ ์ด๋์ผ?