- ์ต๊ทผ LLM ์๊ธฐ ํ๋ จ ๋ฐฉ๋ฒ๋ก ์ ์ฃผ๋ก LLM์ด ์๋ต์ ์์ฑํ๊ณ ์ฌ๋ฐ๋ฅธ ์ถ๋ ฅ ๋ต๋ณ์ด ํฌํจ๋ ๊ฒ์ ํํฐ๋งํ์ฌ ํ๋ จ ๋ฐ์ดํฐ๋ก ์ฌ์ฉํ๋ ๊ฒ์ ์์กดํฉ๋๋ค.
2. "This approach often yields a low-quality fine-tuning training set (e.g., incorrect plans or intermediate reasoning)."
- ์ด๋ฌํ ์ ๊ทผ ๋ฐฉ์์ ์ข ์ข ์ ํ์ง์ ๋ฏธ์ธ ์กฐ์ ํ๋ จ ์ธํธ(์: ์๋ชป๋ ๊ณํ ๋๋ ์ค๊ฐ ์ถ๋ก )๋ฅผ ์์ฑํฉ๋๋ค.
3. "In this paper, we develop a reinforced self-training approach, called ReST-MCTS*, based on integrating process reward guidance with tree search MCTS* for collecting higher-quality reasoning traces as well as per-step value to train policy and reward models."
- ๋ณธ ๋ ผ๋ฌธ์์๋ ์ ์ฑ ๋ฐ ๋ณด์ ๋ชจ๋ธ์ ํ๋ จํ๊ธฐ ์ํด ๋จ๊ณ๋ณ ๊ฐ์น๋ฅผ ํฌํจํ์ฌ ๋ ๋์ ํ์ง์ ์ถ๋ก ํ์ ์ ์์งํ๊ธฐ ์ํด ๊ณผ์ ๋ณด์ ์ง์นจ๊ณผ ํธ๋ฆฌ ํ์ MCTS*๋ฅผ ํตํฉํ ReST-MCTS*๋ผ๋ ๊ฐํ๋ ์๊ธฐ ํ๋ จ ์ ๊ทผ ๋ฐฉ์์ ๊ฐ๋ฐํฉ๋๋ค.
4. "ReST-MCTS* circumvents the per-step manual annotation typically used to train process rewards by tree-search-based reinforcement learning: Given oracle final correct answers, ReST-MCTS* is able to infer the correct process rewards by estimating the probability this step can help lead to the correct answer."
- ReST-MCTS*๋ ์ผ๋ฐ์ ์ผ๋ก ๊ณผ์ ๋ณด์์ ํ๋ จํ๊ธฐ ์ํด ์ฌ์ฉ๋๋ ๋จ๊ณ๋ณ ์๋ ์ฃผ์์ ํธ๋ฆฌ ํ์ ๊ธฐ๋ฐ ๊ฐํ ํ์ต์ผ๋ก ์ฐํํฉ๋๋ค: ์ต์ข ์ฌ๋ฐ๋ฅธ ๋ต์ด ์ฃผ์ด์ง๋ฉด, ReST-MCTS*๋ ์ด ๋จ๊ณ๊ฐ ์ฌ๋ฐ๋ฅธ ๋ต์ผ๋ก ์ด๋ ๊ฐ๋ฅ์ฑ์ ์ถ์ ํ์ฌ ์ฌ๋ฐ๋ฅธ ๊ณผ์ ๋ณด์์ ์ถ๋ก ํ ์ ์์ต๋๋ค.
5. "These inferred rewards serve dual purposes: they act as value targets for further refining the process reward model and also facilitate the selection of high-quality traces for policy model self-training."
- ์ด๋ฌํ ์ถ๋ก ๋ ๋ณด์์ ์ด์ค ๋ชฉ์ ์ ์ํํฉ๋๋ค: ๊ณผ์ ๋ณด์ ๋ชจ๋ธ์ ๋์ฑ ์ ์ ํ๊ธฐ ์ํ ๊ฐ์น ๋ชฉํ๋ก ์์ฉํ๋ฉฐ, ๋ํ ์ ์ฑ ๋ชจ๋ธ ์๊ธฐ ํ๋ จ์ ์ํ ๊ณ ํ์ง ํ์ ์ ํ์ ์ฉ์ดํ๊ฒ ํฉ๋๋ค.
6. "We first show that the tree-search policy in ReST-MCTS* achieves higher accuracy compared with prior LLM reasoning baselines such as Best-of-N and Tree-of-Thought, within the same search budget."
- ์ฐ๋ฆฌ๋ ๋จผ์ ReST-MCTS*์ ํธ๋ฆฌ ํ์ ์ ์ฑ ์ด ๋์ผํ ํ์ ์์ฐ ๋ด์์ Best-of-N ๋ฐ Tree-of-Thought์ ๊ฐ์ ์ด์ LLM ์ถ๋ก ๊ธฐ์ค๋ณด๋ค ๋ ๋์ ์ ํ๋๋ฅผ ๋ฌ์ฑํจ์ ๋ณด์ฌ์ค๋๋ค.
7. "We then show that by using traces searched by this tree-search policy as training data, we can continuously enhance the three language models for multiple iterations, and outperform other self-training algorithms such as ReSTEM and Self-Rewarding LM."
- ๊ทธ๋ฐ ๋ค์ ์ด ํธ๋ฆฌ ํ์ ์ ์ฑ ์ ์ํด ๊ฒ์๋ ํ์ ์ ํ๋ จ ๋ฐ์ดํฐ๋ก ์ฌ์ฉํจ์ผ๋ก์จ ์ฌ๋ฌ ๋ฒ์ ๋ฐ๋ณต์ ํตํด ์ธ ๊ฐ์ง ์ธ์ด ๋ชจ๋ธ์ ์ง์์ ์ผ๋ก ํฅ์์ํฌ ์ ์์ผ๋ฉฐ, ReSTEM ๋ฐ Self-Rewarding LM๊ณผ ๊ฐ์ ๋ค๋ฅธ ์๊ธฐ ํ๋ จ ์๊ณ ๋ฆฌ์ฆ์ ๋ฅ๊ฐํจ์ ๋ณด์ฌ์ค๋๋ค.
๋๊ธ 1