https://arxiv.org/abs/2406.03816

1. "Recent methodologies in LLM self-training mostly rely on LLM generating responses and filtering those with correct output answers as training data."

- ์ตœ๊ทผ LLM ์ž๊ธฐ ํ›ˆ๋ จ ๋ฐฉ๋ฒ•๋ก ์€ ์ฃผ๋กœ LLM์ด ์‘๋‹ต์„ ์ƒ์„ฑํ•˜๊ณ  ์˜ฌ๋ฐ”๋ฅธ ์ถœ๋ ฅ ๋‹ต๋ณ€์ด ํฌํ•จ๋œ ๊ฒƒ์„ ํ•„ํ„ฐ๋งํ•˜์—ฌ ํ›ˆ๋ จ ๋ฐ์ดํ„ฐ๋กœ ์‚ฌ์šฉํ•˜๋Š” ๊ฒƒ์— ์˜์กดํ•ฉ๋‹ˆ๋‹ค.


2. "This approach often yields a low-quality fine-tuning training set (e.g., incorrect plans or intermediate reasoning)."

- ์ด๋Ÿฌํ•œ ์ ‘๊ทผ ๋ฐฉ์‹์€ ์ข…์ข… ์ €ํ’ˆ์งˆ์˜ ๋ฏธ์„ธ ์กฐ์ • ํ›ˆ๋ จ ์„ธํŠธ(์˜ˆ: ์ž˜๋ชป๋œ ๊ณ„ํš ๋˜๋Š” ์ค‘๊ฐ„ ์ถ”๋ก )๋ฅผ ์ƒ์„ฑํ•ฉ๋‹ˆ๋‹ค.


3. "In this paper, we develop a reinforced self-training approach, called ReST-MCTS*, based on integrating process reward guidance with tree search MCTS* for collecting higher-quality reasoning traces as well as per-step value to train policy and reward models."

- ๋ณธ ๋…ผ๋ฌธ์—์„œ๋Š” ์ •์ฑ… ๋ฐ ๋ณด์ƒ ๋ชจ๋ธ์„ ํ›ˆ๋ จํ•˜๊ธฐ ์œ„ํ•ด ๋‹จ๊ณ„๋ณ„ ๊ฐ€์น˜๋ฅผ ํฌํ•จํ•˜์—ฌ ๋” ๋†’์€ ํ’ˆ์งˆ์˜ ์ถ”๋ก  ํ”์ ์„ ์ˆ˜์ง‘ํ•˜๊ธฐ ์œ„ํ•ด ๊ณผ์ • ๋ณด์ƒ ์ง€์นจ๊ณผ ํŠธ๋ฆฌ ํƒ์ƒ‰ MCTS*๋ฅผ ํ†ตํ•ฉํ•œ ReST-MCTS*๋ผ๋Š” ๊ฐ•ํ™”๋œ ์ž๊ธฐ ํ›ˆ๋ จ ์ ‘๊ทผ ๋ฐฉ์‹์„ ๊ฐœ๋ฐœํ•ฉ๋‹ˆ๋‹ค.


4. "ReST-MCTS* circumvents the per-step manual annotation typically used to train process rewards by tree-search-based reinforcement learning: Given oracle final correct answers, ReST-MCTS* is able to infer the correct process rewards by estimating the probability this step can help lead to the correct answer."

- ReST-MCTS*๋Š” ์ผ๋ฐ˜์ ์œผ๋กœ ๊ณผ์ • ๋ณด์ƒ์„ ํ›ˆ๋ จํ•˜๊ธฐ ์œ„ํ•ด ์‚ฌ์šฉ๋˜๋Š” ๋‹จ๊ณ„๋ณ„ ์ˆ˜๋™ ์ฃผ์„์„ ํŠธ๋ฆฌ ํƒ์ƒ‰ ๊ธฐ๋ฐ˜ ๊ฐ•ํ™” ํ•™์Šต์œผ๋กœ ์šฐํšŒํ•ฉ๋‹ˆ๋‹ค: ์ตœ์ข… ์˜ฌ๋ฐ”๋ฅธ ๋‹ต์ด ์ฃผ์–ด์ง€๋ฉด, ReST-MCTS*๋Š” ์ด ๋‹จ๊ณ„๊ฐ€ ์˜ฌ๋ฐ”๋ฅธ ๋‹ต์œผ๋กœ ์ด๋Œ ๊ฐ€๋Šฅ์„ฑ์„ ์ถ”์ •ํ•˜์—ฌ ์˜ฌ๋ฐ”๋ฅธ ๊ณผ์ • ๋ณด์ƒ์„ ์ถ”๋ก ํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.


5. "These inferred rewards serve dual purposes: they act as value targets for further refining the process reward model and also facilitate the selection of high-quality traces for policy model self-training."

- ์ด๋Ÿฌํ•œ ์ถ”๋ก ๋œ ๋ณด์ƒ์€ ์ด์ค‘ ๋ชฉ์ ์„ ์ˆ˜ํ–‰ํ•ฉ๋‹ˆ๋‹ค: ๊ณผ์ • ๋ณด์ƒ ๋ชจ๋ธ์„ ๋”์šฑ ์ •์ œํ•˜๊ธฐ ์œ„ํ•œ ๊ฐ€์น˜ ๋ชฉํ‘œ๋กœ ์ž‘์šฉํ•˜๋ฉฐ, ๋˜ํ•œ ์ •์ฑ… ๋ชจ๋ธ ์ž๊ธฐ ํ›ˆ๋ จ์„ ์œ„ํ•œ ๊ณ ํ’ˆ์งˆ ํ”์  ์„ ํƒ์„ ์šฉ์ดํ•˜๊ฒŒ ํ•ฉ๋‹ˆ๋‹ค.


6. "We first show that the tree-search policy in ReST-MCTS* achieves higher accuracy compared with prior LLM reasoning baselines such as Best-of-N and Tree-of-Thought, within the same search budget."

- ์šฐ๋ฆฌ๋Š” ๋จผ์ € ReST-MCTS*์˜ ํŠธ๋ฆฌ ํƒ์ƒ‰ ์ •์ฑ…์ด ๋™์ผํ•œ ํƒ์ƒ‰ ์˜ˆ์‚ฐ ๋‚ด์—์„œ Best-of-N ๋ฐ Tree-of-Thought์™€ ๊ฐ™์€ ์ด์ „ LLM ์ถ”๋ก  ๊ธฐ์ค€๋ณด๋‹ค ๋” ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋‹ฌ์„ฑํ•จ์„ ๋ณด์—ฌ์ค๋‹ˆ๋‹ค.


7. "We then show that by using traces searched by this tree-search policy as training data, we can continuously enhance the three language models for multiple iterations, and outperform other self-training algorithms such as ReSTEM and Self-Rewarding LM."

- ๊ทธ๋Ÿฐ ๋‹ค์Œ ์ด ํŠธ๋ฆฌ ํƒ์ƒ‰ ์ •์ฑ…์— ์˜ํ•ด ๊ฒ€์ƒ‰๋œ ํ”์ ์„ ํ›ˆ๋ จ ๋ฐ์ดํ„ฐ๋กœ ์‚ฌ์šฉํ•จ์œผ๋กœ์จ ์—ฌ๋Ÿฌ ๋ฒˆ์˜ ๋ฐ˜๋ณต์„ ํ†ตํ•ด ์„ธ ๊ฐ€์ง€ ์–ธ์–ด ๋ชจ๋ธ์„ ์ง€์†์ ์œผ๋กœ ํ–ฅ์ƒ์‹œํ‚ฌ ์ˆ˜ ์žˆ์œผ๋ฉฐ, ReSTEM ๋ฐ Self-Rewarding LM๊ณผ ๊ฐ™์€ ๋‹ค๋ฅธ ์ž๊ธฐ ํ›ˆ๋ จ ์•Œ๊ณ ๋ฆฌ์ฆ˜์„ ๋Šฅ๊ฐ€ํ•จ์„ ๋ณด์—ฌ์ค๋‹ˆ๋‹ค.