Making Large Language Models Better Reasoners with AlignmentReasoning is a cognitive process of using evidence to reach a soundconclusion. The reasoning capability is essential for large language models(LLMs) to serve as the brain of the artificial general intelligence agent.Recent studies reveal that fine-tuning LLMs on data with the chain of thought(COT) reasoning process can significantly enhance their reasoning capabilities.However, we find that the fine-tuned LLMs suffer from an \textit{AssessmentMisalignment} problem, i.e., they frequently assign higher scores to subparCOTs, leading to potential limitations in their reasoning abilities. To addressthis problem, we introduce an \textit{Alignment Fine-Tuning (AFT)} paradigm,which involves three steps: 1) fine-tuning LLMs with COT training data; 2)generating multiple COT responses for each question, and categorizing them intopositive and negative ones based on whether they achieve the correct answer; 3)calibrating the scores of positive and negative responses given by LLMs with anovel constraint alignment loss. Specifically, the constraint alignment losshas two ives: a) Alignment, which guarantees that positive scores surpassnegative scores to encourage answers with high-quality COTs; b) Constraint,which keeps the negative scores confined to a reasonable range to prevent themodel degradation. Beyond just the binary positive and negative feedback, theconstraint alignment loss can be seamlessly adapted to the ranking situationswhen ranking feedback is accessible. Furthermore, we also delve deeply intorecent ranking-based alignment methods, such as DPO, RRHF, and PRO, anddiscover that the constraint, which has been overlooked by these approaches, isalso crucial for their performance. Extensive experiments on four reasoningbenchmarks with both binary and ranking feedback demonstrate the effectivenessof AFT.arxiv.org


์ถ”๋ก ์€ ์ธ๊ณต ์ผ๋ฐ˜ ์ง€๋Šฅ ์—์ด์ „ํŠธ์˜ ์ค‘์š”ํ•œ ํ”„๋กœ์„ธ์Šค์ด๋ฉฐ, ์ตœ๊ทผ ์—ฐ๊ตฌ์— ๋”ฐ๋ฅด๋ฉด COT(์‚ฌ๊ณ  ์‚ฌ์Šฌ) ์ถ”๋ก  ํ”„๋กœ์„ธ์Šค๋ฅผ ํ†ตํ•ฉํ•œ ๋ฐ์ดํ„ฐ๋กœ LLM(๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ)์„ ๋ฏธ์„ธ ์กฐ์ •ํ•˜๋ฉด ์ถ”๋ก  ๋Šฅ๋ ฅ์„ ํ–ฅ์ƒ์‹œํ‚ฌ ์ˆ˜ ์žˆ๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ ์ด๋Ÿฌํ•œ ๋ฏธ์„ธ ์กฐ์ • ํ”„๋กœ์„ธ์Šค๋Š” LLM์ด ์ˆ˜์ค€ ์ดํ•˜์˜ COT์— ๋” ๋†’์€ ์ ์ˆ˜๋ฅผ ํ• ๋‹นํ•˜์—ฌ ์ถ”๋ก  ๊ธฐ๋Šฅ์„ ์ œํ•œํ•˜๋Š” ํ‰๊ฐ€ ๋ถˆ์ผ์น˜ ๋ฌธ์ œ๋กœ ์ด์–ด์ง€๋Š” ๊ฒฝ์šฐ๊ฐ€ ๋งŽ์Šต๋‹ˆ๋‹ค. ์ด ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด ์—ฐ๊ตฌ์›๋“ค์€ ์„ธ ๋‹จ๊ณ„๋ฅผ ํฌํ•จํ•˜๋Š” AFT(Alignment Fine-Tuning) ํŒจ๋Ÿฌ๋‹ค์ž„์„ ๋„์ž…ํ–ˆ์Šต๋‹ˆ๋‹ค. ์ฒซ์งธ, LLM์€ COT ๊ต์œก ๋ฐ์ดํ„ฐ๋กœ ๋ฏธ์„ธ ์กฐ์ •๋ฉ๋‹ˆ๋‹ค. ๋‘˜์งธ, ๊ฐ ์งˆ๋ฌธ์— ๋Œ€ํ•ด ์—ฌ๋Ÿฌ COT ์‘๋‹ต์ด ์ƒ์„ฑ๋˜๊ณ  ์ •ํ™•์„ฑ์— ๋”ฐ๋ผ ๊ธ์ •์  ๋˜๋Š” ๋ถ€์ •์ ์œผ๋กœ ๋ถ„๋ฅ˜๋ฉ๋‹ˆ๋‹ค. ๋งˆ์ง€๋ง‰์œผ๋กœ, ๊ธ์ •์ ์ธ ๋ฐ˜์‘๊ณผ ๋ถ€์ •์ ์ธ ๋ฐ˜์‘์˜ ์ ์ˆ˜๋Š” ๊ณ ํ’ˆ์งˆ COT๊ฐ€ ๋” ๋†’์€ ์ ์ˆ˜๋ฅผ ๋ฐ›๊ณ  ๋ชจ๋ธ ์ €ํ•˜๋ฅผ ๋ฐฉ์ง€ํ•˜๋Š” ์ƒˆ๋กœ์šด ์ œ์•ฝ ์กฐ๊ฑด ์ •๋ ฌ ์†์‹ค์„ ์‚ฌ์šฉํ•˜์—ฌ ๋ณด์ •๋ฉ๋‹ˆ๋‹ค. ์ œ์•ฝ ์กฐ๊ฑด ์ •๋ ฌ ์†์‹ค์€ ์ˆœ์œ„ ์‹œ๋‚˜๋ฆฌ์˜ค์—๋„ ์ ์šฉํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. AFT์˜ ํšจ์œจ์„ฑ์€ ๋ฐ”์ด๋„ˆ๋ฆฌ ๋ฐ ์ˆœ์œ„ ํ”ผ๋“œ๋ฐฑ์„ ๋ชจ๋‘ ์‚ฌ์šฉํ•˜๋Š” 4๊ฐ€์ง€ ์ถ”๋ก  ๋ฒค์น˜๋งˆํฌ์— ๋Œ€ํ•œ ๊ด‘๋ฒ”์œ„ํ•œ ์‹คํ—˜์„ ํ†ตํ•ด ์ž…์ฆ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๋˜ํ•œ ์ด ์—ฐ๊ตฌ์—์„œ๋Š” ์ตœ๊ทผ ์ˆœ์œ„ ๊ธฐ๋ฐ˜ ์ •๋ ฌ ๋ฐฉ๋ฒ•์„ ํƒ์ƒ‰ํ•˜๊ณ  ์„ฑ๊ณผ์— ๋Œ€ํ•œ ์ œ์•ฝ์˜ ์ค‘์š”์„ฑ์„ ๊ฐ•์กฐํ•ฉ๋‹ˆ๋‹ค.

- dc official App