์ดˆ์ง€๋Šฅ์œผ๋กœ ํ–ฅํ•˜๋Š” ๊ธธ, GPT-4 ํ›ˆ๋ จ์‹œํ‚ค๋Š” CriticGPT
OpenAI๋Š” GPT-4์˜ ์˜ค๋ฅ˜๋ฅผ ์ฐพ๊ธฐ ์œ„ํ•ด CriticGPT๋ผ๋Š” ๋ชจ๋ธ ๊ฐœ๋ฐœ


CriticGPT๋Š” GPT-4 ๋ฒ ์ด์Šค๋กœ ChatGPT์˜ ์ฝ”๋“œ๋ฅผ ๋น„ํ‰ํ•ด์ฃผ๋Š” GPT
CriticGPT๋Š” GPT-4์˜ ์‘๋‹ต์„ ๊ฒ€ํ† ํ•˜๊ณ  ์˜ค๋ฅ˜๋ฅผ ์ง€์ ํ•˜๋Š” ์—ญํ• ์„ ํ•จ
์—ฐ๊ตฌ ๊ฒฐ๊ณผ, CriticGPT๊ฐ€ ์‚ฌ๋žŒ๊ณผ ํ˜‘๋ ฅํ–ˆ์„ ๋•Œ ์˜ค๋ฅ˜๋ฅผ ๋” ์ •ํ™•ํ•˜๊ฒŒ ์ฐพ์•„๋ƒ„
CriticGPT๋Š” ์‘๋‹ต์˜ ๋…ผ๋ฆฌ์  ์ผ๊ด€์„ฑ์„ ๊ฒ€ํ† ํ•˜๊ณ , ๋ชจํ˜ธํ•œ ํ‘œํ˜„์„ ์ง€์ ํ•˜๋ฉฐ, ์ •๋ณด์˜ ์ •ํ™•์„ฑ์„ ํ‰๊ฐ€ํ•จ


ํ•˜์ง€๋งŒ ๊ธด ํ…์ŠคํŠธ๋‚˜ ๋ณต์žกํ•œ ์˜ค๋ฅ˜๋ฅผ ์ฒ˜๋ฆฌํ•˜๋Š” ๋ฐ๋Š” ํ•œ๊ณ„๊ฐ€ ์žˆ์Œ
OpenAi์ธก์€ ์•ž์œผ๋กœ ์ด ๋ชจ๋ธ์„ ๊ฐœ์„ ํ•ด ๋‚˜๊ฐˆ ๊ณ„ํš์ด๋ผ ๋ฐํž˜

์ตœ๊ทผ Frontier AI ์—ฐ๊ตฌ์†Œ๋“ค์—์„œ ์ œ์ผ ๋งŽ์ด ๋…ผ์˜๋˜๋Š” ์ฃผ์ œ ์ค‘ ํ•˜๋‚˜๋Š” ๊ฒฐ๊ตญ ์–ด๋А ์ˆœ๊ฐ„ ๊ฐ์ข… ๋ณ‘๋ชฉ์„ ๋šซ๊ณ  AGI โ†’ ASI๋กœ ๊ฐ€๊ธฐ ์œ„ํ•ด AI๋กœ AI๋ฅผ ๊ฐ€๋ฅด์น˜๋Š” ๋ฃจํ”„๋ฅผ ์—ฐ๊ฒฐํ•ด์•ผ ํ•œ๋‹ค๋Š” ์ฃผ์ œ


์˜คํ”ˆAI์˜ ๋ณธ๋ฌธ์—๋„ ๋‚˜์™€์žˆ๋“ฏ ์ด๋Š” ์•ž์œผ๋กœ ์ดˆ์ง€๋Šฅ์˜ ์‹œ๋Œ€์— '์ธ๊ฐ„๋ณด๋‹ค ๋” ๋†’์€ ์ง€๋Šฅ์„ ์–ด๋–ป๊ฒŒ ๊ด€๋ฆฌ๊ฐ๋…ํ•  ๊ฒƒ์ธ๊ฐ€'์— ๋Œ€ํ•œ ์ฃผ์ œ์— ๋Œ€ํ•œ ์—ฐ๊ตฌ
*This is a step towards being able to evaluate outputs from advanced AI systems that can be difficult for people to rate without better tools.*

์ง€๊ธˆ๊นŒ์ง€ ๊ด‘๋ฒ”์œ„ํ•˜๊ฒŒ ํ˜์‹ ์„ ๊ฐ€๋Šฅ์ผ€ ํ–ˆ๋˜ RLHF๋Š” ํ•œ๊ณ„์— ๋ถ€๋”ชํž ๊ฒƒ์ด๊ธฐ ๋•Œ๋ฌธ CriticGPT๋Š” ๊ฒฐ๊ณผ์ ์œผ๋กœ ์‚ฌ๋žŒ๋งŒ ํ–ˆ๋˜ ๊ฒƒ๋ณด๋‹ค ๋” ๋‚˜์€ ๊ฒฐ๊ณผ๋ฅผ ๋ƒˆ๊ณ  ์˜คํ”ˆAI๋Š” ํ–ฅํ›„ ์ด๋ฅผ ํ™•์žฅํ•˜์—ฌ ์‹ค์ œ๋กœ ๋ชจ๋ธ์— ์ ์šฉํ•˜๊ฒ ๋‹ค๋Š” ๋ง์„ ๋‚จ๊น€
*In order to align AI systems that are increasingly complex, weโ€™ll need better tools. In our research on CriticGPT, we found that applying RLHF to GPT-4 has promise to help humans produce better RLHF data for GPT-4. We are planning to scale this work further and put it into practice.*



https://openai.com/index/finding-gpt4s-mistakes-with-gpt-4/