๊ฒฐ๋ก :

์ด ์—ฐ๊ตฌ์—์„œ๋Š” LLM์„ ์œ„ํ•œ 8๋น„ํŠธ ๊ต์œก์„ ์‚ดํŽด๋ด…๋‹ˆ๋‹ค.

8๋น„ํŠธ ์ง‘๋‹จ ํ†ต์‹ , ์˜ตํ‹ฐ๋งˆ์ด์ € ๋ฐ ๋ถ„์‚ฐ ๋ณ‘๋ ฌ ํ›ˆ๋ จ์„ ์ ์ง„์ ์œผ๋กœ ํ†ตํ•ฉํ•˜๋Š” ์ƒˆ๋กœ์šด FP8 ํ˜ผํ•ฉ ์ •๋ฐ€๋„ ํ›ˆ๋ จ ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์†Œ๊ฐœํ•ฉ๋‹ˆ๋‹ค.

์šฐ๋ฆฌ๊ฐ€ ์•„๋Š” ํ•œ, ์ด๋Š” FP8 ์ปดํ“จํŒ…, ์Šคํ† ๋ฆฌ์ง€ ๋ฐ ํ†ต์‹ ์„ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด๋ชจ๋ธ ํ›ˆ๋ จ์˜ ์ „์ฒด ์ง„ํ–‰ ๊ณผ์ •์— ์นจํˆฌ์‹œํ‚ค๋Š” ์ฒซ ๋ฒˆ์งธ ์ž‘์—…์ž…๋‹ˆ๋‹ค.

๊ด‘๋ฒ”์œ„ํ•œ ์‹คํ—˜์„ ํ†ตํ•ด ์ œ์•ˆ๋œ ๋ฐฉ๋ฒ•์ด ๋‹ค์–‘ํ•œ ๊ทœ๋ชจ์˜ GPT ๋ชจ๋ธ ํ›ˆ๋ จ ๋งฅ๋ฝ์—์„œ ํ†ต์‹  ์˜ค๋ฒ„ํ—ค๋“œ๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ์ค„์ด๊ณ  ๋ฉ”๋ชจ๋ฆฌ ํ™œ์šฉ๋„๋ฅผ ์ค„์ด๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค.

ํ–ฅํ›„ ์ž‘์—…์—์„œ๋Š” FP8 GPT ๋ชจ๋ธ์˜ ํฌ๊ธฐ์™€ ๊ต์œก ๋‹จ๊ณ„๋ฅผ ํ™•์žฅ(scale up)ํ•˜๊ณ  8๋น„ํŠธ ํ˜ผํ•ฉ ์ •๋ฐ€๋„ ์ฒด๊ณ„๋ฅผ ์‚ฌ์šฉํ•˜์—ฌ ์ถ”๊ฐ€๋กœ ๊ต์œกํ•  ๊ณ„ํš์ž…๋‹ˆ๋‹ค.

๋˜ํ•œ ์ œ์•ˆ๋œ FP8 ๋ฐฉ์‹์„ ์‚ฌ์šฉํ•˜์—ฌ ๋‹ค์ค‘ ๋ชจ๋“œ ๋Œ€ํ˜• ๋ชจ๋ธ(multi-modal large models)์„ ํ›ˆ๋ จํ•˜๊ณ 

์Šค๋งˆํŠธํฐ๊ณผ ๊ฐ™์€ ๋‹ค์–‘ํ•œ ์—์ง€ ์žฅ์น˜์—์„œ LLM์˜ ๋‚ฎ์€ ๋น„ํŠธ ๋ฐฐํฌ๋ฅผ ํƒ์ƒ‰ํ•ฉ๋‹ˆ๋‹ค .



(์›๋ฌธ)

In this work, we explore 8-bit training for LLMs. We introduce a new FP8 mixed-precision training framework, which incorporates 8-bit collective communication, optimizer, and distributed parallel training in an incremental manner. To our best knowledge, this is the first work infiltrating FP8 compute, storage and communication into the whole progress of large language model training. Extensive experiments demonstrate the proposed method effectively diminishes communication overhead and curtails memory utilization in the context of GPT model training at various scales. In future work, we plan to scale up the size and training steps of the FP8 GPT models and further train them with our 8-bit mixed-precision scheme. Moreover, we will also use the proposed FP8 scheme to train multi-modal large models, and explore low-bit deployment of LLMs on various edge devices, such as smart phones.