๊ธฐ์ค์ ์ด ํ๋ จ์ ๋ค์ด๊ฐ ์ปดํจํ
์ด 100์์ฌํ๋กญ์ค(1e26) ์ด์์ธ๋ฐ ํ์ฌ ๊ฐ์ฅ ํฐ ์คํ์์ค ๋ชจ๋ธ์ธ llama2๊ฐ 5e24์ด๊ณ , gpt3๊ฐ 3e23์
gpt4๋ ์ด ๊ธฐ์ค์ ๋ค์ด๊ฐ์ง ์ ํํ ์์๋ ์๋๋ฐ ์ ๋ฏธ๋๋ gpt5๋ ๋ฌด์กฐ๊ฑด ๋ค์ด๊ฐ๋ค๊ณ ๋ณด๋ฉด ๋ ๋ฏ
https://x.com/soumithchintala/status/1719118791773208713?s=61&t=vw-Q25I5Cfyobm-eAwkHQQ
Regulation starts at roughly two orders of magnitude larger than a ~70B Transformer trained on 2T tokens -- which is ~5e24.
โ Soumith Chintala (@soumithchintala) October 30, 2023
Note: increasing the size of the dataset OR the size of the transformer increases training flops.
The (rumored) size of GPT-4 is regulated. https://t.co/bhay1ElDAp
https://github.com/amirgholami/ai_and_memory_wall#nlp-modelsGitHub - amirgholami/ai_and_memory_wall: AI and Memory Wall blog postAI and Memory Wall blog post. Contribute to amirgholami/ai_and_memory_wall development by creating an account on GitHub.github.com
๋๊ธ 0