AI ๋ชจ๋ธ์์ "๊ทธ๊ฒ"์ ๋ฐ์ดํฐ์
์ด๋ค.
์ ๋ ๊ฑฐ์ 1๋
๊ฐ OpenAI์์ ๊ทผ๋ฌดํด ์์ต๋๋ค. ๊ทธ ์๊ฐ ๋์ ์ ๋ ๋ง์ ์์ฑ ๋ชจ๋ธ์ ํ๋ จ์์ผฐ์ต๋๋ค. ์ฌ์ค์ ๋๊ตฌ๋ ๊ทธ๋ ๊ฒ ๋ง์ด ํ๋ จ์ํฌ ๊ถ๋ฆฌ๊ฐ ์์ ์ ๋๋ก ๋ง์ด์.ย ์ฌ๋ฌ ๋ชจ๋ธ ๊ตฌ์ฑ๊ณผ ํ์ดํผํ๋ผ๋ฏธํฐ๋ฅผ ์กฐ์ ํ๋ฉด์ ๊ด์ฐฐํ ์๊ฐ ๋์, ๋ชจ๋ ํ๋ จ ์คํ ์ฌ์ด์ ์ ์ฌ์ฑ์ด ์๋ค๋ ๊ฒ์ด ์ ์ ๋๋ ทํด์ก์ต๋๋ค.
์ด ๋ชจ๋ธ๋ค์ด ๊ทธ๋ค์ ๋ฐ์ดํฐ์
์ ๋๋๋๋ก ์ ๊ตํ๊ฒ ๊ทผ์ฌํํ๊ณ ์๋ค๋ ์ฌ์ค์ด ์ ์ ๋ถ๋ช
ํด์ง๊ณ ์์ต๋๋ค. ์ด๊ฒ์ด ์๋ฏธํ๋ ๋ฐ๋ ๊ทธ๋ค์ด ๊ฐ๋ ๊ณ ์์ด๊ฐ ๋ฌด์์ธ์ง๋ฅผ ๋ฐฐ์ฐ๋ ๊ฒ๋ฟ๋ง ์๋๋ผ, ์ฌ๋๋ค์ด ์์ฃผ ์ฐ๋ ์ฌ์ง์ด๋ ํํ ์ฐ๋ ๋จ์ด์ ๊ฐ์ ์ค์ํ์ง ์์ ๋ถํฌ ์ฌ์ด์ ๋น๋๋ฅผ ๋ฐฐ์ด๋ค๋ ๊ฒ์
๋๋ค.
์ด๋ ์ถฉ๋ถํย ๊ฐ์ค์น์ ํ๋ จ ์๊ฐ์ ๊ฐ์ง ๋ชจ๋ ๋ชจ๋ธ๋ค์ด ๊ฐ์ ๋ฐ์ดํฐ์
์ผ๋ก ์ถฉ๋ถํ ์ค๋ ํ๋ จ๋๋ฉด ๊ฑฐ์ ๋ชจ๋ ๋์ผํ ์ง์ ์ผ๋ก ์๋ ดํ๋ค๋ ๊ฒ์ผ๋ก ๋ํ๋ฉ๋๋ค. ์ถฉ๋ถํ ํฐ diffusion conv-unets๋ ViT ์์ฑ๊ธฐ์ ๋์ผํ ์ด๋ฏธ์ง๋ฅผ ์์ฑํฉ๋๋ค. AR ์ํ๋ง์ diffusion๊ณผ ๋์ผํ ์ด๋ฏธ์ง๋ฅผ ์์ฑํฉ๋๋ค.
์ด๋ ๋๋ผ์ด ๊ด์ฐฐ์
๋๋ค! ์ด๊ฒ์ ๋ชจ๋ธ ํ๋์ด ์ํคํ
์ฒ, ํ์ดํผํ๋ผ๋ฏธํฐ ๋๋ ์ต์ ํ ์ ํ์ ์ํด ๊ฒฐ์ ๋์ง ์๋๋ค๋ ๊ฒ์ ์๋ฏธํฉ๋๋ค. ๊ทธ๊ฒ์ ๋น์ ์ ๋ฐ์ดํฐ์
์ ์ํด ๊ฒฐ์ ๋ฉ๋๋ค, ๊ทธ ๋ฐ์ ๋ค๋ฅธ ๊ฒ์ ์์ต๋๋ค. ๋ค๋ฅธ ๋ชจ๋ ๊ฒ์ ๊ทธ ๋ฐ์ดํฐ์
์ ํจ์จ์ ์ผ๋ก ๊ทผ์ฌํํ๊ธฐ ์ํด ๊ณ์ฐ์ ์ ๋ฌํ๋ ์๋จ์ ๋ถ๊ณผํฉ๋๋ค.
๊ทธ๋ฌ๋ฏ๋ก ์ฐ๋ฆฌ๊ฐ '๋๋ค', '์ฑGPT', '๋ฐ๋', 'ํด๋ก๋'๋ฅผ ์ธ๊ธํ ๋, ๋ชจ๋ธ ๊ฐ์ค์น๋ฅผ ์ธ๊ธํ๋ ๊ฒ์ด ์๋๋๋ค. ๊ทธ๊ฒ์ ๋ฐ์ดํฐ์
์ ๋งํ๋ ๊ฒ์
๋๋ค.
The โitโ in AI models is the dataset.
Posted on June 10, 2023 by jbetker
Iโve been at OpenAI for almost a year now. In that time, Iโve trained a lot of generative models. More than anyone really has any right to train. As Iโve spent these hours observing the effects of tweaking various model configurations and hyperparameters, one thing that has struck me is the similarities in between all the training runs.
Itโs becoming awfully clear to me that these models are truly approximating their datasets to an incredible degree. What that means is not only that they learn what it means to be a dog or a cat, but the interstitial frequencies between distributions that donโt matter, like what photos humans are likely to take or words humans commonly write down.
What this manifests as is โ trained on the same dataset for long enough, pretty much every model with enough weights and training time converges to the same point. Sufficiently large diffusion conv-unets produce the same images as ViT generators. AR sampling produces the same images as diffusion.
This is a surprising observation! It implies that model behavior is not determined by architecture, hyperparameters, or optimizer choices. Itโs determined by your dataset, nothing else. Everything else is a means to an end in efficiently delivery compute to approximating that dataset.
Then, when you refer to โLambdaโ, โChatGPTโ, โBardโ, or โClaudeโ then, itโs not the model weights that you are referring to. Itโs the dataset.
๊ฐ๋ ๊ธ์ ์ฌ๋ผ์จ ๊ธ์ด ์ค์ํ ๋ด์ฉ์ด๋ผ๊ณ ์๊ฐ๋๋๋ฐ ํ๊ธ๋ฒ์ญ์ด ์ฝ๊ฐ ๋ํดํ ๋ถ๋ถ์ด ์์ด์ ํด๋ก๋+GPT4 ์กฐํฉ์ผ๋ก ์ฌ๋ฒ์ญ ํด ๋ดค๋ค.
https://gall.dcinside.com/mgallery/board/view/?id=thesingularity&no=459454
๋ฉํ llama๊ฐ ์ฌ๊ธฐ์์ ์ฐฉ์ํ ๋ฏ ์ถ๊ณ ์ด๊ฑฐ ์ด ์ฌ๋ ์ํธ๋งํํ ์ณ๋ง๊ณ ์์๋ฏ linkedin ๋ณด๋ ์์ง๋ ๊ทผ๋ฌด์ค์ด๋๋ฐ
๋ชจ๋ธ์ด ์ํคํ ์ฒ๋ ํ๋ผ๋ฏธํฐ ์กฐ์ ์ผ๋ก ์ฐฝ๋ฐ์ ์ธ ๊ตฌ์กฐ๋ฅผ ํ์ฑํ๋ ๊ฒ ์๋๋ผ ๋ฐ์ดํฐ์ ์ ๋ด์ฌ๋ ๋ฌด์ธ๊ฐ๋ก ์๋ ดํ ๋ฟ์ด๋ผ๋ฉด ํฉ์ฑ๋ฐ์ดํฐ์ ์กด์ฌ๊ฐ ์ด๋ค ์ํฅ์ ๋ฏธ์น ์ง ๊ถ๊ธํ๋ค ๊ทธ๊ฒ ์ญ์ ๊ทธ์ ์๋ ด์ ๊ฐ์ํ๋ ์ฉ๋์ ์ง๋์ง ์์์ง ์๋๋ฉด ์๋ก์ด ๋ฌด์ธ๊ฐ์ ๋๋ฌํ๊ฒ ๋๋๊ฒ์ธ์ง.
์์ฝ ์ข
๊ทธ๋ฌ๋๊น ํด๋จธ๋ ธ์ด๋ ๋ง๋ค์ด์ ๋ฐ์ดํฐ ์์ฐฝ ๋ชจ์์ผ ๋๋๊ฑฐ
๋ณ๋ก ์๋ก์ด ์๊ธฐ๋ ์๋๋ . ๋ชจ๋ธ์ DOF๊ฐ ์ถฉ๋ถํ ํฌ๋ฉด ๋น์ฐํ ๋ฐ์ดํฐ์ ์ ์ ๊ทผ์ฌํ๊ฒ ์ง.
โEverything else is a means to an end in efficiently delivery compute to approximating that dataset. Then, when you refer to โLambdaโ, โChatGPTโ, โBardโ, or โClaudeโ then, itโs not the model weights that you are referring to. Itโs the dataset.โ