Dwarkesh Patel (00:06:02 - 00:06:12):
만약 우리가 인간 수준의 지능에 도달하기 전에 스케일링 정체기가 온다면, 돌이켜보면 어떻게 설명할 수 있을까요? 만약 그렇게 판명된다면 어떤 결과가 나올 가능성이 있다고 생각하시나요?
Dario Amodei (00:06:12 - 00:08:57):
저는 근본 이론의 문제와 실질적인 이슈를 구분할 것입니다. 우리가 직면할 수 있는 실질적인 문제 중 하나는 데이터가 부족해질 수 있다는 것입니다. 여러 이유로 그런 일은 일어나지 않을 것 같지만, 순진하게 본다면 우리는 데이터 부족에서 멀지 않습니다. 즉, scaling curve를 계속 그리기에는 데이터가 부족할 수 있습니다. 또 다른 시나리오는, 우리가 사용 가능한 모든 compute를 다 써버렸는데 그것으로는 충분하지 않았고, 그 이후로는 진전이 더디다는 것입니다. 저는 이 둘 중 어느 것도 일어날 것이라고 생각하지 않지만, 그럴 가능성은 있습니다.
근본적인 관점에서 볼 때, 개인적으로 scaling law가 그냥 멈출 것 같지는 않습니다. 만약 그렇게 된다면, 또 다른 이유는 우리가 정확히 올바른 architecture를 가지고 있지 않기 때문일 수 있습니다. 만약 우리가 LSTM이나 RNN으로 시도했다면 slope가 달라졌을 것입니다. 여전히 그곳에 도달할 수는 있겠지만, transformer가 가진 과거를 멀리 attend할 수 있는 능력이 없다면 represent하기 매우 어려운 것들이 있습니다. 만약 우리가 그냥 벽에 부딪히고 그것이 architecture 때문이 아니라면 저는 매우 놀랄 것입니다. 우리는 이미 모델이 할 수 없는 것들이 할 수 있는 것들과 종류가 다르지 않은 것처럼 보이는 지점에 도달했습니다.
몇 년 전에는 그들이 추론할 수 없고, 프로그래밍할 수 없다고 주장할 수 있었을 것입니다. 경계를 그을 수 있었고 아마도 벽에 부딪힐 수도 있다고 말할 수 있었을 것입니다. 저는 우리가 벽에 부딪히지 않을 거라고 생각했고, 다른 몇몇 사람들도 그렇게 생각하지 않았지만, 그때는 그럴 법한 경우였습니다. 지금은 그럴 가능성이 더 적어졌습니다.
그런 일이 일어날 수도 있습니다. 이런 것들은 미친 짓입니다. 우리는 내일 벽에 부딪힐 수도 있습니다. 만약 그렇게 된다면 제 설명은 다음 단어 예측으로 학습시킬 때 손실 함수에 문제가 있다는 것입니다.
만약 정말로 높은 수준에서 프로그래밍하는 법을 배우고 싶다면, 그것은 일부 토큰이 다른 토큰보다 훨씬 더 중요하다는 것을 의미하고 그것들은 충분히 희귀해서 손실 함수가 외관, 즉 엔트로피의 대부분의 비트를 책임지는 것에 지나치게 초점을 맞추게 되고, 대신 정말 필수적인 이 내용에는 초점을 맞추지 않게 됩니다. 그래서 신호가 노이즈에 묻힐 수 있습니다. 저는 여러 이유로 그렇게 전개되지는 않을 것이라고 생각합니다. 하지만 만약 당신이 저에게 - 네, 당신은 2024년 모델을 학습시켰습니다. 그것은 훨씬 더 컸지만 전혀 나아지지 않았고, 모든 architecture를 시도해봤지만 작동하지 않았다고 말한다면, 그것이 제가 내릴 설명일 것입니다.
Dwarkesh Patel (00:06:02 - 00:06:12):
If it turns out that scaling plateaus before we reach human level intelligence, looking back on it, what would be your explanation? What do you think is likely to be the case if that turns out to be the outcome?
Dario Amodei (00:06:12 - 00:08:57):
I would distinguish some problem with the fundamental theory with some practical issue. One practical issue we could have is we could run out of data. For various reasons, I think that's not going to happen but if you look at it very naively we're not that far from running out of data. So it's like we just don't have the data to continue the scaling curves. Another way it could happen is we just use up all of the compute that was available and that wasn't enough and then progress is slow after that. I wouldn't bet on either of those things happening but they could.
From a fundamental perspective, I personally think it's very unlikely that the scaling laws will just stop. If they do, another reason could just be that we don't have quite the right architecture. If we tried to do it with an LSTM or an RNN the slope would be different. It still might be that we get there but there are some things that are just very hard to represent when you don't have the ability to attend far in the past that transformers have. If somehow we just hit a wall and it wasn’t about the architecture I'd be very surprised by that. We're already at the point where to me the things the models can't do don't seem to be different in kind from the things they can do.
You could have made a case a few years ago that they can't reason, they can't program. You could have drawn boundaries and said maybe you'll hit a wall. I didn't think we would hit a wall, a few other people didn't think we would hit a wall, but it was a more plausible case then. It's a less plausible case now.
It could happen. This stuff is crazy. We could hit a wall tomorrow. If that happens my explanation would be there's something wrong with the loss when you train on next word prediction.
If you really want to learn to program at a really high level, it means you care about some tokens much more than others and they're rare enough that the loss function over focuses on the appearance, the things that are responsible for the most bits of entropy, and instead they don't focus on this stuff that's really essential. So you could have the signal drowned out in the noise. I don't think it's going to play out that way for a number of reasons. But if you told me — Yes, you trained your 2024 model. It was much bigger and it just wasn't any better, and you tried every architecture and didn't work, that's the explanation I would reach for.
이 영상 안그래도 며칠 전에 봤는데 아모데이는 스케일링 정체가 올 것이라고 생각하지 않지만, 질문 자체가 '만약'이라는 가정을 하고 억지로 이유를 찾는다면 뭐가 있을까 물어보는 거임
여러 이유로 그런 일은 일어나지 않을 것 같지만, 순진하게 본다면 우리는 데이터 부족에서 멀지 않습니다 (중략) 저는 이 둘 중 어느 것도 일어날 것이라고 생각하지 않지만, 그럴 가능성은 있습니다.
중간에 질문자가 멀티모달 데이터도 있어서 그런 거냐고 묻자 아모데이는 명확하게 답하지 않고 얼버무림
아무것도 모르는 사람들은 대충 보고 그냥 믿어 버린다고
ㅇㅇ 좀 어그로 제목이긴하다. 근데 실제 GPT4나 클로드3랑 대화해보면 현존하는 유용한 지식은 거의 다 씹어먹은 거 같은데 학습시킬 때 좋은 데이터 위주로 먼저 넣었을 거 아냐? 한국어를 예로들면 뉴스, 블로그, 유명커뮤 데이터 다 넣었는데 더 많은 데이터가 필요해서 캣맘카페 데이터를 집어넣으면 성능 개판 되는 건 아닐까? 지금만 해도 트위터 같은데 BOT들이 엄청난 텍스트를 생성하고 있고 앞으로 인간들이 LLM에서 나온 Output을 온라인 상에 엄청나게 복붙한다고 생각하면 데이터 부족은 오지 않을 건데 그게 맞나 싶기도 하고..
하사비스, 아모데이, 수츠케버 등 전부 데이터 고갈 걱정은 안 한다는 관점인 걸로 보여서 나도 걱정은 안함