?si=ckbNcCvoEpizd0u2

인터뷰어: The final question is this. You've been in this field for over a decade, much longer  than many others, and you've seen different landmarks like ImageNet and Transformers. What  do you think the next landmark will look like?

셰인 레그: I think the next landmark that people will  think back to and remember is going much more fully multimodal. That will open out the  sort of understanding that you see in language models into a much larger space of possibilities. And when people think back, they'll think about, “Oh, those old fashioned models, they just did  like chat, they just did text.” It just felt like a very narrow thing whereas now they understand  when you talk to them and they understand images and pictures and video and you can show them  things or things like that. And they will have   much more understanding of what's going on. And  it'll feel like the system's kind of opened up into the world in a much more powerful way.

인터뷰어: Do you mind if I ask a follow-up on that? ChatGPT just released their multimodal feature  and you, in DeepMind, you had the Gato paper, where you have this one model where you can throw images, video games and even actions in there. So far it doesn't seem to have percolated as much  as ChatGPT initially from GPT3 or something. What explains that? Is it just  that people haven't learned to use   multimodality? They're not powerful enough yet?

셰인 레그: I think it's early days. I think you will see understanding images and things more and  more. But I think it's early days in this   transition is when you start really digesting  a lot of video and other things like that, that the systems will start having a much more  grounded understanding of the world and all kinds   of other aspects. And then when that works well,  that will open up naturally lots and lots of new applications and all sorts of new possibilities  because you're not confined to text chat anymore.

인터뷰어: New avenues of training data as well, right?

셰인 레그: Yeah, new training data and all kinds of different applications that aren't just purely textual  anymore. And what are those applications? Well, probably a lot of them we can't even imagine  at the moment because there are just so many   possibilities once you can start dealing with all  sorts of different modalities in a consistent way.


인터뷰어: 마지막 질문입니다. 다른 사람들보다 훨씬 더 오랜 기간 동안 이 분야에 종사해 오셨고, 이미지넷이나 트랜스포머와 같은 다양한 랜드마크를 보셨죠. 다음 랜드마크는 어떤 모습일 것이라고 생각하시나요?

셰인 레그: 사람들이 기억하고 회상할 다음 랜드마크는 훨씬 더 완전한 멀티모달이 될 것이라고 생각합니다. 그러면 언어 모델에서 볼 수 있는 이해가 훨씬 더 넓은 가능성의 공간으로 확장될 것입니다. 그리고 사람들은 "아, 그 옛날 모델들은 채팅이나 문자만 했었지"라고 회상하게 될 것입니다. 하지만 이제는 이미지와 사진, 동영상도 이해할 수 있고, 여러분이 무언가를 보여줄 수도 있습니다. 그리고 무슨 일이 일어나고 있는지 훨씬 더 잘 이해할 수 있게 됩니다. 그리고 시스템이 훨씬 더 강력한 방식으로 세상에 개방된 것처럼 느껴질 것입니다.

인터뷰어: 그 부분에 대한 후속 질문을 해도 될까요? ChatGPT는 방금 멀티모달 기능을 출시했고, 딥마인드에서는 이미지, 비디오 게임, 심지어 액션까지 넣을 수 있는 가토 페이퍼라는 모델이 있습니다. 지금까지는 GPT3나 다른 것에서 ChatGPT가 처음에 나온 것만큼 많이 퍼지지 않은 것 같습니다. 그 이유는 무엇일까요? 사람들이 멀티 모달리티를 사용하는 법을 배우지 않았기 때문일까요? 아직 충분히 강력하지 않나요?

셰인 레그: 아직은 초기 단계라고 생각합니다. 이미지와 사물을 점점 더 많이 이해하게 될 것이라고 생각합니다. 하지만 많은 비디오와 다른 것들을 실제로 소화하기 시작하면 시스템이 세상과 다른 모든 측면에 대해 훨씬 더 근거를 바탕으로 이해하기 시작할 것입니다. 그리고 그것이 잘 작동하면 더 이상 텍스트 채팅에 국한되지 않기 때문에 자연스럽게 수많은 새로운 애플리케이션과 모든 종류의 새로운 가능성이 열릴 것입니다.

인터뷰어: 데이터를 학습할 수 있는 새로운 길도 열리겠죠?

셰인 레그: 네, 새로운 학습 데이터와 더 이상 텍스트에만 국한되지 않는 모든 종류의 다양한 애플리케이션이 있습니다. 어떤 애플리케이션이 있나요? 다양한 양식을 일관된 방식으로 처리할 수 있게 되면 가능성이 무궁무진해지기 때문에 현재로서는 상상조차 할 수 없는 것들이 많을 겁니다.









데미스 하사비스 인터뷰

And just going back to your first question about the generative models, I do think we are right at the beginning of an incredible new era that’s going to play out over the next five, 10 years. Not only in advancing science with AI but in terms of the types of products we can build to improve people’s everyday lives, billions of people in their everyday lives, and help them to be more efficient and to enrich their lives. And I think what we’re seeing today with these chatbots is literally just scratching the surface. There are a lot more types of AI than generative AI. Generative AI is now the “in” thing, but I think that planning and deep reinforcement learning and problem-solving and reasoning, those kinds of capabilities are going to come back in the next wave after this, along with the current capabilities of the current systems. So I think, in a year or two’s time, if we were to talk again, we are going to be talking about entirely new types of products and experiences and services with never-seen-before capabilities. And I’m very excited about building those things, actually. And that’s one of the reasons I’m very excited about leading Google DeepMind now in this new era and focusing on building these AI-powered next-generation products.

제너레이티브 모델에 대한 첫 번째 질문으로 돌아가서, 저는 우리가 향후 5년, 10년 동안 펼쳐질 놀라운 새 시대의 시작점에 있다고 생각합니다. AI로 과학을 발전시키는 것뿐만 아니라 사람들의 일상생활을 개선하고, 수십억 명의 일상생활을 개선하고, 더 효율적이고 풍요로운 삶을 살 수 있도록 도울 수 있는 제품 유형을 구축할 수 있습니다. 그리고 현재 우리가 보고 있는 챗봇은 말 그대로 표면을 살짝 긁어본 것에 불과하다고 생각합니다. 제너레이티브 AI보다 훨씬 더 많은 유형의 AI가 있습니다. 지금은 제너레이티브 AI가 '대세'이지만, 계획 수립과 심층 강화 학습, 문제 해결 및 추론 등 이러한 종류의 기능은 현재 시스템의 현재 기능과 함께 다음 물결에 다시 등장할 것이라고 생각합니다. 따라서 1~2년 후에 다시 이야기하게 된다면 이전에는 볼 수 없었던 기능을 갖춘 완전히 새로운 유형의 제품과 경험, 서비스에 대해 이야기하게 될 것이라고 생각합니다. 그리고 저는 그런 것들을 구축하는 것에 대해 매우 기대하고 있습니다. 이것이 바로 제가 이 새로운 시대에 구글 딥마인드를 이끌며 AI 기반의 차세대 제품을 만드는 데 집중하게 된 것을 매우 기쁘게 생각하는 이유 중 하나입니다.