?si=E2K7w0zhnXW9sYSK



1:14:30~1:33:37

OpenAI 개발자 두 명이 나오는데, 이날 여기서 시연하는 여러 가지 데모가 중요한 게 아니라 데모 앞, 뒤에 한 말을 더 주목할만함.



We'll talk about multimodal stuff, but I think it's important to start off with where we are today. And I think as we all know, people who have been building in the AI space for the last 6 12 18 months, 2023 has really been the year of chatbots and I think it's been incredible to see how much people have actually been able to do, how much value you can create in the world with like just a simple chatbot. And it's it still blows my mind to think about how rudimentary these systems are and how much more value there's going to be created in the next year and in the next decade. And that's why I'm excited for 2024 which I think is really going to be the I don't know if I can trademark this but the year of multimodal models. OpenAI has a ton of multimodal capabilities that are in the works. Some folks might have already tried some of these in ChatGPT and the IOS app or the the web app today, things like Vision taking in images describing them. We'll show that later on. Also the ability to generate images we've had this historically with DallE 2 but DallE 3 really if folks have tried, it takes things to the next level. So excited to to show some of that today as well.
멀티모달에 대해 이야기할 것이지만, 먼저 현재 우리가 어디에 와 있는지부터 살펴보는 것이 중요하다고 생각합니다. 우리 모두 알다시피 지난 6년 12개월 18개월 동안 AI 분야에서 구축해 온 사람들에게 2023년은 정말 챗봇의 해였으며, 사람들이 실제로 얼마나 많은 일을 할 수 있었는지, 간단한 챗봇으로 세상에 얼마나 많은 가치를 창출할 수 있는지 보는 것은 놀라운 일이라고 생각합니다. 그리고 이러한 시스템이 얼마나 초보적인 수준인지, 그리고 내년과 향후 10년 동안 얼마나 더 많은 가치가 창출될지 생각하면 여전히 놀랍습니다. 그래서 저는 2024년이 정말 멀티모달 모델의 해가 될 것이라고 생각하며, 이런 표현을 쓸 수 있을지 모르겠지만 기대가 됩니다. (중간에 개드립 생략) OpenAI에는 현재 개발 중인 수많은 멀티모달 기능이 있습니다. 이미 ChatGPT와 IOS 앱 또는 웹 앱에서 이미지로 설명하는 비전과 같은 기능을 사용해 보신 분들도 계실 것입니다. 나중에 보여드리겠습니다. 또한, 이미지 생성 기능은 이미 DallE 2에서도 사용 가능했지만, DallE 3는 한 단계 더 발전한 기능입니다. 오늘도 그 중 일부를 보여드리게 되어 매우 기쁩니다.


So if you think of uh the way that multimodal capabilities are working right now, it's a little bit of a setup of islands where we have DallE that takes text and generates images. We have whisper that takes in audio and generates text trans. We have GPT with vision capabilities that takes images and text and can reason over both at the same time. But right now this these are all very the seperate things. However you can think of text as a connective tissue between all of these models and there's a lot of interesting things that we can build right now using that Paradigm. But what we're actually really excited for is a future in which there's Unity between all these modalities and this is where we're going. This is not where we're today. But you can think of models in the same way that like GPT can consume images and text simultaneously. Maybe in the future we'll consume may more modalities and we'll output even more modalities to reason about them in the at the same time. However we're not there yet. And so today Logan are going to show you just like some, some uh architecture patterns and some ways in which you can mimic this kind of situation with what we have available today. And some of the patterns that you can start to think about as we move towards this future in which models can a reason way beyond text.
현재 멀티모달 기능이 작동하는 방식을 생각해보면, 섬들과 같은 형태입니다. 텍스트를 받아 이미지를 생성하는 DallE가 있고 오디오를 받아 텍스트 트랜스크립트를 생성하는 위스퍼가 있습니다. 이미지와 텍스트를 동시에 인식하고 추론할 수 있는 비전 기능을 갖춘 GPT도 있습니다. 하지만 현재로서는 이 모든 것이 매우 별개의 것입니다. 텍스트를 이러한 모든 모델 간의 연결 조직으로 생각할 수 있으며, 이러한 패러다임을 사용하여 지금 당장 구축할 수 있는 흥미로운 것들이 많이 있습니다. 하지만 실제로 우리가 정말 기대하는 것은 이러한 모든 양식이 통합되는 미래이며, 이것이 바로 우리가 가고자 하는 방향입니다. 지금 우리가 있는 곳은 이런 곳이 아닙니다. 하지만 GPT와 같이 이미지와 텍스트를 동시에 사용할 수 있는 것과 같은 방식으로 모델을 생각할 수 있습니다. 어쩌면 미래에는 더 많은 양식을 소비하고 동시에 더 많은 양식을 추론할 수 있게 될지도 모릅니다. 하지만 아직 거기까지는 가지 못했습니다. 그래서 오늘 로건은 몇 가지 아키텍처 패턴과 현재 우리가 사용할 수 있는 것으로 이런 상황을 모방할 수 있는 몇 가지 방법을 보여드리려고 합니다. 그리고 모델이 텍스트를 뛰어넘어 추론할 수 있는 미래로 나아갈 때 생각해 볼 수 있는 몇 가지 패턴을 소개합니다.


I were making these demos today, waiting till the the last minute as as always it it was really interesting to see that like really much of the work of making multimodal systems today is like how do you hook everything up together and connect the different modalities and again as he said using text as sort of the the bridge between different modalities. But it's going to be super interesting to see like how much developer efficiency gains there are when you no longer have to do that. And you really just have like a single model that can do text in text out(video at some point), speech in speech out at some point so it'll be super cool to to see when that's possible and uh making demos even easier and simpler.
저는 오늘 이 데모를 만들면서 항상 그렇듯이 마지막 순간까지 기다리면서 오늘날 멀티 모달 시스템을 만드는 작업의 대부분은 어떻게 모든 것을 서로 연결하고 서로 다른 모달리티를 연결할 수 있는지, 그리고 그가 말한 것처럼 텍스트를 일종의 다리로 사용하여 서로 다른 모달리티를 연결하는 것이 정말 흥미로웠습니다. 그러나 더 이상 그렇게 할 필요가 없을 때 개발자의 효율성이 얼마나 향상되는지 보는 것은 매우 흥미로울 것입니다. 그리고 텍스트 인 텍스트 아웃(어느 시점에서는 비디오), 음성 인 음성 아웃(어느 시점에서는 음성)을 할 수 있는 단일 모델만 있으면 언제 가능할지, 그리고 데모를 더욱 쉽고 간단하게 만들 수 있는지 보는 것은 정말 멋질 것입니다.


(데모 끝난 후)


Start thinking multimodal, that's something net new, that's happening these days. And if you have any crazy ideas that you think 'wow it would be really cool if technology could do this' we'll probably be able together. And the products that you'll be able to build 6 months from now, a year from now are going to be incredible so start having this in mind as people who are building AI products and people who are building companies. Think of text as a as a connecting tissue right now. And I think this is a very powerful concept and that's going to continue to be the case for the near future uh and there are many powerful patterns that are yet to be explored when it comes to multimodal stuff especially when it comes to to uh doing things with images uh so really excited to uh soon get in the hands of all of you guys and and to see why you all build with this I think it's uh it's really exciting uh to see a AI start to venture into the visual world.
멀티모달에 대해 생각해 보세요. 멀티모달은 요즘 새롭게 떠오르고 있는 개념입니다. 그리고 '기술이 이걸 할 수 있다면 정말 멋질 것 같다'고 생각하는 미친 아이디어가 있다면 우리가 함께 할 수 있을 것입니다. 그리고 6개월 후, 1년 후 여러분이 만들 수 있는 제품은 놀라울 것이므로 AI 제품을 만드는 사람이나 회사를 만드는 사람으로서 이를 염두에 두시기 바랍니다. 지금 텍스트를 연결 조직이라고 생각하세요. 그리고 저는 이것이 매우 강력한 개념이라고 생각하며 가까운 미래(near future)에도 계속 그럴 것이라고 생각합니다. 특히 이미지로 작업을 수행하는 데있어 멀티 모달 작업과 관련하여 아직 탐색되지 않은 강력한 패턴이 많이 있습니다. 곧 여러분 모두의 손에 들어가서 여러분 모두가 이것을 사용하여 구축하는 이유를 보게되어 정말 흥분됩니다. AI가 시각적 세계로 모험을 시작하는 것을 보는 것은 정말 흥분됩니다.






요약하자면 향후 6~12개월에 외부 개발자들이 만들 수 있는, 이날 데모로 살짝 보여준 '텍스트를 모달리티 간의 연결 조직으로 이용한, 따로 떨어진 섬들 같은' 멀티모달 패라다임은, 모든 모달리티가 싱글 모델에 통합되는 더 먼 미래의 모델이 아닌, 가까운 미래까지 한동안 거쳐가는 패러다임이라는 것임.