물론 트랜스포머 아키텍처가 스케일확장하면 지능이 증가한단게 밝혀졌지만 데이터를 언어밖에 못먹잖아
댓글 9
멀티모달 나온지가 언젠데
익명(1.239)2022-03-03 20:49
complex sequence learning may be the key that unlocks all the rest. This would explain the surprising capacities we see in large language models — which, in the end, are nothing but complex sequence learners. Attention, in turn, has proven to be the key mechanism for achieving complex sequence learning in neural nets — as suggested by the title of the paper introducing the Transformer model whose
익명(125.191)2022-03-03 20:50
답글
successors power today’s LLMs: Attention is all you need.
익명(125.191)2022-03-03 20:50
답글
요약: 언어모델 할 정도면 다 됨
익명(118.44)2022-03-03 20:52
답글
No serious impediment stands in the way of AI researchers training next-generation models on combinations of text with images, sound, and video, and indeed this kind of work is already underway.¹⁷ Such models will also eventually be trained using the active experience of robots in real or simulated worlds, which may play with sand and water and engage in other exploratory “Montessori style
익명(125.191)2022-03-03 20:54
답글
style learning
익명(125.191)2022-03-03 20:54
답글
그리고 GPT3에서 100배 확장이 아니라 2857배 확장임
익명(125.191)2022-03-03 20:56
답글
얼마나 똑똑해질지 궁금하다. 기존 언어모델의 맹점인 간단한 연산까지 할 지능이 생길수 있을지가 관건인듯
익명(121.153)2022-03-03 20:59
답글
계산을 하는 언어모델이 있다??!
그게 돼? 솔직히 아직도 의구스러움. 규모만 키워서 되는 문제인가
멀티모달 나온지가 언젠데
complex sequence learning may be the key that unlocks all the rest. This would explain the surprising capacities we see in large language models — which, in the end, are nothing but complex sequence learners. Attention, in turn, has proven to be the key mechanism for achieving complex sequence learning in neural nets — as suggested by the title of the paper introducing the Transformer model whose
successors power today’s LLMs: Attention is all you need.
요약: 언어모델 할 정도면 다 됨
No serious impediment stands in the way of AI researchers training next-generation models on combinations of text with images, sound, and video, and indeed this kind of work is already underway.¹⁷ Such models will also eventually be trained using the active experience of robots in real or simulated worlds, which may play with sand and water and engage in other exploratory “Montessori style
style learning
그리고 GPT3에서 100배 확장이 아니라 2857배 확장임
얼마나 똑똑해질지 궁금하다. 기존 언어모델의 맹점인 간단한 연산까지 할 지능이 생길수 있을지가 관건인듯
계산을 하는 언어모델이 있다??! 그게 돼? 솔직히 아직도 의구스러움. 규모만 키워서 되는 문제인가