2024년 5월에 gpt4가 튜링테스트 통과했다는 논문이 나옴
실험 방법은 일반적인 튜링테스트 방법임
세부적으론 인간이 상대가 누군지 알려주지 않고 대화를 5분 하게한 후
사람인지 아닌지 판단하게 하고 얼마나 확신하는지 질문함
근데 사람들이 GPT-4랑 대화하고 반 이상이 얘가 사람이라고 착각한거임.
근데 잘 살펴보면
GPT-4는 54%, 사람은 67% 였음
즉 “튜링테스트를 통과했다”를 사람인지 아닌지 분간 못하기 시작했다 기준을 50% 로 잡으면 통과한거고
사람보다 더 사람같다 를 기준으로 잡으면 13% 딸린 거임.
이 논문이 요즘 피터틸이나 유명인사들이 튜링테스트 언급하는 이유임
관심있으면 함 보셈
https://arxiv.org/abs/2405.08007People cannot distinguish GPT-4 from a human in a Turing testWe evaluated 3 systems (ELIZA, GPT-3.5 and GPT-4) in a randomized, controlled, and preregistered Turing test. Human participants had a 5 minute conversation with either a human or an AI, and judged whether or not they thought their interlocutor was human. GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%). The results provide the first robust empirical demonstration that any artificial system passes an interactive 2-player Turing test. The results have implications for debates around machine intelligence and, more urgently, suggest that deception by current AI systems may go undetected. Analysis of participants' strategies and reasoning suggests that stylistic and socio-emotional factors play a larger role in passing the Turing test than traditional notions of intelligence.arxiv.org
실험 방법은 일반적인 튜링테스트 방법임
세부적으론 인간이 상대가 누군지 알려주지 않고 대화를 5분 하게한 후
사람인지 아닌지 판단하게 하고 얼마나 확신하는지 질문함
근데 사람들이 GPT-4랑 대화하고 반 이상이 얘가 사람이라고 착각한거임.
근데 잘 살펴보면
GPT-4는 54%, 사람은 67% 였음
즉 “튜링테스트를 통과했다”를 사람인지 아닌지 분간 못하기 시작했다 기준을 50% 로 잡으면 통과한거고
사람보다 더 사람같다 를 기준으로 잡으면 13% 딸린 거임.
이 논문이 요즘 피터틸이나 유명인사들이 튜링테스트 언급하는 이유임
관심있으면 함 보셈
https://arxiv.org/abs/2405.08007People cannot distinguish GPT-4 from a human in a Turing testWe evaluated 3 systems (ELIZA, GPT-3.5 and GPT-4) in a randomized, controlled, and preregistered Turing test. Human participants had a 5 minute conversation with either a human or an AI, and judged whether or not they thought their interlocutor was human. GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%). The results provide the first robust empirical demonstration that any artificial system passes an interactive 2-player Turing test. The results have implications for debates around machine intelligence and, more urgently, suggest that deception by current AI systems may go undetected. Analysis of participants' strategies and reasoning suggests that stylistic and socio-emotional factors play a larger role in passing the Turing test than traditional notions of intelligence.arxiv.org
기능주의적으로 인간한테 따라가야...