https://arxiv.org/abs/2406.02061Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language ModelsLarge Language Models (LLMs) are often described as being instances of foundation models - that is, models that transfer strongly across various tasks and conditions in few-show or zero-shot manner, while exhibiting scaling laws that predict function improvement when increasing the pre-training scale. These claims of excelling in different functions and tasks rely on measurements taken across various sets of standardized benchmarks showing high scores for such models. We demonstrate here a dramatic breakdown of function and reasoning capabilities of state-of-the-art models trained at the largest available scales which claim strong function, using a simple, short, conventional common sense problem formulated in concise natural language, easily solvable by humans. The breakdown is dramatic, as models also express strong overconfidence in their wrong solutions, while providing often non-sensical "reasoning"-like explanations akin to confabulations to justify and backup the validity of their clearly failed responses, making them sound plausible. Various standard interventions in an attempt to get the right solution, like various type of enhanced prompting, or urging the models to reconsider the wrong solutions again by multi step re-evaluation, fail. We take these initial observations to the scientific and technological community to stimulate urgent re-assessment of the claimed capabilities of current generation of LLMs, Such re-assessment also requires common action to create standardized benchmarks that would allow proper detection of such basic reasoning deficits that obviously manage to remain undiscovered by current state-of-the-art evaluation procedures and benchmarks. Code for reproducing experiments in the paper and raw experiments data can be found at https://github.com/LAION-AI/AIWarxiv.org
기본적인 상식 질문을 꼬아낼경우 프롬프트를 수정하더라도 높은 확률로 답변에 실패함
인간과 달리 지식의 부족 문제가 아니란걸 감안하면 상당히 부정적인 결과
ㅋㅋㅋㅋㅋㅋ
싹다 노동시켜 싹다 알바시켜
왜일까
그게 LLM이니깐
ㅋㅋㅋㅋㅋ
이것도 결국 돌파할듯 ai 붐 일어난지 얼마 안됐으니까
(gpt4가 나온지 1년이 넘었지만 제자리인데도)
1년이면 뭐 과학 기준으론 잠깐이지
팩트) gpt4랑 gemini 1.5 pro는 벤치마크에서 큰 차이가 난다
과학 기준으로는 수십년도 잠깐인데? 20년도 잠깐임 ㅋㅋㅋ
과학기준으론 잠깐 ㅋㅋㅋ 갑자기 선형충이 되버렸네 - dc App
지능이 아니라 합성기니까
확률학적 앵무새
신르쿤 1승 냥냥하죠?
사람으로 치면 융통성이 없다고 봐야하나
암기는 잘하는데 응용을 못하는 타입 근데 암기력이 월등해서 문제들을 통으로 외움
통계기계
싹다 쿠팡뛰어 싹다 배민뛰어
흠
얀르쿤 역시 ai의 아버지다
겨울이.... 온다...
특이점은 안온다
호들갑 싹 빼니 현실이 보이네.,. 그래 이래야지..
역시 얀르쿤의 혜안
슬슬 트랜스포머같이 시대를 뒤흔들 논문이 나올 때가 됐는데
아이 앰 옵티머스 오토봇 롤 아웃
LLM = 앵무새
얀르쿤 또 1승 ㅋㅋㅋ
llm은 결국 혼자 생각못하니까 agi랑 거리가 멀다니까 ㄹㅇ;
혼자 생각할 수 있는 인공지능 같은건 영원히 불가능함
왜 불가능해 사람 뇌가 있는데
결국 스케일로 때려 넣는것도 양질의 데이터가 뒷받침 안해주면 오히려 스케일이 높아질수록 붕괴하게 되있다는것 아무리 칩 때려박아도 데이터에 기업들이 집착하는 이유가 있음
구직 . 좋아요 . 알바라도 할께요
대 르 쿤