https://www.nature.com/articles/s41586-025-09422-z?utm_source=substack&utm_medium=emailDeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning - NatureA new artificial intelligence model, DeepSeek-R1, is introduced, demonstrating that the reasoning abilities of large language models can be incentivized through pure reinforcement learning, removing the need for human-annotated demonstrations.www.nature.comㅇㅇ
학계와 빅테크의 차이를 알 수 있는 부분
리뷰가 존나오래걸려서 R1제로 논문이 이제야 나오노 ㅋㅋㅋㅋㅋ
오 네이처
학계는 이제 끝이군
원래 저널은 리뷰 오래걸리긴 함.. 컨퍼는 그래도 빠르게 돌아가는 편임. 물론 그래봤자 프론티어는 빅테크지만