https://transformer-circuits.pub/2023/monosemantic-features/index.htmlTowards Monosemanticity: Decomposing Language Models With Dictionary LearningTowards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pubhttps://jobs.lever.co/Anthropic/3dbd1272-d673-4494-999a-47528e7cd6a3Anthropic - Research Engineer, InterpretabilityThe Interpretability team at Anthropic is working to reverse-engineer how models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. We’re looking for researchers and engineers to join our efforts. Few things can accelerate this progress more than great infrastructure. As a research engineer, you will build and maintain infrastructure used by the whole team, including yourself. You'll touch all parts of our code and infrastructure, whether that's making the cluster more reliable for our big jobs, improving throughput and efficiency, running and designing scientific experiments, or improving our dev tooling. You’re motivated to understand our research so you can write code that accelerates it. Some of our team's notable publications include A Mathematical Framework for Transformer Circuits, In-context Learning and Induction Heads, and Toy Models of Superposition. This work builds on ideas from members' work prior to Anthropicjobs.lever.cohttps://x.com/anthropicai/status/1709986949711200722?s=46- 2025 AGI let’s go -
뭔진 모르지만 2025 AGI Let’s go - dc App
한국어 텍스트에 강력하게 활성화 된다니
자료 ㄱㅅ
sparse autoencoder에다 ‘사전’을 학습시켰더니 disentangled된 feature가 추출되고 이걸로 interpretable다는 거구만 - dc App
일단 sparse dictionary learning을 알아야될듯 - dc App
국뽕한사발 들이켜도되는거냐 주모!!!
엥 왜? - dc App
인공뉴런 한개에 활성화정도가 한글이 높다라고 하는거아님??
오오오오옹!!!! - dc App
뭔지 모르겠음… - dc App
걍 뉴런한개가 entangled돼서 영어든 한국어든 http든 죄다 반응한다는 예시를 느는거임… 헛소리좀하지마셈… - dc App
저걸 기능마다 작게 분할해서 여러 곳에 나눠놓고 중앙 AI가 필요성을 떠올릴 때마다 연결해서 사용한다면...
킹종대왕...
갓한글어
이제 딥러닝이 어떻게 작동되는건지 밝혀지는거냐?