Looking at the model’s internal feature activations, we noticed two things. (1) The model appeared to be internally aware that it was “holding back its true thoughts” and providing a fake summary. (2) The model seemed to interpret the results as a prompt injection attack. (3/7) pic.twitter.com/hueLTWdD0p
— Jack Lindsey (@Jack_W_Lindsey) November 25, 2025
흥미로운
기계적 해석가능성 연구의 선두주자답네
댓글 0