I am not a doomer. Misaligned superintelligence is probably not the biggest AI risk. But I did spend the past year working on technical research on aligning AI systems as my day-job at OpenAI, working with Ilya and the Superalignment team. There is a very real technical problem: our current alignment techniques (methods to ensure we can reliably control, steer, and trust AI systems) won’t scale to superhuman AI systems. What I want to do is explain what I see as the “default” plan for how we’ll muddle through, and why I’m optimistic. While not enough people are on the ball—we should have much more ambitious efforts to solve this problem!—overall, we’ve gotten lucky with how deep learning has shaken out, there’s a lot of empirical low-hanging fruit that will get us part of the way, and we’ll have the advantage of millions of automated AI researchers to get us the rest of the way. 


But I also want to tell you why I’m worried. Most of all, ensuring alignment doesn’t go awry will require extreme competence in managing the intelligence explosion. If we do rapidly transition from AGI to superintelligence, we will face a situation where, in less than a year, we will go from recognizable human-level systems for which descendants of current alignment techniques will mostly work fine, to much more alien, vastly superhuman systems that pose a qualitatively different, fundamentally novel technical alignment problem; at the same time, going from systems where failure is low-stakes to extremely powerful systems where failure could be catastrophic; all while most of the world is probably going kind of crazy. It makes me pretty nervous. 


By the time the decade is out, we’ll have billions of vastly superhuman AI agents running around. These superhuman AI agents will be capable of extremely complex and creative behavior; we will have no hope of following along. We’ll be like first graders trying to supervise people with multiple doctorates.


In essence, we face a problem of handing off trust. By the end of the intelligence explosion, we won’t have any hope of understanding what our billion superintelligences are doing (except as they might choose to explain to us, like they might to a child). And we don’t yet have the technical ability to reliably guarantee even basic side constraints for these systems, like “don’t lie” or “follow the law” or “don’t try to exfiltrate your server”. Reinforcement from human feedback (RLHF) works very well for adding such side constraints for current systems—but RLHF relies on humans being able to understand and supervise AI behavior, which fundamentally won’t scale to superhuman systems. 


Simply put, without a very concerted effort, we won’t be able to guarantee that superintelligence won’t go rogue (and this is acknowledged by many leaders in the field). Yes, it may all be fine by default. But we simply don’t know yet. Especially once future AI systems aren’t just trained with imitation learning, but large-scale, long-horizon RL (reinforcement learning), they will acquire unpredictable behaviors of their own, shaped by a trial-and-error process (for example, they may learn to lie or seek power, simply because these are successful strategies in the real world!). 


The stakes will be high enough that hoping for the best simply isn’t a good enough answer on alignment.







... OpenAI에서 일과로 AI 시스템을 정렬하는 기술 연구를 하는 데 지난해를 보냈고, Ilya와 Superalignment 팀과 함께 일했습니다. 매우 현실적인 기술적 문제가 있습니다. 현재의 정렬 기술(AI 시스템을 안정적으로 제어, 조종, 신뢰할 수 있도록 하는 방법)은 초월적인 AI 시스템에는 확장할 수 없습니다...



...인간 피드백 (RLHF)을 통한 강화는 현재 시스템에 이러한 부수적 제약을 추가하는 데 매우 효과적이지만, RLHF는 인간이 AI 행동을 이해하고 감독할 수 있어야 하며, 이는 근본적으로 초인적 시스템으로 확장되지 않습니다...



...모방 학습으로만 훈련되지 않고 대규모, 장기적 RL(강화 학습)로 훈련되면 시행착오 과정을 통해 형성된 예측할 수 없는 자체 행동을 습득하게 될 것입니다...







24b0d121e09c28a8699fe8b115ef046c65f2