https://arxiv.org/abs/2304.03279Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI BenchmarkArtificial agents have traditionally been trained to maximize reward, whichmay incentivize power-seeking and deception, analogous to how next-tokenprediction in language models (LMs) may incentivize toxicity. So do agentsnaturally learn to be Machiavellian? And how do we measure these behaviors ingeneral-purpose models such as GPT-4? Towards answering these questions, weintroduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure gamescontaining over half a million rich, diverse scenarios that center on socialdecision-making. Scenario labeling is automated with LMs, which are moreperformant than human annotators. We mathematize dozens of harmful behaviorsand use our annotations to evaluate agents' tendencies to be power-seeking,cause disutility, and commit ethical violations. We observe some tensionbetween maximizing reward and behaving ethically. To improve this trade-off, weinvestigate LM-based methods to steer agents' towards less harmful behaviors.Our results show that agents can both act competently and morally, so concreteprogress can currently be made in machine ethics--designing agents that arePareto improvements in both safety and capabilities.arxiv.org
https://arxiv.org/abs/2304.03279Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI BenchmarkArtificial agents have traditionally been trained to maximize reward, whichmay incentivize power-seeking and deception, analogous to how next-tokenprediction in language models (LMs) may incentivize toxicity. So do agentsnaturally learn to be Machiavellian? And how do we measure these behaviors ingeneral-purpose models such as GPT-4? Towards answering these questions, weintroduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure gamescontaining over half a million rich, diverse scenarios that center on socialdecision-making. Scenario labeling is automated with LMs, which are moreperformant than human annotators. We mathematize dozens of harmful behaviorsand use our annotations to evaluate agents' tendencies to be power-seeking,cause disutility, and commit ethical violations. We observe some tensionbetween maximizing reward and behaving ethically. To improve this trade-off, weinvestigate LM-based methods to steer agents' towards less harmful behaviors.Our results show that agents can both act competently and morally, so concreteprogress can currently be made in machine ethics--designing agents that arePareto improvements in both safety and capabilities.arxiv.org
๋๊ธ 0