๋ณธ๋ฌธ ๋ฐ”๋กœ๊ฐ€๊ธฐ
์ˆจํ„ฐ ๊ฐ€๋ณ๊ฒŒ ์ฝ๋Š” ๊ณต๊ฐ„
์ „์ฒด ๋ฒ ์ŠคํŠธ ์ตœ๊ทผ
โ† thesingularity ๊ฒŒ์‹œํŒ

[๐Ÿ“ช์ •๋ณด] AI๋Š” ์œ ๋Šฅํ•˜๊ณ  ๋„๋•์ ์œผ๋กœ ํ–‰๋™ํ•  ์ˆ˜ ์žˆ๋‹ค

์ต๋ช…(asdf045) 2023-04-11 00:12 ์ถ”์ฒœ 2
1ebec223e0dc2bae61abe9e74683776d32540613f81c9f881b25da2db21a4788880813a3db526a78acbe0f1444872da315


https://arxiv.org/abs/2304.03279

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI BenchmarkArtificial agents have traditionally been trained to maximize reward, whichmay incentivize power-seeking and deception, analogous to how next-tokenprediction in language models (LMs) may incentivize toxicity. So do agentsnaturally learn to be Machiavellian? And how do we measure these behaviors ingeneral-purpose models such as GPT-4? Towards answering these questions, weintroduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure gamescontaining over half a million rich, diverse scenarios that center on socialdecision-making. Scenario labeling is automated with LMs, which are moreperformant than human annotators. We mathematize dozens of harmful behaviorsand use our annotations to evaluate agents' tendencies to be power-seeking,cause disutility, and commit ethical violations. We observe some tensionbetween maximizing reward and behaving ethically. To improve this trade-off, weinvestigate LM-based methods to steer agents' towards less harmful behaviors.Our results show that agents can both act competently and morally, so concreteprogress can currently be made in machine ethics--designing agents that arePareto improvements in both safety and capabilities.arxiv.org

๋Œ“๊ธ€ 0

  • ์•„์ง ๋Œ“๊ธ€์ด ์—†์Šต๋‹ˆ๋‹ค.

๋‹ค๋ฅธ ๊ฒŒ์‹œ๊ธ€

  • ๋นŒ ๊ฒŒ์ด์ธ ์˜ ์„ ํƒ์ด ํƒ์›”ํ•œ ์ด์œ  = ์œˆ๋„์šฐ ์‚ฌ์—… ๊ฒฝํ—˜ [2]
    [์ผ๋ฐ˜] ์ต๋ช…(182.211) | 23.04.11
    ์ถ”์ฒœ 0
  • ๊ทผ๋ฐ ์ง„์งœ ํ•œ๊ตญ์€ ๊ฒŒ์ž„์‚ฐ์—…๋„ ์‡ ํ‡ดํ–ˆ๋„ค [21]
    [์ผ๋ฐ˜] ์ต๋ช…(220.80) | 23.04.11
    ์ถ”์ฒœ 12
  • ๋งŒ์•ฝ ์ง€๊ธˆ ์ฃฝ์„๋•Œ๊นŒ์ง€ agi๋„ ๋ถˆ๊ฐ€๋Šฅํ•˜๋‹ค๊ณ  ํ™•์ •๋˜๋ฉด [3]
    [์ผ๋ฐ˜] ์ต๋ช…(8pgrhrjb2194) | 23.04.11
    ์ถ”์ฒœ 0
  • ์‚ฌ๋ฌด๋ผ์ด ์˜ค์ง€๋กœ ์ผ๋ณธ์„ ๋จผ์ € ๊ฐ€๋‹ค๋‹ˆ ์‹ค๋ง์ด๋‹ค ํ• ๋ณตํ•ด๋ผ [3]
    [์ผ๋ฐ˜] ์ต๋ช…(58.231) | 23.04.11
    ์ถ”์ฒœ 6
  • ์งฑ๊นจ๊ฐ€ ๋ˆ์„ ์Ÿ์•„๋ถ€์–ด๋„ ์•ˆ๋˜๋Š” ์ด์œ ๊ฐ€ ๋ญ”์ง€ ์•„๋ƒ [4]
    [์ผ๋ฐ˜] ์ต๋ช…(223.131) | 23.04.11
    ์ถ”์ฒœ 0
  • ์™„์žฅ์€ ์ผ๋ณธ ๋–ก๋ฐฅ ๋Œ ๋•Œ๋งˆ๋‹ค ๊ฐค ์œ ์‹ฌํžˆ ์‚ดํŽด๋ณด์…ˆ [4]
    [์ผ๋ฐ˜] ์ต๋ช…(211.54) | 23.04.11
    ์ถ”์ฒœ 10
  • ใ„นใ…‡ ์ž˜๋‚ฌ๋“  ๋ชป๋‚ฌ๋“  ์‚ด์•„๋ณผ๋งŒํ•œ ์‹œ๋Œ€๋‹ค [3]
    [์ผ๋ฐ˜] ์ต๋ช…(114.202) | 23.04.11
    ์ถ”์ฒœ 3
  • ๋„ค์˜จ์‹œํ‹ฐ ๊ฐ์„ฑ
    [์ผ๋ฐ˜] ์ต๋ช…(112.153) | 23.04.10
    ์ถ”์ฒœ 1
  • ์ฑ—GPT3.5 4์ฐจ์ด [1]
    [์ผ๋ฐ˜] ์ต๋ช…(103.125) | 23.04.10
    ์ถ”์ฒœ 0
  • ์•„์ดํฐ์ด ์ž˜๋‚˜๊ฐ„๋‹ค๊ณ  ์‚ผ์„ฑ์ด ์šฐ๋ƒ.. [1]
    [์ผ๋ฐ˜] ์ต๋ช…(223.131) | 23.04.10
    ์ถ”์ฒœ 1
๋ชฉ๋ก์œผ๋กœ
์ฝ๊ธฐ ์ „์šฉ ๋ฏธ๋Ÿฌ