1. κΉμ κ°ννμ΅(RL)μ λ°μ μ μΈκ° μ°κ΅¬μλ€μ΄ μλμΌλ‘ μ€κ³ν μκ³ λ¦¬μ¦μ κΈ°λ°νμμΌλ, μ΅κ·Όμλ λ©ν νμ΅ μ λ°μ΄νΈ κ·μΉμ ν΅ν΄ λ€μν RL μμ μμ μ μνν μ μλ μκ³ λ¦¬μ¦μ λ°κ²¬νλ κ²μ΄ κ°λ₯νλ€λ κ²μ΄ λ°νμ‘λ€.
2. Unsupervised Environment Design (UED)μμ μκ°μ λ°μ, μ°λ¦¬λ λ©ν νμ΅λ μ΅μ νκΈ°μ ννλ₯Ό κ·ΉλννκΈ° μν΄ μλμΌλ‘ κ΅μ‘ κ³Όμ μ μμ±νλ μλ‘μ΄ μ κ·Όλ²κ³Ό "μκ³ λ¦¬μ¦μ νν"λΌλ μλ‘μ΄ νν κ·Όμ¬λ₯Ό μ μνλ€.
3. μ€νμ ν΅ν΄, μ°λ¦¬μ λ°©λ² GROOVEλ LPGλ³΄λ€ μ°μν μΌλ°νλ₯Ό λ¬μ±νλ©°, μ΄ μ κ·Όλ²μ μ€μ μΈκ³μ λ€μν νκ²½μ ν΄κ²°ν μ μλ μ§μ ν μΌλ° RL μκ³ λ¦¬μ¦μ λ°κ²¬μ ν₯ν ν κ±Έμμ΄λΌκ³ λ―Ώλλ€.
https://arxiv.org/abs/2310.02782Discovering General Reinforcement Learning Algorithms with Adversarial Environment DesignThe past decade has seen vast progress in deep reinforcement learning (RL) on the back of algorithms manually designed by human researchers. Recently, it has been shown that it is possible to meta-learn update rules, with the hope of discovering algorithms that can perform well on a wide range of RL tasks. Despite impressive initial results from algorithms such as Learned Policy Gradient (LPG), there remains a generalization gap when these algorithms are applied to unseen environments. In this work, we examine how characteristics of the meta-training distribution impact the generalization performance of these algorithms. Motivated by this analysis and building on ideas from Unsupervised Environment Design (UED), we propose a novel approach for automatically generating curricula to maximize the regret of a meta-learned optimizer, in addition to a novel approximation of regret, which we name algorithmic regret (AR). The result is our method, General RL Optimizers Obtained Via Environment Design (GROOVE). In a series of experiments, we show that GROOVE achieves superior generalization to LPG, and evaluate AR against baseline metrics from UED, identifying it as a critical component of environment design in this setting. We believe this approach is a step towards the discovery of truly general RL algorithms, capable of solving a wide range of real-world environments.arxiv.org
μ€λ λ λ μΈκ° κ³μ μμμ§λ€
βwe believe this approach is a step towards the discovery of truly general RLβ μΊ¬ γ·γ· - dc App