https://arxiv.org/abs/2306.13812v2Loss of Plasticity in Deep Continual LearningModern deep-learning systems are specialized to problem settings in whichtraining occurs once and then never again, as opposed to continual-learningsettings in which training occurs continually. If deep-learning systems areapplied in a continual learning setting, then it is well known that they mayfail to remember earlier examples. More fundamental, but less well known, isthat they may also lose their ability to learn on new examples, a phenomenoncalled loss of plasticity. We provide direct demonstrations of loss ofplasticity using the MNIST and ImageNet datasets repurposed for continuallearning as sequences of tasks. In ImageNet, binary classification performancedropped from 89\% accuracy on an early task down to 77\%, about the level of alinear network, on the 2000th task. Loss of plasticity occurred with a widerange of deep network architectures, optimizers, activation functions, batchnormalization, dropout, but was substantially eased by $L^2$-regularization,particularly when combined with weight perturbation. Further, we introduce anew algorithm -- continual backpropagation -- which slightly modifiesconventional backpropagation to reinitialize a small fraction of less-usedunits after each example and appears to maintain plasticity indefinitely.arxiv.org๋ฆฌ์ฒ๋ ์ํผ์ฐ์์ญ์ ํ
๋๊ธ 1