Skip to main navigation Skip to search Skip to main content

AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning

  • Pengxiang Wang*
  • , Hongbo Bo
  • , Jun Hong
  • , Weiru Liu
  • , Kedian Mu
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference Contribution (Conference Proceeding)

40 Downloads (Pure)

Abstract

Catastrophic forgetting is a major issue in task-incremental learning, where a neural network loses what it has learned in previous tasks after being trained on new tasks. A number of architecture-based approaches have been proposed to address this issue. However, the architecture-based approaches suffer from another issue on network capacity when the network learns long sequences of tasks. As the network is trained on an increasing number of new tasks in a long sequence of tasks, more parameters become static to prevent the network from forgetting what it has learned in previous tasks. In this paper, we propose an adaptive task-based hard attention mechanism which allows adaptive updates to static parameters by taking into account the information about previous tasks on both the importance of these parameters to previous tasks and the current network capacity. We develop a new neural network architecture incorporating our proposed Adaptive Hard Attention to the Task (AdaHAT) mechanism. AdaHAT extends an existing architecture-based approach, Hard Attention to the Task (HAT), to learn long sequences of tasks in an incremental manner. We conduct experiments on a number of datasets and compare AdaHAT with a number of baselines, including HAT. Our experimental results show that AdaHAT achieves better average performance over tasks than these baselines, especially on long sequences of tasks, demonstrating the benefits from balancing the trade-off between stability and plasticity of a network when learning such sequences of tasks. Our code is available at github.com/pengxiang-wang/continual-learning-arena.
Original languageEnglish
Title of host publicationMachine Learning and Knowledge Discovery in Databases
Subtitle of host publication Research Track
EditorsAlbert Bifet, Jesse Davis, Tomas Krilavičius, Meelis Kull, Eirini Ntoutsi, Indrė Žliobaitė
PublisherSpringer
Chapter9
Pages 143–160
Number of pages18
Volume14943
ISBN (Electronic)9783031703522
ISBN (Print)9783031703515
DOIs
Publication statusPublished - 22 Aug 2024
Event2024 Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD) -
Duration: 9 Sept 202413 Sept 2024

Publication series

NameLecture Notes in Computer Science
PublisherSpringer
Volume14943
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference2024 Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD)
Period9/09/2413/09/24

Fingerprint

Dive into the research topics of 'AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning'. Together they form a unique fingerprint.

Cite this