Javascript must be enabled to continue!
ADAPTIVE KNOWLEDGE REGULARIZATION FOR CONTINUAL LEARNING IN TRANSFORMER ARCHITECTURES
View through CrossRef
The subject matter of the article is development of latent representation regularization mechanism in transformer-based architecture under conditions of continuous learning with domain shifts. Modern language models achieve high quality in static learning scenarios, but they remain limited in long-term operation cases, incremental adaptation to new domains and lacking resistance to catastrophic forgetting, especially in the absence of access to previously observed data. This paper explores the possibility of overcoming these limitations by combining uncertainty inspired regularization and forgetting attention mechanisms into one transformer architecture. The goal of the study is to design, implement and validate transformer-based architecture with multi-level representation regularization mechanism that can help transformer-based language models efficiently adapt to alternative data distribution while retaining previously acquired knowledge. The proposed approach aims to achieve an equilibrium between model adaptability and stability in continual learning without requiring complete model retraining or legacy data retention. The tasks to be solved in this study include: formalization of an unified conceptual latent representations regularization method that combines Bayesian uncertainty inspired latent representations regularization with head-wise attention scaling in attention mechanism; implementation of this method into the transformer model; creating the experimental case of continuous language modeling with sequential domain shifts; give quantitative estimation of model forgetting and stability in prediction quality; compare the proposed model to classical naive fine-tuning, LoRA and parameter regularization methods. The conclusions demonstrate that the proposed method achieves lower forgetting, lower perplexity on previously learned domains and a better stability–plasticity trade-off than naive fine-tuning, LoRA and Elastic Weight Consolidation, while requiring comparable computational resources. The scientific novelty of proposed approach consists in development of layer-selective latent regularization framework for continual language modeling which integrates an attention with forgetting mechanism with preserving domain-invariant representations through statistical alignment in latent space for reducing forgetting in continual learning scenarios. Unlike existing approaches to continual learning that consider model regularization either on parameter or memory level (by using previous data), the proposed approach moves regularization into representation space, where it uses both direct regularization via proposed uncertainty-based regularization and indirect regularization via attention with forget gate, ensuring the models’ possibility of stable continual language modeling in non-stationary environments.
National Aerospace University - Kharkiv Aviation Institute
Title: ADAPTIVE KNOWLEDGE REGULARIZATION FOR CONTINUAL LEARNING IN TRANSFORMER ARCHITECTURES
Description:
The subject matter of the article is development of latent representation regularization mechanism in transformer-based architecture under conditions of continuous learning with domain shifts.
Modern language models achieve high quality in static learning scenarios, but they remain limited in long-term operation cases, incremental adaptation to new domains and lacking resistance to catastrophic forgetting, especially in the absence of access to previously observed data.
This paper explores the possibility of overcoming these limitations by combining uncertainty inspired regularization and forgetting attention mechanisms into one transformer architecture.
The goal of the study is to design, implement and validate transformer-based architecture with multi-level representation regularization mechanism that can help transformer-based language models efficiently adapt to alternative data distribution while retaining previously acquired knowledge.
The proposed approach aims to achieve an equilibrium between model adaptability and stability in continual learning without requiring complete model retraining or legacy data retention.
The tasks to be solved in this study include: formalization of an unified conceptual latent representations regularization method that combines Bayesian uncertainty inspired latent representations regularization with head-wise attention scaling in attention mechanism; implementation of this method into the transformer model; creating the experimental case of continuous language modeling with sequential domain shifts; give quantitative estimation of model forgetting and stability in prediction quality; compare the proposed model to classical naive fine-tuning, LoRA and parameter regularization methods.
The conclusions demonstrate that the proposed method achieves lower forgetting, lower perplexity on previously learned domains and a better stability–plasticity trade-off than naive fine-tuning, LoRA and Elastic Weight Consolidation, while requiring comparable computational resources.
The scientific novelty of proposed approach consists in development of layer-selective latent regularization framework for continual language modeling which integrates an attention with forgetting mechanism with preserving domain-invariant representations through statistical alignment in latent space for reducing forgetting in continual learning scenarios.
Unlike existing approaches to continual learning that consider model regularization either on parameter or memory level (by using previous data), the proposed approach moves regularization into representation space, where it uses both direct regularization via proposed uncertainty-based regularization and indirect regularization via attention with forget gate, ensuring the models’ possibility of stable continual language modeling in non-stationary environments.
Related Results
A Mixed Regularization Method for Ill-Posed Problems
A Mixed Regularization Method for Ill-Posed Problems
In this paper we propose a mixed regularization method for ill-posed problems. This method combines iterative regularization methods and continuous regularization methods effective...
Automatic Load Sharing of Transformer
Automatic Load Sharing of Transformer
Transformer plays a major role in the power system. It works 24 hours a day and provides power to the load. The transformer is excessive full, its windings are overheated which lea...
Continual Learning of Large Language Models: A Comprehensive Survey
Continual Learning of Large Language Models: A Comprehensive Survey
The challenge of effectively and efficiently adapting statically pre-trained Large Language Models (LLMs) to ever-evolving data distributions remains predominant. When tailored for...
High frequency modeling of power transformers under transients
High frequency modeling of power transformers under transients
This thesis presents the results related to high frequency modeling of power transformers. First, a 25kVA distribution transformer under lightning surges is tested in the laborator...
Continual Learning: Overcoming Catastrophic Forgetting for Adaptive AI Systems
Continual Learning: Overcoming Catastrophic Forgetting for Adaptive AI Systems
Continual learning is a fundamental challenge in artificial intelligence (AI) that aims to enable models to learn from a continuous stream of data while retaining previously acqui...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Study of Transformer Lifetime Due to Loading Process on 20 KV Distribution Line
Study of Transformer Lifetime Due to Loading Process on 20 KV Distribution Line
Power transformer is very important in electric power system due to its function to raise or lower the voltage according to its designation. On the power side, the power transforme...
ANALISIS PENGARUH MASA OPERASIONAL TERHADAP PENURUNAN KAPASITAS TRANSFORMATOR DISTRIBUSI DI PT PLN (PERSERO)
ANALISIS PENGARUH MASA OPERASIONAL TERHADAP PENURUNAN KAPASITAS TRANSFORMATOR DISTRIBUSI DI PT PLN (PERSERO)
One cause the interruption of transformer is loading that exceeds the capabilities of the transformer. The state of continuous overload will affect the age of the transformer and r...

