Javascript must be enabled to continue!
Modeling Unlearning and Relearning with Multi-agent Q-Learning Systems
View through CrossRef
We model unlearning by simulating a Q-agent (using the reinforcement learning Qlearning algorithm), representing a real-world learner, playing the game of Nim against different adversarial agents to learn the optimal Nim strategy. When the Q-agent plays against sub-optimal agents, its percentage of optimal moves is decreased, analogous to a person forgetting (“unlearning”) what they have learned previously. To mitigate the effect of this “unlearning”, we experimented with modulating the Q-learning so that minimal learning occurs with untrusted opponents. This trust-based modulation is modeled by observing opponent moves that are different from those that a Q-agent has learned. This model parallels human trust which tends to increase with those whom one agrees with. With this modulated learning, we observe that a Q-agent with a baseline optimal strategy is able to robustly retain previously learned strategy, in some cases achieving a 0.3 difference in accuracy from the unlearning model. We then ran a three-phase simulation where the Qagent played against optimal agents in the first phase, sub-optimal agents in the second “unlearning” phase, and optimal or random agents in the third phase. We found that even after unlearning, the Q-agent was quickly able to relearn most of its knowledge about the optimal strategy for Nim.
Academy & Industry Research Collaboration Center
Title: Modeling Unlearning and Relearning with Multi-agent Q-Learning Systems
Description:
We model unlearning by simulating a Q-agent (using the reinforcement learning Qlearning algorithm), representing a real-world learner, playing the game of Nim against different adversarial agents to learn the optimal Nim strategy.
When the Q-agent plays against sub-optimal agents, its percentage of optimal moves is decreased, analogous to a person forgetting (“unlearning”) what they have learned previously.
To mitigate the effect of this “unlearning”, we experimented with modulating the Q-learning so that minimal learning occurs with untrusted opponents.
This trust-based modulation is modeled by observing opponent moves that are different from those that a Q-agent has learned.
This model parallels human trust which tends to increase with those whom one agrees with.
With this modulated learning, we observe that a Q-agent with a baseline optimal strategy is able to robustly retain previously learned strategy, in some cases achieving a 0.
3 difference in accuracy from the unlearning model.
We then ran a three-phase simulation where the Qagent played against optimal agents in the first phase, sub-optimal agents in the second “unlearning” phase, and optimal or random agents in the third phase.
We found that even after unlearning, the Q-agent was quickly able to relearn most of its knowledge about the optimal strategy for Nim.
Related Results
UNLEARNING UNSUSTAINABILITY
UNLEARNING UNSUSTAINABILITY
There is an increased urge to facilitate a transformation of the Dutch food to address pressing sustainability challenges. At present, these calls for transformation are most often...
Unlearning in AI: Techniques and Frameworks for Data Deletion in Pretrained Models Under Legal and Ethical Constraints
Unlearning in AI: Techniques and Frameworks for Data Deletion in Pretrained Models Under Legal and Ethical Constraints
Abstract: The rapid expansion of the AI revolution has been propelled by a focus on large-scale pretrained models, which have enabled significant advancements across diverse tasks ...
A survey on large language models unlearning: taxonomy, evaluations, and future directions
A survey on large language models unlearning: taxonomy, evaluations, and future directions
Abstract
Following the introduction of data privacy regulations and “the right to be forgotten”, large language models (LLMs) unlearning has emerged as a promisin...
Evaluation Metrics for Machine Unlearning
Evaluation Metrics for Machine Unlearning
The evaluation of machine unlearning has become increasingly significant as machine learning systems face growing demands for privacy, security, and regulatory compliance. This pap...
Exploring linkages between unlearning and human resource development: Revisiting unlearning cases
Exploring linkages between unlearning and human resource development: Revisiting unlearning cases
AbstractThe purpose of this study was to review unlearning cases and to identify and suggest what roles human resource development (HRD) can play in the unlearning process. By adop...
Intentional unlearning practices in postmassified university systems: Reformation for the metamodern era
Intentional unlearning practices in postmassified university systems: Reformation for the metamodern era
A crucial aspect of the learning cycle, unlearning has recently received more attention in academic discussions about the future of higher education. In an attempt to improve equal...
Route Learning and Transport of Resources during Colony Relocation in Australian Desert Ants
Route Learning and Transport of Resources during Colony Relocation in Australian Desert Ants
Abstract
Many ant species are able to respond to dramatic changes in local conditions by relocating the entire colony to a new location. While we...
Organizational unlearning: A risky food safety strategy?
Organizational unlearning: A risky food safety strategy?
Abstract
Strategically unlearning specific knowledge, behaviors, and practices facilitates product and process innovation, business model evolution, and new marke...

