Javascript must be enabled to continue!
Reward centering in multi-agent reinforcement learning
View through CrossRef
Reward centering (RC) has recently emerged as a simple yet effective paradigm to stabilize training in single-agent reinforcement learning. However, its potential in multi-agent reinforcement learning (MARL) remains unexplored. In this paper, we first construct a reward centering value decomposition framework for MARL, in which global reward signals can be dynamically adjusted to effectively mitigate instability in value function estimation. Then a reward centering policy optimization framework for MARL is also generated, which aims to reduce the instability of policy updates. Based on these two reward centering frameworks, we propose five novel RC-based MARL methods. Extensive experimental results show that these methods can substantially outperform their baselines. Particularly, RC-QMIX and RC-MAPPO methods are best choices for off-policy and on-policy MARL scenarios respectively. This implies that our reward centering mechanisms can efficiently mitigate reward bias and improve the robustness of MARL methods. Our work provides a simple yet general mechanism for enhancing MARL methods and opens a perspective for reward engineering in complex multi-agent scenarios.
Title: Reward centering in multi-agent reinforcement learning
Description:
Reward centering (RC) has recently emerged as a simple yet effective paradigm to stabilize training in single-agent reinforcement learning.
However, its potential in multi-agent reinforcement learning (MARL) remains unexplored.
In this paper, we first construct a reward centering value decomposition framework for MARL, in which global reward signals can be dynamically adjusted to effectively mitigate instability in value function estimation.
Then a reward centering policy optimization framework for MARL is also generated, which aims to reduce the instability of policy updates.
Based on these two reward centering frameworks, we propose five novel RC-based MARL methods.
Extensive experimental results show that these methods can substantially outperform their baselines.
Particularly, RC-QMIX and RC-MAPPO methods are best choices for off-policy and on-policy MARL scenarios respectively.
This implies that our reward centering mechanisms can efficiently mitigate reward bias and improve the robustness of MARL methods.
Our work provides a simple yet general mechanism for enhancing MARL methods and opens a perspective for reward engineering in complex multi-agent scenarios.
Related Results
An examination of how reward associations differentially facilitate and impair Stroop performance
An examination of how reward associations differentially facilitate and impair Stroop performance
Behavioral performance is improved when the color of a Stroop stimulus is tied to a potential reward but is impaired when the irrelevant word meaning is reward related. The facilit...
Reward does not facilitate visual perceptual learning until sleep occurs
Reward does not facilitate visual perceptual learning until sleep occurs
ABSTRACTA growing body of evidence indicates that visual perceptual learning (VPL) is enhanced by reward provided during training. Another line of studies has shown that sleep foll...
An examination of how reward associations facilitate and impair Stroop performance
An examination of how reward associations facilitate and impair Stroop performance
Rewarded stimuli are prioritized by the attentional system. Behavioral performance is improved when the task-relevant dimension is tied to a potential reward but is impaired when t...
Examining the effects of reward and punishment on incidental learning
Examining the effects of reward and punishment on incidental learning
<p>Reward has been shown to improve multiple forms of learning. However, many of these studies do not distinguish whether reward directly benefits learning or if learning is ...
Explicit reward stabilizes motor output by attenuating sensory prediction error driven learning
Explicit reward stabilizes motor output by attenuating sensory prediction error driven learning
ABSTRACT
Motor adaptation driven by sensory prediction errors (SPEs) is often regarded as an automatic, implicit process that operates independen...
Differential and temporally dynamic involvement of primate amygdala nuclei in face animacy and reward information processing
Differential and temporally dynamic involvement of primate amygdala nuclei in face animacy and reward information processing
Abstract
Decision-making is influenced by both expected reward and social factors, such as who offered the outcomes. Thus, although a reward might originally be ind...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
STRENGTH OF BUTT WELDED BUTT JOINT OF REINFORCEMENT OF CLASS A500C
STRENGTH OF BUTT WELDED BUTT JOINT OF REINFORCEMENT OF CLASS A500C
The paper presents the results of experimental studies of the strength of cross-shaped welded joints of types К1-Кт and К3-Рр [1] of thermomechanically hardened reinforcement of cl...

