Javascript must be enabled to continue!
Multi-Agent Natural Actor-Critic Reinforcement Learning Algorithms
View through CrossRef
AbstractMulti-agent actor-critic algorithms are an important part of the Reinforcement Learning (RL) paradigm. We propose three fully decentralized multi-agent natural actor-critic (MAN) algorithms in this work. The objective is to collectively find a joint policy that maximizes the average long-term return of these agents. In the absence of a central controller and to preserve privacy, agents communicate some information to their neighbors via a time-varying communication network. We prove convergence of all the three MAN algorithms to a globally asymptotically stable set of the ODE corresponding to actor update; these use linear function approximations. We show that the Kullback–Leibler divergence between policies of successive iterates is proportional to the objective function’s gradient. We observe that the minimum singular value of the Fisher information matrix is well within the reciprocal of the policy parameter dimension. Using this, we theoretically show that the optimal value of the deterministic variant of the MAN algorithm at each iterate dominates that of the standard gradient-based multi-agent actor-critic (MAAC) algorithm. To our knowledge, it is the first such result in multi-agent reinforcement learning (MARL). To illustrate the usefulness of our proposed algorithms, we implement them on a bi-lane traffic network to reduce the average network congestion. We observe an almost 25% reduction in the average congestion in 2 MAN algorithms; the average congestion in another MAN algorithm is on par with the MAAC algorithm. We also consider a generic 15 agent MARL; the performance of the MAN algorithms is again as good as the MAAC algorithm.
Springer Science and Business Media LLC
Title: Multi-Agent Natural Actor-Critic Reinforcement Learning Algorithms
Description:
AbstractMulti-agent actor-critic algorithms are an important part of the Reinforcement Learning (RL) paradigm.
We propose three fully decentralized multi-agent natural actor-critic (MAN) algorithms in this work.
The objective is to collectively find a joint policy that maximizes the average long-term return of these agents.
In the absence of a central controller and to preserve privacy, agents communicate some information to their neighbors via a time-varying communication network.
We prove convergence of all the three MAN algorithms to a globally asymptotically stable set of the ODE corresponding to actor update; these use linear function approximations.
We show that the Kullback–Leibler divergence between policies of successive iterates is proportional to the objective function’s gradient.
We observe that the minimum singular value of the Fisher information matrix is well within the reciprocal of the policy parameter dimension.
Using this, we theoretically show that the optimal value of the deterministic variant of the MAN algorithm at each iterate dominates that of the standard gradient-based multi-agent actor-critic (MAAC) algorithm.
To our knowledge, it is the first such result in multi-agent reinforcement learning (MARL).
To illustrate the usefulness of our proposed algorithms, we implement them on a bi-lane traffic network to reduce the average network congestion.
We observe an almost 25% reduction in the average congestion in 2 MAN algorithms; the average congestion in another MAN algorithm is on par with the MAAC algorithm.
We also consider a generic 15 agent MARL; the performance of the MAN algorithms is again as good as the MAAC algorithm.
Related Results
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
STRENGTH OF BUTT WELDED BUTT JOINT OF REINFORCEMENT OF CLASS A500C
STRENGTH OF BUTT WELDED BUTT JOINT OF REINFORCEMENT OF CLASS A500C
The paper presents the results of experimental studies of the strength of cross-shaped welded joints of types К1-Кт and К3-Рр [1] of thermomechanically hardened reinforcement of cl...
Fiber reinforcement as an alternative to the compressed zone linear reinforcement and the flexible concrete elements stretched zone prestressing
Fiber reinforcement as an alternative to the compressed zone linear reinforcement and the flexible concrete elements stretched zone prestressing
Abstract
The results of a numerical experiment in the framework of a theoretical study of the strength and crack resistance of the reinforced concrete beams availabl...
FLDQN: Cooperative Multi-Agent Federated Reinforcement Learning for Solving Travel Time Minimization Problems in Dynamic Environments Using SUMO Simulation
FLDQN: Cooperative Multi-Agent Federated Reinforcement Learning for Solving Travel Time Minimization Problems in Dynamic Environments Using SUMO Simulation
The increasing volume of traffic has led to severe challenges, including traffic congestion, heightened energy consumption, increased air pollution, and prolonged travel times. Add...
Conflict-Based Search for Optimal Multi-Agent Pathfinding
Conflict-Based Search for Optimal Multi-Agent Pathfinding
We are assumed a customary of mediators in the multi-agent pathfinding problem (MAPF), each of which has its own start and goal positions. The objective is to discovery paths for e...
Decentralized Multi-Agent Deep Reinforcement Learning: A Competitive-Game Perspective
Decentralized Multi-Agent Deep Reinforcement Learning: A Competitive-Game Perspective
Abstract
Deep reinforcement learning (DRL) has been widely studied in single agent learning but require further development and understanding in the multi-agent field. As ...
Polysaccharides in Asparagus and Asparagus Juice
Polysaccharides in Asparagus and Asparagus Juice
The polysaccharides in asparagus are additionally peremptory to incorporate into this area on cancer prevention agent and calming medical advantages. Polysaccharides are an excepti...
Hybrid Fuzzy PID with Soft Actor-Critic Reinforcement Learning for Wind Turbine Pitch Control
Hybrid Fuzzy PID with Soft Actor-Critic Reinforcement Learning for Wind Turbine Pitch Control
This paper proposes a novel hybrid intelligent control framework integrating a fuzzy supervisory Proportional-Integral-Derivative (PID) controller with Soft Actor-Critic (SAC) rein...

