Javascript must be enabled to continue!
Adaptive Replay Strategies Stabilize Multi-Agent Reinforcement Learning for Differential-Drive Robot Coordination
View through CrossRef
Background: Coordinating differential-drive mobile robots for landmark coverage is challenging due to non-holonomic dynamics, clutter, and sparse rewards. Standard multi-agent RL pipelines often show unstable learning and inconsistent completion in this setting. Methodology: We adopt a centralized-training, decentralized-execution actor-critic without inter-agent commu nication. Our replay-centric design combines a tagged buffer that up-samples goal-reaching transitions and an offline replay initialization that seeds early learning with curated trajectories. Dynamic task assignment uses the Hungarian algorithm during training and evaluation, and we benchmark against uniform replay and established variants. Results: In a cluttered six-robot arena, the approach improves training stability relative to uniform replay. Convergence is faster and requires fewer updates to reach consistent success. Coverage efficiency increases as landmarks are reached earlier across runs. Collisions per episode decrease without adding communication or architectural changes. Multi seed evaluations show gains that persist with narrow confidence intervals. Train-evaluation gaps shrink on unseen maps, indicating improved generalization. Ablations attribute complementary ben efits to the tagged and offline components. Performance remains competitive with prioritized and hindsight replay baselines under matched budgets. Computational overhead is small because sam pling logic changes while network sizes do not. Conclusions: Focus ing on replay design substantially stabilizes multi-agent learning for differential-drive coordination. The pipeline integrates cleanly with standard CTDE implementations and supports practical deployment in coverage tasks.
Universitas Muhammadiyah Yogyakarta
Title: Adaptive Replay Strategies Stabilize Multi-Agent Reinforcement Learning for Differential-Drive Robot Coordination
Description:
Background: Coordinating differential-drive mobile robots for landmark coverage is challenging due to non-holonomic dynamics, clutter, and sparse rewards.
Standard multi-agent RL pipelines often show unstable learning and inconsistent completion in this setting.
Methodology: We adopt a centralized-training, decentralized-execution actor-critic without inter-agent commu nication.
Our replay-centric design combines a tagged buffer that up-samples goal-reaching transitions and an offline replay initialization that seeds early learning with curated trajectories.
Dynamic task assignment uses the Hungarian algorithm during training and evaluation, and we benchmark against uniform replay and established variants.
Results: In a cluttered six-robot arena, the approach improves training stability relative to uniform replay.
Convergence is faster and requires fewer updates to reach consistent success.
Coverage efficiency increases as landmarks are reached earlier across runs.
Collisions per episode decrease without adding communication or architectural changes.
Multi seed evaluations show gains that persist with narrow confidence intervals.
Train-evaluation gaps shrink on unseen maps, indicating improved generalization.
Ablations attribute complementary ben efits to the tagged and offline components.
Performance remains competitive with prioritized and hindsight replay baselines under matched budgets.
Computational overhead is small because sam pling logic changes while network sizes do not.
Conclusions: Focus ing on replay design substantially stabilizes multi-agent learning for differential-drive coordination.
The pipeline integrates cleanly with standard CTDE implementations and supports practical deployment in coverage tasks.
Related Results
Designing a robot to evaluate group formations
Designing a robot to evaluate group formations
Robots are making their way in environments inhabited by people. Whether in domestic or public crowded environments, robots should take into consideration social norms and behavior...
PELATIHAN PERANCANGAN ROBOT BERODA DENGAN DETEKTOR TEPI MEJA PADA SEKOLAH SMA TARSISIUS 1 DAN SMA TRI RATNA
PELATIHAN PERANCANGAN ROBOT BERODA DENGAN DETEKTOR TEPI MEJA PADA SEKOLAH SMA TARSISIUS 1 DAN SMA TRI RATNA
Wheeled robot is a robot which movement is managed by the rotation of Direct Current (DC) motors. These motors are connected to wheels. Wheeled robot usually is used as a teaching ...
Sistem Kendali Hybrid Fuzzy-Pid pada Kinematika Robot Berkaki 4 Menggunakan Sensor Gyroscope
Sistem Kendali Hybrid Fuzzy-Pid pada Kinematika Robot Berkaki 4 Menggunakan Sensor Gyroscope
<p><em>Legged robots have attracted the attention of researchers because of their superior adaptation to complex environments compared to wheeled robots. Legged robots ...
Post-learning replay of hippocampal-striatal activity is biased by reward-prediction signals
Post-learning replay of hippocampal-striatal activity is biased by reward-prediction signals
Abstract
Neural activity encoding recent experiences is replayed during sleep and rest to promote consolidation of memories. However, precisely which features of ex...
Evaluating hippocampal replay without a ground truth
Evaluating hippocampal replay without a ground truth
AbstractDuring rest and sleep, memory traces replay in the brain. The dialogue between brain regions during replay is thought to stabilize labile memory traces for long-term storag...
Theta-band phase locking during encoding leads to coordinated entorhinal-hippocampal replay
Theta-band phase locking during encoding leads to coordinated entorhinal-hippocampal replay
Abstract
Precisely timed interactions between hippocampal and cortical neurons during replay epochs are thought to support memory consolidation. ...
Decision Making in Human-Robot Interaction
Decision Making in Human-Robot Interaction
Processus décisionnels pour l'interaction homme-robot
Un intérêt croissant est aujourd'hui porté sur les robots capables de conduire des activités de collaboration ...

