Archive/A Belief-Driven Hybrid Reinforcement Learning Framework for Decentralized Multi-Robot Navigation Under Partial Observability
A Belief-Driven Hybrid Reinforcement Learning Framework for Decentralized Multi-Robot Navigation Under Partial Observability
Vineetha Malathi, Pramod Sreedharan, Rthuraj Puthiyaveedu Puthiyaveedu Rajesh et al.
28 de julio de 2026
en

Abstract

Decentralized multi-robot navigation is difficult when robots must act from local observations without centralized coordination or explicit inter-robot communication. A belief-driven hybrid reinforcement learning framework is evaluated for planar multi-robot navigation under partial observability. Each robot builds a compact local state from its position, waypoint target, sector-based proximity readings, and a decaying occupancy belief that summarizes recent obstacle evidence. A Deep Deterministic Policy Gradient (DDPG) actor produces continuous velocity proposals, and a lightweight geometric safety-blending layer combines this command with goal-seeking and reactive avoidance vectors before execution. The simulation was revised to use e-puck-compatible heading-limited forward motion rather than side-slip motion. The framework is intentionally solver-free at runtime and does not introduce online constrained optimization or new communication mechanisms. The evaluation reports a controlled five-seed study using seeds 101–105 and a 10-seed stress suite covering scalability, symmetric crossing, corridor, and dense dynamic-obstacle cases. In the controlled nominal evaluation, full three-robot completion occurred in all five runs, with 100.0% mean success, 142.2 mean steps, and no recorded collision timestep. In the hybrid stress suite, nominal, four-robot swap, five-robot crossing, symmetric-deadlock, and corridor cases achieved full success in all 10 seeds. Dense dynamic obstacles were the main failure case, with 5/10 full-success runs, 5 robot timeouts, and 10.1 mean collision events per run. These results support the feasibility of the hybrid structure in moderate tested conditions while showing that dense moving obstacles remain a practical limitation. Formal safety guarantees, matched benchmark comparisons, physical robot validation, and wider randomization remain areas requiring future work.

IPC Classification

H04

Keywords

belief-drivenhybridreinforcementlearningframeworkdecentralizedmulti-robotnavigationpartialobservabilityroboticsdifficultwhenrobotsmustlocalobservationswithoutcentralizedcoordinationexplicitinter-robotcommunicationevaluated
Citar esta publicación

€ 4.00