Archive/PPO-GAT-Follow: Graph-Attention Reinforcement Learning for Robust Robot Person Following in Dense Crowds
PPO-GAT-Follow: Graph-Attention Reinforcement Learning for Robust Robot Person Following in Dense Crowds
Xinyu Zhou, Yongliang Shi, Songhao Piao et al.
24 juillet 2026
en

Abstract

Robot person following (RPF) in dense crowds requires a mobile robot to maintain an appropriate relative position with respect to a moving target while avoiding surrounding pedestrians and satisfying rear-following and social constraints. This paper proposes PPO-GAT-Follow, an interaction-aware reinforcement learning framework for dense-crowd RPF under geometric visibility loss with available target-relative pose estimates. The follower, target pedestrian, and surrounding pedestrians are represented as graph nodes, and a graph attention encoder models their local interactions. A task-oriented reward mechanism jointly accounts for target maintenance, visibility preservation, collision avoidance, proximity-aware social compliance, rear position maintenance, post-arrival stabilization, and action stability. Experiments are conducted in IR-SIM under fixed-route and random-route settings, with comparisons against MPC, DWA, SFM, and an adapted SARL baseline. In the fixed-route setting with 12 background pedestrians, PPO-GAT-Follow achieves a task success rate of 98.8% and a collision rate of 1.1%, improving task success by 10.9 percentage points over MPC. In the random-route setting at the training density, it achieves 83.1% task success and an SPL of 0.815, outperforming MPC by 18.3 percentage points in task success; at this density, it also surpasses SARL in the main task-level metrics. Zero-shot evaluations across crowd densities, together with structural and reward ablations, reward weight sensitivity analysis, tolerance shift tests, multi-seed training, and stress testing under target pose noise and heterogeneous pedestrian dynamics, further demonstrate the effectiveness and reliability of the proposed framework. Gazebo-based validation also demonstrates system integration feasibility with localization, point cloud-based surrounding pedestrian perception, tracking, and UWB-like target-relative pose input. Nevertheless, visual target identification, re-identification, and perception-level occlusion recovery remain outside the scope of the present validation.

IPC Classification

G06

Keywords

ppo-gat-followgraph-attentionreinforcementlearningrobustrobotpersonfollowingdensecrowdssensorsrequiresmobilemaintainappropriaterelativepositionrespectmovingtargetwhileavoidingsurroundingpedestrians
Citer cette publication

€ 4.00