Abstract
The increasing frequency and intensity of wildfires has created an urgent demand for scalable and autonomous wildfire response systems. While recent advances in multi-agent reinforcement learning (MARL) have demonstrated promise for collaborative uncrewed aerial vehicle (UAV)-based wildfire suppression, most existing approaches rely on simplified fire propagation dynamics and highly centralised learning architectures that are difficult to deploy in realistic operational settings. This paper presents a decentralised MARL framework for wildfire suppression that combines stochastic wildfire propagation, wind-driven spread dynamics, and communication-aware multi-agent coordination. The proposed framework extends an existing probabilistic wildfire environment through the incorporation of wind speed and directional effects, producing highly asymmetric and stochastic wildfire behaviour that more closely resembles real wildfire propagation. A decentralised Deep Q-Network (DQN) architecture is then introduced in which UAV agents learn independently through individual replay buffers. To mitigate the sparse-learning challenges introduced by decentralisation, selective experience sharing based on the SUPER algorithm is incorporated, enabling agents to exchange only high-value experiences under realistic communication constraints. Experimental results demonstrate that selective communication significantly improves containment performance and learning efficiency while preserving decentralised execution. The work highlights both the feasibility and challenges of realistic UAV swarm coordination for wildfire suppression, particularly the trade-offs between communication bandwidth, environmental stochasticity, and collaborative performance.
IPC Classification
Keywords
€ 4.00