Abstract
Multimodal object detection leveraging RGB and infrared imagery has become essential for robust all-weather perception in unmanned aerial vehicle (UAV) applications. However, existing methods still struggle with effective cross-modal feature fusion, spatial misalignment between modalities, and scale variation of objects in aerial views. In this paper, we propose CMGFDet, a Cross-Modal Gated Fusion Network with Multi-Receptive Field Aggregation designed for RGB–infrared aerial object detection. Our framework introduces three coordinated modules: (1) a Cross-Modal Feature Fusion Network (CMFFN) that employs a gated attention mechanism to selectively aggregate complementary information from both modalities during encoding; (2) a Global–Local Attention Module (GLAM) that performs hierarchical cross-modal feature alignment by jointly modelling global channel statistics and local spatial correlations in the decoder; and (3) a Multi-Receptive Field Aggregation Network (MRFAN) that captures multi-scale contextual information through parallel depthwise convolutions with diverse kernel sizes. Additionally, we incorporate a deep supervision strategy and a composite loss function to enhance training efficiency. Extensive experiments on four public benchmarks (DroneVehicle, RGBTDronePerson, VEDAI, and VTUAV) show that CMGFDet improves the previous best mAP@0.5 by 1.6%, 2.2%, 1.9%, and 2.2%, respectively. The implementation code will be released upon acceptance to support reproducibility.
IPC Classification
Keywords
€ 4.00