Archive/Spatial Leakage in Classifying NASA FIRMS Thermal Anomalies as Wildfire Incidents: A Leakage-Controlled Evaluation of Radiometric, Temporal, and Spatiotemporal Features
Spatial Leakage in Classifying NASA FIRMS Thermal Anomalies as Wildfire Incidents: A Leakage-Controlled Evaluation of Radiometric, Temporal, and Spatiotemporal Features
Armin Soltan, Alberto González-Martínez
22 juillet 2026
en

Abstract

NASA’s Fire Information for Resource Management System (FIRMS) provides near-real-time thermal anomaly detections from VIIRS, but not all detections correspond to wildfire incidents: industrial heat, agricultural burning, and sensor artifacts produce false alarms that contribute to alert fatigue for emergency-management analysts. We study whether contextual machine learning (ML) features improve wildfire-incident classification from FIRMS detections, and—more importantly—whether reported gains survive leakage-controlled evaluation. We construct a labeled dataset by matching 521,395 VIIRS SNPP detections across CONUS in 2024 to 3766 NIFC 2024 wildfire perimeters, yielding 131,771 (25.3%) wildfire-matched and 389,624 candidate non-wildfire detections spanning 1067 distinct wildfire incidents. We benchmark five operational baselines and six classifiers under four validation regimes (random, event-aware, 5° spatial-block, and temporal holdout) with and without raw geographic coordinates. A naive random split inflates LightGBM to F1 =0.985, but a leakage-controlled event-aware split reduces it to F1 =0.767, and a spatial-block holdout to F1 =0.627. Feature attribution shows geographic coordinates account for 88.9% of model gain—the summed share of LightGBM’s total split-gain attributed to the three coordinate features within the full-feature model; removing coordinates improves spatial-block generalization from F1 =0.627 to 0.818, demonstrating that raw coordinates drive memorization of where 2024 fires occurred rather than transferable discrimination. We further show that spatiotemporal clustering must be causal: a model using full-partition clustering appears strong (F1 =0.908) but leaks future detections, whereas a properly causal trailing-window version ties plain LightGBM in-distribution (F1 =0.762). Combining causal clustering with no raw coordinates is the most robust configuration under spatial transfer (spatial-block F1 =0.868 vs. 0.627 for the coordinate model). Bootstrap 95% confidence intervals show these gaps far exceed statistical uncertainty, and sensitivity analyses show the conclusions are robust to the spatial-block size and to the clustering-window choice. Under natural class prevalence (14%), precision falls to 0.69, and results are sensitive to the labeling buffer. All ML models nonetheless far exceed FIRMS high-confidence thresholding (F1 =0.128). We argue that spatial leakage—not raw accuracy—is the central methodological issue for FIRMS wildfire-incident classification, and recommend coordinate-free, causal spatiotemporal-clustering features evaluated under spatial holdout. The system is intended as an analyst-prioritization decision-support layer, not autonomous incident confirmation.

IPC Classification

G06A01

Keywords

spatialleakageclassifyingnasafirmsthermalanomalieswildfireincidentsleakage-controlledevaluationradiometrictemporalspatiotemporalfeaturesgeohazardsfireinformationresourcemanagementsystemprovidesnear-real-timeanomaly
Citer cette publication

€ 4.00