Archive/ChangePixel: Pixel-Level Evidence-Grounded Disaster Change Narration via Single-Backbone Transfer
ChangePixel: Pixel-Level Evidence-Grounded Disaster Change Narration via Single-Backbone Transfer
Qinyu Zhou, Ben Yang, Xinyan Wei et al.
29 juillet 2026
en

Abstract

Remote sensing change captioning aims to describe disaster-related changes from bi-temporal imagery, yet existing methods typically produce image-level captions without explicit regional evidence, limiting interpretability and weakening the link between generated language and actual changed areas. We present ChangePixel, a single-backbone framework that upgrades remote sensing change captioning into pixel-level, evidence-grounded change narration without introducing new manual grounding labels. ChangePixel incorporates three lightweight modules: a Bi-Temporal Change-Aware Transfer Adapter (BCTA) that converts shared pre- and post-event visual features into change-aware grounding representations, a Change Region Grounding Planner (CRGP) that localizes a compact set of informative changed regions before narration begins, and a Weak Evidence Alignment Bridge (WAB) that converts released change captions into phrase-to-region weak supervision. Through this design, the model jointly produces a global change caption and region-level evidence in the form of pixel masks paired with corresponding local change phrases. Experiments on the Remote Sensing Change Caption (RSCC) dataset and LEVIR-CC demonstrate that ChangePixel provides caption quality (ROUGE 19.52/ST5-SCS 76.91 on RSCC; CIDEr-D 56.82 on LEVIR-CC under zero-shot transfer) that is competitive with general-purpose vision–language models (VLMs) while adding pixel-level spatial evidence to change narration; additionally, evidence localization is quantified on LEVIR-MCI through semantic change-mask metrics, reaching Change mIoU 33.8 (15.3 points higher than a non-learned pixel-difference floor of 18.5), whereas phrase-to-region alignment is assessed qualitatively pending a dedicated grounding benchmark. The proposed framework offers a practical path from coarse image-level captioning to evidence-grounded disaster understanding.

IPC Classification

G06

Keywords

changepixelpixel-levelevidence-groundeddisasterchangenarrationsingle-backbonetransferremotesensingcaptioningaimsdescribedisaster-relatedchangesbi-temporalimageryexistingtypicallyproduceimage-levelcaptionswithoutexplicit
Citer cette publication

€ 4.00