Archive/Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models
Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models
Egor Ushakov, Sofya Zimina, Arsenii Litvinov et al.
31 de julho de 2026
en

Abstract

The performance of AI solutions in mammography is largely determined by data quality, preprocessing methods, and augmentation strategies. However, systematic evaluation of these factors for models trained on aggregated multicenter datasets remains underexplored. This article presents a comparative assessment of the effects of different stages of the training pipeline on the final diagnostic accuracy. Using a pooled dataset (VinDr-Mammo, INBreast, CMMD, CBIS-DDSM), we evaluated each pipeline step—from filtering to architecture selection (EfficientNet-B3, CLIP). External testing was conducted on the MosMed database. Among the tested preprocessing steps, filtering the darkest 5% of images proved most effective. For EfficientNet-B3, optimal geometric and photometric augmentations increased test AUROC on the prepared MosMed test set from 0.844 to 0.900. Domain-specific pretraining and high resolution yielded the best performance: Mammo-CLIP achieved an AUROC of 0.949 ± 0.012, and EfficientNet-B3 reached 0.934 ± 0.013. Overall, this study developed a standardized pipeline that includes sequential data filtering and harmonization, augmentation optimization, and architecture selection. This approach ensures reliable and reproducible results for automated mammogram classification.

IPC Classification

G06A61

Keywords

optimizingmammographypipelinedatafilteringvision-languagemodelsinformaticsperformancesolutionslargelydeterminedqualitypreprocessingaugmentationstrategieshoweversystematicevaluationthesefactorstrainedaggregatedmulticenter
Referencie esta publicação

€ 4.00