Abstract
Accurate segmentation of abdominal organs in Computed Tomography (CT) underpins radiotherapy planning, surgical planning, and disease monitoring. Existing benchmarks rank architectures by a single aggregate Dice score, without per-organ statistical testing or boundary-sensitive metrics, even though models are chosen organ by organ for clinical use. We benchmark ten architectures spanning convolutional, attention-based, transformer, and state–space (Mamba) families on the AMOS CT dataset under one identical nnU-Net-style pipeline; we report per-organ Dice, 95-percentile Hausdorff Distance (HD95), and Normalised Surface Dice, with pairwise significance tested on an independent external dataset (TotalSegmentator). A competitive cluster of convolutional and Mamba models leads; rankings are stable on large organs but reshuffle by 10–13% on the small, geometrically complex ones, and boundary fidelity separates the models into tiers that the Dice ranking hides. This ordering largely holds on the external set (Spearman ρ=0.84). Selecting a model on aggregate Dice alone is therefore unsafe for organ-specific clinical tasks: per-organ overlap and boundary metrics should be the primary acceptance criteria for selecting a model before clinical deployment.
IPC Classification
Keywords
€ 4.00