Archive/Explainable Machine Learning for Predicting Complicated Appendicitis and Expected Hospital Length of Stay in Children: An Exploratory Single-Center Study
Explainable Machine Learning for Predicting Complicated Appendicitis and Expected Hospital Length of Stay in Children: An Exploratory Single-Center Study
Ahmad Turki, Enas Raml, Jury Emad Aboulola
24 de julio de 2026
en

Abstract

Background and Objectives: Pediatric appendicitis remains a common surgical emergency, but early risk stratification of complicated appendicitis and expected hospital length of stay (LOS) remains challenging. This study developed and internally evaluated an explainable machine learning framework for pediatric appendicitis using routinely available clinical, laboratory, and radiological variables. Materials and Methods: This retrospective single-cohort model-development study included 152 pediatric patients with acute appendicitis treated at King Abdulaziz University Hospital, Jeddah, Saudi Arabia, between January 2019 and December 2024. Two prediction tasks were evaluated: complicated appendicitis classification and expected LOS regression. Prediction was performed after initial clinical assessment, first laboratory testing, and initial diagnostic imaging, but before surgery or definitive conservative treatment. Operative findings, histopathological findings, postoperative variables, treatment-response variables, actual LOS-related variables, and final disease-severity labels were excluded as predictors. Model performance was evaluated using repeated nested cross-validation, bootstrap 95% confidence intervals, calibration analysis, benchmark comparisons, LOS sensitivity analyses, and explainability analysis using feature importance and SHAP values. Results: Complicated appendicitis was present in 57 patients (37.5%), and median LOS was 3 days (IQR: 2–6). For complicated appendicitis prediction, the repeated nested cross-validation framework achieved an AUC of 0.788 (95% CI: 0.710–0.856), accuracy of 0.750 (95% CI: 0.684–0.816), sensitivity of 0.684 (95% CI: 0.559–0.796), specificity of 0.789 (95% CI: 0.705–0.868), and Brier score of 0.183 (95% CI: 0.158–0.210). The framework showed comparable performance to a parsimonious logistic regression model (AUC: 0.807). For expected LOS prediction, the regression framework achieved an MAE of 2.150 days (95% CI: 1.632–2.846), MdAE of 1.305 days (95% CI: 1.004–1.493), RMSE of 4.329 days (95% CI: 2.290–6.297), and R2 of 0.292 (95% CI: 0.213–0.481). LOS prediction showed lower MAE, MdAE, and RMSE than median and mean LOS null baseline models. Explainability analysis identified symptom duration, lymphocyte percentage, vomiting, temperature, and age as important predictors of complicated appendicitis, while radiological perforation, symptom duration, radiological collection, CRP, WBC count, NLR, lymphocyte percentage, and appendicolith by radiology contributed to expected LOS prediction. Conclusions: Explainable machine learning methods showed potential for internally validated prediction of complicated appendicitis and expected LOS using post-assessment pre-treatment data. However, this was a retrospective, single-cohort development study with a modest sample size and no external validation cohort. The findings should therefore be interpreted as exploratory and should not be considered evidence of clinical generalizability or readiness for implementation. Larger prospective multicenter studies with external validation, prediction-interval estimation, and clinical utility assessment are required before clinical use.

IPC Classification

G06A61C07

Keywords

explainablemachinelearningpredictingcomplicatedappendicitisexpectedhospitallengthstaychildrenexploratorysingle-centermedicinabackgroundobjectivespediatricremainscommonsurgicalemergencyearlyriskstratification
Citar esta publicación

€ 4.00