Archive/Computational Phenotyping of Autism-Related Behaviors: A Cross-Cultural Machine Learning Study in Bangladesh
Computational Phenotyping of Autism-Related Behaviors: A Cross-Cultural Machine Learning Study in Bangladesh
Saimourya Surabhi, Kaitlyn Dunlap, Parnian Azizian et al.
27 juillet 2026
en

Abstract

Background: Digital behavioral phenotyping of autism spectrum disorder (ASD) offers a promising approach for developing more scalable diagnostic frameworks across diverse global contexts. Machine learning (ML) models show promise for ASD diagnosis using behavioral videos, but critical questions remain regarding whether models trained on data from one country work in another, and how the background of the raters affects the accuracy. Our work addresses these questions by testing whether ML models can accurately diagnose ASD across different populations and rater groups. Methods: This work evaluates the performance of a supervised ML framework for binary classification of ASD versus non-ASD [speech, language and communication disorders (SLC) + neurotypical (NT)] in a cohort of 227 children in Bangladesh. We first assessed the cross-domain model transferability of a clinical-instrument-trained logistic regression model (LR-9) on behavioral ratings that were based on videos of Bangladeshi children interacting with caregivers and toys at two major child development centers in Dhaka, Bangladesh. We then trained five diverse classifiers (Logistic Regression, Random Forest, XGBoost, SVM, and RuleFit) on the full annotated Bangladeshi dataset. Using SHAP-based consensus elbow feature selection, we identified a compact set of features that maintained the performance. Finally, we developed ensemble models to improve predictive stability. Results: The LR-9 model, originally trained on U.S. clinical instrument data, was evaluated on video-based behavioral ratings from 214 Bangladeshi children. When tested on Bangladeshi clinician ratings, the LR-9 model achieved a sensitivity of 86.1% (95% CI: [0.78–0.93]) and AUC of 0.79 (95% CI: [0.73–0.86]). The distinction across rater groups was between trained raters (clinicians and students) and crowd workers, who showed lower sensitivity 28.5% (95% CI: [0.21, 0.39]). When tested on the aggregated ratings from all groups, the model achieved an AUC of 0.78 (95% CI: [0.72–0.84]). Inter-rater reliability followed the same pattern: individual agreement was fair (Krippendorff’s α = 0.26), but the multi-rater consensus was reliable (ICC(1,k) = 0.84), with Bangladeshi clinicians showing the highest agreement (α = 0.34) and crowd workers the lowest (α = 0.20). We then trained new models directly on the Bangladeshi ratings. All model types achieved similar AUC values (0.86–0.89), with overlapping confidence intervals. Using just 8–11 key behaviors kept the similar performance while cutting the features by 66–75%. Combining ensembles gave similar results (e.g., Bayesian averaging: AUC 0.88 [0.78, 0.95]) but with more stable predictions. Conclusion: This study provides evidence that mobile video-based ASD diagnosis can achieve comparable performance (AUC: 0.89 [0.76, 0.96]) to models trained on clinical instrument data. This work contributes to the development of broader adaptable autism detection tools, bypassing the dependence on traditional clinical instrument data.

IPC Classification

G06H04A61

Keywords

computationalphenotypingautism-relatedbehaviorscross-culturalmachinelearningbangladeshbiomedinformaticsbackgrounddigitalbehavioralautismspectrumdisorderofferspromisingapproachdevelopingmorescalablediagnosticframeworksacross
Citer cette publication

€ 4.00