Abstract
Identifying probiotic microorganisms requires linking genomic content to experimentally reported host-interaction traits. A machine-learning framework predicted seven probiotic phenotypes (acid resistance, bile resistance, adhesion, antimicrobial activity, immunomodulation, antioxidant activity, and antiproliferative potential) from functional COG categories and biosynthetic gene cluster annotations in 1183 genomes (780 probiotic and 403 non-probiotic). CatBoost, Random Forest, XGBoost, LightGBM, and Logistic Regression were evaluated with stratified cross-validation, leave-one-out cross-validation, and an 80:20 holdout split. Antioxidant and antiproliferative models performed best, with cross-validation F1-scores of 0.92 and 0.97 and holdout F1-scores of 0.70 and 0.72. SHAP analysis identified model-derived genomic associations rather than causal mechanisms, including an inverse association between T3PKS abundance and antioxidant classification and associations between lipid transport/metabolism features and antiproliferative classification. The framework supports high-throughput in silico prioritization of candidate probiotic strains for phenotype-specific validation. Interpretation remains limited by heterogeneous source annotations, label noise, and potential false negatives in the non-probiotic dataset.
IPC Classification
Keywords
€ 4.00