Archive/Machine Learning Classification of High-Turbidity Exceedance Using Water Quality and Meteorological Predictors
Machine Learning Classification of High-Turbidity Exceedance Using Water Quality and Meteorological Predictors
Saleh H. Alhathloul, Yazeed Algurainy
30 juillet 2026
en

Abstract

High-turbidity exceedance events are important indicators of abnormal coastal water quality conditions, but their relatively low frequency compared with normal observations makes reliable classification challenging. This study developed a scenario-based machine learning framework to classify high-turbidity exceedance using water quality and meteorological predictors from an automated seawater monitoring station along the Arabian Gulf. After quality control screening and temporal matching with NASA POWER meteorological variables, turbidity was converted into a binary target using a 2 NTU threshold, resulting in 4408 low-turbidity observations (85.3%) and 758 high-turbidity observations (14.7%). Four predictor scenarios were evaluated: meteorological-only selected variables, combined selected variables, water-quality-only all variables, and combined all variables. Random Forest, XGBoost, a Support Vector Machine, and K-Nearest Neighbors were tested under three class-imbalance treatments: undersampling, oversampling, and SMOTE. The meteorological-only scenario showed moderate classification ability, with ROC-AUC values of 0.81–0.86, but relatively low precision, indicating limited control of false alarms. Adding water quality predictors substantially improved classification performance, with the combined selected-variable scenario achieving ROC-AUC values of 0.96–0.98. Under oversampling, Random Forest achieved an accuracy = 0.95, precision = 0.83, specificity = 0.97, weighted F1-score = 0.95, and ROC-AUC = 0.98. The water-quality-only scenario also performed strongly, with accuracy values of 0.90–0.95 and ROC-AUC values of 0.95–0.96, confirming that direct water quality measurements carried most of the predictive signal. The combined all-variable scenario produced similarly high performance, with ROC-AUC values of 0.96–0.98, while variable selection reduced redundancy without loss of predictive skill. Computational analysis showed that the top-ranked XGBoost model under the combined all-variable oversampling scenario required only 0.08 s for training and 0.003 ms per sample for prediction. Overall, the results demonstrate that integrating selected water quality and meteorological predictors provides an accurate, efficient, and operationally practical framework for high-turbidity early-warning classification.

IPC Classification

G06H01

Keywords

machinelearningclassificationhigh-turbidityexceedancewaterqualitymeteorologicalpredictorsappliedscienceseventsimportantindicatorsabnormalcoastalconditionsrelativelyfrequencycomparednormalobservationsmakesreliable
Citer cette publication

€ 4.00