Archive/CTGAN-Based Data Augmentation and XGBoost–LSTM Strength Prediction of CSG
CTGAN-Based Data Augmentation and XGBoost–LSTM Strength Prediction of CSG
Guanghui Li, Yupeng Zhang, Qingqing Tian et al.
22 juillet 2026
en

Abstract

Cementitious sand and gravel (CSG) is commonly used in construction engineering; however, its mix proportion design is complex, and traditional physical experiments face limitations such as long cycles, high costs, and susceptibility to external factors when obtaining high-quality sample data. In this study, a foundational dataset was first acquired through physical experiments: 100 sets of CSG specimens with different mix proportions (cement content 40, 50, 60, 70 kg/m3; water-to-binder ratio 1.0, 1.2, 1.4; sand ratio 0.1, 0.2, 0.3, 0.4; fly ash content 20, 30, 40, 50 kg/m3) were prepared. After 28 days of standard curing, compressive strength and splitting tensile strength tests were conducted using a WAW-1000 electro-hydraulic servo universal testing machine, yielding 100 sets of real mechanical property data. The coefficients of variation for all test groups were below 10%, confirming the reliability and repeatability of the experimental data. On this basis, a data augmentation method based on Conditional Tabular Generative Adversarial Networks (CTGAN) is proposed. Through adversarial training between the generator and the discriminator, the model learns the multi-dimensional distribution characteristics of the original CSG data and generates 100 synthetic samples, which are then merged with the original data to expand the dataset to 200 samples. The quality of the synthetic data is evaluated using Wasserstein distance and correlation matrix heatmaps. Furthermore, a hybrid XGBoost–LSTM prediction model is proposed—XGBoost is used for feature construction to capture nonlinear interactions among mix proportion variables, and the constructed features are then fed into an LSTM network for sequential learning and regression prediction. The results show that the CTGAN-generated data are highly consistent with the original data in terms of kernel density distributions and variable correlations, with Wasserstein distance significantly superior to four comparative methods: Bootstrap, SMOTE, GaussianCopula, and TVAE. After augmentation, the XGBoost–LSTM model achieves a coefficient of determination (R2) of 0.9897 for compressive strength prediction (vs. 0.9793 before augmentation) and 0.9801 for splitting tensile strength (vs. 0.9882 before augmentation, a slight decrease). The mean absolute percentage errors (MAPE) are 4.49% and 4.11%, and the root mean square errors (RMSE) are 0.201 and 0.049, respectively; both error metrics are reduced compared with those before augmentation. Compared with baseline models including XGBoost, LSTM, Random Forest (RF), and Support Vector Regression (SVR), the XGBoost–LSTM model exhibits the best performance across all evaluation metrics, and Wilcoxon signed-rank tests confirm that the performance differences are statistically significant (p < 0.05). The proposed method of CTGAN-based data augmentation combined with the XGBoost-LSTM hybrid model provides an effective solution to the problem of insufficient CSG sample data and offers a reference for data enhancement and performance prediction of other small-sample materials.

IPC Classification

G06H04C07B60

Keywords

ctgan-baseddataaugmentationxgboostlstmstrengthpredictionmaterialscementitioussandgravelcommonlyusedconstructionengineeringhoweverproportiondesigncomplextraditionalphysicalexperimentsfacelimitations
Citer cette publication

€ 4.00