Archive/Explainable Text-Based Computational Framework for Psycho-Emotional Risk Classification: Calibration, Interpretability, and Encoder Comparison on Public Datasets
Explainable Text-Based Computational Framework for Psycho-Emotional Risk Classification: Calibration, Interpretability, and Encoder Comparison on Public Datasets
Orazmukhamed Bekmurat, Vassiliy Serbin, Aliya Aizhanova et al.
21 de julho de 2026
en

Abstract

The early detection of psycho-emotional risks remains challenging despite its growing importance. Most existing models rely on a single data type, mainly questionnaires or text, and operate as “black boxes”, limiting practical use. Psycho-emotional states are multidimensional, reflected in textual, structured, and temporal digital signals. Ignoring this complexity may result in information loss and lower prediction accuracy. This paper presents an explainable text-based computational framework for psycho-emotional risk classification using public datasets. Record-level structured information was used only when it was available within the same original observation and was not created by matching records across datasets. The study did not integrate observations across datasets or perform cross-dataset fusion; these aspects are considered potential directions for future research. Implemented in Python using PyTorch, the model was evaluated on two open-text datasets. The BERT-based configuration achieved accuracy of 0.593 ± 0.009 and an ROC-AUC of 0.860 ± 0.006, while RoBERTa-base improved the performance to 0.666 ± 0.005 accuracy and a 0.896 ± 0.004 ROC-AUC under five-fold cross-validation. The results demonstrate that classification quality depends on encoder selection, preprocessing, dataset characteristics, duplicate handling, calibration, and the way that available record-level information is represented. Therefore, the findings should be interpreted as record-level computational classification results rather than the clinical validation of psycho-emotional risk assessment.

IPC Classification

G06A61

Keywords

explainabletext-basedcomputationalframeworkpsycho-emotionalriskclassificationcalibrationinterpretabilityencodercomparisonpublicdatasetsmultimodaltechnologiesinteractionearlydetectionrisksremainschallengingdespitegrowingimportance
Referencie esta publicação

€ 4.00