Archive/SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition
SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition
Alpamis Kutlimuratov, Kongratbay Sharipov, Piratdin Allayarov et al.
30 de julho de 2026
en

Abstract

Recognizing facial emotions automatically from images/videos (FER) still represents a difficult problem for emotion computing, mainly due to variations in the face pose, lighting, occlusion, facial features, and expression intensity in the wild. Recent CNN–Transformer-based hybrid models like POSTER have leveraged local feature learning, landmark guidance, and global dependency modeling to achieve strong performance. Yet these methods give the main focus to spatial and contextual representations while not really going deep into adaptive channel-wise feature importance over multi-scale representations. As different feature channels represent emotions in varying degrees, it is likely that by treating all feature channels equally, one would limit the ability of the learned features to discriminate effectively. To overcome this weakness, this article presents a ResNet-18–Transformer landmark-guided module called SE-POSTER that fuses lightweight Squeeze-and-Excitation (SE) attention modules into the multi-scale feature pyramid of the baseline POSTER architecture. The proposed method carries out feature channel recalibration adaptively at the level of features before Transformer-based global attention modeling, thus allowing the network to focus on emotionally informative feature channels and suppress less relevant responses. The inclusion of SE attention in the network enhances fine, mid, and global levels of feature representations at a very low cost in terms of computation. On the basis of the RAF-DB, FERPlus, and AffectNet datasets, enormous experiments prove that the SE-POSTER framework proposed is capable of steadily boosting recognition accuracy relative to the baseline POSTER and several state-of-the-art FER methods. Especially, the proposed model delivers 92.78% accuracy on RAF-DB while it also shows better robustness and generalization capability under difficult real-world conditions. Moreover, additional ablation studies reveal that multi-level channel recalibration is effective in improving discriminative emotional feature learning.

IPC Classification

G06H04

Keywords

se-posterchannel-enhancedlandmarkguidedtransformerfacialemotionrecognitioninformaticsrecognizingemotionsautomaticallyimagesvideosstillrepresentsdifficultproblemcomputingmainlyvariationsfaceposelighting
Referencie esta publicação

€ 4.00