Archive/Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition
Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition
Kabul Khudaybergenov, Avazjon Marakhimov
28 de julio de 2026
en

Abstract

Skeleton-based human action recognition is an important problem in applied vision systems, yet many existing approaches depend on a single skeleton descriptor or a single feature-learning mechanism. This restriction can weaken the representation of local body kinematics, long-range joint relations, and temporal dependencies within an action sequence. To address these limitations, this paper proposes HILF-GT (Hybrid Invariant Latent Feature Graph Transformer), a hybrid Graph Convolutional Network (GCN)-Transformer framework based on multiple spatio-temporal invariant latent features. The representation module constructs complementary structured tensors from skeleton graphs, inter-joint distances, adjacent-frame joint displacements, and inter-limb angles. Instead of transforming these descriptors into image-like maps for separate Convolutional Neural Network (CNN)-based classification, HILF-GT keeps their graph and temporal organization during learning. A local GCN branch models skeleton-aware kinematic patterns, whereas a graph-aware Transformer branch uses biased self-attention and cross-attention to capture dependencies among distant joints, frames, and latent-feature streams. A Perceiver-style latent bottleneck is further introduced to reduce the memory cost of global attention over frame-joint tokens. Experiments were conducted on four standard benchmark datasets, including NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UTD-MHAD. The proposed method achieved 93.1% and 97.20% accuracy on the NTU-RGB+D 60 Cross-Subject and Cross-View protocols, 88.15% and 90.20% on the NTU-RGB+D 120 Cross-Subject and Cross-Setup protocols, 98.50% on NW-UCLA, and 97.50% on UTD-MHAD.

IPC Classification

G06H04

Keywords

hybridinvariantlatentfeaturegraphtransformerskeleton-basedhumanactionrecognitioninformationimportantproblemappliedvisionsystemsmanyexistingapproachesdependsingleskeletondescriptorfeature-learning
Citar esta publicación

€ 4.00