Abstract
While generative models have become a standard approach for addressing the semantic-to-visual gap in Generalized Zero-Shot Learning (GZSL), existing architectures often struggle with two persistent limitations: cross-modal interference during condition fusion and severe overfitting to the visual distributions of seen classes. To address these bottlenecks, this paper introduces SemanticFlowNet, a framework based on Decoupled Semantic Flow Matching. Specifically, we propose a Decoupled Multi-modal Conditioning mechanism that relies on channel-wise concatenation of temporal encodings, semantic attributes, and visual contexts, which preserves the orthogonal subspaces of each modality and reduces interference. Additionally, we integrate a Dropout-enhanced Adaptive Layer Normalization (AdaLN) module to perturb the rigid memorization of seen classes, utilizing stochastic dropout within the state evolution to simulate the distributional variance of unseen domains. Finally, a Time-Aware Dynamic Reconstruction Penalty is introduced to enforce progressively stricter semantic alignment as the generative ordinary differential equation (ODE) trajectory converges to the target manifold. Evaluations on the CUB, SUN, and AWA2 benchmarks demonstrate the effectiveness of the proposed framework. Notably, SemanticFlowNet achieves a harmonic mean of 77.90% on the CUB dataset in the single-seed full-model setting, providing a competitive baseline for generative GZSL applications.
IPC Classification
Keywords
€ 4.00