Abstract
Existing e-commerce recommender systems, like graph-based retrieval and user behavior-driven Collaborative Filtering, usually take either textual or visual features for item recommendations to users. Traditional hierarchical large language models (HLLM) also suffer from high computational overhead for long user sequences and insufficient cross-modal semantic alignment. In this paper, we propose a multimodal hierarchical large language model (MHLLM) recommender system that leverages user behavior and integrates large language models to enhance text–image search precision and commercial value. Our MHLLM is a two-stage V-shaped model. The model decouples multimodal feature modeling and user behavior modeling: the first stage leverages Item-LLM and Item-CLIP to extract text semantics and cross-modal visual-text features respectively; the second stage adopts a learnable dynamic gating mechanism to adaptively fuse dual-modal features and uses User-LLM to model user preferences for personalized recommendation. We further design a multi-phase training strategy and lightweight feature projection structure to solve the gradient vanishing problem in joint training and reduce computational cost. Experimental results show that our MHLLM significantly outperforms traditional systems, improving key metrics like Recall@5 and NDCG@5 over 20% in most cases. Additionally, by efficient leveraging LLM and CLIP, MHLLM reduces the repeated inference overhead of User-LLM by over 30% via feature caching. Ablation experiments verify the effectiveness of core components such as dynamic gating, gating smoothing loss and task-specific prompts. The proposed model balances recommendation accuracy and deployment efficiency and has practical application value for large-scale e-commerce image-text retrieval and recommendation scenarios.
IPC Classification
Keywords
€ 4.00