Yuheng Fu

dblp:365/8182 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0005-5452-2351ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Artificial intelligence
1 paper
Graph learning · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › sequential recommendation › side information-enhanced sequential recommendation
multimodal sequential recommendation
1.012026
Multi-Way Cascade-Attention Network for Multi-Modal Sequential Recommendation · IEEE Trans. Multim. 2026
Recommender systems
sequential recommendation
1.012026
Multi-Way Cascade-Attention Network for Multi-Modal Sequential Recommendation · IEEE Trans. Multim. 2026
Machine learning › Graph learning
graph neural network
0.312026
Multi-Way Cascade-Attention Network for Multi-Modal Sequential Recommendation · IEEE Trans. Multim. 2026

Methods — techniques the papers use, named apart from their topics

self-attention · 2.0graph propagation · 2.0cross-attention · 2.0cascade attention · 2.0
YearPublicationVenuePosition
2026 Multi-Way Cascade-Attention Network for Multi-Modal Sequential Recommendation
abstract
Sequential recommendation has become a hot topic, which aims to predict the desired items for each user based on his/her historical actions. The mainstream advancements for this task focus on modeling user behaviors in a pure item ID-based manner, which often fail to provide satisfactory results due to data sparsity and cold-start issues. Recently, several studies that leverage multi-modal information (i.e., multi-modal sequential recommendation) have shed light on alleviating such issues. However, we argue that three limitations are still not well addressed: 1) they usually extract ID modality features of an item with an one-hot encoding, which does not include any semantic information; 2) they fail to effectively mitigate the semantic gap issue and explicitly explore the asynchronous interplay between any two modalities; and 3) during the model prediction stage, they neglect the significance of adaptively fusing multi-modal embeddings for each user. To address such defects, we propose a novel framework for multi-modal sequential recommendation, namely, a multi-way cascade-attention network (MCN). Specifically, we apply a lightweight graph propagation network to derive informative representations of the ID-oriented modality, explicitly encoding collaborative signals in the user-item interaction graph. Next, we develop a multi-way cascade-attention module (CAM) to accomplish user behavior sequence alignment across different modality spaces. Each CAM consists of a cross-attention block followed by a series of self-attention blocks. The former encodes the asynchronous interplay between two modalities, while the latter captures intra-modal temporal dependencies. Finally, we design a modality-aware attentive strategy to dynamically fuse the user's dynamic interests across different modality spaces. Our extensive experiments on four public datasets demonstrate the superiority of MCN over recent state-of-the-art recommenders.
Bin Wu 0019, Long Chen 0016, Yuheng Fu, Yunshan Ma 0002, Mingliang Xu 0001, Tat-Seng Chua
IEEE Trans. Multim.3