Desheng Cai

dblp:205/7620 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-7464-9115ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-modal Bipartite Graph Structure Learning with Information Bottleneck for Micro-video Recommendation
abstract
Graph-based recommender systems have become prevalent in micro-video recommendation by modeling user-item interactions as a bipartite graph. However, these methods face two inherent limitations: (1) their reliance on a fixed, pre-defined graph structure makes them susceptible to noisy interactions, and (2) the multi-modal representations they learn often contain redundant information that is not discriminative enough for the recommendation task. To overcome these issues, we propose a novel Multi-modal Bipartite Graph Structure Learning network (MBGSL), which leverages the information bottleneck principle for robust micro-video recommendation. Specifically, MBGSL first learns adaptive graph structures from multi-modal content (e.g., visual, acoustic, textual) through dedicated graph learners to mitigate noise. Then, it applies an intra-modality information bottleneck to learn minimal sufficient representations within each modality and an inter-modality information bottleneck to capture distinctive information across modalities, thereby eliminating redundancy. Furthermore, the model incorporates collaborative signals through a contrastive learning objective to guide the graph structure learning process. Extensive experiments on three real-world datasets demonstrate that MBGSL achieves state-of-the-art performance, significantly surpassing existing baselines.
Ying He 0008, Desheng Cai, Shengsheng Qian, Quan Fang, Yinwei Wei, Changsheng Xu
WWW2
2026 Anchor node guided global-local graph neural networks for multimedia recommendation
Qi Ren, Desheng Cai
J. Vis. Commun. Image Represent.2
2026 Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
abstract
Multimedia recommendation systems leverage user–item interactions and multimodal information to capture user preferences, enabling more accurate and personalized recommendations. Despite notable advancements, existing approaches still face two critical limitations: first, shallow modality fusion often relies on simple concatenation, failing to exploit rich synergic intra- and inter-modal relationships; second, asymmetric feature treatment—where users are only characterized by interaction IDs while items benefit from rich multimodal content—hinders the learning of a shared semantic space. To address these issues, we propose a C ross-Modal R ecursive A ttention N etwork with Dual Graph E mbedding (CRANE) . To tackle shallow fusion, we design a core Recursive Cross-Modal Attention (RCA) mechanism that iteratively refines modality features based on cross-correlations in a joint latent space, effectively capturing high-order intra- and inter-modal dependencies. For symmetric multimodal learning, we explicitly construct users’ multimodal profiles by aggregating features of their interacted items. Furthermore, CRANE integrates a symmetric dual graph framework—comprising a heterogeneous user–item interaction graph and a homogeneous item–item semantic graph—unified by a self-supervised contrastive learning objective to fuse behavioral and semantic signals. Despite these complex modeling capabilities, CRANE maintains high computational efficiency. Theoretical and empirical analyses confirm its scalability and high practical efficiency, achieving faster convergence on small datasets and superior performance ceilings on large-scale ones. Comprehensive experiments on four public real-world datasets validate an average 5% improvement in key metrics over state-of-the-art baselines. Our code is publicly available at https://github.com/MKC-Lab/CRANE .
Ji Dai, Quan Fang, Jun Hu 0016, Desheng Cai, Yang Yang 0122
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Cross-View Sample-Enriched Graph Contrastive Learning Network for Personalized Micro-video Recommendation
abstract
Micro-video recommendation has attracted extensive research attention with the increasing popularity of micro-video sharing platforms. Recently, graph contrastive learning (GCL) is adopted for enhancing the performance of graph neural network based micro-video recommendation. However, these GCL methods may suffer from the following problems: (1) they fail to fully exploit the potential of contrastive learning for ignoring or misjudging highly similar samples, and (2) the complementary recommendation effects between graph structure information and multi-modal feature information are not effectively utilized. In this paper, we propose a novel Cross-View Sample-Enriched Graph Contrastive Learning Network (CSGCL) for micro-video recommendation. Specifically, we build a collaborative learning view and a semantic learning view to learn node representations. For the collaborative learning view, we leverage similar nodes at the structure level to construct an effective collaborative contrastive objective. For the semantic learning view, we derive the k-nearest neighbor graph generated from multi-modal features as the semantic graphs and build a semantic contrastive objective for learning high-quality micro-video representations. Finally, a cross-view contrastive objective is designed to consider the mutually complementary recommendation effects by maximizing the agreement between the two above views. Extensive experiments on three real-world datasets demonstrate that the proposed model outperforms the baselines.
Ying He 0008, Gong-Qing Wu, Desheng Cai, Xuegang Hu
ICMR3
2023 Meta-path based graph contrastive learning for micro-video recommendation
Ying He 0008, Gong-Qing Wu, Desheng Cai, Xuegang Hu
Expert Syst. Appl.3
2023 Heterogeneous Graph Contrastive Learning Network for Personalized Micro-Video Recommendation
abstract
Personalized micro-video recommendation has attracted a lot of research attention with the growing popularity of micro-video sharing platforms. Many efforts have been made to consider micro-video recommendation as a matching task and shown promising performance, while they only focus on simple features or multi-modal attribute information. Recently, Graph Neural Networks (GNNs) have been employed in many recommendation tasks and achieved impressive success. However, these GNN-based methods may suffer from the following limitations: (1) fail to capture the heterogeneity of nodes in user-video bipartite graphs; (2) ignore the non-local (global) semantic correlation information remained in heterogeneous graphs. In this paper, we present a novel approach, Heterogeneous Graph Contrastive Learning Network (HGCL), for personalized micro-video recommendation. To consider heterogeneity in user-video bipartite graphs, we first introduce a heterogeneous graph encoder network for a high-quality representation learning of users and micro-videos. Specifically, we design a random surfing model to generate node-type specific homogeneous graphs to preserve the heterogeneity. Then we propose a graph contrastive learning framework to achieve representation learning on each node-type specific homogeneous graph by maximizing the mutual information between local patches of a graph and the global representation of the entire graph. Finally, a type-crossing objective function is proposed to jointly integrate the node embeddings from different node types to facilitate high-quality representation learning. Experimental results on real-world datasets in the micro-video recommendation task validate the performance of our method, compared with state-of-the-art baseline algorithms.
Desheng Cai, Shengsheng Qian, Quan Fang, Jun Hu 0016, Wenkui Ding, Changsheng Xu
IEEE Trans. Multim.1
2023 User Cold-Start Recommendation via Inductive Heterogeneous Graph Neural Network
abstract
Recently, user cold-start recommendations have attracted a lot of attention from industry and academia. In user cold-start recommendation systems, the user attribute information is often used by existing approaches to learn user preferences due to the unavailability of user action data. However, most existing recommendation methods often ignore the sparsity of user attributes in cold-start recommendation systems. To tackle this limitation, this article proposes a novel Inductive Heterogeneous Graph Neural Network (IHGNN) model, which utilizes the relational information in user cold-start recommendation systems to alleviate the sparsity of user attributes. Our model converts new users, items, and associated multimodal information into a Modality-aware Heterogeneous Graph (M-HG) that preserves the rich and heterogeneous relationship information among them. Specifically, to utilize rich and heterogeneous relational information in an M-HG for enriching the sparse attribute information of new users, we design a strategy based on random walk operations to collect associated neighbors of new users by multiple times sampling operation. Then, a well-designed multiple hierarchical attention aggregation model consisting of the intra- and inter-type attention aggregating module is proposed, focusing on useful connected neighbors and neglecting meaningless and noisy connected neighbors to generate high-quality representations for user cold-start recommendations. Experimental results on three real datasets demonstrate that the IHGNN outperforms the state-of-the-art baselines.
Desheng Cai, Shengsheng Qian, Quan Fang, Jun Hu 0016, Changsheng Xu
ACM Trans. Inf. Syst.1
2022 Dual Confidence Learning Network for Open-World Time Series Classification
Junwei Lv, Ying He 0008, Xuegang Hu, Desheng Cai, Yuqi Chu, Jun Hu 0016
DASFAA (2)4
2022 Adaptive Anti-Bottleneck Multi-Modal Graph Learning Network for Personalized Micro-video Recommendation
abstract
Micro-video recommendation has attracted extensive research attention with the increasing popularity of micro-video sharing platforms. There exists a substantial amount of excellent efforts made to the micro-video recommendation task. Recently, homogeneous (or heterogeneous) GNN-based approaches utilize graph convolutional operators (or meta-path based similarity measures) to learn meaningful representations for users and micro-videos and show promising performance for the micro-video recommendation task. However, these methods may suffer from the following problems: (1) fail to aggregate information from distant or long-range nodes; (2) ignore the varying intensity of users' preferences for different items in micro-video recommendations; (3) neglect the similarities of multi-modal contents of micro-videos for recommendation tasks. In this paper, we propose a novel Adaptive Anti-Bottleneck Multi-Modal Graph Learning Network for personalized micro-video recommendation. Specifically, we design a collaborative representation learning module and a semantic representation learning module to fully exploit user-video interaction information and the similarities of micro-videos, respectively. Furthermore, we utilize an anti-bottleneck module to automatically learn the importance weights of short-range and long-range neighboring nodes to obtain more expressive representations of users and micro-videos. Finally, to consider the varying intensity of users' preferences for different micro-videos, we design and optimize an adaptive recommendation loss to train our model in an end-to-end manner. We evaluate our method on three real-world datasets and the results demonstrate that the proposed model outperforms the baselines.
Desheng Cai, Shengsheng Qian, Quan Fang, Jun Hu 0016, Changsheng Xu
ACM Multimedia1
2022 Heterogeneous Hierarchical Feature Aggregation Network for Personalized Micro-Video Recommendation
abstract
Micro-video recommendation has attracted extensive research attention with the increasing popularity of micro-video sharing platforms. Traditional approaches consider micro-video recommendation as a matching task and ignore the rich relationships among users and micro-videos from various modalities (e.g., visual, acoustic, and textual). Recently, GNN-based approaches show promising performance for the micro-video recommendation task. However, they mainly focus on the homogeneous graph which includes only one type of nodes or relations, and cannot be applied to the heterogeneous graph which consists of users, micro-videos, and related multi-modal information. In this paper, a novel Heterogeneous Hierarchical Feature Aggregation Network (HHFAN) is proposed for personalized micro-video recommendation. Our goal is to explore the highly complicated relationship information among users, micro-videos and related multi-modal information from a modality-aware Heterogeneous Information Graph (M-HIG), and thus generate high-quality user and micro-video embeddings for recommendation. The proposed model consists of two key components: (1) In data structure level, we build a heterogeneous graph and utilize a random walk based sampling strategy to sample neighbors for users and micro-videos. (2) In representation learning level, we design a hierarchical feature aggregation network including the intra- and inter-type feature aggregation networks to better capture the complex structure and rich semantic information in the heterogeneous graph. We evaluate our method on two real-world datasets and the results demonstrate that the proposed model outperforms the baseline methods.
Desheng Cai, Shengsheng Qian, Quan Fang, Changsheng Xu
IEEE Trans. Multim.1
2021 Attentive interaction-driven entity resolution over multi-source web information
Ying He 0008, Gong-Qing Wu, Desheng Cai, Shengjie Hu, Xianyu Bao, Xuegang Hu
Neurocomputing3
2019 Content-aware attributed entity embedding for synonymous named entity discovery
Desheng Cai, Gong-Qing Wu
Neurocomputing1