VLDB 2026 Research / reviewers in the wild / expert
Chunxu Shen
dblp:229/1311
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0005-0361-1709ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRECISE: Pre-training and Fine-tuning Sequential Recommenders with Collaborative and Semantic InformationabstractRecommendation platforms commonly offer diverse content scenarios for users to interact with. Pre-training models are the most commonly used approach in recommendation systems to capture users' full-domain interests. Traditional ID-based pre-training models mainly capture user interests by leveraging collaborative signals. However, a prevalent drawback of those systems is the incapacity to handle cold-start scenarios. With the recent advent of large language models, there has been a significant increase in research efforts exploiting LLMs to extract semantic information for items. However, text-based recommendations highly rely on elaborate feature engineering and often fail to capture collaborative similarities. Chonggang Song, Chunxu Shen, Yaoming Wu, Lingling Yi |
CIKM | 2 |
| 2025 | Improving LLMs for Recommendation with Out-Of-Vocabulary TokensabstractCharacterizing users and items through vector representations is crucial for various tasks in recommender systems. Recent approaches attempt to apply Large Language Models (LLMs) in recommendation through a question&answer format, where real items (eg, Item No.2024) are represented with compound words formed from in-vocabulary tokens (eg, “item“, “20“, “24“). However, these tokens are not suitable for representing items, as their meanings are shaped by pre-training on natural language tasks, limiting the model’s ability to capture user-item relationships effectively. In this paper, we explore how to effectively characterize users and items in LLM-based recommender systems from the token construction view. We demonstrate the necessity of using out-of-vocabulary (OOV) tokens for the characterization of items and users, and propose a well-constructed way of these OOV tokens. By clustering the learned representations from historical user-item interactions, we make the representations of user/item combinations share the same OOV tokens if they have similar properties. This construction allows us to capture user/item relationships well (memorization) and preserve the diversity of descriptions of users and items (diversity). Furthermore, integrating these OOV tokens into the LLM’s vocabulary allows for better distinction between users and items and enhanced capture of user-item relationships during fine-tuning on downstream tasks. Our proposed framework outperforms existing state-of-the-art methods across various downstream recommendation tasks. Ting-Ji Huang, Chunxu Shen, Kai-Qi Liu, De-Chuan Zhan, Han-Jia Ye |
ICML | 3 |
| 2025 | Hierarchical Graph Information Bottleneck for Multi-Behavior RecommendationabstractIn real-world recommendation scenarios, users typically engage with platforms through multiple types of behavioral interactions. Multi-behavior recommendation algorithms aim to leverage various auxiliary user behaviors to enhance prediction for target behaviors of primary interest (e.g., buy), thereby overcoming performance limitations caused by data sparsity in target behavior records. Current state-of-the-art approaches typically employ hierarchical design following either cascading (e.g., view$\rightarrow$cart$\rightarrow$buy) or parallel (unified$\rightarrow$behavior$\rightarrow$specific components) paradigms, to capture behavioral relationships. However, these methods still face two critical challenges: (1) severe distribution disparities across behaviors, and (2) negative transfer effects caused by noise in auxiliary behaviors. In this paper, we propose a novel model-agnostic Hierarchical Graph Information Bottleneck (HGIB) framework for multi-behavior recommendation to effectively address these challenges. Following information bottleneck principles, our framework optimizes the learning of compact yet sufficient representations that preserve essential information for target behavior prediction while eliminating task-irrelevant redundancies. To further mitigate interaction noise, we introduce a Graph Refinement Encoder (GRE) that dynamically prunes redundant edges through learnable edge dropout mechanisms. We conduct comprehensive experiments on three real-world public datasets, which demonstrate the superior effectiveness of our framework. Beyond these widely used datasets in the academic community, we further expand our evaluation on several real industrial scenarios and conduct an online A/B testing, showing again a significant improvement in multi-behavior recommendations. The source code of our proposed HGIB is available at https://github.com/zhy99426/HGIB. Hengyu Zhang 0001, Chunxu Shen, Xiangguo Sun, Jie Tan 0001, Yanchao Tan, Yu Rong 0001, Hong Cheng 0001, Lingling Yi |
RecSys | 2 |
| 2025 | Adaptive Graph Integration for Cross-Domain Recommendation via Heterogeneous Graph CoordinatorsabstractIn the digital era, users typically interact with diverse items across multiple domains (e.g., e-commerce, streaming platforms, and social networks), generating intricate heterogeneous interaction graphs. Leveraging multi-domain data can improve recommendation systems by enriching user insights and mitigating data sparsity in individual domains. However, integrating such multi-domain knowledge for cross-domain recommendation remains challenging due to inherent disparities in user behavior and item characteristics and the risk of negative transfer, where irrelevant or conflicting information from the source domains adversely impacts the target domain's performance. To tackle these challenges, we propose HAGO, a novel framework with Heterogeneous Adaptive Graph coOrdinators, which dynamically integrates multi-domain graphs into a cohesive structure. HAGO adaptively adjusts the connections between coordinators and multi-domain graph nodes to enhance beneficial inter-domain interactions while alleviating negative transfer. Furthermore, we introduce a universal multi-domain graph pre-training strategy alongside HAGO to collaboratively learn high-quality node representations across domains. Being compatible with various graph-based models and pre-training techniques, HAGO demonstrates broad applicability and effectiveness. Extensive experiments show that our framework outperforms state-of-the-art methods in cross-domain recommendation scenarios, underscoring its potential for real-world applications. The source code is available at https://github.com/zhy99426/HAGO. Hengyu Zhang 0001, Chunxu Shen, Xiangguo Sun, Jie Tan 0001, Yu Rong 0001, Chengzhi Piao, Hong Cheng 0001, Lingling Yi |
SIGIR | 2 |
| 2021 | Contextualize Knowledge Bases with Transformer for End-to-end Task-Oriented Dialogue SystemsabstractIncorporating knowledge bases (KB) into endto-end task-oriented dialogue systems is challenging, since it requires to properly represent the entity of KB, which is associated with its KB context and dialogue context.The existing works represent the entity with only perceiving a part of its KB context, which can lead to the less effective representation due to the information loss, and adversely favor KB reasoning and response generation.To tackle this issue, we explore to fully contextualize the entity representation by dynamically perceiving all the relevant entities and dialogue history.To achieve this, we propose a COntextaware Memory Enhanced Transformer framework (COMET), which treats the KB as a sequence and leverages a novel Memory Mask to enforce the entity to only focus on its relevant entities and dialogue history, while avoiding the distraction from the irrelevant entities.Through extensive experiments, we show that our COMET framework can achieve superior performance over the state of the arts. Yanjie Gou, Yinjie Lei, Lingqiao Liu, Yong Dai 0001, Chunxu Shen |
EMNLP (1) | 5 |
| 2019 | Discriminative Correlation Filter Network for Robust Landmark Tracking in Ultrasound Guided Intervention
Chunxu Shen, Jishuai He, Yibin Huang, Jian Wu 0012 |
MICCAI (5) | 1 |
| 2019 | Siamese Spatial Pyramid Matching Network with Location Prior for Anatomical Landmark Tracking in 3-Dimension Ultrasound Sequence
Jishuai He, Chunxu Shen, Yibin Huang, Jian Wu 0012 |
PRCV (2) | 2 |
| 2018 | An Online Learning Approach for Robust Motion Tracking in Liver Ultrasound Sequence
Chunxu Shen, Huabei Shi, Yibin Huang, Jian Wu 0012 |
PRCV (3) | 1 |