VLDB 2026 Research / reviewers in the wild / expert
Fan Li 0017
dblp:73/237-17
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0000-0889-4684ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Recommender systems · 79% Machine learning and data management · 11% Information retrieval · 10% | |
| Artificial intelligence
2 papers |
Representation and self-supervised learning · 79% Efficient and distributed learning · 21% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
video recommendation |
1.7 | 2 | 2025 | Leveraging Label Distributions as Anchors to Enhance Video Recommendation · KDD (2) 2025 Calibrating Video Watch-time Predictions with Credible Prototype Alignment · ICML 2025 |
Recommender systems › video recommendation
watch-time prediction |
1.7 | 2 | 2025 | Leveraging Label Distributions as Anchors to Enhance Video Recommendation · KDD (2) 2025 Calibrating Video Watch-time Predictions with Credible Prototype Alignment · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation matching
feature alignment |
0.9 | 1 | 2025 | Aligning and Balancing ID and Multimodal Representations for Recommendation · KDD (2) 2025 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.9 | 1 | 2025 | Aligning and Balancing ID and Multimodal Representations for Recommendation · KDD (2) 2025 |
Recommender systems › click-through rate prediction
calibration |
0.9 | 1 | 2025 | Calibrating Video Watch-time Predictions with Credible Prototype Alignment · ICML 2025 |
Machine learning and data management
label distribution learning |
0.9 | 1 | 2025 | Leveraging Label Distributions as Anchors to Enhance Video Recommendation · KDD (2) 2025 |
Recommender systems
multimodal recommendation |
0.9 | 1 | 2025 | Aligning and Balancing ID and Multimodal Representations for Recommendation · KDD (2) 2025 |
Recommender systems
diversified recommendation |
0.8 | 1 | 2024 | Contextual Distillation Model for Diversified Recommendation · KDD 2024 |
Information retrieval
reranking |
0.8 | 1 | 2024 | Contextual Distillation Model for Diversified Recommendation · KDD 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.2 | 1 | 2024 | Contextual Distillation Model for Diversified Recommendation · KDD 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2024 | Contextual Distillation Model for Diversified Recommendation · KDD 2024 |
Methods — techniques the papers use, named apart from their topics
wasserstein distance · 1.7multimodal large language model · 1.7gradient modulation · 1.7optimal transport · 1.7knowledge distillation · 1.5contrastive learning · 1.5attention mechanism · 1.5vector quantized variational autoencoder · 0.9variational autoencoder · 0.9LASSO · 0.9maximal marginal relevance · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibrating Video Watch-time Predictions with Credible Prototype AlignmentabstractAccurately predicting user watch-time is crucial for enhancing user stickiness and retention in video recommendation systems. Existing watch-time prediction approaches typically involve transformations of watch-time labels for prediction and subsequent reversal, ignoring both the natural distribution properties of label and the instance representation confusion that results in inaccurate predictions. In this paper, we propose ProWTP, a two-stage method combining prototype learning and optimal transport for watch-time regression prediction, suitable for any deep recommendation model. Specifically, we observe that the watch-ratio (the ratio of watch-time to video duration) within the same duration bucket exhibits a multimodal distribution. To facilitate incorporation into models, we use a hierarchical vector quantised variational autoencoder (HVQ-VAE) to convert the continuous label distribution into a high-dimensional discrete distribution, serving as credible prototypes for calibrations. Based on this, ProWTP views the alignment between prototypes and instance representations as a Semi-relaxed Unbalanced Optimal Transport (SUOT) problem, where the marginal constraints of prototypes are relaxed. And the corresponding optimization problem is reformulated as a weighted Lasso problem for solution. Moreover, ProWTP introduces the assignment and compactness losses to encourage instances to cluster closely around their respective prototypes, thereby enhancing the prototype-level distinguishability. Finally, we conducted extensive offline experiments on two industrial datasets, demonstrating our consistent superiority in real-world application. Chao Cui, Shisong Tang, Fan Li 0017, Jiechao Gao, Hechang Chen |
ICML | 3 |
| 2025 | Aligning and Balancing ID and Multimodal Representations for RecommendationabstractLarge-scale recommendation systems mainly rely on sparse ID features, struggling with data sparsity. It's important to use multimodal information to assist ID learning for better performance. However, there exists two challenges: (1) distribution discrepancy between multimodal and ID makes direct integration prone to user-item mismatch; (2) slower convergence of multimodal representations compared to ID, causing optimization imbalance under a unified objective, which limits the potential of multimodal representations. In this paper, we comprehensively investigate the two problems and proposes a framework named AB-Rec to align and balance ID and multimodal representations learning for recommendation. We design three alignment tasks to fine-tune a pre-trained multimodal large language model (MLLM), which is then utilized to generate a unified multimodal representation for each item. AB-Rec aligns the distributions of ID and multimodal representations by minimizing the in-batch Wasserstein distance, and maximizes the distance between the two types of representations for the same item to avoid representation collapse. To solve the optimization imbalance, we propose a gradient modulation method that adaptively controls the optimization process by monitoring the contribution differences between ID and multimodal representations. Finally, we conduct extensive offline experiments on four datasets and an A/B test on an online video platform, demonstrating the effectiveness and scalability of our proposed method. Binrui Wu, Shisong Tang, Fan Li 0017, Chang Meng, Jingyu Xiao, Jiechao Gao |
KDD (2) | 3 |
| 2025 | Leveraging Label Distributions as Anchors to Enhance Video RecommendationabstractIn video recommendation systems, accurately predicting watch time is crucial for enhancing user engagement and retention. Traditional methods typically apply label transformations or mitigate duration bias to improve performance but overlook that erroneous instance representations are the primary cause of significant prediction errors. Moreover, these approaches predominantly rely on point perdition, limiting their robustness. To address these challenges, we propose LDA, a novel prediction paradigm that optimizes instance representations by explicitly leveraging label distributions as anchors within the model, enabling more accurate and robust predictions. Our analysis reveals that watch ratio across different duration groups exhibit distinct multi-peak distributions, reflecting the strong aggregation of user behavior. Based on this finding, we employ Vector Quantized Variational Auto-encoder (VQ-VAE) to convert the continuous watch ratio distribution into representative anchors that capture these multi-peak characteristics within each duration group. Subsequently, we project both instance representations and anchors into a common space and utilize Optimal Transport (OT) to generate pseudo-labels aligned with the anchor distribution, allowing instances to obtain structured coordinates within this space during training. Finally, we derive optimized instance representations for watch time prediction by aggregating anchor vectors through weighted integration. Extensive offline experiments on two datasets and large-scale online A/B testing on a short-video platform with over 300 million DAUs demonstrate the consistent superiority of LDA in watch time prediction. Chao Cui, Shisong Tang, Fan Li 0017, Huafeng Cao, Jiechao Gao, Hechang Chen |
KDD (2) | 4 |
| 2025 | Adaptive Gradient Masking for Balancing ID and MLLM-based Representations in RecommendationabstractIn large-scale recommendation systems, multimodal (MM) content is increasingly introduced to enhance the generalization of ID features.
The rise of Multimodal Large Language Models (MLLMs) enables the construction of unified user and item representations.
However, the semantic distribution gap between MM and ID representations leads to \textit{convergence inconsistency} during joint training: the ID branch converges quickly, while the MM branch requires more epochs, thus limiting overall performance.
To address this, we propose a two-stage framework including MM representation learning and joint training optimization.
First, we fine-tune the MLLM to generate unified user and item representations, and introduce collaborative signals by post-aligning user ID representations to alleviate semantic differences.
Then, we propose an Adaptive Gradient Masking (AGM) training strategy to dynamically regulate parameter updates between ID and MLLM branches.
AGM estimates the contribution of each representation with mutual information, and applies non-uniform gradient masking at the sub-network level to balance optimization.
We provide theoretical analysis of AGM's effectiveness and further introduce an unbiased variant, AGM*, to enhance training stability.
Experiments on offline and online A/B tests validate the effectiveness of our approach in mitigating convergence inconsistency and improving performance. Yidong Wu, Binrui Wu, Fan Li 0017, Jiechao Gao |
NeurIPS | 4 |
| 2024 | Contextual Distillation Model for Diversified RecommendationabstractThe diversity of recommendation is equally crucial as accuracy in improving user experience. Existing studies, e.g., Determinantal Point Process (DPP) and Maximal Marginal Relevance (MMR), employ a greedy paradigm to iteratively select items that optimize both accuracy and diversity. However, prior methods typically exhibit quadratic complexity, limiting their applications to the re-ranking stage and are not applicable to other recommendation stages with a larger pool of candidate items, such as the pre-ranking and ranking stages. In this paper, we propose Contextual Distillation Model (CDM), an efficient recommendation model that addresses diversification, suitable for the deployment in all stages of industrial recommendation pipelines. Specifically, CDM utilizes the candidate items in the same user request as context to enhance the diversification of the results. We propose a contrastive context encoder that employs attention mechanisms to model both positive and negative contexts. For the training of CDM, we compare each target item with its context embedding and utilize the knowledge distillation framework to learn the win probability of each target item under the MMR algorithm, where the teacher is derived from MMR outputs. During inference, ranking is performed through a linear combination of the recommendation and student model scores, ensuring both diversity and efficiency. We perform offline evaluations on two industrial datasets and conduct online A/B test of CDM on the short-video platform KuaiShou. The considerable enhancements observed in both recommendation quality and diversity, as shown by metrics, provide strong superiority for the effectiveness of CDM. Fan Li 0017, Xu Si, Shisong Tang, Dingmin Wang, Kunyan Han, Guorui Zhou, Yang Song 0008, Hechang Chen |
KDD | 1 |