VLDB 2026 Research / reviewers in the wild / expert
Haochen Sui
dblp:378/5409
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0000-0405-7301ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 84% Data mining · 16% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 77% Trustworthy machine learning · 23% | |
| Computer networks
1 paper |
Network measurement and analytics · 77% Content delivery and video streaming · 23% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.9 | 1 | 2025 | Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems · KDD (2) 2025 |
Recommender systems
click-through rate prediction |
0.9 | 1 | 2025 | Contrastive Prototype Framework for Calibrating Video Recommendation · ACM Multimedia 2025 |
Recommender systems
cross-domain recommendation |
0.9 | 1 | 2025 | Prototype-Guided Representation Projection for Multi-Domain Multi-Task Recommendation · ACM Multimedia 2025 |
Recommender systems
debiased recommendation |
0.9 | 1 | 2025 | Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems · KDD (2) 2025 |
Data mining
representation learning |
0.9 | 1 | 2025 | Prototype-Guided Representation Projection for Multi-Domain Multi-Task Recommendation · ACM Multimedia 2025 |
Recommender systems › debiased recommendation
selection bias |
0.9 | 1 | 2025 | Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems · KDD (2) 2025 |
Recommender systems
video recommendation |
0.9 | 1 | 2025 | Contrastive Prototype Framework for Calibrating Video Recommendation · ACM Multimedia 2025 |
Machine learning › Trustworthy machine learning
calibration |
0.3 | 1 | 2025 | Contrastive Prototype Framework for Calibrating Video Recommendation · ACM Multimedia 2025 |
Recommender systems
mixture of experts |
0.3 | 1 | 2025 | Prototype-Guided Representation Projection for Multi-Domain Multi-Task Recommendation · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
prototype learning · 2.6structural causal model · 1.7monte carlo algorithm · 1.7likelihood maximization · 1.7contrastive learning · 1.7causal intervention · 1.7timesieve · 0.9robustness analysis · 0.9orthogonal loss · 0.9optimal transport · 0.9mixture of experts · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Addressing Correlated Latent Exogenous Variables in Debiased Recommender SystemsabstractRecommendation systems (RS) aim to provide personalized content, but they face a challenge in unbiased learning due to selection bias, where users only interact with items they prefer. This bias leads to a distorted representation of user preferences, which hinders the accuracy and fairness of recommendations. To address the issue, various methods such as error imputation based, inverse propensity scoring, and doubly robust techniques have been developed. Despite the progress, from the structural causal model perspective, previous debiasing methods in RS assume the independence of the exogenous variables. In this paper, we release this assumption and propose a learning algorithm based on likelihood maximization to learn a prediction model. We first discuss the correlation and difference between unmeasured confounding and our scenario, then we propose a unified method that effectively handles latent exogenous variables. Specifically, our method models the data generation process with latent exogenous variables under mild normality assumptions. We then develop a Monte Carlo algorithm to numerically estimate the likelihood function. Extensive experiments on synthetic datasets and three real-world datasets demonstrate the effectiveness of our proposed method. The code is at https://github.com/WallaceSUI/kdd25-background-variable. Shuqiang Zhang, Yuchao Zhang 0001, Jinkun Chen, Haochen Sui |
KDD (2) | 4 |
| 2025 | From Guesswork to Guarantee: Towards Faithful Multimedia Web Forecasting with TimeSieveabstractThe domain of time series forecasting has gained significant attention due to its critical applications in multimedia-rich web traffic (including video streaming workloads and dynamic content delivery) and cross-platform advertisement click predictions, which are essential for web operations planning. While models like TimeSieve have demonstrated strong capabilities in predicting web visitation metrics, they suffer from critical unfaithfulness issues, including sensitivity to random seeds, input noise, layer noise, and parametric perturbations. To address these limitations, we propose Faithful TimeSieve (FTS), an enhanced framework designed to improve prediction reliability and robustness. Our approach systematically detects and mitigates unfaithfulness in TimeSieve, significantly enhancing its stability and consistency. Experimental results demonstrate that FTS substantially improves the model's faithfulness, setting a new standard for temporal forecasting methods. This advancement not only increases TimeSieve's reliability but also contributes to more robust temporal modeling, particularly crucial for web traffic forecasting where prediction accuracy directly impacts operational decisions. Our work thus represents a significant step toward more dependable time series predictions in web-related applications. Songning Lai, Ninghui Feng, Jiechao Gao, Hao Wang 0220, Haochen Sui, Xin Zou 0001, Wenshuo Chen, Lijie Hu, Hang Zhao 0010, Xuming Hu, Yutao Yue |
ACM Multimedia | 5 |
| 2025 | Contrastive Prototype Framework for Calibrating Video RecommendationabstractOnline video recommendation systems often build binary labels based on play complete rate (i.e., the ratio of watch time to video duration), such as complete play and effective play, using them as implicit feedback for Click-Through Rate (CTR) prediction tasks to gauge user interest. Existing works tend to improve prediction accuracy by designing complex models, overlooking that a key cause of inaccurate predictions is the disorganization of instance representation space. To address this issue, we explore a novel approach using prototype learning to calibrate the instance representation space of deep recommendation models and propose a model-agnostic Contrastive Prototype Framework (CPF). Firstly, CPF partitions the instance space into different subspaces based on duration, then generates positive and negative prototype pairs for each subspace from pre-trained recommendation model. Subsequently, we map the instance representations to the prototype space and calibrate them by reducing the distance to the corresponding prototypes. Ultimately, the prediction is derived from the linear combination of the estimated values associated with each prototype. To prevent disorganization in the prototype space during training, we design contrastive and orthogonality losses to constrain the learning of prototypes. Additionally, we show that how CPF effectively addresses the duration bias from the perspective of causal intervention. Offline experiments on two datasets demonstrate that CPF improves recommendation accuracy over several baseline models in predicting five widely used implicit feedback labels. We have also deployed CPF on a short video platform, validating its effectiveness in real-world scenarios. Fan Li 0032, Jiazhen Huang, Shisong Tang, Huafeng Cao, Haochen Sui, Xiaoyu Kang |
ACM Multimedia | 6 |
| 2025 | Prototype-Guided Representation Projection for Multi-Domain Multi-Task RecommendationabstractMulti-domain and multi-task learning enhance the efficiency and performance of industrial recommendation systems by integrating information from different domains/tasks to model user interests uniformly. However, existing methods suffer from the problem of representation entanglement, which limits the effective handling of commonality and specificity among various domains/tasks. In this paper, we propose a Prototype-guided Representation Projection (PRP) model to address this issue, which explores a novel direction of applying prototype learning to deal with complex domain/task relationships in the recommendation field. To identify inter-domain/task commonality, PRP initially uses a shared Mixture of Experts (MoE) architecture to learn representations for each sample, projecting them into a common prototype space across all domains/tasks. For domain/task specificity, specific feature extraction experts are employed, and sample representations are projected to the corresponding prototype spaces, constrained by an orthogonal loss to ensure the independence of those spaces. Moreover, PRP utilizes Optimal Transport (OT) to guide the correct representation projection within the prototype spaces, employing the linear combination of prototypes as the new sample representation. We conduct offline experiments on two open-source datasets and deploy our approach in an online system for A/B testing. Extensive experimental results consistently demonstrate that our approach outperforms existing methods. Binrui Wu, Haochen Sui, Jiaye Lin, Jiechao Gao, Keyan Jin |
ACM Multimedia | 2 |