Haiyuan Zhao

dblp:317/6939 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-1134-2877ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Towards Unbiased and Real-Time Staytime Prediction for Live Streaming Recommendation
abstract
Live streaming has emerged as a dynamic content format that delivers real-time and interactive experiences to users. Distinguished by the short lifespan and immersive nature of live rooms, live streaming poses two key challenges for recommendation: (1) Timeliness: the model must rapidly identify and promote relevant live rooms to target users within a limited window; and (2) Accurate staytime prediction: since extended watching often reflects content quality and user satisfaction, precisely predicting staytime serves as a critical indicator of recommendation relevance and user engagement. Existing approaches often improve timeliness by repeatedly sending staytime signals to accelerate model learning. However, this introduces label truncation bias, distorting the unbiased estimation of high staytime samples. To reconcile these competing demands, we propose MS3M (Multi-Stream Segmented Staytime Modeling), a novel framework that leverages multiple data streams for faster learning while employing segmented staytime modeling-converting staytime regression into a series of time-segmented classification tasks to ensure unbiased training. Furthermore, to address the sparsity of high staytime samples, MS3M's task-dependent architecture allows high staytime parameters to leverage prior knowledge from low staytime data, significantly improving generalization for long-duration watching behaviors. Extensive offline experiments and online A/B tests on TikTok confirm that MS3M effectively balances timeliness and unbiased learning, leading to substantial gains in recommendation accuracy. The proposed approach currently serves TikTok's live streaming recommendation system, contributing to continuous improvement in user watching experience.
Haiyuan Zhao, Changshuo Zhang, Zhen Ouyang, Bin Yuan 0005, Qinglei Wang, Zuotao Liu
CIKM1
2025 Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
abstract
Previous studies have found that PLM-based retrieval models exhibit a preference for LLM-generated content, assigning higher relevance scores to these documents even when their semantic quality is comparable to human-written ones. This phenomenon, known as source bias, threatens the sustainable development of the information access ecosystem. However, the underlying causes of source bias remain unexplored. In this paper, we explain the process of information retrieval with a causal graph and discover that PLM-based retrievers learn perplexity features for relevance estimation, causing source bias by ranking the documents with low perplexity higher. Theoretical analysis further reveals that the phenomenon stems from the positive correlation between the gradients of the loss functions in language modeling task and retrieval task. Based on the analysis, a causal-inspired inference-time debiasing method is proposed, called **C**ausal **D**iagnosis and **C**orrection (CDC). CDC first diagnoses the bias effect of the perplexity and then separates the bias effect from the overall estimated relevance score. Experimental results across three domains demonstrate the superior debiasing effectiveness of CDC, emphasizing the validity of our proposed explanatory framework. Source codes are available at https://github.com/WhyDwelledOnAi/Perplexity-Trap.
Sunhao Dai, Haiyuan Zhao, Liang Pang 0001, Xiao Zhang 0034, Gang Wang 0056, Zhenhua Dong, Jun Xu 0001, Ji-Rong Wen
ICLR3
2025 Uncertainty-aware evidential learning for legal case retrieval with noisy correspondence
Weicong Qin, Weijie Yu 0003, Kepu Zhang, Haiyuan Zhao, Jun Xu 0001, Ji-Rong Wen
Inf. Sci.4
2024 Counteracting Duration Bias in Video Recommendation via Counterfactual Watch Time
abstract
In video recommendation, an ongoing effort is to satisfy users' personalized information needs by leveraging their logged watch time. However, watch time prediction suffers from duration bias, hindering its ability to reflect users' interests accurately. Existing label-correction approaches attempt to uncover user interests through grouping and normalizing observed watch time according to video duration. Although effective to some extent, we found that these approaches regard completely played records (i.e., a user watches the entire video) as equally high interest, which deviates from what we observed on real datasets: users have varied explicit feedback proportion when completely playing videos. In this paper, we introduce the counterfactual watch time (CWT), the potential watch time a user would spend on the video if its duration is sufficiently long. Analysis shows that the duration bias is caused by the truncation of CWT due to the video duration limitation, which usually occurs on those completely played records. Besides, a Counterfactual Watch Model (CWM) is proposed, revealing that CWT equals the time users get the maximum benefit from video recommender systems. Moreover, a cost-based transform function is defined to transform the CWT into the estimation of user interest, and the model can be learned by optimizing a counterfactual likelihood function defined over observed user watch times. Extensive experiments on three real video recommendation datasets and online A/B testing demonstrated that CWM effectively enhanced video recommendation accuracy and counteracted the duration bias.
Haiyuan Zhao, Guohao Cai, Jieming Zhu, Zhenhua Dong, Jun Xu 0001, Ji-Rong Wen
KDD1
2023 Uncovering ChatGPT's Capabilities in Recommender Systems
abstract
The debut of ChatGPT has recently attracted significant attention from the natural language processing (NLP) community and beyond. Existing studies have demonstrated that ChatGPT shows significant improvement in a range of downstream NLP tasks, but the capabilities and limitations of ChatGPT in terms of recommendations remain unclear. In this study, we aim to enhance ChatGPT’s recommendation capabilities by aligning it with traditional information retrieval (IR) ranking capabilities, including point-wise, pair-wise, and list-wise ranking. To achieve this goal, we re-formulate the aforementioned three recommendation policies into prompt formats tailored specifically to the domain at hand. Through extensive experiments on four datasets from different domains, we analyze the distinctions among the three recommendation policies. Our findings indicate that ChatGPT achieves an optimal balance between cost and performance when equipped with list-wise ranking. This research sheds light on a promising direction for aligning ChatGPT with recommendation tasks. To facilitate further explorations in this area, the full code and detailed original results are open-sourced at https://github.com/rainym00d/LLM4RS.
Sunhao Dai, Ninglu Shao, Haiyuan Zhao, Weijie Yu 0003, Zihua Si, Chen Xu 0010, Zhongxiang Sun, Xiao Zhang 0034, Jun Xu 0001
RecSys3
2023 Uncovering User Interest from Biased and Noised Watch Time in Video Recommendation
abstract
In the video recommendation, watch time is commonly adopted as an indicator of user interest. However, watch time is not only influenced by the matching of users’ interests but also by other factors, such as duration bias and noisy watching. Duration bias refers to the tendency for users to spend more time on videos with longer durations, regardless of their actual interest level. Noisy watching, on the other hand, describes users taking time to determine whether they like a video or not, which can result in users spending time watching videos they do not like. Consequently, the existence of duration bias and noisy watching make watch time an inadequate label for indicating user interest. Furthermore, current methods primarily address duration bias and ignore the impact of noisy watching, which may limit their effectiveness in uncovering user interest from watch time. In this study, we first analyze the generation mechanism of users’ watch time from a unified causal viewpoint. Specifically, we considered the watch time as a mixture of the user’s actual interest level, the duration-biased watch time, and the noisy watch time. To mitigate both the duration bias and noisy watching, we propose Debiased and Denoised watch time Correction (D2Co), which can be divided into two steps: First, we employ a duration-wise Gaussian Mixture Model plus frequency-weighted moving average for estimating the bias and noise terms; then we utilize a sensitivity-controlled correction function to separate the user interest from the watch time, which is robust to the estimation error of bias and noise terms. The experiments on two public video recommendation datasets and online A/B testing indicate the effectiveness of the proposed method.
Haiyuan Zhao, Lei Zhang 0006, Jun Xu 0001, Guohao Cai, Zhenhua Dong, Ji-Rong Wen
RecSys1
2023 Separating Examination and Trust Bias from Click Predictions for Unbiased Relevance Ranking
abstract
Alleviating the examination and trust bias in ranking systems is an important research line in unbiased learning-to-rank (ULTR). Current methods typically use the propensity to correct the biased user clicks and then learn ranking models based on the corrected clicks. Though successes have been achieved, directly modifying the clicks suffers from the inherent high variance because the propensities are usually involved in the denominators of corrected clicks. The problem gets even worse in the situation of mixed examination and trust bias. To address the issue, this paper proposes a novel ULTR method called Decomposed Ranking Debiasing (DRD). DRD is tailored for learning unbiased relevance models with low variance in the existence of examination and trust bias. Unlike existing methods that directly modify the original user clicks, DRD proposes to decompose each click prediction as the combination of a relevance term outputted by the ranking model and other bias terms. The unbiased relevance model, therefore, can be learned by fitting the overall click predictions to the biased user clicks. A joint learning algorithm is developed to learn the relevance and bias models' parameters alternatively. Theoretical analysis showed that, compared with existing methods, DRD has lower variance while retains unbiasedness. Empirical studies indicated that DRD can effectively reduce the variance and outperform the state-of-the-art ULTR baselines.
Haiyuan Zhao, Jun Xu 0001, Xiao Zhang 0034, Guohao Cai, Zhenhua Dong, Ji-Rong Wen
WSDM1