VLDB 2026 Research / reviewers in the wild / expert
Kaiqiao Zhan
dblp:134/4750
· DBLP profile ↗
9ranked-venue papers in the field
0as first author
9since 2021 · last 2026
0009-0000-1642-7840ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 8Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gesture Clustering for Real-Time User Disentanglement in Shared-Account RecommendationabstractShared-account usage is common on short-video platforms, especially on mobile and tablet devices, where a single device is accessed by multiple users. While existing industrial solutions generally focus on behavior sequence purification to disentangle mixed user preferences, such approaches inherently depend on behavior accumulation and therefore lack the capability for real-time user identification. To adapt to online recommendation, utilizing gesture interaction features is a natural and promising option, as they (1) are instantaneous without behavior collection and (2) naturally encode fine-grained user operation habits. Nevertheless, we empirically observe that directly incorporating raw gesture features into recommendation models yields limited gains. Identity-discriminative patterns embedded in gesture signals are largely entangled during the main model training, preventing them from being leveraged as explicit and reliable identity cues. As a result, efficiently utilizing gesture information to provide more distinct identity signals for recommendation models remains a critical challenge. To address this issue, we propose G-CORE (Gesture Clustering for Real-time REcommendation), an unsupervised framework that disentangles gesture representations via clustering before integrating them into the main recommendation model. By providing clearer and more identity-aware signals, G-CORE enables the main model with faster user switching without relying on a volume of behavior accumulation. Through extensive offline experiments and online A/B tests on Kuaishou platform, G-CORE demonstrates its effectiveness in various shared-account scenarios, and has been successfully deployed in the Mobile and Tablet system of the platform. Huiying Hu, Xinlang Yue, Kexin Yi, Lingzhen Xu, Yangyi Fang, Yongqi Liu 0002, Kaiqiao Zhan |
SIGIR | 8 |
| 2026 | Revisiting Collaborative Filtering by Unleashing the Power of Similarity
Xinlang Yue, Yongqi Liu 0002, Kaiqiao Zhan |
SIGIR | 6 |
| 2026 | Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video RecommendationabstractShort-video recommender systems typically optimize ranking models using dense user behavioral signals, such as clicks and watch time. However, these signals are only indirect proxies of user satisfaction and often suffer from noise and bias. Recently, explicit satisfaction feedback collected through questionnaires has emerged as a high-quality direct alignment supervision, but is extremely sparse and easily overwhelmed by abundant behavioral data, making it difficult to incorporate into online recommendation models. To address these challenges, we propose a novel framework which is towards End-to-End Alignment of user Satisfaction via Questionnaire, named EASQ, to enable real-time alignment of ranking models with true user satisfaction. Specifically, we first construct an independent parameter pathway for sparse questionnaire signals by combining a multi-task architecture and a lightweight LoRA module. The multi-task design separates sparse satisfaction supervision from dense behavioral signals, preventing the former from being overwhelmed. The LoRA module pre-inject these preferences in a parameter-isolated manner, ensuring stability in the backbone while optimizing user satisfaction. Furthermore, we employ a DPO-based optimization objective tailored for online learning, which aligns the main model outputs with sparse satisfaction signals in real time. This design enables end-to-end online learning, allowing the model to continuously adapt to new questionnaire feedback while maintaining the stability and effectiveness of the backbone. Extensive offline experiments and large-scale online A/B tests demonstrate that EASQ consistently improves user satisfaction metrics across multiple scenarios. EASQ has been successfully deployed in a production short-video recommendation system, delivering significant and stable business gains. Minzhi Xie, Tiantian He 0005, Zixiu Wang, Lantao Hu, Yongqi Liu 0002, Han Li 0005, Kaiqiao Zhan, Kun Gai |
SIGIR | 10 |
| 2026 | Unifying User Satisfaction and Creator Incentive: A Constrained Optimization Framework for Short-Video RecommendationabstractShort-video recommendation systems typically optimize for user satisfaction. However, allocating exposure to creators at critical growth stages incentivizes long-term content supply despite compromising immediate user engagement. Existing efforts concerning creator interests aim at either improving creator exposure fairness, matching creators with suitable audiences, or leveraging creator behavior to enhance user welfare. Directly maximizing the joint value of user satisfaction and creator incentive at recommendation time, however, remains largely unaddressed. This presents two challenges. First, the two objectives are heterogeneous in nature, making it non-trivial to formulate this joint optimization as a tractable problem. Second, optimizing creator incentive requires globally coordinated decisions across requests, making real-time serving infeasible. To address these challenges, we formulate the joint maximization as a constrained optimization problem that unifies the two heterogeneous objectives. We further derive an efficient online algorithm based on the primal-dual method, which decouples global incentive constraints into real-time decisions with theoretical guarantees. Experiments on a large-scale short-video platform demonstrate consistent improvements in joint user-creator value over existing baselines. Xiaoru Qu, Dingyi Zhang, Zhangxi Yan, Hu Liu 0001, Jian Liang 0002, Kaiqiao Zhan |
SIGIR | 8 |
| 2026 | PushGen: Push Notifications Generation with LLMabstractWe present PushGen, an automated framework for generating high-quality push notifications comparable to human-crafted content. With the rise of generative models, there is growing interest in leveraging LLMs for push content generation. Although LLMs make content generation straightforward and cost-effective, maintaining stylistic control and reliable quality assessment remains challenging, as both directly impact user engagement. To address these issues, PushGen combines two key components: (1) a controllable category prompt technique to guide LLM outputs toward desired styles, and (2) a reward model that ranks and selects generated candidates. Extensive offline and online experiments demonstrate its effectiveness, which has been deployed in large-scale industrial applications, serving hundreds of millions of users daily. Shifu Bie, Jiangxia Cao, Zixiao Luo, Yichuan Zou, Lu Zhang 0084, Linxun Chen, Zhaojie Liu, Guorui Zhou, Kaiqiao Zhan, Kun Gai |
WSDM | 11 |
| 2025 | xMTF: A Formula-Free Model for Reinforcement-Learning-Based Multi-Task Fusion in Recommender SystemsabstractRecommender systems need to optimize various types of user feedback, e.g., clicks, likes, and shares. A typical recommender system handling multiple types of feedback has two components: a multi-task learning (MTL) module, predicting feedback such as click-through rate and like rate; and a multi-task fusion (MTF) module, integrating these predictions into a single score for item ranking. MTF is essential for ensuring user satisfaction, as it directly influences recommendation outcomes. Recently, reinforcement learning (RL) has been applied to MTF tasks to improve long-term user satisfaction. However, existing RL-based MTF methods are formula-based methods, which only adjust limited coefficients within pre-defined formulas. The pre-defined formulas restrict the RL search space and become a bottleneck for MTF. To overcome this, we propose a formula-free MTF framework. We demonstrate that any suitable fusion function can be expressed as a composition of single-variable monotonic functions, as per the Sprecher Representation Theorem. Leveraging this, we introduce a novel learnable monotonic fusion cell (MFC) to replace pre-defined formulas. We call this new MFC-based model eXtreme MTF (xMTF). Furthermore, we employ a two-stage hybrid (TSH) learning strategy to train xMTF effectively. By expanding the MTF search space, xMTF outperforms existing methods in extensive offline and online experiments. Yang Cao 0015, Changhao Zhang, Kaiqiao Zhan, Ben Wang 0006 |
WWW | 4 |
| 2025 | Unleashing the Potential of Two-Tower Models: Diffusion-Based Cross-Interaction for Large-Scale MatchingabstractTwo-tower models are widely adopted in the industrial-scale matching stage across a broad range of application domains, such as content recommendations, advertisement systems, and search engines. This model efficiently handles large-scale candidate item screening by separating user and item representations. However, the decoupling network also leads to a neglect of potential information interaction between the user and item representations. Current state-of-the-art (SOTA) approaches include adding a shallow fully connected layer(i.e., COLD), which is limited by performance and can only be used in the ranking stage. For performance considerations, another approach attempts to capture historical positive interaction information from the other tower by regarding them as the input features(i.e., DAT). Later research showed that the gains achieved by this method are still limited because of lacking the guidance on the next user intent. To address the aforementioned challenges, we propose a "cross-interaction decoupling architecture" within our matching paradigm. This user-tower architecture leverages a diffusion module to reconstruct the next positive intention representation and employs a mixed-attention module to facilitate comprehensive cross-interaction. During the next positive intention generation, we further enhance the accuracy of its reconstruction by explicitly extracting the temporal drift within user behavior sequences. Experiments on two real-world datasets and one industrial dataset demonstrate that our method outperforms the SOTA two-tower models significantly, and our diffusion approach outperforms other generative models in reconstructing item representations. Yihan Wang 0026, Zhexin Han, Kaiqiao Zhan, Ben Wang 0006 |
WWW | 5 |
| 2024 | Enhancing Playback Performance in Video Recommender Systems with an On-Device Gating and Ranking Framework
Zhenghao Qi, Honghuan Wu, Tieyao Zhang, Hao Li 0190, Yimin Tu, Kaiqiao Zhan, Ben Wang 0006 |
CIKM | 8 |
| 2024 | RPAF: A Reinforcement Prediction-Allocation Framework for Cache Allocation in Large-Scale Recommender SystemsabstractModern recommender systems are built upon computation-intensive infrastructure, and it is challenging to perform real-time computation for each request, especially in peak periods, due to the limited computational resources. Recommending by user-wise result caches is widely used when the system cannot afford a real-time recommendation. However, it is challenging to allocate real-time and cached recommendations to maximize the users’ overall engagement. This paper shows two key challenges to cache allocation, i.e., the value-strategy dependency and the streaming allocation. Then, we propose a reinforcement prediction-allocation framework (RPAF) to address these issues. RPAF is a reinforcement-learning-based two-stage framework containing prediction and allocation stages. The prediction stage estimates the values of the cache choices considering the value-strategy dependency, and the allocation stage determines the cache choices for each individual request while satisfying the global budget constraint. We show that the challenge of training RPAF includes globality and the strictness of budget constraints, and a relaxed local allocator (RLA) is proposed to address this issue. Moreover, a PoolRank algorithm is used in the allocation stage to deal with the streaming allocation problem. Experiments show that RPAF significantly improves users’ engagement under computational budget constraints. Shuo Su, Yao Wang 0020, Kaiqiao Zhan, Ben Wang 0006, Kun Gai |
RecSys | 6 |