Shunyu Zhang

dblp:288/1696 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0000-1936-1162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SID-Coord: Coordinating Semantic IDs for ID-based Ranking in Short-Video Search
abstract
Large-scale short-video search ranking models are typically trained on sparse co-occurrence signals over hashed item identifiers (HIDs). While effective at memorizing frequent interactions, such ID-based models struggle to generalize to long-tailed items with limited exposure. This memorization–generalization trade-off remains a longstanding challenge in such industrial systems. We propose SID-Coord, a lightweight Semantic ID framework that incorporates discrete, trainable semantic IDs (SIDs) directly into ID-based ranking models. Instead of treating semantic signals as auxiliary dense features, SID-Coord represents semantics as structured identifiers and coordinates HID-based memorization with SID-based generalization within a unified modeling framework. To enable effective coordination, SID-Coord introduces three components: (1) an attention-based fusion module over hierarchical SIDs to capture multi-level semantics, (2) a target-aware HID–SID gating mechanism that adaptively balances memorization and generalization, and (3) a SID-driven interest alignment module that models the semantic similarity distribution between target items and user histories. SID-Coord can be integrated into existing production ranking systems without modifying the backbone model. Online A/B experiments in a real-world production environment show statistically significant improvements, with a +0.664% gain in long-play rate in search and a +0.369% increase in search playback duration.
Shunyu Zhang, Xiaoze Jiang, Jingwei Zhuo
SIGIR3
2025 Unconstrained Monotonic Calibration of Predictions in Deep Ranking Systems
abstract
Ranking models primarily focus on modeling the relative order of predictions while often neglecting the significance of the accuracy of their absolute values. However, accurate absolute values are essential for certain downstream tasks, necessitating the calibration of the original predictions. To address this, existing calibration approaches typically employ predefined transformation functions with order-preserving properties to adjust the original predictions. Unfortunately, these functions often adhere to fixed forms, such as piece-wise linear functions, which exhibit limited expressiveness and flexibility, thereby constraining their effectiveness in complex calibration scenarios. To mitigate this issue, we propose implementing a calibrator using an Unconstrained Monotonic Neural Network (UMNN), which can learn arbitrary monotonic functions with great modeling power. This approach significantly relaxes the constraints on the calibrator, improving its flexibility and expressiveness while avoiding excessively distorting the original predictions by requiring monotonicity. Furthermore, to optimize this highly flexible network for calibration, we introduce a novel additional loss function termed Smooth Calibration Loss (SCLoss), which aims to fulfill a necessary condition for achieving the ideal calibration state. Extensive offline experiments confirm the effectiveness of our method in achieving superior calibration performance. Moreover, deployment in Kuaishou's large-scale online video ranking system demonstrates that the method's calibration improvements translate into enhanced business metrics. The source code is available at https://github.com/baiyimeng/UMC.
Yimeng Bai, Shunyu Zhang, Yang Zhang 0072, Hu Liu 0001, Wentian Bao, Enyun Yu, Fuli Feng, Wenwu Ou
SIGIR2
2024 Knowledge Enhanced Pre-training for Cross-lingual Dense Retrieval
abstract
In recent years, multilingual pre-trained language models (mPLMs) have achieved significant progress in cross-lingual dense retrieval. However, most mPLMs neglect the importance of knowledge. Knowledge always conveys similar semantic concepts in a language-agnostic manner, while query-passage pairs in cross-lingual retrieval also share common factual information. Motivated by this observation, we introduce KEPT, a novel mPLM that effectively leverages knowledge to learn language-agnostic semantic representations. To achieve this, we construct a multilingual knowledge base using hyperlinks and cross-language page alignment data annotated by Wiki. From this knowledge base, we mine intra- and cross-language pairs by extracting symmetrically linked segments and multilingual entity descriptions. Subsequently, we adopt contrastive learning with the mined pairs to pre-train KEPT. We evaluate KEPT on three widely-used benchmarks, considering both zero-shot cross-lingual transfer and supervised multilingual fine-tuning scenarios. Extensive experimental results demonstrate that KEPT achieves strong multilingual and cross-lingual retrieval performance with significant improvements over existing mPLMs.
Hang Zhang 0029, Yeyun Gong, Dayiheng Liu, Shunyu Zhang, Xingwei He 0003, Jiancheng Lv 0001, Jian Guo 0016
LREC/COLING4
2024 A Self-boosted Framework for Calibrated Ranking
abstract
Scale-calibrated ranking systems are ubiquitous in real-world applications nowadays, which pursue accurate ranking quality and calibrated probabilistic predictions simultaneously.For instance, in the advertising ranking system, the predicted click-through rate (CTR) is utilized for ranking and required to be calibrated for the downstream cost-per-click ads bidding.Recently, multi-objective based methods have been wildly adopted as a standard approach for Calibrated Ranking, which incorporates the combination of two loss functions: a pointwise loss that focuses on calibrated absolute values and a ranking loss that emphasizes relative orderings.However, when applied to industrial online applications, existing multi-objective CR approaches still suffer from two crucial limitations.First, previous methods need to aggregate the full candidate list within a single mini-batch to compute the ranking loss.Such aggregation strategy violates extensive data shuffling which has long been proven beneficial for preventing overfitting, and thus degrades the training effectiveness.Second, existing multi-objective methods apply the two inherently conflicting loss functions on a single probabilistic prediction, which results in a sub-optimal trade-off between calibration and ranking.To tackle the two limitations, we propose a Self-Boosted framework for Calibrated Ranking (SBCR).In SBCR, the predicted ranking scores by the online deployed model are dumped into context features.With these additional context features, each single item can perceive the overall distribution of scores in the whole ranking list, so that the ranking loss can be constructed without the need for sample aggregation.As the deployed model is a few versions older than the training model, the dumped predictions reveal what was failed to learn and keep boosting the model to correct previously mis-predicted items.Moreover, a calibration module is introduced to decouple the point loss and ranking loss.The two losses are applied before and after the calibration module separately, which
Shunyu Zhang, Hu Liu 0001, Wentian Bao, Enyun Yu, Yang Song 0008
KDD1
2024 TIM: Temporal Interaction Model in Notification System
abstract
Modern mobile applications heavily rely on the notification system to acquire daily active users and enhance user engagement.Being able to proactively reach users, the system has to decide when to send notifications to users.Although many researchers have studied optimizing the timing of sending notifications, they only utilized users' contextual features, without modeling users' behavior patterns.Additionally, these efforts only focus on individual notifications, and there is a lack of studies on optimizing the holistic timing of multiple notifications within a period.To bridge these gaps, we propose the Temporal Interaction Model (TIM), which models users' behavior patterns by estimating CTR in every time slot over a day in our short video application Kuaishou.TIM leverages long-term user historical interaction sequence features such as notification receipts, clicks, watch time and effective views, and employs a temporal attention unit (TAU) to extract user behavior patterns.Moreover, we provide an elegant strategy of holistic notifications send time control to improve user engagement while minimizing disruption.We evaluate the effectiveness of TIM through offline experiments and online A/B tests.The results indicate that TIM is a reliable tool for forecasting user behavior, leading to a remarkable enhancement in user engagement without causing undue disturbance.
Huxiao Ji, Linchuan Li, Shunyu Zhang, Cunyi Zhang, Wenwu Ou
ICMR4
2023 Query-dominant User Interest Network for Large-Scale Search Ranking
abstract
Historical behaviors have shown great effect and potential in various prediction tasks, including recommendation and information retrieval. The overall historical behaviors are various but noisy while search behaviors are always sparse. Most existing approaches in personalized search ranking adopt the sparse search behaviors to learn representation with bottleneck, which do not sufficiently exploit the crucial long-term interest. In fact, there is no doubt that user long-term interest is various but noisy for instant search, and how to exploit it well still remains an open problem.
Yong Yuan 0004, Jingyou Hou, Bingqing Ke, Junlin He, Shunyu Zhang, Enyun Yu, Wenwu Ou
CIKM10
2023 Modeling Sequential Sentence Relation to Improve Cross-lingual Dense Retrieval
Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001
ICLR1
2022 Multi-View Document Representation Learning for Open-Domain Dense Retrieval
abstract
Dense retrieval has achieved impressive advances in first-stage retrieval from a largescale document collection, which is built on bi-encoder architecture to produce single vector representation of query and document.However, a document can usually answer multiple potential queries from different views.So the single vector representation of a document is hard to match with multi-view queries, and faces a semantic mismatch problem.This paper proposes a multi-view document representation learning framework, aiming to produce multiview embeddings to represent documents and enforce them to align with different queries.First, we propose a simple yet effective method of generating multiple embeddings through viewers.Second, to prevent multi-view embeddings from collapsing to the same one, we further propose a global-local loss with annealed temperature to encourage the multiple viewers to better align with different potential queries.Experiments show our method outperforms recent works and achieves state-of-the-art results. * Work done during internship at Microsoft Research Asia.Q1: Where can people using iPods on planes view the device's interface?A1: Individual seat-back displays.Q2: What are two airlines that considered implementing iPod connections but did not join the 2007 agreement?A2: KLM and Air France.
Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001
ACL (1)1