Xunfan Cai

dblp:379/4060 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0001-3973-8577ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval
1.012026
Scaling and Stabilizing Large-Scale Embedding-Based Retrieval · SIGIR 2026
Information retrieval
hard negative mining
1.012026
Scaling and Stabilizing Large-Scale Embedding-Based Retrieval · SIGIR 2026
Information retrieval › retrieval models
knowledge distillation for retrieval
1.012026
Scaling and Stabilizing Large-Scale Embedding-Based Retrieval · SIGIR 2026
Information retrieval › ranking
learning to rank
1.012026
Scaling and Stabilizing Large-Scale Embedding-Based Retrieval · SIGIR 2026
Information retrieval › retrieval models
neural retrieval
1.012026
Scaling and Stabilizing Large-Scale Embedding-Based Retrieval · SIGIR 2026

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.0hard negative mining · 1.0cross-batch sampling · 1.0
YearPublicationVenuePosition
2026 Scaling and Stabilizing Large-Scale Embedding-Based Retrieval
abstract
Embedding-based retrieval (EBR) is foundational to large-scale e-commerce search, yet its effectiveness is often constrained by the quality of training signals and the representational capacity of the encoder. Standard dual-encoders suffer from a training-inference gap: they are optimized on narrow candidate pools but must discriminate against hundreds of millions of items during inference. Furthermore, while transitioning to higher-capacity backbones can mitigate this gap, simply replacing a mature model can lead to inconsistent retrieval behavior and a loss of the domain-specific knowledge established in previous iterations. In this paper, we present a unified pipeline deployed at Walmart that addresses both signal quality and model evolution. Our contributions are two-fold: (1) Hybrid Hard Negative Mining: We integrate Online Cross-Batch Sampling to increase negative diversity by an order of magnitude and Hybrid Offline Mining, which combines cross-encoder predictions with metadata heuristics to identify nuanced mismatches. (2) Legacy-Aware Distillation: We transition from DistilBERT to a higher-capacity GTE-base encoder. To ensure a smooth and superior transition, we introduce a Warm-Start Distillation technique that transfers domain-specific expertise from the legacy model to the new backbone. Validated through extensive offline experiments and online A/B testing, the proposed pipeline is deployed in live production, delivering a +7.34% improvement in NDCG@5 and a +0.50% lift in gross revenue.
Zhen Yang 0051, Juexin Lin, Hongwei Shang 0001, Kaihao Li, Feng Liu 0051, Satya Chembolu, Xunfan Cai, Cun Mu, Ciya Liao
SIGIR7