Sachin Yadav 0004

dblp:263/4935-4 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0001-9277-9987ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 BERT-Based Cross-Encoder for Large-Scale Engagement Prediction and Re-ranking in Walmart Search Engine
abstract
Product search systems must not only return relevant items but also understand users' implicit preferences beyond their explicit queries. For instance, when searching for "steak", most users implicitly prefer beef steak over equally relevant alternatives like pork steak. Predicting such engagement preferences presents a more complex challenge than traditional relevance modeling, as it requires capturing nuanced query-item relationships that reflect both relevance and user intent. To capture these nuances, we extend semantic understanding to engagement prediction by learning directly from query-item text with engagement labels as supervision rather than relying on historical engagement statistics as input. Our approach effectively captures users' implicit preferences across diverse query types, from tail queries where historical signals are sparse to broad queries where understanding latent intent is critical. Extensive experiments on Walmart's production search data demonstrate significant improvements over production model with a strong relevance foundation: +1.71% add-to-cart lift in interleaving tests and +0.33% in overall search sessions with add-to-cart. Our model is deployed in the production environment of Walmart.com.
Philip Fu, Ajit Puthenputhussery, Changsung Kang, Cun Mu, Sachin Yadav 0004, Hongwei Shang 0001
SIGIR7
2024 Enhancing Relevance of Embedding-based Retrieval at Walmart
abstract
Embedding-based neural retrieval (EBR) is an effective search retrieval method in product search for tackling the vocabulary gap between customer search queries and products. The initial launch of our EBR system at Walmart yielded significant gains in relevance and add-to-cart rates [1]. However, despite EBR generally retrieving more relevant products for reranking, we have observed numerous instances of relevance degradation. Enhancing retrieval performance is crucial, as it directly influences product reranking and affects the customer shopping experience. Factors contributing to these degradations include false positives/negatives in the training data and the inability to handle query misspellings. To address these issues, we present several approaches to further strengthen the capabilities of our EBR model in terms of retrieval relevance. We introduce a Relevance Reward Model (RRM) based on human relevance feedback. We utilize RRM to remove noise from the training data and distill it into our EBR model through a multi-objective loss. In addition, we present the techniques to increase the performance of our EBR model, such as typo-aware training, and semi-positive generation. The effectiveness of our EBR is demonstrated through offline relevance evaluation, online AB tests, and successful deployments to live production.
Juexin Lin, Sachin Yadav 0004, Feng Liu 0051, Nicholas Rossi, Praveen Reddy Suram, Satya Chembolu, Prijith Chandran, Hrushikesh Mohapatra, Alessandro Magnani, Ciya Liao
CIKM2
2022 Semantic Retrieval at Walmart
abstract
In product search, the retrieval of candidate products before re-ranking is more mission critical and challenging than other search like web search, especially for tail queries, which have a complex and specific search intent. In this paper, we present a hybrid system for e-commerce search deployed at Walmart that combines traditional inverted index and embedding-based neural retrieval to better answer user tail queries. Our system significantly improved the relevance of the search engine, measured by both offline and online evaluations. The improvements were achieved through a combination of different approaches. We present a new technique to train the neural model at scale. and describe how the system was deployed in production with little impact on response time. We highlight multiple learnings and practical tricks that were used in the deployment of this system.
Alessandro Magnani, Feng Liu 0051, Suthee Chaidaroon, Sachin Yadav 0004, Praveen Reddy Suram, Ajit Puthenputhussery, Min Xie 0002, Anirudh Kashi, Ciya Liao
KDD4