Changsung Kang

dblp:94/4225 · DBLP profile ↗
← Back
12ranked-venue papers in the field
3as first author
4since 2021 · last 2026
0009-0007-5305-8256ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 8 (1 first)Data Mining & Knowledge Discovery · 3 (2 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 BERT-Based Cross-Encoder for Large-Scale Engagement Prediction and Re-ranking in Walmart Search Engine
abstract
Product search systems must not only return relevant items but also understand users' implicit preferences beyond their explicit queries. For instance, when searching for "steak", most users implicitly prefer beef steak over equally relevant alternatives like pork steak. Predicting such engagement preferences presents a more complex challenge than traditional relevance modeling, as it requires capturing nuanced query-item relationships that reflect both relevance and user intent. To capture these nuances, we extend semantic understanding to engagement prediction by learning directly from query-item text with engagement labels as supervision rather than relying on historical engagement statistics as input. Our approach effectively captures users' implicit preferences across diverse query types, from tail queries where historical signals are sparse to broad queries where understanding latent intent is critical. Extensive experiments on Walmart's production search data demonstrate significant improvements over production model with a strong relevance foundation: +1.71% add-to-cart lift in interleaving tests and +0.33% in overall search sessions with add-to-cart. Our model is deployed in the production environment of Walmart.com.
Philip Fu, Ajit Puthenputhussery, Changsung Kang, Cun Mu, Sachin Yadav 0004, Hongwei Shang 0001
SIGIR4
2026 Learning to Summarize for Search Relevance with Reinforcement Learning
Nitin Yadav, Changsung Kang, Hongwei Shang 0001
SIGIR2
2025 Large Scale Deployment of BERT Based Cross Encoder Model for Re-Ranking in Walmart Search Engine
abstract
Re-ranking plays a crucial role in product search by reassessing products from the primary retrieval system based on specific engagement and relevance criteria. While transformer-based models like the cross encoder have advanced the relevance of ranking models in recent years, a significant challenge arises from the high latency cost associated with running a cross encoder model at runtime. This challenge becomes more pronounced in the long-tail segment, where conventional techniques like caching prove ineffective. To tackle these issues, our paper introduces a scalable framework featuring a BERT-based cross encoder model for re-ranking, deployed in the Walmart search engine. We employ strategies such as intermediate representations, operator fusion, and vectorization to improve the inference latency of the cross encoder model. Furthermore, we provide a detailed discussion on the runtime implementation, highlighting key learnings and practical tricks that ensured minimal impact on response latency during production. Finally, we present the results of online experiments, including manual evaluation and interleaving test conducted on real-world e-commerce search traffic.
Ajit Puthenputhussery, Changsung Kang, Alessandro Magnani, Tian Zhang 0015, Hongwei Shang 0001, Nitin Yadav, Prijith Chandran, Bhavin Madhani, Yuan-Tai Fu, He Wang 0041, Zbigniew Gasiorek, Salvatore Tornatore, Srikanth Dasaka, Vivek Agrawal, Michael Bowersox, Cun Mu, Ciya Liao
SIGIR2
2025 Meta-Learning to Rank for Sparsely Supervised Queries
abstract
Supervisory signals are a critical resource for training learning to rank models. In many real-world search and retrieval scenarios, these signals may not be readily available or could be costly to obtain for some queries. The examples include domains where labeling requires professional expertise, applications with strong privacy constraints, and user engagement information that are too scarce. We refer to these scenarios as sparsely supervised queries which pose significant challenges to traditional learning to rank models. In this work, we address sparsely supervised queries by proposing a novel meta-learning to rank framework which leverages fast learning and adaption capability of meta-learning. The proposed approach accounts for the fact that different queries have different optimal parameters for their rankers, in contrast to traditional learning to rank models which only learn a global ranking model applied to all the queries. In consequence, the proposed method would yield significant advantages especially when new queries are of different characteristics with the training queries. Moreover, the proposed meta-learning to rank framework is generic and flexible. We conduct a set of comprehensive experiments on both public datasets and a real-world e-commerce dataset. The results demonstrate that the proposed meta-learning approach can significantly enhance the performance of learning to rank models with sparsely labeled queries.
Xuyang Wu 0002, Ajit Puthenputhussery, Hongwei Shang 0001, Changsung Kang, Yi Fang 0008
ACM Trans. Inf. Syst.4
2017 Large-Scale Location Prediction for Web Pages
abstract
Location information of Web pages plays an important role in location-sensitive tasks such as Web search ranking for location-sensitive queries. However, such information is usually ambiguous, incomplete, or even missing, which raises the problem of location prediction for Web pages. Meanwhile, Web pages are massive and often noisy, which pose challenges to the majority of existing algorithms for location prediction. In this paper, we propose a novel and scalable location prediction framework for Web pages based on the query-URL click graph. In particular, we introduce a concept of term location vectors to capture location distributions for all terms and develop an automatic approach to learn the importance of each term location vector for location prediction. Empirical results on a large URL set demonstrate that the proposed framework significantly improves the location prediction accuracy comparing with various representative baselines. We further provide a principled way to incorporate the proposed framework into the search ranking task and experimental results on a commercial search engine show that the proposed method remarkably boosts the ranking performance for location-sensitive queries.
Yuening Hu, Changsung Kang, Jiliang Tang, Dawei Yin 0001, Yi Chang 0001
IEEE Trans. Knowl. Data Eng.2
2016 Learning to Rewrite Queries
abstract
It is widely known that there exists a semantic gap between web documents and user queries and bridging this gap is crucial to advance information retrieval systems. The task of query rewriting, aiming to alter a given query to a rewrite query that can close the gap and improve information retrieval performance, has attracted increasing attention in recent years. However, the majority of existing query rewriters are not designed to boost search performance and consequently their rewrite queries could be sub-optimal. In this paper, we propose a learning to rewrite framework that consists of a candidate generating phase and a candidate ranking phase. The candidate generating phase provides us the flexibility to reuse most of existing query rewriters; while the candidate ranking phase allows us to explicitly optimize search relevance. Experimental results on a commercial search engine demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the important components of the proposed framework.
Jiliang Tang, Hua Ouyang, Changsung Kang, Dawei Yin 0001, Yi Chang 0001
CIKM4
2016 Ranking Relevance in Yahoo Search
abstract
Search engines play a crucial role in our daily lives. Relevance is the core problem of a commercial search engine. It has attracted thousands of researchers from both academia and industry and has been studied for decades. Relevance in a modern search engine has gone far beyond text matching, and now involves tremendous challenges. The semantic gap between queries and URLs is the main barrier for improving base relevance. Clicks help provide hints to improve relevance, but unfortunately for most tail queries, the click information is too sparse, noisy, or missing entirely. For comprehensive relevance, the recency and location sensitivity of results is also critical. In this paper, we give an overview of the solutions for relevance in the Yahoo search engine. We introduce three key techniques for base relevance -- ranking functions, semantic matching features and query rewriting. We also describe solutions for recency sensitive relevance and location sensitive relevance. This work builds upon 20 years of existing efforts on Yahoo search, summarizes the most recent advances and provides a series of practical relevance solutions. The performance reported is based on Yahoo's commercial search engine, where tens of billions of urls are indexed and served by the ranking system.
Dawei Yin 0001, Yuening Hu, Jiliang Tang, Tim Daly Jr., Mianwei Zhou, Hua Ouyang, Changsung Kang, Hongbo Deng, Chikashi Nobata, Jean-Marc Langlois, Yi Chang 0001
KDD8
2016 Learning Query and Document Relevance from a Web-scale Click Graph
abstract
Click-through logs over query-document pairs provide rich and valuable information for multiple tasks in information retrieval. This paper proposes a vector propagation algorithm on the click graph to learn vector representations for both queries and documents in the same semantic space. The proposed approach incorporates both click and content information, and the produced vector representations can directly improve ranking performance for queries and documents that have been observed in the click log. For new queries and documents that are not in the click log, we propose a two-step framework to generate the vector representation, which significantly improves the coverage of our vectors while maintaining the high quality. Experiments on Web-scale search logs from a major commercial search engine demonstrate the effectiveness and scalability of the proposed method. Evaluation results show that NDCG scores are significantly improved against multiple baselines by using the proposed method both as a ranking model and as a feature in a learning-to-rank framework.
Shan Jiang 0001, Yuening Hu, Changsung Kang, Tim Daly Jr., Dawei Yin 0001, Yi Chang 0001, ChengXiang Zhai
SIGIR3
2014 A hierarchical Dirichlet model for taxonomy expansion for search engines
abstract
Emerging trends and products pose a challenge to modern search engines since they must adapt to the constantly changing needs and interests of users. For example, vertical search engines, such as Amazon, eBay, Walmart, Yelp and Yahoo! Local, provide business category hierarchies for people to navigate through millions of business listings. The category information also provides important ranking features that can be used to improve search experience. However, category hierarchies are often manually crafted by some human experts and they are far from complete. Manually constructed category hierarchies cannot handle the ever-changing and sometimes long-tail user information needs. In this paper, we study the problem of how to expand an existing category hierarchy for a search/navigation system to accommodate the information needs of users more comprehensively. We propose a general framework for this task, which has three steps: 1) detecting meaningful missing categories; 2) modeling the category hierarchy using a hierarchical Dirichlet model and predicting the optimal tree structure according to the model; 3) reorganizing the corpus using the complete category structure, i.e., associating each webpage with the relevant categories from the complete category hierarchy. Experimental results demonstrate that our proposed framework generates a high-quality category hierarchy and significantly boosts the retrieval performance.
Changsung Kang, Yi Chang 0001, Jiawei Han 0001
WWW2
2012 Predicting primary categories of business listings for local search
abstract
We consider the problem of identifying primary categories of a business listing among the categories provided by the owner of the business. The category information submitted by business owners cannot be trusted with absolute certainty since they may purposefully add some secondary or irrelevant categories to increase recall in local search results, which makes category search very challenging for local search engines. Thus, identifying primary categories of a business is a crucial problem in local search. This problem can be cast as a multi-label classification problem with a large number of categories. However, the large scale of the problem makes it infeasible to use conventional supervised-learning-based text categorization approaches.
Changsung Kang, Jeehaeng Lee, Yi Chang 0001
CIKM1
2012 Learning to rank with multi-aspect relevance for vertical search
abstract
Many vertical search tasks such as local search focus on specific domains. The meaning of relevance in these verticals is domain-specific and usually consists of multiple well-defined aspects (e.g., text matching and distance in local search). Thus the overall relevance between a query and a document is a tradeoff between multiple relevance aspects. Such a tradeoff can vary for different types of queries or in different contexts. In this paper, we explore these vertical-specific aspects in the learning to rank setting. We propose a novel formulation in which the relevance between a query and a document is assessed with respect to each aspect, forming the multi-aspect relevance. In order to compute a ranking function, we study two types of learning-based approaches to estimate the tradeoff between these relevance aspects: a label aggregation method and a model aggregation method. Since there are only a few aspects, a minimal amount of training data is needed to learn the tradeoff. We conduct both offline and online test experiments on a local search engine and the experimental results show that our proposed multi-aspect relevance formulation is very promising. The two types of aggregation methods perform more effectively than a set of baseline methods including a conventional learning to rank method.
Changsung Kang, Xuanhui Wang, Yi Chang 0001, Belle L. Tseng
WSDM1
2011 Learning to re-rank web search results with multiple pairwise features
abstract
Web search ranking functions are typically learned to rank search results based on features of individual documents, i.e., pointwise features. Hence, the rich relationships among documents, which contain multiple types of useful information, are either totally ignored or just explored very limitedly. In this paper, we propose to explore multiple pairwise relationships between documents in a learning setting to rerank search results. In particular, we use a set of pairwise features to capture various kinds of pairwise relationships and design two machine learned re-ranking methods to effectively combine these features with a base ranking function: a pairwise comparison method and a pairwise function decomposition method. Furthermore, we propose several schemes to estimate the potential gains of our re-ranking methods on each query and selectively apply them to queries with high confidence. Our experiments on a large scale commercial search engine editorial data set show that considering multiple pairwise relationships is quite beneficial and our proposed methods can achieve significant gain over methods which only consider pointwise features or a single type of pairwise relationship.
Changsung Kang, Xuanhui Wang, Ciya Liao, Yi Chang 0001, Belle L. Tseng, Zhaohui Zheng 0001
WSDM1