Yubin Kim 0001

dblp:96/7992 · DBLP profile ↗
← Back
16ranked-venue papers in the field
6as first author
7since 2021 · last 2026
0000-0001-5033-2677ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 13 (5 first)Data Mining & Knowledge Discovery · 2 (1 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 The Third Search Futures Workshop at ECIR'26
Leif Azzopardi, Charles L. A. Clarke, Claudia Hauff, Yubin Kim 0001, Zhaochun Ren, Adam Roegiest, Johanne R. Trippas, Saber Zerhoudi
ECIR (3)4
2026 SIGIR 2026 Workshop on eCommerce (ECOM26)
abstract
The eCommerce search and recommendations space is a unique, dynamic domain within information retrieval (IR), characterized by multimodality and industry-driven challenges. While the basic task of fulfilling a user's information need aligns with web search, the methodologies employed are distinct. On eCommerce platforms, the data available for retrieval and ranking differs significantly, as do the success signals (e.g.\ adding items to a cart, purchasing). The special theme of ECOM26 is User Interaction and Experience: Agentic-driven Trends. Our focus for 2026 is on fostering deeper engagement through interactive discussions, exploring crucial topics such as shifts in user interaction paradigms, and addressing emerging topics such as evaluation metrics for LLMs, multimodality, and the interplay between organic and sponsored search. With our discussion-heavy format and structured facilitation, we aim to spark conversation among all participants, beyond that of the usual interactions between presenters and audience questions.
Dean E. Alvarez, Aditya Chichani, Surya Kallumadi, Yubin Kim 0001, Tracy Holloway King, Andrew Trotman
SIGIR4
2025 First International Workshop on Data Quality-Aware Multimodal Recommendation (DaQuaMRec)
Claudio Pomo, Dietmar Jannach, Yubin Kim 0001, Daniele Malitesta, Alberto Carlo Maria Mancino, Julian J. McAuley, Alessandro B. Melchiorre, Shah Nawaz
RecSys3
2025 SIGIR 2025 Workshop on eCommerce (ECOM25): From Research to Product: Challenges, Lessons, and Opportunities in eCommerce Search and Recommendations
abstract
The eCommerce search and recommendations space is a unique and dynamic domain within information retrieval (IR), characterized by its multimodality and industry-driven challenges.While the basic task of fulfilling a user's information need aligns with web search, the methodologies employed are distinct.On eCommerce platforms (e.g.Alibaba, Amazon, eBay, Etsy, Flipkart, Walmart), the data available for retrieval and ranking differs significantly, as do the success signals (e.g.adding items to a cart, purchasing).Our focus for 2025 is on fostering deeper engagement through interactive discussions, exploring crucial topics such as navigating irreproducibility in research-to-product pipelines, and addressing emerging topics such as evaluation metrics for LLMs, multimodality, and the interplay between organic and sponsored search.With our discussion-heavy format and structured facilitation, we aim to spark conversation among all participants.
Yubin Kim 0001, Tracy Holloway King, Aditya Chichani, Pallavi Gudipati, Andrew Trotman
SIGIR1
2024 SIGIR 2024 Workshop on eCommerce (ECOM24)
Surya Kallumadi, Yubin Kim 0001, Tracy Holloway King, Maarten de Rijke, Vamsi Salaka
SIGIR2
2023 eCom'23: The SIGIR 2023 Workshop on eCommerce
abstract
eCommerce Information Retrieval (IR) is receiving increasing attention in the academic literature and is an essential component of some of the largest web sites (e.g. Airbnb, Alibaba, Amazon, eBay, Facebook, Flipkart, Lowes's, Taobao, Target). SIGIR has for several years seen sponsorship from eCommerce organizations, reflecting the importance of IR research to them. The purpose of this workshop is (1) to bring together researchers and practitioners of eCommerce IR to discuss topics unique to it, (2) to determine how to use eCommerce's unique combination of free text, structured data, and customer behavior data to improve search relevance, and (3) to examine how to build datasets and evaluate algorithms in this domain.
Surya Kallumadi, Yubin Kim 0001, Tracy Holloway King, Shervin Malmasi, Maarten de Rijke, Jacopo Tagliabue
SIGIR2
2022 Applications and Future of Dense Retrieval in Industry
abstract
Large-scale search engines are often designed as tiered systems with at least two layers. The L1 candidate retrieval layer efficiently generates a subset of potentially relevant documents (typically ~1000 documents) from a corpus many orders of magnitude larger in size. L1 systems emphasize efficiency and are designed to maximize recall. The L2 re-ranking layer uses a more computationally expensive, but more accurate model (e.g. learning-to-rank or neural model) to re-rank the candidates generated by L1 in order to maximize precision of the final result list.
Yubin Kim 0001
SIGIR1
2020 Overview of the Health Search and Data Mining (HSDM 2020) Workshop
abstract
We present HSDM, a full-day workshop on Health Search and Data Mining co-located with WSDM 2020's Health Day. This event builds on recent biomedical workshops in the NLP and ML communities but puts a clear emphasis on search and data mining (and their intersection) that is lacking in other venues. The program will include two keynote addresses by key opinion leaders in the clinical, search, and data mining domains. The technical program consists of 6 original research presentations. Finally, we will close with a panel discussion with keynote speakers, PC members, and the audience.
Carsten Eickhoff, Yubin Kim 0001, Ryen W. White
WSDM2
2019 Using Collection Shards to Study Retrieval Performance Effect Sizes
abstract
Despite the bulk of research studying how to more accurately compare the performance of IR systems, less attention is devoted to better understanding the different factors that play a role in such performance and how they interact. This is the case of shards, i.e., partitioning a document collection into sub-parts, which are used for many different purposes, ranging from efficiency to selective search or making test collection evaluation more accurate. In all these cases, there is empirical knowledge supporting the importance of shards, but we lack actual models that allow us to measure the impact of shards on system performance and how they interact with topics and systems. We use the general linear mixed model framework and present a model that encompasses the experimental factors of system, topic, shard, and their interaction effects. This detailed model allows us to more accurately estimate differences between the effect of various factors. We study shards created by a range of methods used in prior work and better explain observations noted in prior work in a principled setting and offer new insights. Notably, we discover that the topic*shard interaction effect, in particular, is a large effect almost globally across all datasets, an observation that, to our knowledge, has not been measured before.
Nicola Ferro 0001, Yubin Kim 0001, Mark Sanderson
ACM Trans. Inf. Syst.2
2017 Learning To Rank Resources
abstract
We present a learning-to-rank approach for resource selection. We develop features for resource ranking and present a training approach that does not require human judgments. Our method is well-suited to environments with a large number of resources such as selective search, is an improvement over the state-of-the-art in resource selection for selective search, and is statistically equivalent to exhaustive search even for recall-oriented metrics such as [email protected], an area in which selective search was lacking.
Zhuyun Dai, Yubin Kim 0001, Jamie Callan
SIGIR2
2017 Efficient distributed selective search
Yubin Kim 0001, Jamie Callan, J. Shane Culpepper, Alistair Moffat
Inf. Retr. J.1
2016 Does Selective Search Benefit from WAND Optimization?
Yubin Kim 0001, Jamie Callan, J. Shane Culpepper, Alistair Moffat
ECIR1
2016 Load-Balancing in Distributed Selective Search
abstract
Simulation and analysis have shown that selective search can reduce the cost of large-scale distributed information retrieval. By partitioning the collection into small topical shards, and then using a resource ranking algorithm to choose a subset of shards to search for each query, fewer postings are evaluated. Here we extend the study of selective search using a fine-grained simulation investigating: selective search efficiency in a parallel query processing environment; the difference in efficiency when term-based and sample-based resource selection algorithms are used; and the effect of two policies for assigning index shards to machines. Results obtained for two large datasets and four large query logs confirm that selective search is significantly more efficient than conventional distributed search. In particular, we show that selective search is capable of both higher throughput and lower latency in a parallel environment than is exhaustive search.
Yubin Kim 0001, Jamie Callan, J. Shane Culpepper, Alistair Moffat
SIGIR1
2016 Using the Crowd to Improve Search Result Ranking and the Search Experience
abstract
Despite technological advances, algorithmic search systems still have difficulty with complex or subtle information needs. For example, scenarios requiring deep semantic interpretation are a challenge for computers. People, on the other hand, are well suited to solving such problems. As a result, there is an opportunity for humans and computers to collaborate during the course of a search in a way that takes advantage of the unique abilities of each. While search tools that rely on human intervention will never be able to respond as quickly as current search engines do, recent research suggests that there are scenarios where a search engine could take more time if it resulted in a much better experience. This article explores how crowdsourcing can be used at query time to augment key stages of the search pipeline. We first explore the use of crowdsourcing to improve search result ranking. When the crowd is used to replace or augment traditional retrieval components such as query expansion and relevance scoring, we find that we can increase robustness against failure for query expansion and improve overall precision for results filtering. However, the gains that we observe are limited and unlikely to make up for the extra cost and time that the crowd requires. We then explore ways to incorporate the crowd into the search process that more drastically alter the overall experience. We find that using crowd workers to support rich query understanding and result processing appears to be a more worthwhile way to make use of the crowd during search. Our results confirm that crowdsourcing can positively impact the search experience but suggest that significant changes to the search process may be required for crowdsourcing to fulfill its potential in search systems.
Yubin Kim 0001, Kevyn Collins-Thompson, Jaime Teevan
ACM Trans. Intell. Syst. Technol.1
2015 How Random Decisions Affect Selective Distributed Search
abstract
Selective distributed search is a retrieval architecture that reduces search costs by partitioning a corpus into topical shards such that only a few shards need to be searched for each query. Prior research created topical shards by using random seed documents to cluster a random sample of the full corpus. The resource selection algorithm might use a different random sample of the corpus. These random components make selective search non-deterministic. This paper studies how these random components affect experimental results. Experiments on two ClueWeb09 corpora and four query sets show that in spite of random components, selective search is stable for most queries.
Zhuyun Dai, Yubin Kim 0001, Jamie Callan
SIGIR2
2010 ProbClean: A probabilistic duplicate detection system
abstract
One of the most prominent data quality problems is the existence of duplicate records. Current data cleaning systems usually produce one clean instance (repair) of the input data, by carefully choosing the parameters of the duplicate detection algorithms. Finding the right parameter settings can be hard, and in many cases, perfect settings do not exist. We propose ProbClean, a system that treats duplicate detection procedures as data processing tasks with uncertain outcomes. We use a novel uncertainty model that compactly encodes the space of possible repairs corresponding to different parameter settings. ProbClean efficiently supports relational queries and allows new types of queries against a set of possible repairs.
George Beskales, Mohamed A. Soliman, Ihab F. Ilyas, Shai Ben-David, Yubin Kim 0001
ICDE5