Kedar Bellare

dblp:42/5376 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
3since 2021 · last 2025
0009-0004-6827-8719ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Maps Ranking Optimization in Airbnb
Malay Haldar, Kedar Bellare, Sherry Chen, Soumyadip Banerjee 0003, Xiaotang Wang, Mustafa Abdool, Huiji Gao, Pavan Tapadia, Li-wei He, Sanjeev Katariya, Stephanie Moyerman
CIKM3
2025 Augmenting Guest Search Results with Recommendations at Airbnb
abstract
Users on Airbnb often perform exhaustive searches with varying conditions to find suitable accommodations. However, overly narrow search criteria can lead to insufficient results, causing frustration and abandonment of search journeys. To address these challenges, we developed flexible pivot recommendations that dynamically augment search results by suggesting alternative dates, relaxing amenity requirements, or adjusting price constraints. These recommendations align with users' broader travel intent, resulting in a measurable improvement in booking rates on the platform.
Philbert Lin, Dishant Ailawadi, Soumyadip Banerjee 0003, Shashank Dabriwal, Hao Li 0196, Kedar Bellare, Li-wei He, Sanjeev Katariya
CIKM7
2024 Learning to Rank for Maps at Airbnb
abstract
As a two-sided marketplace, Airbnb brings together hosts who own listings for rent with prospective guests from around the globe. Results from a guest's search for listings are displayed primarily through two interfaces: (1) as a list of rectangular cards that contain on them the listing image, price, rating, and other details, referred to as list-results (2) as oval pins on a map showing the listing price, called map-results. Both these interfaces, since their inception, have used the same ranking algorithm that orders listings by their booking probabilities and selects the top listings for display. But some of the basic assumptions underlying ranking, built for a world where search results are presented as lists, simply break down for maps. This paper describes how we rebuilt ranking for maps by revising the mathematical foundations of how users interact with search results. Our iterative and experiment-driven approach led us through a path full of twists and turns, ending in a unified theory for the two interfaces. Our journey shows how assumptions taken for granted when designing machine learning algorithms may not apply equally across all user interfaces, and how they can be adapted. The net impact was one of the largest improvements in user experience for Airbnb which we discuss as a series of experimental validations.
Malay Haldar, Kedar Bellare, Sherry Chen, Soumyadip Banerjee 0003, Xiaotang Wang, Mustafa Abdool, Huiji Gao, Pavan Tapadia, Li-wei He, Sanjeev Katariya
KDD3
2015 Errata for "Crowdsourcing Algorithms for Entity Resolution" (PVLDB 7(12): 1071-1082)
abstract
We discovered that there was a duplicate figure in our paper. We accidentally put Figure 13(b) for Figure 12(b). We have provided the correct Figure 12(b) above (See Figure 1). Figure 1 plots the recall of various strategies as a function of the number of questions asked for Places dataset. There was no error in the discussion in our paper (See Section 6.2.1 in our paper for more details).
Norases Vesdapunt, Kedar Bellare, Nilesh N. Dalvi
Proc. VLDB Endow.2
2014 Crowdsourcing Algorithms for Entity Resolution
abstract
In this paper, we study a hybrid human-machine approach for solving the problem of Entity Resolution (ER). The goal of ER is to identify all records in a database that refer to the same underlying entity, and are therefore duplicates of each other. Our input is a graph over all the records in a database, where each edge has a probability denoting our prior belief (based on Machine Learning models) that the pair of records represented by the given edge are duplicates. Our objective is to resolve all the duplicates by asking humans to verify the equality of a subset of edges, leveraging the transitivity of the equality relation to infer the remaining edges (e.g. a = c can be inferred given a = b and b = c ). We consider the problem of designing optimal strategies for asking questions to humans that minimize the expected number of questions asked. Using our theoretical framework, we analyze several strategies, and show that a strategy, claimed as " optimal " for this problem in a recent work, can perform arbitrarily bad in theory. We propose alternate strategies with theoretical guarantees. Using both public datasets as well as the production system at Facebook, we show that our techniques are effective in practice.
Norases Vesdapunt, Kedar Bellare, Nilesh N. Dalvi
Proc. VLDB Endow.2
2013 Discovering Hierarchical Structure for Sources and Entities
abstract
In this paper, we consider the problem of jointly learning hierarchies over a set of sources and entities based on their containment relationship. We model the concept of hierarchy using a set of latent binary features and propose a generative model that assigns those latent features to sources and entities in order to maximize the probability of the observed containment. To avoid fixing the number of features beforehand, we consider a non-parametric approach based on the Indian Buffet Process. The hierarchies produced by our algorithm can be used for completing missing associations and discovering structural bindings in the data. Using simulated and real datasets we provide empirical evidence of the effectiveness of the proposed approach in comparison to the existing hierarchy agnostic approaches.
Aditya Pal, Nilesh N. Dalvi, Kedar Bellare
AAAI3
2013 WOO: A Scalable and Multi-tenant Platform for Continuous Knowledge Base Synthesis
abstract
Search, exploration and social experience on the Web has recently undergone tremendous changes with search engines, web portals and social networks offering a different perspective on information discovery and consumption. This new perspective is aimed at capturing user intents, and providing richer and highly connected experiences. The new battleground revolves around technologies for the ingestion, disambiguation and enrichment of entities from a variety of structured and unstructured data sources - we refer to this process as knowledge base synthesis. This paper presents the design, implementation and production deployment of the Web Of Objects (WOO) system, a Hadoop-based platform tackling such challenges. WOO has been designed and implemented to enable various products in Yahoo! to synthesize knowledge bases (KBs) of entities relevant to their domains. Currently, the implementation of WOO we describe is used by various Yahoo! properties such as Intonow, Yahoo! Local, Yahoo! Events and Yahoo! Search. This paper highlights: (i) challenges that arise in designing, building and operating a platform that handles multi-domain, multi-version, and multi-tenant disambiguation of web-scale knowledge bases (hundreds of millions of entities), (ii) the architecture and technical solutions we devised, and (iii) an evaluation on real-world production datasets.
Kedar Bellare, Carlo Curino, Ashwin Machanavajjhala, Peter Mika, Mandar Rahurkar, Aamod Sane
Proc. VLDB Endow.1
2013 Active Sampling for Entity Matching with Guarantees
Kedar Bellare, Suresh Parthasarathy Iyengar, Aditya G. Parameswaran, Vibhor Rastogi
ACM Trans. Knowl. Discov. Data1
2012 Active sampling for entity matching
abstract
In entity matching, a fundamental issue while training a classifier to label pairs of entities as either duplicates or non-duplicates is the one of selecting informative training examples. Although active learning presents an attractive solution to this problem, previous approaches minimize the misclassification rate (0-1 loss) of the classifier, which is an unsuitable metric for entity matching due to class imbalance (i.e., many more non-duplicate pairs than duplicate pairs). To address this, a recent paper [1] proposes to maximize recall of the classifier under the constraint that its precision should be greater than a specified threshold. However, the proposed technique requires the labels of all n input pairs in the worst-case.
Kedar Bellare, Suresh Parthasarathy Iyengar, Aditya G. Parameswaran, Vibhor Rastogi
KDD1
2011 SampleRank: Training Factor Graphs with Atomic Gradients
Michael L. Wick, Khashayar Rohanimanesh, Kedar Bellare, Aron Culotta, Andrew McCallum
ICML3
2009 Generalized Expectation Criteria for Bootstrapping Extractors using Record-Text Alignment
Kedar Bellare, Andrew McCallum
EMNLP1
2009 Alternating Projections for Learning with Expectation Constraints
Kedar Bellare, Gregory Druck, Andrew McCallum
UAI1
2005 A Conditional Random Field for Discriminatively-trained Finite-state String Edit Distance
Andrew McCallum, Kedar Bellare, Fernando Pereira 0003
UAI2
2004 Generic Text Summarization Using WordNet
Kedar Bellare, Anish Das Sarma, Atish Das Sarma, Navneet Loiwal, Vaibhav Mehta, Ganesh Ramakrishnan, Pushpak Bhattacharyya
LREC1