Saurabh Sohoney

dblp:69/9012 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-6576-7788ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2025 AddressBind: Cross-modal Alignment of Addresses and Geocodes
abstract
Mapping addresses to geolocations accurately is a challenging and important problem, with many real-world applications such as delivery logistics, map building and path finding. High quality embedding of geospatial data (e.g., addresses, geocodes) which is grounded in real world play an important role in success of modeling tasks such as geocoding and address resolution/matching. Existing state-of-the-art (SOTA) approaches [9] have proposed to transform the address embedding space to mimic real world proximity via a triplet loss, but requires triplet engineering which is error prone and difficult to scale. In this work, we propose to embed addresses and geocodes data in the same embedding space to enable late fusion of cross-modal semantics and remove dependency on triplet creation. Our proposed model outperforms SOTA baselines (including Multilingual-E5-Large-Instruct [32], a top model on MTEB leaderboard) by improving geolocation accuracy and geocode outliers across geographies with diverse writing standards. We also observe significant gains in address embeddings quality intrinsically and the approach supports to jointly align more geospatial modalities.
Govind, Sayan Putatunda, Saurabh Sohoney
SIGSPATIAL/GIS3
2024 Address De-duplication using Iterative k-Core Graph Decomposition
abstract
A de-duplicated and complete address catalog is essential for any application or business which needs to manage large volumes of address data such as delivery logistics, first-responder services and government databases. For catalog creation, address data is usually procured from disparate sources, which often vary in quality, coverage, and introduce duplicates or variations of the same physical address. Address de-duplication is therefore a crucial step for creating a clean and unified address catalog. De-duplication is even more challenging at a global scale, due to diversity in address writing styles, which might lack standardized addressing systems and can be multi-lingual. In this paper, we formulate address de-duplication as an unsupervised graph clustering problem and propose SANGAM, a novel adaptation of the k-core graph decomposition algorithm. We evaluate this solution on diverse geographic regions around the world. In comparison to existing methods, we observe improvements on the F-beta measure for three datasets. Our key contributions are: (1) formulating address de-duplication as a graph clustering problem, (2) proposing SANGAM, a robust and generic de-duplication approach, and (3) validating its effectiveness on diverse geographies across three continents - Americas, Africa and Europe. (4) Further, we deploy our solution and show the positive impact on geocode learning, an essential application of our solution.
Sriganesh Balamurugan, Vaasudev Narayanan, Saurabh Sohoney
SIGSPATIAL/GIS3
2024 Accurate Customer Address Matching via Weak Supervision for Geocode Learning
abstract
Determining the precise location of customers is important for an efficient and reliable delivery experience, both for customers and delivery associates. Address text is a primary source of information provided by customers about their location. In this paper, we study the important and challenging task of matching free-form customer address text to determine if two addresses represent the same physical building. We introduce a novel address matching framework that leverages transformer-based encoder to prevent tedious and time-consuming efforts spent on manual feature engineering by the baseline model. Furthermore, our proposed framework employs weak supervision to leverage historic delivery information and generate high-quality labeled data. This reduces the requirement for massive amounts of labeled data, typically needed for transformer-based models. Our experiments on manually curated datasets demonstrate the effective and generic nature of our approach, as we achieve 15.57% improvement in recall at 95% precision, on average, compared to the current baseline model across four geographies. We also introduce delivery point (DP) geocode learning for cold-start addresses as a downstream application of customer address matching. In addition to offline experiments, we performed online A/B experiments for DP geocode learning with our proposed approach and observed delivery precision improved by 8.09% and delivery defects reduced by 11.78% on average across four geographies in comparison to the baseline model.
Arpan Paul, Saket Maheshwary, Saurabh Sohoney
SIGSPATIAL/GIS3
2022 Learning geospatially aware place embeddings via weak-supervision
abstract
Understanding and representing real-world places (physical locations where drivers can deliver packages) is key to successfully and efficiently delivering packages to customer's doorstep. Prerequisite to this is the task of capturing similarity and relatedness between places. Intuitively, places that belong to a same building should have similar characteristics in geospatial as well as textual space. However, these assumptions fail in practice as existing methods use customer address text as a proxy for places. While providing the address text, customers tend to miss-out on key tokens, use vernacular content or place synonyms and do not follow a standard structure making them inherently ambiguous. Thus, modelling the problem from linguistic perspective alone is not sufficient. To overcome these shortcomings, we adapt various state-of-the-art embedding learning techniques to geospatial domain and propose Places-FastText, Places-Bert, and Places-GraphSage. We train these models using weak-supervision by innovatively leveraging different geospatial signals already available from historical delivery data. Our experiments and intrinsic evaluation demonstrate the significance of utilizing these signals and neighborhood information in learning geospatially aware place embeddings. Conclusions are further validated by observing significant improvements in two domain specific tasks viz., Pair-wise Matching ([email protected] improves by 29%) and Candidate Generation (avg [email protected] improves by 10%) as evaluated on UAE addresses.
Vamsi Krishna Penumadu, Nitesh Methani, Saurabh Sohoney
SIGSPATIAL/GIS3
2010 All Words Domain Adapted WSD: Finding a Middle Ground between Supervision and Unsupervision
Mitesh M. Khapra, Anup Kulkarni, Saurabh Sohoney, Pushpak Bhattacharyya
ACL3
2010 Value for Money: Balancing Annotation Effort, Lexicon Building and Accuracy for Multilingual WSD
Mitesh M. Khapra, Saurabh Sohoney, Anup Kulkarni, Pushpak Bhattacharyya
COLING2