Moumita Bhattacharya

dblp:157/0759 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-7836-4504ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Fourth Workshop on Personalization and Recommendations in Search (PaRiS)
abstract
With proliferation of personal computing devices and large number of logged-in experiences, search has evolved to a stage with many different product scenarios where personalization plays a crucial role for relevance quality and user satisfaction.The purpose of this workshop is have a forum where the latest research and advancements specifically on Personalization and Recommendations in Search (PaRiS) can be discussed in conjunction with KDD 2025.This will be the fourth instance of this workshop.We held three very successful instances of this workshop at the SIGIR 2024 [3], WebConf 2023 [2] and WSDM 2022.This year we will especially focus on applications of LLM and Generative AI to enable personalization and recommendations [1] in the context of search, for example, conversational assistants, while continuing to use this workshop to discuss other advances and applications in the context of personalized search and recommendations in the context of search.
Sudarshan Lamkhede, Moumita Bhattacharya
KDD (2)2
2024 Joint Modeling of Search and Recommendations Via an Unified Contextual Recommender (UniCoRn)
abstract
Search and recommendation systems are essential in many services, and they are often developed separately, leading to complex maintenance and technical debt. In this paper, we present a unified deep learning model that efficiently handles key aspects of both tasks.
Moumita Bhattacharya, Vito Ostuni, Sudarshan Lamkhede
RecSys1
2024 Third Workshop on Personalization and Recommendations in Search (PaRiS)
Sudarshan Lamkhede, Hamed Zamani, Moumita Bhattacharya, Hongning Wang
SIGIR3
2022 Deep Search Relevance Ranking in Practice
abstract
Machine learning techniques for developing industry-scale search engines have long been a prominent part of most domains and their online products. Search relevance algorithms are key components of products across different fields, including e-commerce, streaming services, and social networks. In this tutorial, we give an introduction to such large-scale search ranking systems, specifically focusing on deep learning techniques in this area. The topics we cover are the following: (1) Overview of search ranking systems in practice, including classical and machine learning techniques; (2) Introduction to sequential and language models in the context of search ranking; and (3) Knowledge distillation approaches for this area. For each of the aforementioned sessions, we first give an introductory talk and then go over an hands-on tutorial to really hone in on the concepts. We cover fundamental concepts using demos, case studies, and hands-on examples, including the latest Deep Learning methods that have achieved state-of-the-art results in generating the most relevant search results. Moreover, we show example implementations of these methods in python, leveraging a variety of open-source machine-learning/deep-learning libraries as well as real industrial data or open-source data.
Linsey Pang, Wei Liu 0007, Keng-hao Chang, Moumita Bhattacharya, Xianjing Liu, Stephen D. Guo
KDD5
2022 Augmenting Netflix Search with In-Session Adapted Recommendations
abstract
We motivate the need for recommendation systems that can cater to the members’ in-the-moment intent by leveraging their interactions from the current session. We provide an overview of an end-to-end in-session adaptive recommendations system in the context of Netflix Search. We discuss the challenges and potential solutions when developing such a system at production scale.
Moumita Bhattacharya, Sudarshan Lamkhede
RecSys1
2020 Query as Context for Item-to-Item Recommendation
abstract
Recommender Systems is one of the main machine learning applications for an e-commerce platforms such as Etsy, a two sided marketplace. A frequent usage of such system is for item-to-item recommendations that show similar items (also referred as listings) based on the listing a user is currently viewing. Item-to-item recommendations typically take into account information that are only associated with the target listing and other listings in the inventory. However, other contextual information such as user intents, queries and seasonality, are often not taken into account. In this talk, we will present two approaches we developed to utilize additional contextual information in the form of queries in generating item-to-item recommendations. Moreover, we will present our journey in migrating Etsy’s rankers from linear to non-linear models. Additionally, we propose new metrics to evaluate candidate sets that accesses diversity and price spread, while not compromising relevance. The proposed metrics can also be used beyond the current application. Our proposed candidate set generation approach outperforms the model in production as well as yielding significant lift in conversion rate and other engagement metrics as indicated by several A/B tests.
Moumita Bhattacharya, Amey Barapatre
RecSys1
2018 Co-occurrence of medical conditions: Exposing patterns through probabilistic topic modeling of snomed codes
abstract
Patients associated with multiple co-occurring health conditions often face aggravated complications and less favorable outcomes. Co-occurring conditions are especially prevalent among individuals suffering from kidney disease, an increasingly widespread condition affecting 13% of the general population in the US. This study aims to identify and characterize patterns of co-occurring medical conditions in patients employing a probabilistic framework. Specifically, we apply topic modeling in a non-traditional way to find associations across SNOMED-CT codes assigned and recorded in the EHRs of >13,000 patients diagnosed with kidney disease. Unlike most prior work on topic modeling, we apply the method to codes rather than to natural language. Moreover, we quantitatively evaluate the topics, assessing their tightness and distinctiveness, and also assess the medical validity of our results. Our experiments show that each topic is succinctly characterized by a few highly probable and unique disease codes, indicating that the topics are tight. Furthermore, inter-topic distance between each pair of topics is typically high, illustrating distinctiveness. Last, most coded conditions grouped together within a topic, are indeed reported to co-occur in the medical literature. Notably, our results uncover a few indirect associations among conditions that have hitherto not been reported as correlated in the medical literature.
Moumita Bhattacharya, Claudine Jurkovitz, Hagit Shatkay
J. Biomed. Informatics1
2017 Assessing chronic kidney disease from office visit records using hierarchical meta-classification of an imbalanced dataset
abstract
Chronic Kidney Disease (CKD) is an increasingly prevalent condition affecting 13% of the US population. The disease is often a silent condition, making its diagnosis challenging. Identifying CKD stages from standard office visit records can help in early detection of the disease and lead to timely intervention. The dataset we use is highly imbalanced. We propose a hierarchical meta-classification method, aiming to stratify CKD by severity levels, employing simple quantitative non-text features gathered from office visit records, while addressing data imbalance. Our method effectively stratifies CKD severity levels obtaining high average sensitivity, precision and F-measure (~93%). We also conduct experiments in which the dimensionality of the data is significantly reduced to include only the most salient features. Our results show that the good performance of our system is retained even when using the reduced feature sets, as well as under much reduced training sets, indicating that our method is stable and generalizable.
Moumita Bhattacharya, Claudine Jurkovitz, Hagit Shatkay
BIBM1
2017 Identifying articles relevant to drug-drug interaction: Addressing class imbalance
abstract
Interactions between drugs (also known as drug-drug interactions or DDIs), which may cause adverse affects, are of much concern; predicting, anticipating and avoiding them is key for improving patient safety and treatment outcome. Knowledge of DDIs is important for physicians to avoid adverse effects when prescribing two drugs simultaneously. DDIs are often published in the biomedical literature; however, gathering information about DDIs is time consuming given the shear volume of publications. Automatic text classification can speed up access to documents related to DDIs. However, the biomedical literature contains a relatively small number of publications relevant to DDIs, compared to the vast amount of irrelevant publications. This imbalance can lead to incorrect classification. While methods addressing class imbalance have been introduced to correctly identify items in the minority (relevant) class to improve recall, they often misclassify items in the majority (irrelevant) class, which leads to low precision. To reduce the number of irrelevant documents misclassified as relevant (false positive), we develop a two-stage cascade classifier. In each step, we separate publication abstracts that are DDI-relevant from those that are either drug-irrelevant or drug-relevant but DDI-irrelevant. We compare our classifier with other popular learning methods that aim to handle imbalance, applying the methods to a well-curated corpus consisting of DDI-relevant and DDI-irrelevant PubMed abstracts. Our method achieves higher precision and F1 measure than other methods while maintaining similar recall.
Moumita Bhattacharya, Heng-Yi Wu, Pengyuan Li 0001, Lang Li 0001, Hagit Shatkay
BIBM2
2016 Identifying patterns of associated-conditions through topic models of Electronic Medical Records
abstract
Multiple adverse health conditions co-occurring in a patient are typically associated with poor prognosis and increased office or hospital visits. Developing methods to identify patterns of co-occurring conditions can assist in diagnosis. Thus, identifying patterns of association among co-occurring conditions is of growing interest. In this paper, we report preliminary results from a data-driven study, in which we apply a machine learning method, namely, topic modeling, to Electronic Medical Records (EMRs), aiming to identify patterns of associated conditions. Specifically, we use the well-established Latent Dirichlet Allocation (LDA), a method based on the idea that documents can be modeled as a mixture of latent topics, where each topic is a distribution over words. In our study, we adapt the LDA model to identify latent topics in patients' EMRs. We evaluate the performance of our method both qualitatively and quantitatively, and show that the obtained topics indeed align well with distinct medical phenomena characterized by co-occurring conditions.
Moumita Bhattacharya, Claudine Jurkovitz, Hagit Shatkay
BIBM1
2014 Identifying growth-patterns in children by applying cluster analysis to electronic medical records
abstract
Obesity is one of the leading health concerns in the United States. Researchers and health care providers are interested in understanding factors affecting obesity and detecting the likelihood of obesity as early as possible. In this paper, we set out to recognize children who have higher risk of obesity by identifying distinct growth patterns in them. This is done by using clustering methods, which group together children who share similar body measurements over a period of time. The measurements characterizing children within the same cluster are plotted as a function of age. We refer to these plots as growth-pattern curves. We show that distinct growth-pattern curves are associated with different clusters and thus can be used to separate children into the topmost (heaviest), middle, or bottom-most cluster based on early growth measurements.
Moumita Bhattacharya, Deborah Ehrenthal, Hagit Shatkay
BIBM1