VLDB 2026 Research / reviewers in the wild / expert
Ritwik Banerjee
dblp:117/4012
· DBLP profile ↗
13ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Local Privacy Laws in a Globalized WorldabstractPersonal data has emerged as a highly valuable yet sensitive asset that drives business decisions, enables targeted advertising, and generates substantial revenue for companies, while simultaneously facilitating invasive monitoring of users. In recent years, research on digital privacy violations, including undue access, collection, and sharing of user data, has grown significantly. Much of this research adopts the European General Data Protection Regulation (GDPR) as the primary reference framework. This is reasonable, as GDPR was a pioneering legislation, and many of its stipulations are clear and unambiguous. However, we argue that focusing solely on GDPR (and a small set of other Western regulatory frameworks) ignores privacy-related concerns, attitudes, and problems faced by users from other locales, creating a significant research blind spot. Shantanu Sharma 0001, Ethan Myers, Lorenzo De Carli, Ritwik Banerjee, Indrakshi Ray |
CODASPY | 4 |
| 2026 | A Longitudinal, Multinational, and Multilingual Corpus of News Coverage of the Russo-Ukrainian War
Dikshya Mohanty, Taisiia Sabadyn, Jelwin Rodrigues, Chenlu Wang, Abhishek Kalugade, Ritwik Banerjee |
LREC | 6 |
| 2025 | Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding TasksabstractDetecting deviant language such as sexism, or nuanced language such as metaphors or sarcasm, is crucial for enhancing the safety, clarity, and interpretation of social interactions. While existing classifiers deliver strong results on these tasks, they often come with significant computational cost and high data demands. In this work, we propose Class Distillation (ClaD), a novel training paradigm that targets the core challenge: distilling a small, well-defined target class from a highly diverse and heterogeneous background. ClaD integrates two key innovations: (i) a loss function informed by the structural properties of class distributions, based on Mahalanobis distance, and (ii) an interpretable decision algorithm optimized for class separation. Across three benchmark detection tasks – sexism, metaphor, and sarcasm – ClaD outperforms competitive baselines, and even with smaller language models and orders of magnitude fewer parameters, achieves performance comparable to several large language models. These results demonstrate ClaD as an efficient tool for pragmatic language understanding tasks that require gleaning a small target class from a larger heterogeneous background. Chenlu Wang, Weimin Lyu, Ritwik Banerjee |
ACL (1) | 3 |
| 2025 | HDCR: Cross-Lingual Medical Misinformation Detection Through Contrastive Claim-Evidence ReasoningabstractThe rapid dissemination of health information online enables dangerous distortions that threaten public health. We present$\text{HD}^{2} \text{CR}$(health-information distortion detection with contrastive reasoning), a framework for fine-grained detection of medical misinformation. Through systematic analysis of health news patterns, we identify four primary distortion types: over-generalization, exaggeration, under-generalization, and false causality. Our contributions include: a cross-lingual corpus of 72,275 English and Chinese claim-evidence pairs with validated distortion labels; a dual-encoder architecture with contrastive cross-attention that explicitly models semantic divergence between claims and biomedical evidence; and extensive evaluations demonstrating$\text{HD}^{2}{ }^{2} \text{CR}^{\prime}$s superior performance: 93.1% binary$F_{1}$and$87.3 \% 5$-class accuracy, with robust cross-lingual generalization (only 2.9% degradation between Chinese and English). Chaoyuan Zuo, Ritwik Banerjee |
BIBM | 2 |
| 2025 | Large-Scale Biomedical Expert Finding for Health Claim Verification: A PubMed-based Retrieval FrameworkabstractVerifying health claims amid rampant misinformation requires identifying qualified experts - a manual process that cannot scale. We address this challenge with a computational framework that automatically identifies biomedical researchers to evaluate health claims by analyzing their PubMed publication profiles. We establish the first benchmark for this cross-genre retrieval task, linking 93,404 health claims to 153,147 biomedical experts. Our two-stage neural pipeline addresses the semantic heterogeneity between informal health claims and formal research literature. Systematic evaluation reveals a striking finding: domain-specific models achieve 84.2% Mean Reciprocal Rank, substantially outperforming general-purpose alternatives, including state-of-the-art LLM-based rerankers fine-tuned on domain-specific data. Our findings underscore the necessity of specialized benchmarks for cross-genre information retrieval, and specialized pretraining for biomedical expert identification, while our scalable architecture enables rapid, automated expert matching for evidence-based claim verification in clinical and public health contexts. Chaoyuan Zuo, Chenlu Wang, Ritwik Banerjee |
BIBM | 3 |
| 2025 | Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case StudiesabstractState-of-the-art automatic speech recognition (ASR) models like Whisper, perform poorly on atypical speech, such as that produced by individuals with dysarthria.Past works for atypical speech have mostly investigated fully personalized (or idiosyncratic) models, but modeling strategies that can both generalize and handle idiosyncracy could be more effective for capturing atypical speech.To investigate this, we compare four strategies: (a) normative models trained on typical speech (no personalization), (b) idiosyncratic models completely personalized to individuals, (c) dysarthric-normative models trained on other dysarthric speakers, and (d) dysarthric-idiosyncratic models which combine strategies by first modeling normative patterns before adapting to individual speech.In this case study, we find the dysarthricidiosyncratic model performs better than idiosyncratic approach while requiring less than half as much personalized data (36.43WER with 128 train size vs 36.99 with 256).Further, we found that tuning the speech encoder alone (as opposed to the LM decoder) yielded the best results reducing word error rate from 71% to 32% on average.Our findings highlight the value of leveraging both normative (crossspeaker) and idiosyncratic (speaker-specific) patterns to improve ASR for underrepresented speech populations. 1 * Equal contribution 1 Github: VishnuRaja98/Dysarthric-Speech-Transcription Vishnu Raja, Adithya V. Ganesan, Anand Syamkumar, Ritwik Banerjee, H. Andrew Schwartz |
EMNLP | 4 |
| 2025 | Harnessing Language Models to Analyze Android App Permission FidelityabstractAndroid’s vast app ecosystem (over 2 million apps) poses significant privacy risks, as current methods for inferring permissions from descriptions - keyword matching, traditional natural language processing (NLP), and recurrent neural networks (RNNs) - struggle with accurate inference due to imprecise, ambiguous, or incomplete natural language descriptions. This gap undermines regulatory transparency and user trust, necessitating tools that reconcile stated functionality with actual data practices. We demonstrate that large language models like GPT-4o, applied in a zero-shot inference setting, leverage contextual reasoning to infer permissions competitively, while fine-tuned encoders (BERT, BART) surpass state-of-the-art performance when trained on minimally annotated datasets augmented with paraphrases, achieving $50-70 \%$ gains in weighted and macro $F_{1}$ scores. By enabling precise permission auditing with reduced annotation costs, our work advances scalable, adaptable solutions for privacy compliance across resource-constrained and highstakes environments. Yunik Tamrakar, Ritwik Banerjee, Ethan Myers, Lorenzo De Carli, Indrakshi Ray |
PST | 2 |
| 2024 | From Claim to Evidence: Verifying Chinese Health Claims with Medical Literature
Chaoyuan Zuo, Chenlu Wang, Ritwik Banerjee |
NLPCC (4) | 4 |
| 2023 | Cross-Genre Retrieval for Information Integrity: A COVID-19 Case Study
Chaoyuan Zuo, Chenlu Wang, Ritwik Banerjee |
ADMA (5) | 3 |
| 2020 | Querying Across Genres for Medical Claims in NewsabstractWe present a query-based biomedical information retrieval task across two vastly different genres -newswire and research literaturewhere the goal is to find the research publication that supports the primary claim made in a health-related news article.For this task, we present a new dataset of 5,034 claims from news paired with research abstracts.Our approach consists of two steps: (i) selecting the most relevant candidates from a collection of 222k research abstracts, and (ii) re-ranking this list.We compare the classical IR approach using BM25 with more recent transformerbased models.Our results show that crossgenre medical IR is a viable task, but incorporating domain-specific knowledge is crucial. Chaoyuan Zuo, Narayan Acharya, Ritwik Banerjee |
EMNLP (1) | 3 |
| 2015 | Internet Outages, the Eyewitness Accounts: Analysis of the Outages Mailing List
Ritwik Banerjee, Abbas Razaghpanah, Luis Chiang, Akassh Mishra, Vyas Sekar, Yejin Choi 0001, Phillipa Gill |
PAM | 1 |
| 2014 | Keystroke Patterns as Prosody in Digital Writings: A Case Study with Deceptive Reviews and EssaysabstractIn this paper, we explore the use of keyboard strokes as a means to access the real-time writ-ing process of online authors, analogously to prosody in speech analysis, in the context of deception detection. We show that differences in keystroke patterns like editing maneuvers and duration of pauses can help distinguish be-tween truthful and deceptive writing. Empiri-cal results show that incorporating keystroke-based features lead to improved performance in deception detection in two different do-mains: online reviews and essays. 1 Ritwik Banerjee, Song Feng 0002, Yejin Choi 0001 |
EMNLP | 1 |
| 2012 | Characterizing Stylistic Elements in Syntactic Structure
Song Feng 0002, Ritwik Banerjee, Yejin Choi 0001 |
EMNLP-CoNLL | 2 |