VLDB 2026 Research / reviewers in the wild / expert
Tathagata Raha
dblp:271/4295
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Deep learning architectures and training · 64% Language models and text generation · 19% Trustworthy machine learning · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence-to-sequence generation |
1.0 | 1 | 2026 | Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation · ACL (1) 2026 |
Medical and health informatics › biomedical natural language processing
medical language model |
0.9 | 1 | 2025 | Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency · EMNLP 2025 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
factual consistency |
0.3 | 1 | 2026 | Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
fairness and bias |
0.3 | 1 | 2025 | Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
bias analysis · 1.7large language model · 1.0cross-examination · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text GenerationabstractTathagata Raha, Clement Christophe, Nada Saadi, Hamza A Javed, Marco AF Pimentel, Ronnie Rajan, Praveenkumar Kanithi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tathagata Raha, Clément Christophe, Nada Saadi, Hamza Javed, Marco A. F. Pimentel, Ronnie Rajan, Praveenkumar Kanithi |
ACL (1) | 1 |
| 2025 | Building Trust in Clinical LLMs: Bias Analysis and Dataset TransparencyabstractSvetlana Maslenkova, Clement Christophe, Marco AF Pimentel, Tathagata Raha, Muhammad Umar Salman, Ahmed Al Mahrooqi, Avani Gupta, Shadab Khan, Ronnie Rajan, Praveenkumar Kanithi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Svetlana Maslenkova, Clément Christophe, Marco A. F. Pimentel, Tathagata Raha, Muhammad Umar Salman, Ahmed Al-Mahrooqi, Avani Gupta, Shadab Khan, Ronnie Rajan, Praveen K. Kanithi |
EMNLP | 4 |
| 2023 | Neural Models for Factual Inconsistency Classification with ExplanationsabstractFactual consistency is one of the most important requirements when editing high quality documents. It is extremely important for automatic text generation systems like summarization, question answering, dialog modeling, and language modeling. Still, automated factual inconsistency detection is rather under-studied. Existing work has focused on (a) finding fake news keeping a knowledge base in context, or (b) detecting broad contradiction (as part of natural language inference literature). However, there has been no work on detecting and explaining types of factual inconsistencies in text, without any knowledge base in context. In this paper, we leverage existing work in linguistics to formally define five types of factual inconsistencies. Based on this categorization, we contribute a novel dataset, FICLE (Factual Inconsistency CLassification with Explanation), with $$\sim $$ 8K samples where each sample consists of two sentences (claim and context) annotated with type and span of inconsistency. When the inconsistency relates to an entity type, it is labeled as well at two levels (coarse and fine-grained). Further, we leverage this dataset to train a pipeline of four neural models to predict inconsistency type with explanations, given a (claim, context) sentence pair. Explanations include inconsistent claim fact triple, inconsistent context span, inconsistent claim component, coarse and fine-grained inconsistent entity types. The proposed system first predicts inconsistent spans from claim and context; and then uses them to predict inconsistency types and inconsistent entity types (when inconsistency is due to entities). We experiment with multiple Transformer-based natural language classification as well as generative models, and find that DeBERTa performs the best. Our proposed methods provide a weighted F1 of $$\sim $$ 87% for inconsistency type classification across the five classes. We make the code and dataset publicly available ( https://github.com/blitzprecision/FICLE ). Tathagata Raha, Mukund Choudhary, Abhinav Menon, KV Aditya Srivatsa, Manish Gupta 0001, Vasudeva Varma |
ECML/PKDD (3) | 1 |
| 2022 | Leveraging Mental Health Forums for User-level Depression Detection on Social MediaabstractThe number of depression and suicide risk cases on social media platforms is ever-increasing, and the lack of depression detection mechanisms on these platforms is becoming increasingly apparent. A majority of work in this area has focused on leveraging linguistic features while dealing with small-scale datasets. However, one faces many obstacles when factoring into account the vastness and inherent imbalance of social media content. In this paper, we aim to optimize the performance of user-level depression classification to lessen the burden on computational resources. The resulting system executes in a quicker, more efficient manner, in turn making it suitable for deployment. To simulate a platform agnostic framework, we simultaneously replicate the size and composition of social media to identify victims of depression. We systematically design a solution that categorizes post embeddings, obtained by fine-tuning transformer models such as RoBERTa, and derives user-level representations using hierarchical attention networks. We also introduce a novel mental health dataset to enhance the performance of depression categorization. We leverage accounts of depression taken from this dataset to infuse domain-specific elements into our framework. Our proposed methods outperform numerous baselines across standard metrics for the task of depression detection in text. Sravani Boinepelli, Tathagata Raha, Harika Abburi, Pulkit Parikh, Niyati Chhaya, Vasudeva Varma |
LREC | 2 |