EDBT 2026 Demo / reviewers in the wild / expert
Ananya Sai
dblp:222/7937 · also Ananya B. Sai
· DBLP profile ↗
7ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-1953-6214ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 38% Machine translation · 16% Question answering and dialogue systems · 13% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
machine translation evaluation |
0.7 | 1 | 2023 | IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian Languages · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis
multilingual NLP |
0.7 | 1 | 2023 | Bi-Phone: Modeling Inter Language Phonetic Influences in Text · ACL (1) 2023 |
Natural language and speech › Speech recognition and synthesis
phonetic modeling |
0.7 | 1 | 2023 | Bi-Phone: Modeling Inter Language Phonetic Influences in Text · ACL (1) 2023 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.5 | 1 | 2021 | Perturbation CheckLists for Evaluating NLG Evaluation Metrics · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue evaluation
dialogue response evaluation |
0.4 | 1 | 2019 | Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue Responses · AAAI 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2019 | Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue Responses · AAAI 2019 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.3 | 1 | 2018 | ElimiNet: A Model for Eliminating Options for Reading Comprehension with Multiple Choice Questions · IJCAI 2018 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
multiple-choice question answering |
0.3 | 1 | 2018 | ElimiNet: A Model for Eliminating Options for Reading Comprehension with Multiple Choice Questions · IJCAI 2018 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
neural question answering |
0.3 | 1 | 2018 | ElimiNet: A Model for Eliminating Options for Reading Comprehension with Multiple Choice Questions · IJCAI 2018 |
Methods — techniques the papers use, named apart from their topics
transliteration · 0.7phonetic influence modeling · 0.7human evaluation · 0.5checklists · 0.5linear system theory · 0.4adversarial attack · 0.4neural network · 0.3ensemble · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in LegislationabstractAtharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya B. Sai, John J Nay, Tanmay Rajpurohit, Ashwin Kalyan, Balaraman Ravindran. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Atharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya Sai, John J. Nay, Tanmay Rajpurohit, Ashwin Kalyan, Balaraman Ravindran |
ACL (1) | 4 |
| 2023 | IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesabstractAnanya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, Raj Dabre. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ananya Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Mitesh M. Khapra, Raj Dabre |
ACL (1) | 1 |
| 2023 | Bi-Phone: Modeling Inter Language Phonetic Influences in TextabstractAbhirut Gupta, Ananya B. Sai, Richard Sproat, Yuri Vasilevski, James Ren, Ambarish Jash, Sukhdeep Sodhi, Aravindan Raghuveer. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Abhirut Gupta, Ananya Sai, Richard Sproat, Yuri Vasilevski, James S. Ren, Ambarish Jash, Sukhdeep S. Sodhi, Aravindan Raghuveer |
ACL (1) | 2 |
| 2021 | Perturbation CheckLists for Evaluating NLG Evaluation MetricsabstractNatural Language Generation (NLG) evaluation is a multifaceted task requiring assessment of multiple desirable criteria, e.g., fluency, coherency, coverage, relevance, adequacy, overall quality, etc. Across existing datasets for 6 NLG tasks, we observe that the human evaluation scores on these multiple criteria are often not correlated.For example, there is a very low correlation between human scores on fluency and data coverage for the task of structured data to text generation.This suggests that the current recipe of proposing new automatic evaluation metrics for NLG by showing that they correlate well with scores assigned by humans for a single criteria (overall quality) alone is inadequate.Indeed, our extensive study involving 25 automatic evaluation metrics across 6 different tasks and 18 different evaluation criteria shows that there is no single metric which correlates well with human scores on all desirable criteria, for most NLG tasks.Given this situation, we propose CheckLists for better design and evaluation of automatic metrics.We design templates which target a specific criteria (e.g., coverage) and perturb the output such that the quality gets affected only along this specific criteria (e.g., the coverage drops).We show that existing evaluation metrics are not robust against even such simple perturbations and disagree with scores assigned by humans to the perturbed output.The proposed templates thus allow for a fine-grained assessment of automatic evaluation metrics exposing their limitations and will facilitate better design, analysis and evaluation of such metrics. 1 Ananya Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, Mitesh M. Khapra |
EMNLP (1) | 1 |
| 2020 | Improving Dialog Evaluation with a Multi-reference Adversarial Dataset and Large Scale PretrainingabstractThere is an increasing focus on model-based dialog evaluation metrics such as ADEM, RUBER, and the more recent BERT-based metrics. These models aim to assign a high score to all relevant responses and a low score to all irrelevant responses. Ideally, such models should be trained using multiple relevant and irrelevant responses for any given context. However, no such data is publicly available, and hence existing models are usually trained using a single relevant response and multiple randomly selected responses from other contexts (random negatives). To allow for better training and robust evaluation of model-based metrics, we introduce the DailyDialog++ dataset, consisting of (i) five relevant responses for each context and (ii) five adversarially crafted irrelevant responses for each context. Using this dataset, we first show that even in the presence of multiple correct references, n-gram based metrics and embedding based metrics do not perform well at separating relevant responses from even random negatives. While model-based metrics perform better than n-gram and embedding based metrics on random negatives, their performance drops substantially when evaluated on adversarial examples. To check if large scale pretraining could help, we propose a new BERT-based evaluation metric called DEB, which is pretrained on 727M Reddit conversations and then finetuned on our dataset. DEB significantly outperforms existing models, showing better correlation with human judgments and better performance on random negatives (88.27% accuracy). However, its performance again drops substantially when evaluated on adversarial responses, thereby highlighting that even large-scale pretrained evaluation models are not robust to the adversarial examples in our dataset. The dataset 1 and code 2 are publicly available. Ananya Sai, Akash Kumar Mohankumar, Siddhartha Arora, Mitesh M. Khapra |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue ResponsesabstractAutomatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM (Lowe et al. 2017) formulated the automatic evaluation of dialogue systems as a learning problem and showed that such a model was able to predict responses which correlate significantly with human judgements, both at utterance and system level. Their system was shown to have beaten word-overlap metrics such as BLEU with large margins. We start with the question of whether an adversary can game the ADEM model. We design a battery of targeted attacks at the neural network based ADEM evaluation system and show that automatic evaluation of dialogue systems still has a long way to go. ADEM can get confused with a variation as simple as reversing the word order in the text! We report experiments on several such adversarial scenarios that draw out counterintuitive scores on the dialogue responses. We take a systematic look at the scoring function proposed by ADEM and connect it to linear system theory to predict the shortcomings evident in the system. We also devise an attack that can fool such a system to rate a response generation system as favorable. Finally, we allude to future research directions of using the adversarial attacks to design a truly automated dialogue evaluation system. Ananya Sai, Mithun Das Gupta, Mitesh M. Khapra, Mukundhan Srinivasan |
AAAI | 1 |
| 2018 | ElimiNet: A Model for Eliminating Options for Reading Comprehension with Multiple Choice QuestionsabstractThe task of Reading Comprehension with Multiple Choice Questions, requires a human (or machine) to read a given {passage, question} pair and select one of the n given options. The current state of the art model for this task first computes a question-aware representation for the passage and then selects the option which has the maximum similarity with this representation. However, when humans perform this task they do not just focus on option selection but use a combination of elimination and selection. Specifically, a human would first try to eliminate the most irrelevant option and then read the passage again in the light of this new information (and perhaps ignore portions corresponding to the eliminated option). This process could be repeated multiple times till the reader is finally ready to select the correct option. We propose ElimiNet, a neural network-based model which tries to mimic this process. Specifically, it has gates which decide whether an option can be eliminated given the {passage, question} pair and if so it tries to make the passage representation orthogonal to this eliminated option (akin to ignoring portions of the passage corresponding to the eliminated option). The model makes multiple rounds of partial elimination to refine the passage representation and finally uses a selection module to pick the best option. We evaluate our model on the recently released large scale RACE dataset and show that it outperforms the current state of the art model on 7 out of the 13 question types in this dataset. Further, we show that taking an ensemble of our elimination-selection based method with a selection based method gives us an improvement of 3.1% over the best-reported performance on this dataset. Soham Parikh, Ananya Sai, Preksha Nema, Mitesh M. Khapra |
IJCAI | 2 |