VLDB 2026 Research / reviewers in the wild / expert
Nikita Nangia
dblp:199/2397
· DBLP profile ↗
9ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Question answering and dialogue systems · 26% Language models and text generation · 23% Trustworthy machine learning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 7 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
natural language understanding |
0.9 | 3 | 2021 | SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems · NeurIPS 2019 Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark · ACL (1) 2019 What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? · ACL/IJCNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.6 | 1 | 2022 | What Makes Reading Comprehension Questions Difficult? · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems
question difficulty estimation |
0.6 | 1 | 2022 | What Makes Reading Comprehension Questions Difficult? · ACL (1) 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
information gathering |
0.5 | 1 | 2021 | What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? · ACL/IJCNLP (1) 2021 |
Machine learning › Trustworthy machine learning
fairness |
0.4 | 1 | 2020 | CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
social bias evaluation |
0.4 | 1 | 2020 | CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
masked language modeling |
0.1 | 1 | 2020 | CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
crowdsourcing · 1.9survey · 1.3benchmark construction · 0.8manual annotation · 0.6leaderboard evaluation · 0.4BERT fine-tuning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language ModelsabstractMargaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat |
NAACL (Long Papers) | 20 |
| 2023 | What Do NLP Researchers Believe? Results of the NLP Community MetasurveyabstractJulian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman |
ACL (1) | 8 |
| 2022 | What Makes Reading Comprehension Questions Difficult?abstractFor a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-future state-of-the-art systems.However, we do not yet know how best to select text sources to collect a variety of challenging examples.In this study, we crowdsource multiple-choice reading comprehension questions for passages taken from seven qualitatively distinct sources, analyzing what attributes of passages contribute to the difficulty and question types of the collected examples.To our surprise, we find that passage source, length, and readability measures do not significantly affect question difficulty.Through our manual annotation of seven reasoning types, we observe several trends between passage sources and reasoning types, e.g., logical reasoning is more often required in questions written for technical passages.These results suggest that when creating a new benchmark dataset, selecting a diverse set of passages can help ensure a diverse range of question types, but that passage difficulty need not be a priority. Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. Bowman |
ACL (1) | 2 |
| 2022 | QuALITY: Question Answering with Long Input Texts, Yes!abstractRichard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman |
NAACL-HLT | 4 |
| 2021 | What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?abstractNikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman |
ACL/IJCNLP (1) | 1 |
| 2020 | CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsabstractWarning: This paper contains explicit statements of offensive stereotypes and may be upsetting.Pretrained language models, especially masked language models (MLMs) have seen success across many NLP tasks.However, there is ample evidence that they use the cultural biases that are undoubtedly present in the corpora they are trained on, implicitly creating harm with biased representations.To measure some forms of social bias in language models against protected demographic groups in the US, we introduce the Crowdsourced Stereotype Pairs benchmark (CrowS-Pairs).CrowS-Pairs has 1508 examples that cover stereotypes dealing with nine types of bias, like race, religion, and age.In CrowS-Pairs a model is presented with two sentences: one that is more stereotyping and another that is less stereotyping.The data focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups.We find that all three of the widelyused MLMs we evaluate substantially favor sentences that express stereotypes in every category in CrowS-Pairs.As work on building less biased models advances, this dataset can be used as a benchmark to evaluate progress. Nikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. Bowman |
EMNLP (1) | 1 |
| 2019 | Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE BenchmarkabstractThe GLUE benchmark (Wang et al., 2019b) is a suite of language understanding tasks which has seen dramatic progress in the past year, with average performance moving from 70.0 at launch to 83.9, state of the art at the time of writing (May 24, 2019).Here, we measure human performance on the benchmark, in order to learn whether significant headroom remains for further progress.We provide a conservative estimate of human performance on the benchmark through crowdsourcing: Our annotators are non-experts who must learn each task from a brief set of instructions and 20 examples.In spite of limited training, these annotators robustly outperform the state of the art on six of the nine GLUE tasks and achieve an average score of 87.1.Given the fast pace of progress however, the headroom we observe is quite limited.To reproduce the datapoor setting that our annotators must learn in, we also train the BERT model (Devlin et al., 2019) in limited-data regimes, and conclude that low-resource sentence classification remains a challenge for modern neural network approaches to text understanding.How do you prepare for a job interview?How do I prepare for my first job interview? 1 Table 6: Another ten randomly sampled examples from QQP's development set.Pairs of sentences with a label of 1 are marked as paraphrases in QQP. Nikita Nangia, Samuel R. Bowman |
ACL (1) | 1 |
| 2019 | SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding SystemsabstractIn the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, offers a single-number metric that summarizes progress on a diverse set of such tasks, but performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research. In this paper we present SuperGLUE, a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, a software toolkit, and a public leaderboard. SuperGLUE is available at https://super.gluebenchmark.com. Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman |
NeurIPS | 3 |
| 2018 | A Broad-Coverage Challenge Corpus for Sentence Understanding through InferenceabstractAdina Williams, Nikita Nangia, Samuel Bowman. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Adina Williams, Nikita Nangia, Samuel R. Bowman |
NAACL-HLT | 2 |