Nikita Nangia

dblp:199/2397 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Question answering and dialogue systems · 26% Language models and text generation · 23% Trustworthy machine learning · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
natural language understanding
0.932021
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems · NeurIPS 2019
Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark · ACL (1) 2019
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.612022
What Makes Reading Comprehension Questions Difficult? · ACL (1) 2022
Natural language and speech › Question answering and dialogue systems
question difficulty estimation
0.612022
What Makes Reading Comprehension Questions Difficult? · ACL (1) 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
information gathering
0.512021
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? · ACL/IJCNLP (1) 2021
Machine learning › Trustworthy machine learning
fairness
0.412020
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models · EMNLP (1) 2020
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
social bias evaluation
0.412020
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models · EMNLP (1) 2020
Natural language and speech › Language models and text generation
masked language modeling
0.112020
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

crowdsourcing · 1.9survey · 1.3benchmark construction · 0.8manual annotation · 0.6leaderboard evaluation · 0.4BERT fine-tuning · 0.4
YearPublicationVenuePosition
2025 SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models
abstract
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
NAACL (Long Papers)20
2023 What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
abstract
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
ACL (1)8
2022 What Makes Reading Comprehension Questions Difficult?
abstract
For a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-future state-of-the-art systems.However, we do not yet know how best to select text sources to collect a variety of challenging examples.In this study, we crowdsource multiple-choice reading comprehension questions for passages taken from seven qualitatively distinct sources, analyzing what attributes of passages contribute to the difficulty and question types of the collected examples.To our surprise, we find that passage source, length, and readability measures do not significantly affect question difficulty.Through our manual annotation of seven reasoning types, we observe several trends between passage sources and reasoning types, e.g., logical reasoning is more often required in questions written for technical passages.These results suggest that when creating a new benchmark dataset, selecting a diverse set of passages can help ensure a diverse range of question types, but that passage difficulty need not be a priority.
Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. Bowman
ACL (1)2
2022 QuALITY: Question Answering with Long Input Texts, Yes!
abstract
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman
NAACL-HLT4
2021 What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
abstract
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman
ACL/IJCNLP (1)1
2020 CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
abstract
Warning: This paper contains explicit statements of offensive stereotypes and may be upsetting.Pretrained language models, especially masked language models (MLMs) have seen success across many NLP tasks.However, there is ample evidence that they use the cultural biases that are undoubtedly present in the corpora they are trained on, implicitly creating harm with biased representations.To measure some forms of social bias in language models against protected demographic groups in the US, we introduce the Crowdsourced Stereotype Pairs benchmark (CrowS-Pairs).CrowS-Pairs has 1508 examples that cover stereotypes dealing with nine types of bias, like race, religion, and age.In CrowS-Pairs a model is presented with two sentences: one that is more stereotyping and another that is less stereotyping.The data focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups.We find that all three of the widelyused MLMs we evaluate substantially favor sentences that express stereotypes in every category in CrowS-Pairs.As work on building less biased models advances, this dataset can be used as a benchmark to evaluate progress.
Nikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. Bowman
EMNLP (1)1
2019 Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark
abstract
The GLUE benchmark (Wang et al., 2019b) is a suite of language understanding tasks which has seen dramatic progress in the past year, with average performance moving from 70.0 at launch to 83.9, state of the art at the time of writing (May 24, 2019).Here, we measure human performance on the benchmark, in order to learn whether significant headroom remains for further progress.We provide a conservative estimate of human performance on the benchmark through crowdsourcing: Our annotators are non-experts who must learn each task from a brief set of instructions and 20 examples.In spite of limited training, these annotators robustly outperform the state of the art on six of the nine GLUE tasks and achieve an average score of 87.1.Given the fast pace of progress however, the headroom we observe is quite limited.To reproduce the datapoor setting that our annotators must learn in, we also train the BERT model (Devlin et al., 2019) in limited-data regimes, and conclude that low-resource sentence classification remains a challenge for modern neural network approaches to text understanding.How do you prepare for a job interview?How do I prepare for my first job interview? 1 Table 6: Another ten randomly sampled examples from QQP's development set.Pairs of sentences with a label of 1 are marked as paraphrases in QQP.
Nikita Nangia, Samuel R. Bowman
ACL (1)1
2019 SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
abstract
In the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, offers a single-number metric that summarizes progress on a diverse set of such tasks, but performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research. In this paper we present SuperGLUE, a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, a software toolkit, and a public leaderboard. SuperGLUE is available at https://super.gluebenchmark.com.
Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman
NeurIPS3
2018 A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
abstract
Adina Williams, Nikita Nangia, Samuel Bowman. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Adina Williams, Nikita Nangia, Samuel R. Bowman
NAACL-HLT2