Alicia Parrish

dblp:248/7544 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-1054-0516ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 44% Language models and text generation · 32% Generative modeling · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
safety evaluation
1.522025
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models · NeurIPS 2025
DICES Dataset: Diversity in Conversational AI Evaluation for Safety · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk · NeurIPS 2025
Natural language and speech › Language models and text generation › alignment
pluralistic alignment
0.912025
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems
conversational agents
0.712023
DICES Dataset: Diversity in Conversational AI Evaluation for Safety · NeurIPS 2023
Machine learning › Trustworthy machine learning
Data-centric AI
0.712023
DataPerf: Benchmarks for Data-Centric AI Development · NeurIPS 2023
Machine learning › Trustworthy machine learning
fairness
0.712023
DICES Dataset: Diversity in Conversational AI Evaluation for Safety · NeurIPS 2023
Natural language and speech › Language models and text generation
linguistic generalization
0.412019
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › pre-trained language model › knowledge probing
linguistic knowledge probing
0.412019
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation
pre-trained language model
0.412019
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

survey · 1.3risk management process · 0.9metaevaluation scoring · 0.9human evaluation · 0.9LLM judgment · 0.9dataset construction · 0.7benchmarking · 0.7analysis methods · 0.4
YearPublicationVenuePosition
2025 Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
abstract
Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes impacting benchmark bias, variance, coverage, or people's capacity to understand benchmark evidence. Using the National Institute of Standards and Technology's risk management process as a foundation, this research iteratively analyzed 26 popular benchmarks, identifying 57 potential failure modes and 196 corresponding mitigation strategies. The mitigations reduce failure likelihood and/or severity, providing a frame for evaluating "benchmark risk," which is scored to provide a metaevaluation benchmark: BenchRisk. Higher scores indicate benchmark users are less likely to reach an incorrect or unsupported conclusion about an LLM. All 26 scored benchmarks present significant risk within one or more of the five scored dimensions (comprehensiveness, intelligibility, consistency, correctness, and longevity), which points to important open research directions for the field of LLM benchmarking. The BenchRisk workflow allows for comparison between benchmarks; as an open-source tool, it also facilitates the identification and sharing of risks and their mitigations.
Sean McGregor, Vassil Tashev, Armstrong Foundjem, Aishwarya Ramasethu, Sadegh AlMahdi Kazemi Zarkouei, Chris Knotz, Kongtao Chen, Alicia Parrish, Anka Reuel, Heather Frase
NeurIPS8
2025 Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
abstract
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems.Content Warning: The paper includes sensitive content that may be harmful.
Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang 0006, Mark Diaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo
NeurIPS7
2024 GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
abstract
Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex Taylor, Mark Diaz, Ding Wang, Gregory Serapio-García. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex S. Taylor, Mark Diaz, Ding Wang 0006, Gregory Serapio-García
NAACL-HLT5
2023 What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
abstract
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
ACL (1)3
2023 DICES Dataset: Diversity in Conversational AI Evaluation for Safety
abstract
Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This requirement overly simplifies the natural subjectivity present in many tasks, and obscures the inherent diversity in human perceptions and opinions about many content items. Preserving the variance in content and diversity in human perceptions in datasets is often quite expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is socio-culturally situated in this context. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographics information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. The DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of safety for conversational AI. We further describe a set of metrics that show how rater diversity influences safety perception across different geographic regions, ethnicity groups, age groups, and genders. The goal of the DICES dataset is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems.
Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher Homan, Alicia Parrish, Gregory Serapio-García, Vinodkumar Prabhakaran, Ding Wang 0006
NeurIPS5
2023 DataPerf: Benchmarks for Data-Centric AI Development
abstract
Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry.
Mark Mazumder, Colby R. Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Frederick Diamos, Gregory Frederick Diamos, Lynn He, Alicia Parrish, Hannah Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett D. Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller 0001, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung-Levy, Newsha Ardalani, Praveen K. Paritosh, Ce Zhang 0001, James Zou 0001, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi
NeurIPS9
2022 QuALITY: Question Answering with Long Input Texts, Yes!
abstract
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman
NAACL-HLT2
2021 NOPE: A Corpus of Naturally-Occurring Presuppositions in English
abstract
Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021.
Alicia Parrish, Sebastian Schuster 0001, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen
CoNLL1
2020 BLiMP: The Benchmark of Linguistic Minimal Pairs for English
abstract
We introduce The Benchmark of Linguistic Minimal Pairs (BLiMP),1 a challenge set for evaluating the linguistic knowledge of language models (LMs) on major grammatical phenomena in English. BLiMP consists of 67 individual datasets, each containing 1,000 minimal pairs—that is, pairs of minimally different sentences that contrast in grammatical acceptability and isolate specific phenomenon in syntax, morphology, or semantics. We generate the data according to linguist-crafted grammar templates, and human aggregate agreement with the labels is 96.4%. We evaluate n-gram, LSTM, and Transformer (GPT-2 and Transformer-XL) LMs by observing whether they assign a higher probability to the acceptable sentence in each minimal pair. We find that state-of-the-art models identify morphological contrasts related to agreement reliably, but they struggle with some subtle semantic and syntactic phenomena, such as negative polarity items and extraction islands.
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics2
2020 Erratum: "BLiMP: The Benchmark of Linguistic Minimal Pairs for English"
abstract
We correct wrongly reported results on BLiMP.
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng 0013, Sheng-Fu Wang, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics2
2019 Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs
abstract
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alex Warstadt, Ioana Grosu, Wei Peng 0013, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, Samuel R. Bowman
EMNLP/IJCNLP (1)10