VLDB 2026 Research / reviewers in the wild / expert
Aida Mostafazadeh Davani
dblp:248/7592
· DBLP profile ↗
13ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-7013-1810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Moral Foundations Reddit CorpusabstractMoral framing and sentiment can affect a variety of online and offline behaviors, including donation, environmental action, political engagement, and protest. Various computational methods in Natural Language Processing (NLP) have been used to detect moral sentiment from textual data, but achieving strong performance in such subjective tasks requires large, hand-annotated datasets. Previous corpora annotated for moral sentiment have proven valuable, and have generated new insights both within NLP and across the social sciences, but have been limited to Twitter. To facilitate improving our understanding of the role of moral rhetoric, we present the Moral Foundations Reddit Corpus, a collection of 16,123 English Reddit comments that have been curated from 12 distinct subreddits, hand-annotated by at least three trained annotators for 8 categories of moral sentiment (i.e., Care, Proportionality, Equality, Purity, Authority, Loyalty, Thin Morality, Implicit/Explicit Morality) based on the updated Moral Foundations Theory (MFT) framework. We evaluate baselines using large language models (Llama3-8B, Ministral-8B) in zero-shot, few-shot, and PEFT (Parameter-Efficient Fine-Tuning) settings, comparing their performance to fine-tuned encoder-only models like BERT (Bidirectional Encoder Representations from Transformers). The results show that LLMs continue to lag behind fine-tuned encoders on this subjective task, underscoring the ongoing need for human-annotated moral corpora for AI alignment evaluation. Keywords: moral sentiment annotation, moral values, moral foundations theory, multi-label text classification, large language models, benchmark dataset, evaluation and alignment resource Jackson Trager, Alireza S. Ziabari, Elnaz Rahmati, Aida Mostafazadeh Davani, Preni Golazizian, Farzan Karimi-Malekabadi, Ali Omrani, Zhihe Li, Brendan Kennedy 0001, Georgios Chochlakis, Nils Karl Reimer, Melissa Reyes, Kelsey Cheng, Mellow Wei, Christina Merrifield, Arta Khosravi, Evans Alvarez, Morteza Dehghani |
LREC | 4 |
| 2025 | A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI EvaluationsabstractSocietal stereotypes are at the center of a myriad of responsible AI interventions targeted at reducing the generation and propagation of potentially harmful outcomes.While these efforts are much needed, they tend to be fragmented and often address different parts of the issue without adopting a unified or holistic approach to social stereotypes and how they impact various parts of the machine learning pipeline.As a result, current interventions fail to capitalize on the underlying mechanisms that are common across different types of stereotypes, and to anchor on particular aspects that are relevant in certain cases.In this paper, we draw on social psychological research and build on NLP data and methods, to propose a unified framework to operationalize stereotypes in generative AI evaluations.Our framework identifies key components of stereotypes that are crucial in AI evaluation, including the target group, associated attribute, relationship characteristics, perceiving group, and context.We also provide considerations and recommendations for its responsible use.CONTENT WARNING: This paper contains examples of stereotypes that may be offensive. Aida Mostafazadeh Davani, Sunipa Dev, Héctor Pérez-Urbina, Vinodkumar Prabhakaran |
EMNLP | 1 |
| 2025 | Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image ModelsabstractCurrent text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems.Content Warning: The paper includes sensitive content that may be harmful. Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang 0006, Mark Diaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo |
NeurIPS | 8 |
| 2024 | D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and EvaluationabstractWhile human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection.Recent studies that critically examine this issue are often focused on Western contexts, and solely document differences across age, gender, or racial groups.Consequently, NLP research on subjectivity have failed to consider that individuals within demographic groups may hold diverse values, which influence their perceptions beyond group norms.To effectively incorporate these considerations into NLP pipelines, we need datasets with extensive parallel annotations from a variety of social and cultural groups.In this paper we introduce the D3CODE dataset: a large-scale cross-cultural dataset of parallel annotations for offensive language in over 4.5K English sentences annotated by a pool of more than 4k annotators, balanced across gender and age, from across 21 countries, representing eight geo-cultural regions.The dataset captures annotators' moral values along six moral foundations: care, equality, proportionality, authority, loyalty, and purity.Our analyses reveal substantial regional variations in annotators' perceptions that are shaped by individual moral values, providing crucial insights for developing pluralistic, culturally sensitive NLP models. Aida Mostafazadeh Davani, Mark Diaz, Dylan K. Baker, Vinodkumar Prabhakaran |
EMNLP | 1 |
| 2024 | GRASP: A Disagreement Analysis Framework to Assess Group Associations in PerspectivesabstractVinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex Taylor, Mark Diaz, Ding Wang, Gregory Serapio-García. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex S. Taylor, Mark Diaz, Ding Wang 0006, Gregory Serapio-García |
NAACL-HLT | 4 |
| 2023 | SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative ModelsabstractAkshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, Sunipa Dev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Akshita Jha, Aida Mostafazadeh Davani, Chandan K. Reddy, Shachi Dave, Vinodkumar Prabhakaran, Sunipa Dev |
ACL (1) | 2 |
| 2023 | Hate Speech Classifiers Learn Normative Social StereotypesabstractAbstract Social stereotypes negatively impact individuals’ judgments about different groups and may have a critical role in understanding language directed toward marginalized groups. Here, we assess the role of social stereotypes in the automated detection of hate speech in the English language by examining the impact of social stereotypes on annotation behaviors, annotated datasets, and hate speech classifiers. Specifically, we first investigate the impact of novice annotators’ stereotypes on their hate-speech-annotation behavior. Then, we examine the effect of normative stereotypes in language on the aggregated annotators’ judgments in a large annotated corpus. Finally, we demonstrate how normative stereotypes embedded in language resources are associated with systematic prediction errors in a hate-speech classifier. The results demonstrate that hate-speech classifiers reflect social stereotypes against marginalized groups, which can perpetuate social inequalities when propagated at scale. This framework, combining social-psychological and computational-linguistic methods, provides insights into sources of bias in hate-speech moderation, informing ongoing debates regarding machine learning fairness. Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy 0001, Morteza Dehghani |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective AnnotationsabstractAbstract Majority voting and averaging are common approaches used to resolve annotator disagreements and derive single ground truth labels from multiple annotations. However, annotators may systematically disagree with one another, often reflecting their individual biases and values, especially in the case of subjective tasks such as detecting affect, aggression, and hate speech. Annotator disagreements may capture important nuances in such tasks that are often ignored while aggregating annotations to a single ground truth. In order to address this, we investigate the efficacy of multi-annotator models. In particular, our multi-task based approach treats predicting each annotators’ judgements as separate subtasks, while sharing a common learned representation of the task. We show that this approach yields same or better performance than aggregating labels in the data prior to training across seven different binary classification tasks. Our approach also provides a way to estimate uncertainty in predictions, which we demonstrate better correlate with annotation disagreements than traditional methods. Being able to model uncertainty is especially useful in deployment scenarios where knowing when not to make a prediction is important. Aida Mostafazadeh Davani, Mark Diaz, Vinodkumar Prabhakaran |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | On Transferability of Bias Mitigation Effects in Language Model Fine-TuningabstractXisen Jin, Francesco Barbieri, Brendan Kennedy, Aida Mostafazadeh Davani, Leonardo Neves, Xiang Ren. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xisen Jin, Francesco Barbieri, Brendan Kennedy 0001, Aida Mostafazadeh Davani, Leonardo Neves, Xiang Ren 0001 |
NAACL-HLT | 4 |
| 2020 | Contextualizing Hate Speech Classifiers with Post-hoc ExplanationabstractHate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways.Such biases manifest in false positives when these identifiers are present, due to models' inability to learn the contexts which constitute a hateful usage of identifiers.We extract post-hoc explanations from fine-tuned BERT classifiers to detect bias towards identity terms.Then, we propose a novel regularization technique based on these explanations that encourages models to learn from the context of group identifiers in addition to the identifiers themselves.Our approach improved over baselines in limiting false positives on out-of-domain data while maintaining or improving in-domain performance. Brendan Kennedy 0001, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, Xiang Ren 0001 |
ACL | 3 |
| 2020 | Hatred is in the Eye of the Annotator: Hate Speech Classifiers Learn Human-Like Social Stereotypes
Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy 0001, Shreya Havaldar, Morteza Dehghani |
CogSci | 1 |
| 2019 | Subtle differences in language experience moderate performance on language-based cognitive tests
Maury Courtland, Aida Mostafazadeh Davani, Melissa Reyes, Leigh Yeh, Jun Yen Leung, Brendan Kennedy 0001, Morteza Dehghani, Jason D. Zevin |
CogSci | 2 |
| 2019 | Reporting the Unreported: Event Extraction for Analyzing the Local Representation of Hate CrimesabstractAida Mostafazadeh Davani, Leigh Yeh, Mohammad Atari, Brendan Kennedy, Gwenyth Portillo Wightman, Elaine Gonzalez, Natalie Delong, Rhea Bhatia, Arineh Mirinjian, Xiang Ren, Morteza Dehghani. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Aida Mostafazadeh Davani, Leigh Yeh, Mohammad Atari, Brendan Kennedy 0001, Gwenyth Portillo-Wightman, Elaine Gonzalez, Natalie Delong, Rhea Bhatia, Arineh Mirinjian, Xiang Ren 0001, Morteza Dehghani |
EMNLP/IJCNLP (1) | 1 |