VLDB 2026 Research / reviewers in the wild / expert
Mark Diaz
dblp:194/9668 · also Mark Díaz
· DBLP profile ↗
17ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-0167-9839ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMsabstractLLMs becoming increasingly personalized to users’ language style raises both excitement and concerns for minority users such as Black American English (BAE) speakers. Yet, previous work has predominantly focused on user perceptions of out-of-context BAE statements by LLMs rather than naturalistic multi-turn interactions, and has ignored such systems’ effects on users’ self-perception. In this work, we examine the effects that multi-turn interactions with speech and text BAE-producing LLMs have on BAE speakers’ perceptions of the LLM and of themselves. We observe a significant change in participant self-esteem following the interactions, and notable qualitative differences between BAE-LLM and Standard American English (SAE) LLM interactions. We also observe significant effects of BAE-usage on user perception of the model within speech-based interactions. Our findings suggest that the effects of BAE-usage by an LLM agent on model- and self-perception among BAE-speaking users are complex and widely varied. Mikayla Campbell, Joel Mire, Mark Diaz, Maarten Sap |
CHI | 3 |
| 2026 | How Tech Workers Contend with Hazards of Humanlikeness in Generative AIabstractGenerative AI’s humanlike qualities are driving its rapid adoption in professional domains. However, this anthropomorphic appeal raises concerns from HCI and responsible AI scholars about potential hazards and harms, such as overtrust in system outputs. To investigate how technology workers navigate these humanlike qualities and anticipate emergent harms, we conducted focus groups with 30 professionals across six job functions (ML engineering, product policy, UX research and design, product management, technology writing, and communications). Our findings reveal an unsettled knowledge environment surrounding humanlike generative AI, where workers’ varying perspectives illuminate a range of potential risks for individuals, knowledge work fields, and society. We argue that workers require comprehensive support, including clearer conceptions of “humanlikeness” to effectively mitigate these risks. To aid in mitigation strategies, we provide a conceptual map articulating the identified hazards and their connection to conflated notions of “humanlikeness.” Mark Diaz, Renee Shelby, Eric Corbett, Andrew Smart |
CHI | 1 |
| 2025 | Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image ModelsabstractCurrent text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems.Content Warning: The paper includes sensitive content that may be harmful. Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang 0006, Mark Diaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo |
NeurIPS | 6 |
| 2024 | SoUnD Framework: Analyzing (So)cial Representation in (Un)structured (D)ataabstractDecisions about how to responsibly collect, use and document data often rely upon understanding how people are represented in data. Yet, the unlabeled nature and scale of data used in foundation model development poses a direct challenge to systematic analyses of downstream risks, such as representational harms. We provide a framework designed to help RAI practitioners more easily plan and structure analyses of how people are represented in unstructured data and identify downstream risks. The framework is organized into groups of analyses that map to 3 basic questions: 1) Who is represented in the data, 2) What content is in the data, and 3) How are the two associated. We use the framework to analyze human representation in two commonly used datasets: the Common Crawl web corpus (C4) of 356 billion tokens, and the LAION-400M dataset of 400 million text-image pairs, both developed in the English language. We illustrate how the framework informs action steps for hypothetical teams faced with data use, development, and documentation decisions. Ultimately, the framework structures human representation analyses and maps out analysis planning considerations, goals, and risk mitigation actions at different stages of dataset and model development. Mark Diaz, Sunipa Dev, Emily Reif, Remi Denton, Vinodkumar Prabhakaran |
AIES (1) | 1 |
| 2024 | What Makes An Expert? Reviewing How ML Researchers Define "Expert"abstractHuman experts are often engaged in the development of machine learning systems to collect and validate data, consult on algorithm development, and evaluate system performance. At the same time, who counts as an ‘expert’ and what constitutes ‘expertise’ is not always explicitly defined. In this work, we review 112 academic publications that explicitly reference ‘expert’ and ‘expertise’ and that describe the development of machine learning (ML) systems to survey how expertise is characterized and the role experts play. We find that expertise is often undefined and forms of knowledge outside of formal education and professional certification are rarely sought, which has implications for the kinds of knowledge that are recognized and legitimized in ML development. Moreover, we find that expert knowledge tends to be utilized in ways focused on mining textbook knowledge, such as through data annotation. We discuss the ways experts are engaged in ML development in relation to deskilling, the social construction of expertise, and implications for responsible AI development. We point to a need for reflection and specificity in justifications of domain expert engagement, both as a matter of documentation and reproducibility, as well as a matter of broadening the range of recognized expertise. Mark Diaz, Angela D. R. Smith |
AIES (1) | 1 |
| 2024 | The Illusion of Artificial InclusionabstractHuman participants play a central role in the development of modern artificial intelligence (AI) technology, in psychological science, and in user research. Recent advances in generative AI have attracted growing interest to the possibility of replacing human participants in these domains with AI surrogates. We survey several such “substitution proposals” to better understand the arguments for and against substituting human participants with modern generative AI. Our scoping review indicates that the recent wave of these proposals is motivated by goals such as reducing the costs of research and development work and increasing the diversity of collected data. However, these proposals ignore and ultimately conflict with foundational values of work with human participants: representation, inclusion, and understanding. This paper critically examines the principles and goals underlying human participation to help chart out paths for future work that truly centers and empowers participants. William Agnew, A. Stevie Bergman, Jennifer Chien, Mark Diaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, Kevin R. McKee |
CHI | 4 |
| 2024 | D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and EvaluationabstractWhile human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection.Recent studies that critically examine this issue are often focused on Western contexts, and solely document differences across age, gender, or racial groups.Consequently, NLP research on subjectivity have failed to consider that individuals within demographic groups may hold diverse values, which influence their perceptions beyond group norms.To effectively incorporate these considerations into NLP pipelines, we need datasets with extensive parallel annotations from a variety of social and cultural groups.In this paper we introduce the D3CODE dataset: a large-scale cross-cultural dataset of parallel annotations for offensive language in over 4.5K English sentences annotated by a pool of more than 4k annotators, balanced across gender and age, from across 21 countries, representing eight geo-cultural regions.The dataset captures annotators' moral values along six moral foundations: care, equality, proportionality, authority, loyalty, and purity.Our analyses reveal substantial regional variations in annotators' perceptions that are shaped by individual moral values, providing crucial insights for developing pluralistic, culturally sensitive NLP models. Aida Mostafazadeh Davani, Mark Diaz, Dylan K. Baker, Vinodkumar Prabhakaran |
EMNLP | 2 |
| 2024 | STAR: SocioTechnical Approach to Red Teaming Language ModelsabstractLaura Weidinger, John F J Mellor, Bernat Guillén Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Stevie Bergman, Mikel D. Rodriguez, Verena Rieser, William Isaac. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Laura Weidinger, John Mellor, Bernat Guillen Pegueroles, Nahema Marchal, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Stevie Bergman, Mikel Rodriguez, Verena Rieser, William Isaac 0001 |
EMNLP | 8 |
| 2024 | GRASP: A Disagreement Analysis Framework to Assess Group Associations in PerspectivesabstractVinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex Taylor, Mark Diaz, Ding Wang, Gregory Serapio-García. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex S. Taylor, Mark Diaz, Ding Wang 0006, Gregory Serapio-García |
NAACL-HLT | 7 |
| 2023 | DICES Dataset: Diversity in Conversational AI Evaluation for SafetyabstractMachine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This requirement overly simplifies the natural subjectivity present in many tasks, and obscures the inherent diversity in human perceptions and opinions about many content items. Preserving the variance in content and diversity in human perceptions in datasets is often quite expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is socio-culturally situated in this context. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographics information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. The DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of safety for conversational AI. We further describe a set of metrics that show how rater diversity influences safety perception across different geographic regions, ethnicity groups, age groups, and genders. The goal of the DICES dataset is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems. Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher Homan, Alicia Parrish, Gregory Serapio-García, Vinodkumar Prabhakaran, Ding Wang 0006 |
NeurIPS | 3 |
| 2023 | PaLM: Scaling Language Modeling with PathwaysabstractLarge language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel |
J. Mach. Learn. Res. | 59 |
| 2022 | Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective AnnotationsabstractAbstract Majority voting and averaging are common approaches used to resolve annotator disagreements and derive single ground truth labels from multiple annotations. However, annotators may systematically disagree with one another, often reflecting their individual biases and values, especially in the case of subjective tasks such as detecting affect, aggression, and hate speech. Annotator disagreements may capture important nuances in such tasks that are often ignored while aggregating annotations to a single ground truth. In order to address this, we investigate the efficacy of multi-annotator models. In particular, our multi-task based approach treats predicting each annotators’ judgements as separate subtasks, while sharing a common learned representation of the task. We show that this approach yields same or better performance than aggregating labels in the data prior to training across seven different binary classification tasks. Our approach also provides a way to estimate uncertainty in predictions, which we demonstrate better correlate with annotation disagreements than traditional methods. Being able to model uncertainty is especially useful in deployment scenarios where knowing when not to make a prediction is important. Aida Mostafazadeh Davani, Mark Diaz, Vinodkumar Prabhakaran |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | Addressing Age-Related Bias in Sentiment AnalysisabstractRecent studies have identified various forms of bias in language-based models, raising concerns about the risk of propagating social biases against certain groups based on sociodemographic factors (e.g., gender, race, geography). In this study, we analyze the treatment of age-related terms across 15 sentiment analysis models and 10 widely-used GloVe word embeddings and attempt to alleviate bias through a method of processing model training data. Our results show significant age bias is encoded in the outputs of many sentiment analysis algorithms and word embeddings, and we can alleviate this bias by manipulating training data. Mark Diaz, Isaac L. Johnson, Amanda Lazar, Anne Marie Piper, Darren Gergle |
IJCAI | 1 |
| 2019 | Whose Walkability?: Challenges in Algorithmically Measuring Subjective ExperienceabstractThe Walk Score is a patented algorithm for measuring the walkability of a given geographic area. In addition to its use in real estate, the accompanying API is used in a range of research in public health and urban development. This study explores how neighborhood residents differently understand the notion of walkability as well as the extent to which their personal definitions of neighborhood walkability are reflected in the Walk Score's underlying algorithm. We find that, while the Walk Score generally aligns with residents' priorities around walkability, significant subjective aspects that influence walking behavior are not reflected in the score, raising the need to consider implications for using algorithmic tools like the Walk Score in certain research contexts. We discuss the challenge of measuring subjective experience and how designers might begin to address it. We call for qualitative evaluations of algorithmic tools to help determine appropriate contexts of use. Mark Diaz, Nicholas Diakopoulos |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2019 | "The cavalry ain't coming in to save us": Supporting Capacities and Relationships through Civic TechabstractCities are increasingly integrating sensing and information and communication technologies to improve municipal services, civic engagement, and quality of life for residents. Although these civic technologies have the potential to affect economic, social, and environmental factors, there has been less focus on residents of lower income communities' involvement in civic technology design. Based on two public forums held in underserved communities, we describe residents' perceptions of civic technologies in their communities and challenges that limit the technologies' viability. We found that residents viewed civic technology as a tool that should strengthen existing community assets by providing an avenue to connect assets and build upon them. We describe how an asset-based approach can move us toward designing civic technology that develops stronger relationships among community-led initiatives, and between the community and local government--- rather than a data-driven approach to civic tech that focuses on transactions between residents and city services. Jessa Dickinson, Mark Diaz, Christopher A. Le Dantec, Sheena Lewis Erete |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2018 | Addressing Age-Related Bias in Sentiment AnalysisabstractComputational approaches to text analysis are useful in understanding aspects of online interaction, such as opinions and subjectivity in text. Yet, recent studies have identified various forms of bias in language-based models, raising concerns about the risk of propagating social biases against certain groups based on sociodemographic factors (e.g., gender, race, geography). In this study, we contribute a systematic examination of the application of language models to study discourse on aging. We analyze the treatment of age-related terms across 15 sentiment analysis models and 10 widely-used GloVe word embeddings and attempt to alleviate bias through a method of processing model training data. Our results demonstrate that significant age bias is encoded in the outputs of many sentiment analysis algorithms and word embeddings. We discuss the models' characteristics in relation to output bias and how these models might be best incorporated into research. Mark Diaz, Isaac L. Johnson, Amanda Lazar, Anne Marie Piper, Darren Gergle |
CHI | 1 |
| 2017 | Going Gray, Failure to Hire, and the Ick Factor: Analyzing How Older Bloggers Talk about AgeismabstractAgeism is a pervasive, and often invisible, form of discrimination. Though it can affect people of all ages, older adults in particular face age-related stereotypes and bias in their everyday lives. In this paper, we describe the ways in which older bloggers articulate a collective narrative on ageism as it appears in their lives, develop a community with anti-ageist interests, and discuss strategies to navigate and change societal views and institutions. Bloggers criticize stereotypical notions that focus exclusively on losses that occur with age and advocate a view that takes into account the complexity and positive aspects of older adulthood. This paper contributes a unique case of online collective action among older adults while drawing on their online discourse as a way of understanding what ageism means for CSCW. Amanda Lazar, Mark Diaz, Robin Brewer, Chelsea Kim, Anne Marie Piper |
CSCW | 2 |