William Barr Held

dblp:400/6173 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-1410-951XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021
YearPublicationVenuePosition
2026 Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
abstract
The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly.To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while reducing costs and data redundancy.Analyzing 10 subset selection methods with 18 audio models across 40 tasks covering major LAM evaluation dimensions, we show that subsets of just 50 examples (0.3% of data) can achieve over 0.93 Pearson correlation with full benchmark scores.To understand how well these scores align with what practitioners ultimately care about-user satisfaction-we collect 776 human preference ratings from realistic voice assistant conversations, finding that both subsets and full benchmark achieve only 0.85 correlation with human.To better predict preferences, we trained regression models on these selected subsets, achieving 0.98 correlation-outperforming regression models trained on both random subsets and the full benchmark.This demonstrates that in regression modeling, well-curated subsets outpredict the full benchmark, showing quality over quantity.We opensource these regression-weighted subsets as the HUMANS benchmark, an efficient proxy for LAM evaluation that captures both benchmark performance and user preferences.
Woody Haosheng Gan, William Barr Held, Diyi Yang
ACL (1)2
2025 Distilling an End-to-End Voice Assistant Without Instruction Training Data
abstract
Voice assistants, such as Siri and Google Assistant, typically model audio and text separately, resulting in lost speech information and increased complexity. Recent efforts to address this with end-to-end Speech Large Language Models (speech-in, text-out) trained with supervised finetuning (SFT) have led to models “forgetting” capabilities from text-only LLMs. Our work proposes an alternative paradigm for training Speech LLMs without instruction data, using the response of a text-only LLM to transcripts as self-supervision. Importantly, this process can be performed without annotated responses. We show that our Distilled Voice Assistant (DiVA) generalizes to Spoken Question Answering, Classification, and Translation. Furthermore, DiVA better matches user preferences, achieving a 72% win rate compared with state-of-the-art models like Qwen 2 Audio, despite using >100x less training compute.
William Barr Held, Weiyan Shi 0001, Minzhi Li, Michael J. Ryan, Diyi Yang
ACL (1)1
2025 Mind the Gap: Static and Interactive Evaluations of Large Audio Models
abstract
Minzhi Li, William Held, Michael J. Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, Diyi Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Minzhi Li, William Barr Held, Michael J. Ryan, Kunat Pipatanakul, P. P. Manakul, Diyi Yang
ACL (1)2
2025 SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
abstract
Michael J. Ryan, Omar Shaikh, Aditri Bhagirath, Daniel Frees, William Held, Diyi Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Michael J. Ryan, Omar Shaikh, Aditri Bhagirath, Daniel Frees, William Barr Held, Diyi Yang
ACL (1)5
2025 Culture Cartography: Mapping the Landscape of Cultural Knowledge
abstract
To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training.How do we find knowledge that is (1) salient to ingroup users, but (2) unknown to LLMs?The most common solutions are single-initiative: either researchers define challenging questions that users passively answer (traditional annotation), or users actively produce data that researchers structure as benchmarks (knowledge extraction).The process would benefit from mixed-initiative collaboration, where users guide the process to meaningfully reflect their cultures, and LLMs steer the process to meet the researcher's goals.We propose CULTURE CARTOGRAPHY as a methodology that operationalizes this mixed-initiative vision.Here, an LLM initializes annotation with questions for which it has low-confidence answers, making explicit both its prior knowledge and the gaps therein.This allows a human respondent to fill these gaps and steer the model towards salient topics through direct edits.We implement CULTURE CARTOGRAPHY as a tool called CULTURE EXPLORER.Compared to a baseline where humans answer LLMproposed questions, we find that CULTURE EX-PLORER more effectively produces knowledge that strong models like DeepSeek R1, Llama-4 and GPT-4o are missing, even with web search.Fine-tuning on this data boosts the accuracy of Llama models by up to 19.2% on related culture benchmarks.
Caleb Ziems, William Barr Held, Jane Dwivedi-Yu, Amir Goldberg, David Grusky, Diyi Yang
EMNLP2
2024 Unintended Impacts of LLM Alignment on Global Representation
abstract
Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO).Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness.However, human preferences are not universal, and aligning to specific preference sets may have unintended effects.We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide.Our results show that current alignment procedures create disparities between English dialects and global opinions.We find alignment improves capabilities in several languages.We conclude by discussing design decisions that led to these unintended impacts and recommendations for more equitable preference tuning.We make our code and data publicly available on Github 1 .
Michael J. Ryan, William Barr Held, Diyi Yang
ACL (1)2
2024 Can Large Language Models Transform Computational Social Science?
abstract
Abstract Large language models (LLMs) are capable of successfully performing many language processing tasks zero-shot (without training data). If zero-shot LLMs can also reliably classify and explain social phenomena like persuasiveness and political ideology, then LLMs could augment the computational social science (CSS) pipeline in important ways. This work provides a road map for using LLMs as CSS tools. Towards this end, we contribute a set of prompting best practices and an extensive evaluation pipeline to measure the zero-shot performance of 13 language models on 25 representative English CSS benchmarks. On taxonomic labeling tasks (classification), LLMs fail to outperform the best fine-tuned models but still achieve fair levels of agreement with humans. On free-form coding tasks (generation), LLMs produce explanations that often exceed the quality of crowdworkers’ gold references. We conclude that the performance of today’s LLMs can augment the CSS research pipeline in two ways: (1) serving as zero-shot data annotators on human annotation teams, and (2) bootstrapping challenging creative generation tasks (e.g., explaining the underlying attributes of a text). In summary, LLMs are posed to meaningfully participate in social science analysis in partnership with humans.
Caleb Ziems, William Barr Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang 0001, Diyi Yang
Comput. Linguistics2
2023 DAMP: Doubly Aligned Multilingual Parser for Task-Oriented Dialogue
abstract
William Held, Christopher Hidey, Fei Liu, Eric Zhu, Rahul Goel, Diyi Yang, Rushin Shah. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
William Barr Held, Christopher Hidey, Eric Zhu, Rahul Goel, Diyi Yang, Rushin Shah
ACL (1)1
2023 On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
abstract
Warning: This paper contains several toxic and offensive statements.Generating a Chain of Thought (CoT) has been shown to consistently improve large language model (LLM) performance on a wide range of NLP tasks.However, prior work has mainly focused on logical reasoning tasks (e.g.arithmetic, commonsense QA); it remains unclear whether improvements hold for more diverse types of reasoning, especially in socially situated contexts.Concretely, we perform a controlled evaluation of zero-shot CoT across two socially sensitive domains: harmful questions and stereotype benchmarks.We find that zeroshot CoT reasoning in sensitive domains significantly increases a model's likelihood to produce harmful or undesirable output, with trends holding across different prompt formats and model variants.Furthermore, we show that harmful CoTs increase with model size, but decrease with improved instruction following.Our work suggests that zero-shot CoT should be used with caution on socially important tasks, especially when marginalized groups or sensitive topics are involved.
Omar Shaikh, Hongxin Zhang 0005, William Barr Held, Michael S. Bernstein, Diyi Yang
ACL (1)3
2023 Multi-VALUE: A Framework for Cross-Dialectal English NLP
abstract
Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users.Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts.Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE).We introduce a suite of resources for evaluating and achieving English dialect invariance.The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features.Multi-VALUE maps SAE to synthetic forms of each dialect.First, we use this system to stress tests question answering, machine translation, and semantic parsing.Stress tests reveal significant performance disparities for leading models on nonstandard dialects.Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems.Finally, we partner with native speakers of Chicano and Indian English to release new goldstandard variants of the popular CoQA task.To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/.
Caleb Ziems, William Barr Held, Jingfeng Yang 0001, Jwala Dhamala, Rahul Gupta 0001, Diyi Yang
ACL (1)2
2023 Shapley Head Pruning: Identifying and Removing Interference in Multilingual Transformers
abstract
Multilingual transformer-based models demonstrate remarkable zero and few-shot transfer across languages by learning and reusing language-agnostic features.However, as a fixed-size model acquires more languages, its performance across all languages degrades.Those who attribute this interference phenomenon to limited model capacity address the problem by adding additional parameters, despite evidence that transformer-based models are overparameterized.In this work, we show that it is possible to reduce interference by instead identifying and pruning language-specific attention heads.First, we use Shapley Values, a credit allocation metric from coalitional game theory, to identify attention heads that introduce interference.Then, we show that pruning such heads from a fixed model improves performance for a target language on both sentence classification and structural prediction.Finally, we provide insights on language-agnostic and language-specific attention heads using attention visualization. 1
William Barr Held, Diyi Yang
EACL1
2023 DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic Rules
abstract
Existing large language models (LLMs) that mainly focus on Standard American English (SAE) often lead to significantly worse performance when being applied to other English dialects.While existing mitigations tackle discrepancies for individual target dialects, they assume access to high-accuracy dialect identification systems.The boundaries between dialects are inherently flexible, making it difficult to categorize language into discrete predefined categories.In this work, we propose DADA (Dialect Adaptation via Dynamic Aggregation), a modular approach to imbue SAE-trained models with multi-dialectal robustness by composing adapters which handle specific linguistic features.The compositional architecture of DADA allows for both targeted adaptation to specific dialect variants and simultaneous adaptation to various dialects.We show that DADA is effective for both single task and instruction finetuned language models, offering an extensible and interpretable framework for adapting existing LLMs to different English dialects.
William Barr Held, Diyi Yang
EMNLP2
2023 Task-Agnostic Low-Rank Adapters for Unseen English Dialects
abstract
Large Language Models (LLMs) are trained on corpora disproportionally weighted in favor of Standard American English.As a result, speakers of other dialects experience significantly more failures when interacting with these technologies.In practice, these speakers often accommodate their speech to be better understood.Our work shares the belief that language technologies should be designed to accommodate the diversity in English dialects and not the other way around.However, prior works on dialect struggle with generalizing to evolving and emerging dialects in a scalable manner.To fill this gap, our method, Hyper-LoRA, leverages expert linguistic knowledge to enable resource-efficient adaptation via hypernetworks.By disentangling dialect-specific and cross-dialectal information, HyperLoRA improves generalization to unseen dialects in a task-agnostic fashion.Not only is HyperLoRA more scalable in the number of parameters, but it also achieves the best or most competitive performance across 5 dialects in a zero-shot setting.In this way, our approach facilitates access to language technology for billions of English dialect speakers who are traditionally underrepresented.
Zedian Xiao, William Barr Held, Diyi Yang
EMNLP2
2021 Focus on what matters: Applying Discourse Coherence Theory to Cross Document Coreference
abstract
Performing event and entity coreference resolution across documents vastly increases the number of candidate mentions, making it intractable to do the full n 2 pairwise comparisons.Existing approaches simplify by considering coreference only within document clusters, but this fails to handle inter-cluster coreference, common in many applications.As a result cross-document coreference algorithms are rarely applied to downstream tasks.We draw on an insight from discourse coherence theory: potential coreferences are constrained by the reader's discourse focus.We model the entities/events in a reader's focus as a neighborhood within a learned latent embedding space which minimizes the distance between mentions and the centroids of their gold coreference clusters.We then use these neighborhoods to sample only hard negatives to train a fine-grained classifier on mention pairs and their local discourse features.Our approach 1 achieves state-of-the-art results for both events and entities on the ECB+, Gun Violence, Football Coreference, and Cross-Domain Cross-Document Coreference corpora.Furthermore, training on multiple corpora improves average performance across all datasets by 17.2 F1 points, leading to a robust coreference resolution model for use in downstream tasks where link distribution is unknown.
William Barr Held, Dan Iter, Daniel Jurafsky
EMNLP (1)1