Caleb Ziems

dblp:252/5058 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-2792-6298ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Culture Cartography: Mapping the Landscape of Cultural Knowledge
abstract
To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training.How do we find knowledge that is (1) salient to ingroup users, but (2) unknown to LLMs?The most common solutions are single-initiative: either researchers define challenging questions that users passively answer (traditional annotation), or users actively produce data that researchers structure as benchmarks (knowledge extraction).The process would benefit from mixed-initiative collaboration, where users guide the process to meaningfully reflect their cultures, and LLMs steer the process to meet the researcher's goals.We propose CULTURE CARTOGRAPHY as a methodology that operationalizes this mixed-initiative vision.Here, an LLM initializes annotation with questions for which it has low-confidence answers, making explicit both its prior knowledge and the gaps therein.This allows a human respondent to fill these gaps and steer the model towards salient topics through direct edits.We implement CULTURE CARTOGRAPHY as a tool called CULTURE EXPLORER.Compared to a baseline where humans answer LLMproposed questions, we find that CULTURE EX-PLORER more effectively produces knowledge that strong models like DeepSeek R1, Llama-4 and GPT-4o are missing, even with web search.Fine-tuning on this data boosts the accuracy of Llama models by up to 19.2% on related culture benchmarks.
Caleb Ziems, William Barr Held, Jane Dwivedi-Yu, Amir Goldberg, David Grusky, Diyi Yang
EMNLP1
2024 Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles
abstract
Julia Kruk, Michela Marchini, Rijul Magu, Caleb Ziems, David Muchlinski, Diyi Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Julia Kruk, Michela Marchini, Rijul Magu, Caleb Ziems, David Muchlinski, Diyi Yang
ACL (1)4
2024 Can Large Language Models Transform Computational Social Science?
abstract
Abstract Large language models (LLMs) are capable of successfully performing many language processing tasks zero-shot (without training data). If zero-shot LLMs can also reliably classify and explain social phenomena like persuasiveness and political ideology, then LLMs could augment the computational social science (CSS) pipeline in important ways. This work provides a road map for using LLMs as CSS tools. Towards this end, we contribute a set of prompting best practices and an extensive evaluation pipeline to measure the zero-shot performance of 13 language models on 25 representative English CSS benchmarks. On taxonomic labeling tasks (classification), LLMs fail to outperform the best fine-tuned models but still achieve fair levels of agreement with humans. On free-form coding tasks (generation), LLMs produce explanations that often exceed the quality of crowdworkers’ gold references. We conclude that the performance of today’s LLMs can augment the CSS research pipeline in two ways: (1) serving as zero-shot data annotators on human annotation teams, and (2) bootstrapping challenging creative generation tasks (e.g., explaining the underlying attributes of a text). In summary, LLMs are posed to meaningfully participate in social science analysis in partnership with humans.
Caleb Ziems, William Barr Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang 0001, Diyi Yang
Comput. Linguistics1
2023 NormBank: A Knowledge Bank of Situational Social Norms
abstract
We present NORMBANK, a knowledge bank of 155k situational norms.This resource is designed to ground flexible normative reasoning for interactive, assistive, and collaborative AI systems.Unlike prior commonsense resources, NORMBANK grounds each inference within a multivalent sociocultural frame, which includes the setting (e.g., restaurant), the agents' contingent roles (waiter, customer), their attributes (age, gender), and other physical, social, and cultural constraints (e.g., the temperature or the country of operation).In total, NORMBANK contains 63k unique constraints from a taxonomy that we introduce and iteratively refine here.Constraints then apply in different combinations to frame social norms.Under these manipulations, norms are non-monotonic -one can cancel an inference by updating its frame even slightly.Still, we find evidence that neural models can help reliably extend the scope and coverage of NORMBANK.We further demonstrate the utility of this resource with a series of transfer experiments.For data and code, see
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Y. Halevy, Diyi Yang
ACL (1)1
2023 Multi-VALUE: A Framework for Cross-Dialectal English NLP
abstract
Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users.Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts.Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE).We introduce a suite of resources for evaluating and achieving English dialect invariance.The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features.Multi-VALUE maps SAE to synthetic forms of each dialect.First, we use this system to stress tests question answering, machine translation, and semantic parsing.Stress tests reveal significant performance disparities for leading models on nonstandard dialects.Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems.Finally, we partner with native speakers of Chicano and Indian English to release new goldstandard variants of the popular CoQA task.To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/.
Caleb Ziems, William Barr Held, Jingfeng Yang 0001, Jwala Dhamala, Rahul Gupta 0001, Diyi Yang
ACL (1)1
2023 Impressions: Visual Semiotics and Aesthetic Impact Understanding
abstract
Is aesthetic impact different from beauty?Is visual salience a reflection of its capacity for effective communication?We present Impressions, 1 a novel dataset through which to investigate the semiotics of images, and how specific visual features and design choices can elicit specific emotions, thoughts and beliefs.We posit that the impactfulness of an image extends beyond formal definitions of aesthetics, to its success as a communicative act, where style contributes as much to meaning formation as the subject matter.However, prior image captioning datasets are not designed to empower state-of-the-art architectures to model potential human impressions or interpretations of images.To fill this gap, we design an annotation task heavily inspired by image analysis techniques in the Visual Arts to collect 1,440 image-caption pairs and 4,320 unique annotations exploring impact, pragmatic image description, impressions, and aesthetic design choices.We show that existing multimodal image captioning and conditional generation models struggle to simulate plausible human responses to images.However, this dataset significantly improves their ability to model impressions and aesthetic evaluations of images through fine-tuning and few-shot adaptation.
Julia Kruk, Caleb Ziems, Diyi Yang
EMNLP2
2023 CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
abstract
Annotated data plays a critical role in Natural Language Processing (NLP) in training models and evaluating their performance.Given recent developments in Large Language Models (LLMs), models such as ChatGPT demonstrate zero-shot capability on many text-annotation tasks, comparable with or even exceeding human annotators.Such LLMs can serve as alternatives for manual annotation, due to lower costs and higher scalability.However, limited work has leveraged LLMs as complementary annotators, nor explored how annotation work is best allocated among humans and LLMs to achieve both quality and cost objectives.We propose CoAnnotating, a novel paradigm for Human-LLM co-annotation of unstructured texts at scale.Under this framework, we utilize uncertainty to estimate LLMs' annotation capability.Our empirical study shows CoAnnotating to be an effective means to allocate work from results on different datasets, with up to 21% performance improvement over random baseline.For code implementation, see https: //github.com/SALT-NLP/CoAnnotating.
Minzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan, Nancy F. Chen, Zhengyuan Liu, Diyi Yang
EMNLP3
2022 VALUE: Understanding Dialect Disparity in NLU
abstract
English Natural Language Understanding (NLU) systems have achieved great performances and even outperformed humans on benchmarks like GLUE and SuperGLUE.However, these benchmarks contain only textbook Standard American English (SAE).Other dialects have been largely overlooked in the NLP community.This leads to biased and inequitable NLU systems that serve only a sub-population of speakers.To understand disparities in current models and to facilitate more dialect-competent NLU systems, we introduce the VernAcular Language Understanding Evaluation (VALUE) benchmark, a challenging variant of GLUE that we created with a set of lexical and morphosyntactic transformation rules.In this initial release (V.1), we construct rules for 11 features of African American Vernacular English (AAVE), and we recruit fluent AAVE speakers to validate each feature transformation via linguistic acceptability judgments in a participatory design manner.Experiments show that these new dialectal features can lead to a drop in model performance.
Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, Diyi Yang
ACL (1)1
2022 Inducing Positive Perspectives with Text Reframing
abstract
Sentiment transfer is one popular example of a text style transfer task, where the goal is to reverse the sentiment polarity of a text.With a sentiment reversal comes also a reversal in meaning.We introduce a different but related task called positive reframing in which we neutralize a negative point of view and generate a more positive perspective for the author without contradicting the original meaning.Our insistence on meaning preservation makes positive reframing a challenging and semantically rich task.To facilitate rapid progress, we introduce a large-scale benchmark, POSITIVE PSY-CHOLOGY FRAMES, with 8,349 sentence pairs and 12,755 structured annotations to explain positive reframing in terms of six theoreticallymotivated reframing strategies.Then we evaluate a set of state-of-the-art text style transfer models, and conclude by discussing key challenges and directions for future work.To download the data, see https://github. com/GT-SALT/positive-frames
Caleb Ziems, Minzhi Li, Anthony Zhang, Diyi Yang
ACL (1)1
2022 The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
abstract
Content Warning: some examples in this paper may be offensive or upsetting.Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user's trust in the moral integrity of the system.Moral deviations are difficult to mitigate because moral judgments are not universal, and there may be multiple competing judgments that apply to a situation simultaneously.In this work, we introduce a new resource, not to authoritatively resolve moral ambiguities, but instead to facilitate systematic understanding of the intuitions, values and moral judgments reflected in the utterances of dialogue systems.The MORAL INTEGRITY CORPUS, MIC , is such a resource, which captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs).Each RoT reflects a particular moral conviction that can explain why a chatbot's reply may appear acceptable or problematic.We further organize RoTs with a set of 9 moral and social attributes and benchmark performance for attribute classification.Most importantly, we show that current neural language models can automatically generate new RoTs that reasonably describe previously unseen interactions, but they still struggle with certain scenarios.Our findings suggest that MIC will be a useful resource for understanding and language models' implicit moral assumptions and flexibly benchmarking the integrity of conversational agents.
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Y. Halevy, Diyi Yang
ACL (1)1
2021 Racism is a virus: anti-asian hate and counterspeech in social media during the COVID-19 crisis
abstract
The spread of COVID-19 has sparked racism and hate on social media targeted towards Asian communities. However, little is known about how racial hate spreads during a pandemic and the role of counterspeech in mitigating this spread. In this work, we study the evolution and spread of anti-Asian hate speech through the lens of Twitter. We create COVID-HATE, the largest dataset of anti-Asian hate and counterspeech spanning 14 months, containing over 206 million tweets, and a social network with over 127 million nodes. By creating a novel hand-labeled dataset of 3,355 tweets, we train a text classifier to identify hateful and counterspeech tweets that achieves an average macro-F1 score of 0.832. Using this dataset, we conduct longitudinal analysis of tweets and users. Analysis of the social network reveals that hateful and counterspeech users interact and engage extensively with one another, instead of living in isolated polarized communities. We find that nodes were highly likely to become hateful after being exposed to hateful content in the year 2020. Notably, counterspeech messages discourage users from turning hateful, potentially suggesting a solution to curb hate on web and social media platforms. Data and code is available at http://claws.cc.gatech.edu/covid.
Bing He 0002, Caleb Ziems, Sandeep Soni, Naren Ramakrishnan, Diyi Yang, Srijan Kumar
ASONAM2
2021 Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
abstract
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, Diyi Yang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, Diyi Yang
EMNLP (1)2
2020 Aggressive, Repetitive, Intentional, Visible, and Imbalanced: Refining Representations for Cyberbullying Classification
Caleb Ziems, Ymir Vigfusson, Fred Morstatter
ICWSM1