EDBT 2026 Demo / reviewers in the wild / expert
Robert D. Hawkins
dblp:168/8718 · also Robert X. D. Hawkins
· DBLP profile ↗
51ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0001-9089-8544ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 11 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 37 · 9 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsabstractRecent evaluations show that large language models (LLMs) frequently fail to challenge users' harmful beliefs in domains ranging from medical advice to social reasoning.We present a unifying analysis through the lens of pragmatics: these safety failures can be understood and addressed as LLMs exhibiting excessive accommodation and insufficient epistemic vigilance.We show that the pragmatic factors affecting accommodation and epistemic vigilance in humans (at-issueness, linguistic encoding, and source reliability) influence LLM behaviors in similar ways.We demonstrate how these factors explain performance differences across three safety benchmarks that test models' ability to challenge harmful beliefs, spanning misinformation (Cancer-Myth, SAGE-Eval) and sycophancy (ELEPHANT).This pragmatic lens further motivates prompting interventions, such as adding the phrase "wait a minute", that drastically improve performance on these difficult benchmarks by shifting pragmatic cues.Our results have practical implications for benchmark design and underscore the importance of pragmatics for understanding model behavior and improving performance. Myra Cheng, Robert D. Hawkins, Daniel Jurafsky |
ACL (1) | 2 |
| 2026 | Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical TasksabstractA quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner’s hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world. Kiyosu Maeda, William P. McCarthy, Ching-Yi Tsai, Jeffrey Mu, Robert D. Hawkins, Judith E. Fan, Parastoo Abtahi |
CHI | 6 |
| 2025 | How do we get to know someone? Diagnostic questions for inferring personal traits
Erik Brockbank, Tobias Gerstenberg, Judith E. Fan, Robert D. Hawkins |
CogSci | 4 |
| 2025 | A Computational Account of Epistemic Vigilance: Learning from Selective Truths through Bayesian Reasoning
Robert D. Hawkins, Charley M. Wu, Michael Franke |
CogSci | 2 |
| 2025 | Amortizing Structure Discovery with Generative Flow Networks
Alex Guerra, Aditya Palaparthi, Sarah-Jane Leslie, Robert D. Hawkins |
CogSci | 4 |
| 2025 | Two paths to variation in semantic judgments: How ambiguity and conceptual diversity drive individual differences in meaning
Raja Marjieh, Nori Jacoby, Robert D. Hawkins |
CogSci | 4 |
| 2025 | Minding the Politeness Gap in Cross-cultural Communication
Yuka Machino, Max H. Siegel, Matthias Hofer 0002, Josh Tenenbaum, Robert D. Hawkins |
CogSci | 5 |
| 2025 | Learning to communicate a shared wavelength
Wasita Mahaphanit, Benjamin Keller, Robert D. Hawkins, Jonathan Phillips, Luke Chang |
CogSci | 3 |
| 2025 | Function shapes form: Compositionality emerges from communicative needs, not environmental structure alone
Jessica Mankewitz, Robert D. Hawkins |
CogSci | 2 |
| 2025 | Dynamics of topic exploration in conversation
Helen Schmidt, Claire Bergey, Changyi Zhou, Chelsea Helion, Robert D. Hawkins |
CogSci | 5 |
| 2025 | Polite Speech Generation in Humans and Language Models
Robert D. Hawkins |
CogSci | 2 |
| 2025 | Overcoming sparse and uneven evidence with natural language communication
Yuliya Zubak, Robert D. Hawkins |
CogSci | 2 |
| 2025 | Comparing human and LLM politeness strategies in free productionabstractPolite speech poses a fundamental alignment challenge for large language models (LLMs).Humans deploy a rich repertoire of linguistic strategies to balance informational and social goals -from positive approaches that build rapport (compliments, expressions of interest) to negative strategies that minimize imposition (hedging, indirectness).We investigate whether LLMs employ a similarly context-sensitive repertoire by comparing human and LLM responses to English-language scenarios in both constrained and open-ended production tasks.We find that larger models (≥70B parameters) successfully replicate key effects from the computational pragmatics literature, and human evaluators prefer LLM-generated responses in open-ended contexts.However, further linguistic analyses reveal that models disproportionately rely on negative politeness strategies to create distance even in positive contexts, potentially leading to misinterpretations.While LLMs thus demonstrate an impressive command of politeness strategies, these systematic differences provide important groundwork for making intentional choices about pragmatic behavior in human-AI communication. Robert D. Hawkins |
EMNLP | 2 |
| 2025 | Core Knowledge Deficits in Multi-Modal Language ModelsabstractWhile Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge—rudimentary cognitive abilities innate to humans from early childhood.
To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We evaluate 230 models with 11 different prompts, leading to a total of 2,530 data points for analysis. Our experiments uncover four key findings, collectively demonstrating core knowledge deficits in MLLMs: they consistently underperform and show reduced, or even absent, scalability on low-level abilities relative to high-level ones.
Finally, we propose Concept Hacking, a novel controlled evaluation method, that reveals MLLMs fail to progress toward genuine core knowledge understanding, but instead rely on shortcut learning as they scale. Project page at https://williamium3000.github.io/core-knowledge/. Yijiang Li, Qingying Gao, Tianwei Zhao, Bingyang Wang, Haiyun Lyu, Robert D. Hawkins, Nuno Vasconcelos, Tal Golan, Dezhi Luo, Hokin Deng |
ICML | 7 |
| 2024 | A longitudinal analysis of children's communicative acts
Claire Bergey, Misha O'Keeffe, Robert D. Hawkins |
CogSci | 3 |
| 2024 | When and why does shared reality generalize?
Wasita Mahaphanit, Christopher Welker, Helen Schmidt, Luke Chang, Robert D. Hawkins |
CogSci | 5 |
| 2024 | Toddlers Actively Sample from Reliable and Unreliable Speakers
Jessica Mankewitz, Robert D. Hawkins, Jenny R. Saffran |
CogSci | 2 |
| 2023 | Advancing Cognitive Science and AI with Cognitive-AI Benchmarking
Felix J. Binder, Logan Matthew Cross, Yoni Friedman, Robert D. Hawkins, Dan Yamins, Judith E. Fan |
CogSci | 4 |
| 2023 | Semantic uncertainty guides the extension of conventions to new referents
Ron Eliav, Anya Ji, Yoav Artzi, Robert D. Hawkins |
CogSci | 4 |
| 2023 | Overinformative Question Answering by Humans and Machines
Polina Tsvilodub, Michael Franke, Robert D. Hawkins, Noah D. Goodman |
CogSci | 3 |
| 2023 | Same situation, different goals: Exploring how people construct stories from streams of experience
Naomi Vaida, Susan T. Fiske, Robert D. Hawkins |
CogSci | 3 |
| 2022 | Two's company but six is a crowd: emergence of conventions in multiparty communication games
Veronica Boyce, Robert D. Hawkins, Noah D. Goodman, Michael C. Frank |
CogSci | 2 |
| 2022 | Identifying concept libraries from language about object structure
Catherine Wong, William P. McCarthy, Gabriel Grand, Yoni Friedman, Josh Tenenbaum, Jacob Andreas, Robert D. Hawkins, Judith E. Fan |
CogSci | 7 |
| 2022 | Mixed-effects transformers for hierarchical adaptationabstractLanguage differs dramatically from context to context.To some degree, large language models like GPT-3 account for such variation by conditioning on strings of initial input text, or prompts.However, prompting can be ineffective when contexts are sparse, out-of-sample, or extra-textual.In this paper, we introduce the mixed-effects transformer (MET), a novel approach for learning hierarchically-structured prefixes-lightweight modules prepended to an input sequence-to account for structured variation in language use.Specifically, we show how the popular class of mixedeffects regression models may be extended to transformer-based architectures using a regularized prefix-tuning procedure with dropout.We evaluate this approach on several domainadaptation benchmarks, finding that it learns contextual variation from minimal data while generalizing well to unseen contexts. Julia White 0001, Noah D. Goodman, Robert D. Hawkins |
EMNLP | 3 |
| 2022 | Abstract Visual Reasoning with Tangram ShapesabstractWe introduce KILOGRAM, a resource for studying abstract visual reasoning in humans and machines.Drawing on the history of tangram puzzles as stimuli in cognitive science, we build a richly annotated dataset that, with >1k distinct stimuli, is orders of magnitude larger and more diverse than prior resources.It is both visually and linguistically richer, moving beyond whole shape descriptions to include segmentation maps and part labels.We use this resource to evaluate the abstract visual reasoning capacities of recent multi-modal models.We observe that pre-trained weights demonstrate limited abstract reasoning, which dramatically improves with fine-tuning.We also observe that explicitly describing parts aids abstract reasoning for both humans and models, especially when jointly encoding the linguistic and visual inputs. Anya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr, Wai Keen Vong, Robert D. Hawkins, Yoav Artzi |
EMNLP | 6 |
| 2022 | Using natural language and program abstractions to instill human inductive biases in machinesabstractStrong inductive biases give humans the ability to quickly learn to perform a variety of tasks. Although meta-learning is a method to endow neural networks with useful inductive biases, agents trained by meta-learning may sometimes acquire very different strategies from humans. We show that co-training these agents on predicting representations from natural language task descriptions and programs induced to generate such tasks guides them toward more human-like inductive biases. Human-generated language descriptions and program induction models that add new learned primitives both contain abstract concepts that can compress description length. Co-training on these representations result in more human-like behavior in downstream meta-reinforcement learning agents than less abstract controls (synthetic language descriptions, program induction without learned primitives), suggesting that the abstraction supported by these representations is key. Sreejan Kumar, Carlos G. Correa, Ishita Dasgupta 0001, Raja Marjieh, Michael Y. Hu, Robert D. Hawkins, Jonathan D. Cohen 0003, Nathaniel D. Daw, Karthik Narasimhan, Thomas L. Griffiths 0001 |
NeurIPS | 6 |
| 2022 | How to talk so AI will learn: Instructions, descriptions, and autonomyabstractFrom the earliest years of our lives, humans use language to express our beliefs and desires. Being able to talk to artificial agents about our preferences would thus fulfill a central goal of value alignment. Yet today, we lack computational models explaining such language use. To address this challenge, we formalize learning from language in a contextual bandit setting and ask how a human might communicate preferences over behaviors. We study two distinct types of language: instructions, which provide information about the desired policy, and descriptions, which provide information about the reward function. We show that the agent's degree of autonomy determines which form of language is optimal: instructions are better in low-autonomy settings, but descriptions are better when the agent will need to act independently. We then define a pragmatic listener agent that robustly infers the speaker's reward function by reasoning about how the speaker expresses themselves. We validate our models with a behavioral experiment, demonstrating that (1) our speaker model predicts human behavior, and (2) our pragmatic listener successfully recovers humans' reward functions. Finally, we show that this form of social learning can integrate with and reduce regret in traditional reinforcement learning. We hope these insights facilitate a shift from developing agents that obey language to agents that learn from it. Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths 0001, Dylan Hadfield-Menell |
NeurIPS | 2 |
| 2021 | Learning Rewards From Linguistic FeedbackabstractWe explore unconstrained natural language feedback as a learning signal for artificial agents. Humans use rich and varied language to teach, yet most prior work on interactive learning from language assumes a particular form of input (e.g., commands). We propose a general framework which does not make this assumption, instead using aspect-based sentiment analysis to decompose feedback into sentiment over the features of a Markov decision process. We then infer the teacher's reward function by regressing the sentiment on the features, an analogue of inverse reinforcement learning. To evaluate our approach, we first collect a corpus of teaching behavior in a cooperative task where both teacher and learner are human. We implement three artificial learners: sentiment-based "literal" and "pragmatic" models, and an inference network trained end-to-end to predict rewards. We then re-run our initial experiment, pairing human teachers with these artificial learners. All three models successfully learn from interactive human feedback. The inference network approaches the performance of the "literal" sentiment model, while the "pragmatic" model nears human performance. Our work provides insight into the information structure of naturalistic linguistic feedback as well as methods to leverage it for reinforcement learning. Theodore R. Sumers, Mark K. Ho, Robert D. Hawkins, Karthik Narasimhan, Thomas L. Griffiths 0001 |
AAAI | 3 |
| 2021 | Respect the code: Speakers expect novel conventions to generalize within but not across social group boundaries
Robert D. Hawkins, Irina Liu, Adele Goldberg 0002, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2021 | Contextual Flexibility Guides Communication in a Cooperative Language Game
Abhilasha Ashok Kumar, Ketika Garg, Robert D. Hawkins |
CogSci | 3 |
| 2021 | Learning to communicate about shared procedural abstractions
William P. McCarthy, Robert D. Hawkins, Cameron Holdaway, Judith E. Fan |
CogSci | 2 |
| 2021 | Extending rational models of communication from beliefs to actions
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths 0001 |
CogSci | 2 |
| 2021 | Open-domain clarification question generation without question examplesabstractAn overarching goal of natural language processing is to enable machines to communicate seamlessly with humans.However, natural language can be ambiguous or unclear.In cases of uncertainty, humans engage in an interactive process known as repair: asking questions and seeking clarification until their uncertainty is resolved.We propose a framework for building a visually grounded questionasking model capable of producing polar (yesno) clarification questions to resolve misunderstandings in dialogue.Our model uses an expected information gain objective to derive informative questions from an off-the-shelf image captioner without requiring any supervised question-answer data.We demonstrate our model's ability to pose questions that improve communicative success in a goal-oriented 20 questions game with synthetic and human answerers. Julia White 0001, Gabriel Poesia, Robert D. Hawkins, Dorsa Sadigh, Noah D. Goodman |
EMNLP (1) | 3 |
| 2020 | Generalizing meanings from partners to populations: Hierarchical inference supports convention formation on networks
Robert D. Hawkins, Noah D. Goodman, Adele Goldberg 0002, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2020 | Parents scaffold the formation of conversational pacts with their children
Ashley C. Leung, Robert D. Hawkins, Daniel Yurovsky |
CogSci | 2 |
| 2020 | Continual Adaptation for Efficient Machine CommunicationabstractTo communicate with new partners in new contexts, humans rapidly form new linguistic conventions.Recent neural language models are able to comprehend and produce the existing conventions present in their training data, but are not able to flexibly and interactively adapt those conventions on the fly as humans do.We introduce an interactive repeated reference task as a benchmark for models of adaptation in communication and propose a regularized continual learning framework that allows an artificial agent initialized with a generic language model to more accurately and efficiently communicate with a partner over time.We evaluate this framework through simulations on COCO and in real-time reference game experiments with human partners. Robert D. Hawkins, Minae Kwon, Dorsa Sadigh, Noah D. Goodman |
CoNLL | 1 |
| 2020 | Investigating representations of verb bias in neural language modelsabstractLanguages typically provide more than one grammatical construction to express certain types of messages. A speaker's choice of construction is known to depend on multiple factors, including the choice of main verb -- a phenomenon known as \emph{verb bias}. Here we introduce DAIS, a large benchmark dataset containing 50K human judgments for 5K distinct sentence pairs in the English dative alternation. This dataset includes 200 unique verbs and systematically varies the definiteness and length of arguments. We use this dataset, as well as an existing corpus of naturally occurring data, to evaluate how well recent neural language models capture human preferences. Results show that larger models perform better than smaller models, and transformer architectures (e.g. GPT-2) tend to out-perform recurrent architectures (e.g. LSTMs) even under comparable parameter and training settings. Additional analyses of internal feature representations suggest that transformers may better integrate specific lexical information with grammatical constructions. Robert D. Hawkins, Takateru Yamakoshi, Thomas L. Griffiths 0001, Adele Goldberg 0002 |
EMNLP (1) | 1 |
| 2019 | Disentangling contributions of visual information and interaction history in the formation of graphical conventions
Robert D. Hawkins, Megumi Sano, Noah D. Goodman, Judith W. Fan |
CogSci | 1 |
| 2019 | Using replication studies to teach research methods in cognitive science
Josh de Leeuw, Janet K. Andrews, Kenneth R. Livingston, Michael Franke, Joshua K. Hartshorne, Robert D. Hawkins, Jordan Wagge |
CogSci | 6 |
| 2019 | Communicating semantic part information in drawings
Kushin Mukherjee, Robert D. Hawkins, Judith W. Fan |
CogSci | 2 |
| 2019 | Shapeglot: Learning Language for Shape DifferentiationabstractIn this work we explore how fine-grained differences between the shapes of common objects are expressed in language, grounded on 2D and/or 3D object representations. We first build a large scale, carefully controlled dataset of human utterances each of which refers to a 2D rendering of a 3D CAD model so as to distinguish it from a set of shape-wise similar alternatives. Using this dataset, we develop neural language understanding (listening) and production (speaking) models that vary in their grounding (pure 3D forms via point-clouds vs. rendered 2D images), the degree of pragmatic reasoning captured (e.g. speakers that reason about a listener or not), and the neural architecture (e.g. with or without attention). We find models that perform well with both synthetic and human partners, and with held out utterances and objects. We also find that these models are capable of zero-shot transfer learning to novel object classes (e.g. transfer from training on chairs to testing on lamps), as well as to real-world images drawn from furniture catalogs. Lesion studies indicate that the neural listeners depend heavily on part-related words and associate these words correctly with visual parts of objects (without any explicit supervision on such parts), and that transfer to novel classes is most successful when known part-related words are available. This work illustrates a practical approach to language grounding, and provides a novel case study in the relationship between object shape and linguistic structure when it comes to object differentiation. Panos Achlioptas, Leonidas J. Guibas, Noah D. Goodman, Judy Fan, Robert D. Hawkins |
ICCV | 5 |
| 2018 | Emerging abstractions: Lexical conventions are shaped by communicative context
Robert D. Hawkins, Michael Franke, Kenny Smith, Noah D. Goodman |
CogSci | 1 |
| 2017 | Convention-formation in iterated reference games
Robert D. Hawkins, Mike Frank, Noah D. Goodman |
CogSci | 1 |
| 2017 | Mentioning atypical properties of objects is communicatively efficient
Elisa Kreiss, Robert D. Hawkins, Judith Degen, Noah D. Goodman |
CogSci | 2 |
| 2017 | Colors in Context: A Pragmatic Neural Model for Grounded Language UnderstandingabstractWe present a model of pragmatic referring expression interpretation in a grounded communication task (identifying colors from descriptions) that draws upon predictions from two recurrent neural network classifiers, a speaker and a listener, unified by a recursive pragmatic reasoning framework. Experiments show that this combined pragmatic model interprets color descriptions more accurately than the classifiers from which it is built, and that much of this improvement results from combining the speaker and listener perspectives. We observe that pragmatic reasoning helps primarily in the hardest cases: when the model must distinguish very similar colors, or when few utterances adequately express the target color. Our findings make use of a newly-collected corpus of human utterances in color reference games, which exhibit a variety of pragmatic behaviors. We also show that the embedded speaker model reproduces many of these pragmatic behaviors. Will Monroe, Robert D. Hawkins, Noah D. Goodman, Christopher Potts |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Animal, dog, or dalmatian? Level of abstraction in nominal referring expressions
Caroline Graf, Judith Degen, Robert D. Hawkins, Noah D. Goodman |
CogSci | 3 |
| 2016 | Conversational expectations account for apparent limits on theory of mind use
Robert D. Hawkins, Noah D. Goodman |
CogSci | 1 |
| 2016 | The Emergence of Conventions
Robert D. Hawkins, Noah D. Goodman, Olga Feher, Kenny Smith, Robert L. Goldstone, Thomas L. Griffiths 0001 |
CogSci | 1 |
| 2015 | Why do you ask? Good questions provoke informative answers
Robert D. Hawkins, Andreas Stuhlmüller, Judith Degen, Noah D. Goodman |
CogSci | 1 |
| 2015 | Emergent Collective Sensing in Human Groups
P. M. Krafft, Robert D. Hawkins, Alex Pentland, Noah D. Goodman, Josh Tenenbaum |
CogSci | 2 |
| 2013 | Real-Time Strategy: Multi-Level Dynamics in an Uncertain Environment
Robert D. Hawkins, Robert L. Goldstone |
CogSci | 1 |