EDBT 2026 Demo / reviewers in the wild / expert
Sean Trott
dblp:179/2729
· DBLP profile ↗
20ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-6003-3731ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMsabstractSean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov, Pamela D. Rivière. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sean Trott, Samuel M. Taylor, Cameron R. Jones, James A. Michaelov, Pamela D. Rivière |
ACL (1) | 1 |
| 2026 | Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical AmbiguityabstractAbstract Despite an in-principle understanding of self-attention matrix operations in Transformer language models (LMs), it remains unclear precisely how these operations map onto interpretable computations or functions—and how or when individual attention heads develop specialized attention patterns. Here, we present a pipeline to systematically probe attention mechanisms, and we illustrate its value by leveraging lexical ambiguity—where a single word has multiple meanings—to isolate attention mechanisms that contribute to word sense disambiguation. We take a “developmental” approach: first, using publicly available Pythia LM checkpoints, we identify inflection points in disambiguation performance for each LM in the suite; in 14M and 410M, we identify heads whose attention to disambiguating words covaries with overall disambiguation performance across development. We then stress-test the robustness of these heads to stimulus perturbations: in 14M, we find limited robustness, but in 410M, we identify multiple heads with surprisingly generalizable behavior. Then, in a causal analysis, we find that ablating the target heads demonstrably impairs disambiguation performance, particularly in 14M. We additionally reproduce developmental analyses of 14M across all of its random seeds. Together, these results suggest: that disambiguation benefits from a constellation of mechanisms, some of which (especially in 14M) are highly sensitive to the position and part-of-speech of the disambiguating cue; and that larger models (410M) may contain heads with more robust disambiguation behavior. They also join a growing body of work that highlights the value of adopting a developmental perspective when probing LM mechanisms. Pamela D. Rivière, Sean Trott |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | The role of language in human and machine intelligence
Gary Lupyan, Sean Trott, Martin Zettersten, Hunter Gentry, Thomas L. Griffiths 0001, Anna A. Ivanova |
CogSci | 2 |
| 2025 | Does Language Stabilize Quantity Representations in Vision Transformers?
Pamela D. Rivière, Oisin Parkinson-Coombs, Cameron R. Jones, Sean Trott |
CogSci | 4 |
| 2025 | Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language ModelsabstractRecent studies show that deep vision-only and language-only models—trained on disjoint modalities—nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of where in each network this convergence emerges, what visual or linguistic cues support it, whether it captures human preferences in many-to-many image-text scenarios, and how aggregating exemplars of the same concept affects alignment. Here, we systematically investigate these questions. We find that alignment peaks in mid-to-late layers of both model types, reflecting a shift from modality-specific to conceptually shared representations. This alignment is robust to appearance-only changes but collapses when semantics are altered (e.g., object removal or word-order scrambling), highlighting that the shared code is truly semantic. Moving beyond the one-to-one image-caption paradigm, a forced-choice “Pick-a-Pic” task shows that human preferences for image-caption matches are mirrored in the embedding spaces across all vision-language model pairs. This pattern holds bidirectionally when multiple captions correspond to a single image, demonstrating that models capture fine-grained semantic distinctions akin to human judgments. Surprisingly, averaging embeddings across exemplars amplifies alignment rather than blurring detail. Together, our results demonstrate that unimodal networks converge on a shared semantic code that aligns with human judgments and strengthens with exemplar aggregation. Zoe Wanying He, Sean Trott, Meenakshi Khosla |
EMNLP | 2 |
| 2025 | Evaluating Contextualized Representations of (Spanish) Ambiguous Words: A New Lexical Resource and Empirical AnalysisabstractPamela D Riviere, Anne L. Beatty-Martínez, Sean Trott. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Pamela D. Rivière, Anne L. Beatty-Martínez, Sean Trott |
NAACL (Long Papers) | 3 |
| 2025 | AI-Augmented Predictions: LLM Assistants Improve Human Forecasting AccuracyabstractLarge language models (LLMs) match and sometimes exceed human performance in many domains. This study explores the potential of LLMs to augment human judgment in a forecasting task. We evaluate the effect on human forecasters of two LLM assistants: one designed to provide high-quality (“superforecasting”) advice, and the other designed to be overconfident and base-rate neglecting, thus providing noisy forecasting advice. We compare participants using these assistants to a control group that received a less advanced model that did not provide numerical predictions or engage in explicit discussion of predictions. Participants ( N \(=\) 991) answered a set of six forecasting questions and had the option to consult their assigned LLM assistant throughout. Our preregistered analyses show that interacting with each of our frontier LLM assistants significantly enhances prediction accuracy by between 24% and 28% compared to the control group. Exploratory analyses showed a pronounced outlier effect in one forecasting item, without which we find that the superforecasting assistant increased accuracy by 41%, compared with 29% for the noisy assistant. We further examine whether LLM forecasting augmentation disproportionately benefits less skilled forecasters, degrades the wisdom-of-the-crowd by reducing prediction diversity, or varies in effectiveness with question difficulty. Our data do not consistently support these hypotheses. Our results suggest that access to a frontier LLM assistant, even a noisy one, can be a helpful decision aid in cognitively demanding tasks compared to a less powerful model that does not provide specific forecasting advice. However, the effects of outliers suggest that further research into the robustness of this pattern is needed. Philipp Schoenegger, Peter S. Park, Ezra Karger, Sean Trott, Philip Tetlock |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2024 | Does reading words help you to read minds? A comparison of humans and LLMs at a recursive mindreading task
Cameron R. Jones, Sean Trott, Ben Bergen 0001 |
CogSci | 2 |
| 2024 | Context-dependent and Dynamic Effects of Distributional and Sensorimotor Distance Measures on EEG
Harshada Vinaya, Sean Trott, Diane Pecher, René Zeelenberg, Seana Coulson |
CogSci | 2 |
| 2024 | Multimodal Language Models Show Evidence of Embodied SimulationabstractMultimodal large language models (MLLMs) are gaining popularity as partial solutions to the “symbol grounding problem” faced by language models trained on text alone. However, little is known about whether and how these multiple modalities are integrated. We draw inspiration from analogous work in human psycholinguistics on embodied simulation, i.e., the hypothesis that language comprehension is grounded in sensorimotor representations. We show that MLLMs are sensitive to implicit visual features like object shape (e.g., “The egg was in the skillet” implies a frying egg rather than one in a shell). This suggests that MLLMs activate implicit information about object shape when it is implied by a verbal description of an event. We find mixed results for color and orientation, and rule out the possibility that this is due to models’ insensitivity to those features in our dataset overall. We suggest that both human psycholinguistics and computational models of language could benefit from cross-pollination, e.g., with the potential to establish whether grounded representations play a functional role in language processing. Cameron R. Jones, Sean Trott |
LREC/COLING | 2 |
| 2024 | Do Multimodal Large Language Models and Humans Ground Language Similarly?abstractAbstract Large Language Models (LLMs) have been criticized for failing to connect linguistic meaning to the world—for failing to solve the “symbol grounding problem.” Multimodal Large Language Models (MLLMs) offer a potential solution to this challenge by combining linguistic representations and processing with other modalities. However, much is still unknown about exactly how and to what degree MLLMs integrate their distinct modalities—and whether the way they do so mirrors the mechanisms believed to underpin grounding in humans. In humans, it has been hypothesized that linguistic meaning is grounded through “embodied simulation,” the activation of sensorimotor and affective representations reflecting described experiences. Across four pre-registered studies, we adapt experimental techniques originally developed to investigate embodied simulation in human comprehenders to ask whether MLLMs are sensitive to sensorimotor features that are implied but not explicit in descriptions of an event. In Experiment 1, we find sensitivity to some features (color and shape) but not others (size, orientation, and volume). In Experiment 2, we identify likely bottlenecks to explain an MLLM’s lack of sensitivity. In Experiment 3, we find that despite sensitivity to implicit sensorimotor features, MLLMs cannot fully account for human behavior on the same task. Finally, in Experiment 4, we compare the psychometric predictive power of different MLLM architectures and find that ViLT, a single-stream architecture, is more predictive of human responses to one sensorimotor feature (shape) than CLIP, a dual-encoder architecture—despite being trained on orders of magnitude less data. These results reveal strengths and limitations in the ability of current MLLMs to integrate language with other modalities, and also shed light on the likely mechanisms underlying human language comprehension. Cameron R. Jones, Ben Bergen 0001, Sean Trott |
Comput. Linguistics | 3 |
| 2024 | Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME)abstractAbstract We address a growing debate about the extent to which large language models (LLMs) produce behavior consistent with Theory of Mind (ToM) in humans. We present EPITOME: a battery of six experiments that tap diverse ToM capacities, including belief attribution, emotional inference, and pragmatic reasoning. We elicit a performance baseline from human participants for each task. We use the dataset to ask whether distributional linguistic information learned by LLMs is sufficient to explain ToM in humans. We compare performance of five LLMs to a baseline of responses from human comprehenders. Results are mixed. LLMs display considerable sensitivity to mental states and match human performance in several tasks. Yet, they commit systematic errors in others, especially those requiring pragmatic reasoning on the basis of mental state information. Such uneven performance indicates that human-level ToM may require resources beyond distributional information. Cameron R. Jones, Sean Trott, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | Distrubutional Semantics Still Can't Account for Affordances
Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, Ben Bergen 0001 |
CogSci | 5 |
| 2022 | Can a pressure against homophones explain phonological neighborhoods?
Sean Trott, Ben Bergen 0001 |
CogSci | 1 |
| 2021 | RAW-C: Relatedness of Ambiguous Words in Context (A New Lexical Resource for English)abstractSean Trott, Benjamin Bergen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sean Trott, Ben Bergen 0001 |
ACL/IJCNLP (1) | 1 |
| 2020 | (Re)construing Meaning in NLPabstractHuman speakers have an extensive toolkit of ways to express themselves.In this paper, we engage with an idea largely absent from discussions of meaning in natural language understanding-namely, that the way something is expressed reflects different ways of conceptualizing or construing the information being conveyed.We first define this phenomenon more precisely, drawing on considerable prior work in theoretical cognitive semantics and psycholinguistics.We then survey some dimensions of construed meaning and show how insights from construal could inform theoretical and practical work in NLP. Sean Trott, Tiago Timponi Torrent, Nancy Chang, Nathan Schneider 0001 |
ACL | 1 |
| 2020 | Effects of Battle and Journey Metaphors on Charitable Donations for Cancer Patients
Alex Liebscher, Sean Trott, Ben Bergen 0001 |
CogSci | 2 |
| 2019 | Prosodic cues signal the intent of potential indirect requests
Sean Trott, Stefanie Reed, Victor Ferreira, Ben Bergen 0001 |
CogSci | 1 |
| 2019 | Sub-morphemic form-meaning systematicity: the impact of onset phones on word concreteness
Sean Trott, Arturs Semenuks, Ben Bergen 0001 |
CogSci | 1 |
| 2016 | Exploiting deep semantics and compositionality of natural language for Human-Robot-InteractionabstractWe are developing a natural language interface for human robot interaction that implements reasoning about deep semantics in natural language. To realize the required deep analysis, we employ methods from cognitive linguistics, namely the modular and compositional framework of Embodied Construction Grammar (ECG) [18]. Using ECG, robots are able to solve fine-grained reference resolution problems and other issues related to deep semantics and compositionality of natural language. This also includes verbal interaction with humans to clarify commands and queries that are too ambiguous to be executed safely. We implement our NLU framework as a ROS package and present proof-of-concept scenarios with different robots, as well as a survey on the state of the art in knowledge-based language HRI. Manfred Eppe, Sean Trott, Jerome A. Feldman |
IROS | 2 |