EDBT 2026 Demo / reviewers in the wild / expert
Cameron R. Jones
dblp:377/9416 · also Cameron Robert Jones
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-6609-8966ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMsabstractSean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov, Pamela D. Rivière. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sean Trott, Samuel M. Taylor, Cameron R. Jones, James A. Michaelov, Pamela D. Rivière |
ACL (1) | 3 |
| 2025 | Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task
Jared Moore, Rasmus Overmark, Ned Cooper, Beba Cibralic, Nick Haber, Cameron R. Jones |
CogSci | 6 |
| 2025 | Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?
Zhiqiang Pi, Annapurna Vadaparty, Ben Bergen 0001, Cameron R. Jones |
CogSci | 4 |
| 2025 | Judging the Judges: Displacing and Inverting the Turing test to Investigate the Interrogator
Ishika Rathi, Ben Bergen 0001, Cameron R. Jones |
CogSci | 3 |
| 2025 | Does Language Stabilize Quantity Representations in Vision Transformers?
Pamela D. Rivière, Oisin Parkinson-Coombs, Cameron R. Jones, Sean Trott |
CogSci | 3 |
| 2024 | Does reading words help you to read minds? A comparison of humans and LLMs at a recursive mindreading task
Cameron R. Jones, Sean Trott, Ben Bergen 0001 |
CogSci | 1 |
| 2024 | Multimodal Language Models Show Evidence of Embodied SimulationabstractMultimodal large language models (MLLMs) are gaining popularity as partial solutions to the “symbol grounding problem” faced by language models trained on text alone. However, little is known about whether and how these multiple modalities are integrated. We draw inspiration from analogous work in human psycholinguistics on embodied simulation, i.e., the hypothesis that language comprehension is grounded in sensorimotor representations. We show that MLLMs are sensitive to implicit visual features like object shape (e.g., “The egg was in the skillet” implies a frying egg rather than one in a shell). This suggests that MLLMs activate implicit information about object shape when it is implied by a verbal description of an event. We find mixed results for color and orientation, and rule out the possibility that this is due to models’ insensitivity to those features in our dataset overall. We suggest that both human psycholinguistics and computational models of language could benefit from cross-pollination, e.g., with the potential to establish whether grounded representations play a functional role in language processing. Cameron R. Jones, Sean Trott |
LREC/COLING | 1 |
| 2024 | Does GPT-4 pass the Turing test?abstractWe evaluated GPT-4 in a public online Turing test.The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the baseline set by human participants (66%).Participants' decisions were based mainly on linguistic style (35%) and socioemotional traits (27%), supporting the idea that intelligence, narrowly conceived, is not sufficient to pass the Turing test.Participant knowledge about LLMs and number of games played positively correlated with accuracy in detecting AI, suggesting learning and practice as possible strategies to mitigate deception.Despite known limitations as a test of intelligence, we argue that the Turing test continues to be relevant as an assessment of naturalistic communication and deception.AI models with the ability to masquerade as humans could have widespread societal consequences, and we analyse the effectiveness of different strategies and criteria for judging humanlikeness. Cameron R. Jones, Ben Bergen 0001 |
NAACL-HLT | 1 |
| 2024 | Do Multimodal Large Language Models and Humans Ground Language Similarly?abstractAbstract Large Language Models (LLMs) have been criticized for failing to connect linguistic meaning to the world—for failing to solve the “symbol grounding problem.” Multimodal Large Language Models (MLLMs) offer a potential solution to this challenge by combining linguistic representations and processing with other modalities. However, much is still unknown about exactly how and to what degree MLLMs integrate their distinct modalities—and whether the way they do so mirrors the mechanisms believed to underpin grounding in humans. In humans, it has been hypothesized that linguistic meaning is grounded through “embodied simulation,” the activation of sensorimotor and affective representations reflecting described experiences. Across four pre-registered studies, we adapt experimental techniques originally developed to investigate embodied simulation in human comprehenders to ask whether MLLMs are sensitive to sensorimotor features that are implied but not explicit in descriptions of an event. In Experiment 1, we find sensitivity to some features (color and shape) but not others (size, orientation, and volume). In Experiment 2, we identify likely bottlenecks to explain an MLLM’s lack of sensitivity. In Experiment 3, we find that despite sensitivity to implicit sensorimotor features, MLLMs cannot fully account for human behavior on the same task. Finally, in Experiment 4, we compare the psychometric predictive power of different MLLM architectures and find that ViLT, a single-stream architecture, is more predictive of human responses to one sensorimotor feature (shape) than CLIP, a dual-encoder architecture—despite being trained on orders of magnitude less data. These results reveal strengths and limitations in the ability of current MLLMs to integrate language with other modalities, and also shed light on the likely mechanisms underlying human language comprehension. Cameron R. Jones, Ben Bergen 0001, Sean Trott |
Comput. Linguistics | 1 |
| 2024 | Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME)abstractAbstract We address a growing debate about the extent to which large language models (LLMs) produce behavior consistent with Theory of Mind (ToM) in humans. We present EPITOME: a battery of six experiments that tap diverse ToM capacities, including belief attribution, emotional inference, and pragmatic reasoning. We elicit a performance baseline from human participants for each task. We use the dataset to ask whether distributional linguistic information learned by LLMs is sufficient to explain ToM in humans. We compare performance of five LLMs to a baseline of responses from human comprehenders. Results are mixed. LLMs display considerable sensitivity to mental states and match human performance in several tasks. Yet, they commit systematic errors in others, especially those requiring pragmatic reasoning on the basis of mental state information. Such uneven performance indicates that human-level ToM may require resources beyond distributional information. Cameron R. Jones, Sean Trott, Ben Bergen 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Distrubutional Semantics Still Can't Account for Affordances
Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, Ben Bergen 0001 |
CogSci | 1 |
| 2021 | The Role of Physical Inference in Pronoun Resolution
Cameron R. Jones, Ben Bergen 0001 |
CogSci | 1 |