Cameron R. Jones

dblp:377/9416 · also Cameron Robert Jones · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-6609-8966ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs
abstract
Sean Trott, Samuel M. Taylor, Cameron Robert Jones, James A. Michaelov, Pamela D. Rivière. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sean Trott, Samuel M. Taylor, Cameron R. Jones, James A. Michaelov, Pamela D. Rivière
ACL (1)3
2025 Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task
Jared Moore, Rasmus Overmark, Ned Cooper, Beba Cibralic, Nick Haber, Cameron R. Jones
CogSci6
2025 Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?
Zhiqiang Pi, Annapurna Vadaparty, Ben Bergen 0001, Cameron R. Jones
CogSci4
2025 Judging the Judges: Displacing and Inverting the Turing test to Investigate the Interrogator
Ishika Rathi, Ben Bergen 0001, Cameron R. Jones
CogSci3
2025 Does Language Stabilize Quantity Representations in Vision Transformers?
Pamela D. Rivière, Oisin Parkinson-Coombs, Cameron R. Jones, Sean Trott
CogSci3
2024 Does reading words help you to read minds? A comparison of humans and LLMs at a recursive mindreading task
Cameron R. Jones, Sean Trott, Ben Bergen 0001
CogSci1
2024 Multimodal Language Models Show Evidence of Embodied Simulation
abstract
Multimodal large language models (MLLMs) are gaining popularity as partial solutions to the “symbol grounding problem” faced by language models trained on text alone. However, little is known about whether and how these multiple modalities are integrated. We draw inspiration from analogous work in human psycholinguistics on embodied simulation, i.e., the hypothesis that language comprehension is grounded in sensorimotor representations. We show that MLLMs are sensitive to implicit visual features like object shape (e.g., “The egg was in the skillet” implies a frying egg rather than one in a shell). This suggests that MLLMs activate implicit information about object shape when it is implied by a verbal description of an event. We find mixed results for color and orientation, and rule out the possibility that this is due to models’ insensitivity to those features in our dataset overall. We suggest that both human psycholinguistics and computational models of language could benefit from cross-pollination, e.g., with the potential to establish whether grounded representations play a functional role in language processing.
Cameron R. Jones, Sean Trott
LREC/COLING1
2024 Does GPT-4 pass the Turing test?
abstract
We evaluated GPT-4 in a public online Turing test.The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the baseline set by human participants (66%).Participants' decisions were based mainly on linguistic style (35%) and socioemotional traits (27%), supporting the idea that intelligence, narrowly conceived, is not sufficient to pass the Turing test.Participant knowledge about LLMs and number of games played positively correlated with accuracy in detecting AI, suggesting learning and practice as possible strategies to mitigate deception.Despite known limitations as a test of intelligence, we argue that the Turing test continues to be relevant as an assessment of naturalistic communication and deception.AI models with the ability to masquerade as humans could have widespread societal consequences, and we analyse the effectiveness of different strategies and criteria for judging humanlikeness.
Cameron R. Jones, Ben Bergen 0001
NAACL-HLT1
2024 Do Multimodal Large Language Models and Humans Ground Language Similarly?
abstract
Abstract Large Language Models (LLMs) have been criticized for failing to connect linguistic meaning to the world—for failing to solve the “symbol grounding problem.” Multimodal Large Language Models (MLLMs) offer a potential solution to this challenge by combining linguistic representations and processing with other modalities. However, much is still unknown about exactly how and to what degree MLLMs integrate their distinct modalities—and whether the way they do so mirrors the mechanisms believed to underpin grounding in humans. In humans, it has been hypothesized that linguistic meaning is grounded through “embodied simulation,” the activation of sensorimotor and affective representations reflecting described experiences. Across four pre-registered studies, we adapt experimental techniques originally developed to investigate embodied simulation in human comprehenders to ask whether MLLMs are sensitive to sensorimotor features that are implied but not explicit in descriptions of an event. In Experiment 1, we find sensitivity to some features (color and shape) but not others (size, orientation, and volume). In Experiment 2, we identify likely bottlenecks to explain an MLLM’s lack of sensitivity. In Experiment 3, we find that despite sensitivity to implicit sensorimotor features, MLLMs cannot fully account for human behavior on the same task. Finally, in Experiment 4, we compare the psychometric predictive power of different MLLM architectures and find that ViLT, a single-stream architecture, is more predictive of human responses to one sensorimotor feature (shape) than CLIP, a dual-encoder architecture—despite being trained on orders of magnitude less data. These results reveal strengths and limitations in the ability of current MLLMs to integrate language with other modalities, and also shed light on the likely mechanisms underlying human language comprehension.
Cameron R. Jones, Ben Bergen 0001, Sean Trott
Comput. Linguistics1
2024 Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME)
abstract
Abstract We address a growing debate about the extent to which large language models (LLMs) produce behavior consistent with Theory of Mind (ToM) in humans. We present EPITOME: a battery of six experiments that tap diverse ToM capacities, including belief attribution, emotional inference, and pragmatic reasoning. We elicit a performance baseline from human participants for each task. We use the dataset to ask whether distributional linguistic information learned by LLMs is sufficient to explain ToM in humans. We compare performance of five LLMs to a baseline of responses from human comprehenders. Results are mixed. LLMs display considerable sensitivity to mental states and match human performance in several tasks. Yet, they commit systematic errors in others, especially those requiring pragmatic reasoning on the basis of mental state information. Such uneven performance indicates that human-level ToM may require resources beyond distributional information.
Cameron R. Jones, Sean Trott, Ben Bergen 0001
Trans. Assoc. Comput. Linguistics1
2022 Distrubutional Semantics Still Can't Account for Affordances
Cameron R. Jones, Tyler A. Chang, Seana Coulson, James A. Michaelov, Sean Trott, Ben Bergen 0001
CogSci1
2021 The Role of Physical Inference in Pronoun Resolution
Cameron R. Jones, Ben Bergen 0001
CogSci1