Tomer D. Ullman

dblp:64/10432 · also Tomer David Ullman · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0003-1722-2382ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 8 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Making Sense of Nonsense
Jennifer Hu 0001, Felix Sosa, Tomer D. Ullman
CogSci3
2025 The Uncanny Valley meets the Humorous Hill: Things are funny when they match a pattern but fall short on quality
Antara Raaghavi Bhattacharya, Jennifer Hu 0001, Tomer D. Ullman
CogSci3
2025 Stumped! Learning to think outside the box in 3-7 year old children
Junyi Chu, Misha O'Keeffe, Silvia Kancong Liu, Elizabeth Baraff Bonawitz, Tomer D. Ullman
CogSci5
2025 Intuitions about prosocial backfiring: Four to seven-year olds' understanding of when helping might cause offense
Kiera Parece, Tomer D. Ullman, Laura Schulz
CogSci2
2025 Language models assign responsibility based on actual rather than counterfactual contributions
Eric J. Bigelow, Tobias Gerstenberg, Tomer D. Ullman, Samuel Gershman
CogSci4
2025 Forking Paths in Neural Text Generation
abstract
Estimating uncertainty in Large Language Models (LLMs) is important for properly evaluating LLMs, and ensuring safety for users. However, prior approaches to uncertainty estimation focus on the final answer in generated text, ignoring intermediate steps that might dramatically impact the outcome. We hypothesize that there exist key forking tokens, such that re-sampling the system at those specific tokens, but not others, leads to very different outcomes. To test this empirically, we develop a novel approach to representing uncertainty dynamics across individual tokens of text generation, and applying statistical models to test our hypothesis. Our approach is highly flexible: it can be applied to any dataset and any LLM, without fine tuning or accessing model weights. We use our method to analyze LLM responses on 7 different tasks across 4 domains, spanning a wide range of typical use cases. We find many examples of forking tokens, including surprising ones such as a space character instead of a colon, suggesting that LLMs are often just a single token away from saying something very different.
Eric J. Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer D. Ullman
ICLR4
2025 One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity
abstract
Sonia Krishna Murthy, Tomer Ullman, Jennifer Hu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sonia K. Murthy, Tomer D. Ullman, Jennifer Hu 0001
NAACL (Long Papers)2
2024 MMToM-QA: Multimodal Theory of Mind Question Answering
abstract
Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, Tianmin Shu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chuanyang Jin, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum, Tianmin Shu
ACL (1)7
2024 Shades of Zero: Distinguishing impossibility from inconceivability
Jennifer Hu 0001, Felix Sosa, Tomer D. Ullman
CogSci3
2024 The Task Task: Creative problem generation in humans and language models
Junyi Chu, Jennifer Hu 0001, Tomer D. Ullman
CogSci3
2024 Exploring Loophole Behavior: A Comparative Study of Autistic and Non-Autistic Populations
Kiera Parece, Sophie Bridgers, Tomer D. Ullman, Laura Schulz
CogSci3
2024 In-Context Learning Dynamics with Random Binary Sequences
abstract
Large language models (LLMs) trained on huge text datasets demonstrate intriguing capabilities, achieving state-of-the-art performance on tasks they were not explicitly trained for. The precise nature of LLM capabilities is often mysterious, and different prompts can elicit different capabilities through in-context learning. We propose a framework that enables us to analyze in-context learning dynamics to understand latent concepts underlying LLMs’ behavioral patterns. This provides a more nuanced understanding than success-or-failure evaluation benchmarks, but does not require observing internal activations as a mechanistic interpretation of circuits would. Inspired by the cognitive science of human randomness perception, we use random binary sequences as context and study dynamics of in-context learning by manipulating properties of context data, such as sequence length. In the latest GPT-3.5+ models, we find emergent abilities to generate seemingly random numbers and learn basic formal languages, with striking in-context learning dynamics where model outputs transition sharply from seemingly random behaviors to deterministic repetition.
Eric J. Bigelow, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Tomer D. Ullman
ICLR5
2024 Approximate planning in spatial search
abstract
How people plan is an active area of research in cognitive science, neuroscience, and artificial intelligence. However, tasks traditionally used to study planning in the laboratory tend to be constrained to artificial environments, such as Chess and bandit problems. To date there is still no agreed-on model of how people plan in realistic contexts, such as navigation and search, where values intuitively derive from interactions between perception and cognition. To address this gap and move towards a more naturalistic study of planning, we present a novel spatial Maze Search Task (MST) where the costs and rewards are physically situated as distances and locations. We used this task in two behavioral experiments to evaluate and contrast multiple distinct computational models of planning, including optimal expected utility planning, several one-step heuristics inspired by studies of information search, and a family of planners that deviate from optimal planning, in which action values are estimated by the interactions between perception and cognition. We found that people's deviations from optimal expected utility are best explained by planners with a limited horizon, however our results do not exclude the possibility that in human planning action values may be also affected by cognitive mechanisms of numerosity and probability perception. This result makes a novel theoretical contribution in showing that limited planning horizon generalizes to spatial planning, and demonstrates the value of our multi-model approach for understanding cognition.
Marta Kryven, Suhyoun Yu, Max Kleiman-Weiner, Tomer D. Ullman, Josh Tenenbaum
PLoS Comput. Biol.4
2023 Skirting the Sacred: Moral Violations Make Intentional Misunderstandings Worse
Kiera Parece, Sophie Bridgers, Laura Schulz, Tomer D. Ullman
CogSci4
2022 Adults' Evaluations of Rote and Reflective Teachers
Ilona Bass, Elizabeth Baraff Bonawitz, Tomer D. Ullman
CogSci3
2022 People's evaluation of programs that drive agents' behavior
Eric J. Bigelow, Tomer D. Ullman
CogSci2
2022 Combining mental simulation and abstract reasoning explains people's reaction time in an intuitive physics task
Felix Sosa, Samuel Gershman, Tomer D. Ullman
CogSci3
2021 Loopholes, a Window into Value Alignment and the Learning of Meaning
Sophie Bridgers, Laura Schulz, Tomer D. Ullman
CogSci3
2021 Unsupervised Discovery of 3D Physical Objects from Video
Yilun Du, Kevin A. Smith 0001, Tomer D. Ullman, Josh Tenenbaum, Jiajun Wu 0001
ICLR3
2021 AGENT: A Benchmark for Core Psychological Reasoning
abstract
For machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants can tell agents from objects, expecting agents to act efficiently to achieve goals given constraints. Despite recent interest in machine agents that reason about other agents, it is not clear if such agents learn or hold the core psychology principles that drive human reasoning. Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four scenarios (goal preferences, action efficiency, unobserved constraints, and cost-reward trade-offs) that probe key concepts of core intuitive psychology. We validate AGENT with human-ratings, propose an evaluation protocol emphasizing generalization, and compare two strong baselines built on Bayesian inverse planning and a Theory of Mind neural network. Our results suggest that to pass the designed tests of core intuitive psychology at human levels, a model must acquire or have built-in representations of how agents plan, combining utility computations and core knowledge of objects and physics.
Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan 0001, Kevin A. Smith 0001, Shari Liu, Dan Gutfreund, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman
ICML9
2021 Temporal and Object Quantification Networks
abstract
We present Temporal and Object Quantification Networks (TOQ-Nets), a new class of neuro-symbolic networks with a structural bias that enables them to learn to recognize complex relational-temporal events. This is done by including reasoning layers that implement finite-domain quantification over objects and time. The structure allows them to generalize directly to input instances with varying numbers of objects in temporal sequences of varying lengths. We evaluate TOQ-Nets on input domains that require recognizing event-types in terms of complex temporal relational patterns. We demonstrate that TOQ-Nets can generalize from small amounts of data to scenarios containing more objects than were present during training and to temporal warpings of input sequences.
Jiayuan Mao, Zhezheng Luo, Chuang Gan 0001, Josh Tenenbaum, Jiajun Wu 0001, Leslie Pack Kaelbling, Tomer D. Ullman
IJCAI7
2021 A Bayesian-Symbolic Approach to Reasoning and Learning in Intuitive Physics
abstract
Humans can reason about intuitive physics in fully or partially observed environments even after being exposed to a very limited set of observations. This sample-efficient intuitive physical reasoning is considered a core domain of human common sense knowledge. One hypothesis to explain this remarkable capacity, posits that humans quickly learn approximations to the laws of physics that govern the dynamics of the environment. In this paper, we propose a Bayesian-symbolic framework (BSP) for physical reasoning and learning that is close to human-level sample-efficiency and accuracy. In BSP, the environment is represented by a top-down generative model of entities, which are assumed to interact with each other under unknown force laws over their latent and observed properties. BSP models each of these entities as random variables, and uses Bayesian inference to estimate their unknown properties. For learning the unknown forces, BSP leverages symbolic regression on a novel grammar of Newtonian physics in a bilevel optimization setup. These inference and regression steps are performed in an iterative manner using expectation-maximization, allowing BSP to simultaneously learn force laws while maintaining uncertainty over entity properties. We show that BSP is more sample-efficient compared to neural alternatives on controlled synthetic datasets, demonstrate BSP's applicability to real-world common sense scenes and study BSP's performance on tasks previously used to study human physical reasoning.
Kai Xu 0016, Akash Srivastava, Dan Gutfreund, Felix Sosa, Tomer D. Ullman, Josh Tenenbaum, Charles Sutton
NeurIPS5
2020 The Origins of Common Sense in Humans and Machines
Kevin A. Smith 0001, Eliza Kosoy, Alison Gopnik, Deepak Pathak, Alan Fern, Josh Tenenbaum, Tomer D. Ullman
CogSci7
2020 The fine structure of surprise in intuitive physics: when, why, and how much?
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman
CogSci7
2020 Look before you leap: Quantitative tradeoffs between peril and reward in action understanding
Nensi Gjata, Tomer D. Ullman, Elizabeth S. Spelke, Shari Liu
CogSci2
2020 Adventures in Flatland: Perceiving Social Interactions Under Physical Dynamics
Tianmin Shu, Marta Kryven, Tomer D. Ullman, Josh Tenenbaum
CogSci3
2019 People's perception of others' risk preferences
Shari Liu, John McCoy, Tomer D. Ullman
CogSci3
2019 Draping an Elephant: Uncovering Children's Reasoning About Cloth-Covered Objects
Tomer D. Ullman, Eliza Kosoy, Ilker Yildirim, Amir Arsalan Soltani, Max H. Siegel, Josh Tenenbaum, Elizabeth S. Spelke
CogSci1
2019 Modeling Expectation Violation in Intuitive Physics with Coarse Probabilistic Object Representations
abstract
From infancy, humans have expectations about how objects will move and interact. Even young children expect objects not to move through one another, teleport, or disappear. They are surprised by mismatches between physical expectations and perceptual observations, even in unfamiliar scenes with completely novel objects. A model that exhibits human-like understanding of physics should be similarly surprised, and adjust its beliefs accordingly. We propose ADEPT, a model that uses a coarse (approximate geometry) object-centric representation for dynamic 3D scene understanding. Inference integrates deep recognition networks, extended probabilistic physical simulation, and particle filtering for forming predictions and expectations across occlusion. We also present a new test set for measuring violations of physical expectations, using a range of scenarios derived from developmental psychology. We systematically compare ADEPT, baseline models, and human expectations on this test set. ADEPT outperforms standard network architectures in discriminating physically implausible scenes, and often performs this discrimination at the same level as people.
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman
NeurIPS7
2018 Models of Human Scientific Discovery
Robert L. Goldstone, Alison Gopnik, Paul Thagard, Tomer D. Ullman
CogSci4
2018 Moral Dynamics: A Computational Model of Moral Judgment
Felix Sosa, Tomer D. Ullman, Samuel Gershman, Josh Tenenbaum, Tobias Gerstenberg
CogSci2
2017 Thinking and Guessing: Bayesian and Empirical Models of How Humans Search
Marta Kryven, Tomer D. Ullman, William Cowan, Josh Tenenbaum
CogSci2
2017 What's worth the effort: Ten-month-old infants infer the value of goals from the costs of actions
Shari Liu, Tomer D. Ullman, Josh Tenenbaum, Elizabeth S. Spelke
CogSci2
2017 Weight matters: The role of physical weight in non-physical language across age and culture
Tomer D. Ullman, Santiago Alonso-Diaz, Stephen Ferrigno, Sarina Zahid, Celeste Kidd
CogSci1
2017 A Compositional Object-Based Approach to Learning Physical Dynamics
Michael Chang 0003, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum
ICLR (Poster)2
2016 Outcome or Strategy? A Bayesian Model of Intelligence Attribution
Marta Kryven, Tomer D. Ullman, William Cowan, Josh Tenenbaum
CogSci2
2016 Coalescing the Vapors of Human Experience into a Viable and Meaningful Comprehension
Tomer D. Ullman, Max H. Siegel, Josh Tenenbaum, Samuel Gershman
CogSci1
2016 The Pragmatics of Spatial Language
Tomer D. Ullman, Yang Xu 0023, Noah D. Goodman
CogSci1
2014 Wins above replacement: Responsibility attributions as counterfactual replacements
Tobias Gerstenberg, Tomer D. Ullman, Max Kleiman-Weiner, David A. Lagnado, Josh Tenenbaum
CogSci2
2014 Learning physical theories from dynamical scenes
Tomer D. Ullman, Andreas Stuhlmüller, Noah D. Goodman, Josh Tenenbaum
CogSci1
2013 Minimal Nativism: How does cognitive development get off the ground?
Tomer D. Ullman, Josh Tenenbaum, Noah D. Goodman, Shimon Ullman, Elizabeth S. Spelke
CogSci1
2012 Computational Models of Intuitive Physics
Peter W. Battaglia, Tomer D. Ullman, Josh Tenenbaum, Adam Sanborn, Kenneth D. Forbus, Tobias Gerstenberg, David A. Lagnado
CogSci2
2012 Probabilistic generative models for counterfactual reasoning and blame attribution
John McCoy, Tomer D. Ullman, Andreas Stuhlmüller, Tobias Gerstenberg, Josh Tenenbaum
CogSci2
2011 Forward Physics: How people learn and generalize novel dynamical models
Tomer D. Ullman, Noah D. Goodman, Josh Tenenbaum
CogSci1
2009 Help or Hinder: Bayesian Models of Social Goal Inference
abstract
Everyday social interactions are heavily influenced by our snap judgments about others goals. Even young infants can infer the goals of intentional agents from observing how they interact with objects and other agents in their environment: e.g., that one agent is helping orhindering anothers attempt to get up a hill or open a box. We propose a model for how people can infer these social goals from actions, based on inverse planning in multiagent Markov decision problems (MDPs). The model infers the goal most likely to be driving an agents behavior by assuming the agent acts approximately rationally given environmental constraints and its model of other agents present. We also present behavioral evidence in support of this model over a simpler, perceptual cue-based alternative.
Tomer D. Ullman, Chris L. Baker, Owen Macindoe, Owain Evans, Noah D. Goodman, Josh Tenenbaum
NIPS1