EDBT 2026 Demo / reviewers in the wild / expert
Tomer D. Ullman
dblp:64/10432 · also Tomer David Ullman
· DBLP profile ↗
45ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0003-1722-2382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 8 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Making Sense of Nonsense
Jennifer Hu 0001, Felix Sosa, Tomer D. Ullman |
CogSci | 3 |
| 2025 | The Uncanny Valley meets the Humorous Hill: Things are funny when they match a pattern but fall short on quality
Antara Raaghavi Bhattacharya, Jennifer Hu 0001, Tomer D. Ullman |
CogSci | 3 |
| 2025 | Stumped! Learning to think outside the box in 3-7 year old children
Junyi Chu, Misha O'Keeffe, Silvia Kancong Liu, Elizabeth Baraff Bonawitz, Tomer D. Ullman |
CogSci | 5 |
| 2025 | Intuitions about prosocial backfiring: Four to seven-year olds' understanding of when helping might cause offense
Kiera Parece, Tomer D. Ullman, Laura Schulz |
CogSci | 2 |
| 2025 | Language models assign responsibility based on actual rather than counterfactual contributions
Eric J. Bigelow, Tobias Gerstenberg, Tomer D. Ullman, Samuel Gershman |
CogSci | 4 |
| 2025 | Forking Paths in Neural Text GenerationabstractEstimating uncertainty in Large Language Models (LLMs) is important for properly evaluating LLMs, and ensuring safety for users. However, prior approaches to uncertainty estimation focus on the final answer in generated text, ignoring intermediate steps that might dramatically impact the outcome. We hypothesize that there exist key forking tokens, such that re-sampling the system at those specific tokens, but not others, leads to very different outcomes. To test this empirically, we develop a novel approach to representing uncertainty dynamics across individual tokens of text generation, and applying statistical models to test our hypothesis. Our approach is highly flexible: it can be applied to any dataset and any LLM, without fine tuning or accessing model weights. We use our method to analyze LLM responses on 7 different tasks across 4 domains, spanning a wide range of typical use cases. We find many examples of forking tokens, including surprising ones such as a space character instead of a colon, suggesting that LLMs are often just a single token away from saying something very different. Eric J. Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer D. Ullman |
ICLR | 4 |
| 2025 | One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversityabstractSonia Krishna Murthy, Tomer Ullman, Jennifer Hu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sonia K. Murthy, Tomer D. Ullman, Jennifer Hu 0001 |
NAACL (Long Papers) | 2 |
| 2024 | MMToM-QA: Multimodal Theory of Mind Question AnsweringabstractChuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, Tianmin Shu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chuanyang Jin, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum, Tianmin Shu |
ACL (1) | 7 |
| 2024 | Shades of Zero: Distinguishing impossibility from inconceivability
Jennifer Hu 0001, Felix Sosa, Tomer D. Ullman |
CogSci | 3 |
| 2024 | The Task Task: Creative problem generation in humans and language models
Junyi Chu, Jennifer Hu 0001, Tomer D. Ullman |
CogSci | 3 |
| 2024 | Exploring Loophole Behavior: A Comparative Study of Autistic and Non-Autistic Populations
Kiera Parece, Sophie Bridgers, Tomer D. Ullman, Laura Schulz |
CogSci | 3 |
| 2024 | In-Context Learning Dynamics with Random Binary SequencesabstractLarge language models (LLMs) trained on huge text datasets demonstrate intriguing capabilities, achieving state-of-the-art performance on tasks they were not explicitly trained for. The precise nature of LLM capabilities is often mysterious, and different prompts can elicit different capabilities through in-context learning. We propose a framework that enables us to analyze in-context learning dynamics to understand latent concepts underlying LLMs’ behavioral patterns. This provides a more nuanced understanding than success-or-failure evaluation benchmarks, but does not require observing internal activations as a mechanistic interpretation of circuits would. Inspired by the cognitive science of human randomness perception, we use random binary sequences as context and study dynamics of in-context learning by manipulating properties of context data, such as sequence length. In the latest GPT-3.5+ models, we find emergent abilities to generate seemingly random numbers and learn basic formal languages, with striking in-context learning dynamics where model outputs transition sharply from seemingly random behaviors to deterministic repetition. Eric J. Bigelow, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Tomer D. Ullman |
ICLR | 5 |
| 2024 | Approximate planning in spatial searchabstractHow people plan is an active area of research in cognitive science, neuroscience, and artificial intelligence. However, tasks traditionally used to study planning in the laboratory tend to be constrained to artificial environments, such as Chess and bandit problems. To date there is still no agreed-on model of how people plan in realistic contexts, such as navigation and search, where values intuitively derive from interactions between perception and cognition. To address this gap and move towards a more naturalistic study of planning, we present a novel spatial Maze Search Task (MST) where the costs and rewards are physically situated as distances and locations. We used this task in two behavioral experiments to evaluate and contrast multiple distinct computational models of planning, including optimal expected utility planning, several one-step heuristics inspired by studies of information search, and a family of planners that deviate from optimal planning, in which action values are estimated by the interactions between perception and cognition. We found that people's deviations from optimal expected utility are best explained by planners with a limited horizon, however our results do not exclude the possibility that in human planning action values may be also affected by cognitive mechanisms of numerosity and probability perception. This result makes a novel theoretical contribution in showing that limited planning horizon generalizes to spatial planning, and demonstrates the value of our multi-model approach for understanding cognition. Marta Kryven, Suhyoun Yu, Max Kleiman-Weiner, Tomer D. Ullman, Josh Tenenbaum |
PLoS Comput. Biol. | 4 |
| 2023 | Skirting the Sacred: Moral Violations Make Intentional Misunderstandings Worse
Kiera Parece, Sophie Bridgers, Laura Schulz, Tomer D. Ullman |
CogSci | 4 |
| 2022 | Adults' Evaluations of Rote and Reflective Teachers
Ilona Bass, Elizabeth Baraff Bonawitz, Tomer D. Ullman |
CogSci | 3 |
| 2022 | People's evaluation of programs that drive agents' behavior
Eric J. Bigelow, Tomer D. Ullman |
CogSci | 2 |
| 2022 | Combining mental simulation and abstract reasoning explains people's reaction time in an intuitive physics task
Felix Sosa, Samuel Gershman, Tomer D. Ullman |
CogSci | 3 |
| 2021 | Loopholes, a Window into Value Alignment and the Learning of Meaning
Sophie Bridgers, Laura Schulz, Tomer D. Ullman |
CogSci | 3 |
| 2021 | Unsupervised Discovery of 3D Physical Objects from Video
Yilun Du, Kevin A. Smith 0001, Tomer D. Ullman, Josh Tenenbaum, Jiajun Wu 0001 |
ICLR | 3 |
| 2021 | AGENT: A Benchmark for Core Psychological ReasoningabstractFor machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants can tell agents from objects, expecting agents to act efficiently to achieve goals given constraints. Despite recent interest in machine agents that reason about other agents, it is not clear if such agents learn or hold the core psychology principles that drive human reasoning. Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four scenarios (goal preferences, action efficiency, unobserved constraints, and cost-reward trade-offs) that probe key concepts of core intuitive psychology. We validate AGENT with human-ratings, propose an evaluation protocol emphasizing generalization, and compare two strong baselines built on Bayesian inverse planning and a Theory of Mind neural network. Our results suggest that to pass the designed tests of core intuitive psychology at human levels, a model must acquire or have built-in representations of how agents plan, combining utility computations and core knowledge of objects and physics. Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan 0001, Kevin A. Smith 0001, Shari Liu, Dan Gutfreund, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
ICML | 9 |
| 2021 | Temporal and Object Quantification NetworksabstractWe present Temporal and Object Quantification Networks (TOQ-Nets), a new class of neuro-symbolic networks with a structural bias that enables them to learn to recognize complex relational-temporal events. This is done by including reasoning layers that implement finite-domain quantification over objects and time. The structure allows them to generalize directly to input instances with varying numbers of objects in temporal sequences of varying lengths. We evaluate TOQ-Nets on input domains that require recognizing event-types in terms of complex temporal relational patterns. We demonstrate that TOQ-Nets can generalize from small amounts of data to scenarios containing more objects than were present during training and to temporal warpings of input sequences. Jiayuan Mao, Zhezheng Luo, Chuang Gan 0001, Josh Tenenbaum, Jiajun Wu 0001, Leslie Pack Kaelbling, Tomer D. Ullman |
IJCAI | 7 |
| 2021 | A Bayesian-Symbolic Approach to Reasoning and Learning in Intuitive PhysicsabstractHumans can reason about intuitive physics in fully or partially observed environments even after being exposed to a very limited set of observations. This sample-efficient intuitive physical reasoning is considered a core domain of human common sense knowledge. One hypothesis to explain this remarkable capacity, posits that humans quickly learn approximations to the laws of physics that govern the dynamics of the environment. In this paper, we propose a Bayesian-symbolic framework (BSP) for physical reasoning and learning that is close to human-level sample-efficiency and accuracy. In BSP, the environment is represented by a top-down generative model of entities, which are assumed to interact with each other under unknown force laws over their latent and observed properties. BSP models each of these entities as random variables, and uses Bayesian inference to estimate their unknown properties. For learning the unknown forces, BSP leverages symbolic regression on a novel grammar of Newtonian physics in a bilevel optimization setup. These inference and regression steps are performed in an iterative manner using expectation-maximization, allowing BSP to simultaneously learn force laws while maintaining uncertainty over entity properties. We show that BSP is more sample-efficient compared to neural alternatives on controlled synthetic datasets, demonstrate BSP's applicability to real-world common sense scenes and study BSP's performance on tasks previously used to study human physical reasoning. Kai Xu 0016, Akash Srivastava, Dan Gutfreund, Felix Sosa, Tomer D. Ullman, Josh Tenenbaum, Charles Sutton |
NeurIPS | 5 |
| 2020 | The Origins of Common Sense in Humans and Machines
Kevin A. Smith 0001, Eliza Kosoy, Alison Gopnik, Deepak Pathak, Alan Fern, Josh Tenenbaum, Tomer D. Ullman |
CogSci | 7 |
| 2020 | The fine structure of surprise in intuitive physics: when, why, and how much?
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
CogSci | 7 |
| 2020 | Look before you leap: Quantitative tradeoffs between peril and reward in action understanding
Nensi Gjata, Tomer D. Ullman, Elizabeth S. Spelke, Shari Liu |
CogSci | 2 |
| 2020 | Adventures in Flatland: Perceiving Social Interactions Under Physical Dynamics
Tianmin Shu, Marta Kryven, Tomer D. Ullman, Josh Tenenbaum |
CogSci | 3 |
| 2019 | People's perception of others' risk preferences
Shari Liu, John McCoy, Tomer D. Ullman |
CogSci | 3 |
| 2019 | Draping an Elephant: Uncovering Children's Reasoning About Cloth-Covered Objects
Tomer D. Ullman, Eliza Kosoy, Ilker Yildirim, Amir Arsalan Soltani, Max H. Siegel, Josh Tenenbaum, Elizabeth S. Spelke |
CogSci | 1 |
| 2019 | Modeling Expectation Violation in Intuitive Physics with Coarse Probabilistic Object RepresentationsabstractFrom infancy, humans have expectations about how objects will move and interact. Even young children expect objects not to move through one another, teleport, or disappear. They are surprised by mismatches between physical expectations and perceptual observations, even in unfamiliar scenes with completely novel objects. A model that exhibits human-like understanding of physics should be similarly surprised, and adjust its beliefs accordingly. We propose ADEPT, a model that uses a coarse (approximate geometry) object-centric representation for dynamic 3D scene understanding. Inference integrates deep recognition networks, extended probabilistic physical simulation, and particle filtering for forming predictions and expectations across occlusion. We also present a new test set for measuring violations of physical expectations, using a range of scenarios derived from developmental psychology. We systematically compare ADEPT, baseline models, and human expectations on this test set. ADEPT outperforms standard network architectures in discriminating physically implausible scenes, and often performs this discrimination at the same level as people. Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
NeurIPS | 7 |
| 2018 | Models of Human Scientific Discovery
Robert L. Goldstone, Alison Gopnik, Paul Thagard, Tomer D. Ullman |
CogSci | 4 |
| 2018 | Moral Dynamics: A Computational Model of Moral Judgment
Felix Sosa, Tomer D. Ullman, Samuel Gershman, Josh Tenenbaum, Tobias Gerstenberg |
CogSci | 2 |
| 2017 | Thinking and Guessing: Bayesian and Empirical Models of How Humans Search
Marta Kryven, Tomer D. Ullman, William Cowan, Josh Tenenbaum |
CogSci | 2 |
| 2017 | What's worth the effort: Ten-month-old infants infer the value of goals from the costs of actions
Shari Liu, Tomer D. Ullman, Josh Tenenbaum, Elizabeth S. Spelke |
CogSci | 2 |
| 2017 | Weight matters: The role of physical weight in non-physical language across age and culture
Tomer D. Ullman, Santiago Alonso-Diaz, Stephen Ferrigno, Sarina Zahid, Celeste Kidd |
CogSci | 1 |
| 2017 | A Compositional Object-Based Approach to Learning Physical Dynamics
Michael Chang 0003, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum |
ICLR (Poster) | 2 |
| 2016 | Outcome or Strategy? A Bayesian Model of Intelligence Attribution
Marta Kryven, Tomer D. Ullman, William Cowan, Josh Tenenbaum |
CogSci | 2 |
| 2016 | Coalescing the Vapors of Human Experience into a Viable and Meaningful Comprehension
Tomer D. Ullman, Max H. Siegel, Josh Tenenbaum, Samuel Gershman |
CogSci | 1 |
| 2016 | The Pragmatics of Spatial Language
Tomer D. Ullman, Yang Xu 0023, Noah D. Goodman |
CogSci | 1 |
| 2014 | Wins above replacement: Responsibility attributions as counterfactual replacements
Tobias Gerstenberg, Tomer D. Ullman, Max Kleiman-Weiner, David A. Lagnado, Josh Tenenbaum |
CogSci | 2 |
| 2014 | Learning physical theories from dynamical scenes
Tomer D. Ullman, Andreas Stuhlmüller, Noah D. Goodman, Josh Tenenbaum |
CogSci | 1 |
| 2013 | Minimal Nativism: How does cognitive development get off the ground?
Tomer D. Ullman, Josh Tenenbaum, Noah D. Goodman, Shimon Ullman, Elizabeth S. Spelke |
CogSci | 1 |
| 2012 | Computational Models of Intuitive Physics
Peter W. Battaglia, Tomer D. Ullman, Josh Tenenbaum, Adam Sanborn, Kenneth D. Forbus, Tobias Gerstenberg, David A. Lagnado |
CogSci | 2 |
| 2012 | Probabilistic generative models for counterfactual reasoning and blame attribution
John McCoy, Tomer D. Ullman, Andreas Stuhlmüller, Tobias Gerstenberg, Josh Tenenbaum |
CogSci | 2 |
| 2011 | Forward Physics: How people learn and generalize novel dynamical models
Tomer D. Ullman, Noah D. Goodman, Josh Tenenbaum |
CogSci | 1 |
| 2009 | Help or Hinder: Bayesian Models of Social Goal InferenceabstractEveryday social interactions are heavily influenced by our snap judgments about others goals. Even young infants can infer the goals of intentional agents from observing how they interact with objects and other agents in their environment: e.g., that one agent is helping orhindering anothers attempt to get up a hill or open a box. We propose a model for how people can infer these social goals from actions, based on inverse planning in multiagent Markov decision problems (MDPs). The model infers the goal most likely to be driving an agents behavior by assuming the agent acts approximately rationally given environmental constraints and its model of other agents present. We also present behavioral evidence in support of this model over a simpler, perceptual cue-based alternative. Tomer D. Ullman, Chris L. Baker, Owen Macindoe, Owain Evans, Noah D. Goodman, Josh Tenenbaum |
NIPS | 1 |