Tobias Gerstenberg

dblp:133/0561 · DBLP profile ↗
← Back
78ranked-venue papers
12as first author
45since 2021 · last 2025
0000-0002-9162-0779ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 12 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 69 · 12 first-author · 36 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Spot the ball: Inferring Hidden Information from Human Behavioral Cues
Neha Balamurugan, Sarah A. Wu, Cristóbal Eyzaguirre, Adam Chun, Tobias Gerstenberg
CogSci5
2025 How do we get to know someone? Diagnostic questions for inferring personal traits
Erik Brockbank, Tobias Gerstenberg, Judith E. Fan, Robert D. Hawkins
CogSci2
2025 Taking others for granted: balancing personal and presentational goals in action selection
Victor Btesh, David A. Lagnado, Tobias Gerstenberg
CogSci3
2025 Toward a Formal Pragmatics of Explanation
Jacqueline Harding, Tobias Gerstenberg, Thomas Icard
CogSci2
2025 Learning From 'What Might Have Been': A Bayesian Model of Learning from Regret
Kate Petrova, James J. Gross, Tobias Gerstenberg
CogSci3
2025 Cause and fault in development
David Rose, Cici Hou, Shaun Nichols, Tobias Gerstenberg, Ellen M. Markman
CogSci4
2025 Leave a trace: Recursive reasoning about deceptive behavior
Verona Teo, Sarah A. Wu, Erik Brockbank, Tobias Gerstenberg
CogSci4
2025 Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
Lionel Wong, Katie Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy J. O'Donnell, Alexander K. Lew, Jacob Andreas, Tyler Brooke-Wilson, Josh Tenenbaum
CogSci6
2025 Language models assign responsibility based on actual rather than counterfactual contributions
Eric J. Bigelow, Tobias Gerstenberg, Tomer D. Ullman, Samuel Gershman
CogSci3
2025 Learning to Plan from Actual and Counterfactual Experiences
Justin Yang, Tobias Gerstenberg
CogSci2
2025 Generics revisited: Analyzing generalizations in children's books and caregivers' speech
Sunny Yu, Alvin Wei Ming Tan, Siying Zhang, Xuhui Miao, Riley Carlson, Tobias Gerstenberg, David Rose
CogSci6
2025 Young children can use counterfactual simulation to reason about task performance
Peter Zhu, Chloe R. Chang, Tobias Gerstenberg, Hyowon Gweon
CogSci3
2025 Causal-PIK: Causality-based Physical Reasoning with a Physics-Informed Kernel
abstract
Tasks that involve complex interactions between objects with unknown dynamics make planning before execution difficult. These tasks require agents to iteratively improve their actions after actively exploring causes and effects in the environment. For these type of tasks, we propose Causal-PIK, a method that leverages Bayesian optimization to reason about causal interactions via a Physics-Informed Kernel to help guide efficient search for the best next action. Experimental results on Virtual Tools and PHYRE physical reasoning benchmarks show that Causal-PIK outperforms state-of-the-art results, requiring fewer actions to reach the goal. We also compare Causal-PIK to human studies, including results from a new user study we conducted on the PHYRE benchmark. We find that Causal-PIK remains competitive on tasks that are very challenging, even for human problem-solvers.
Carlota Parés-Morlans, Michelle Yi, Sarah A. Wu, Rika Antonova, Tobias Gerstenberg, Jeannette Bohg
ICML6
2024 Anticipating the Risks and Benefits of Counterfactual World Simulation Models (Extended Abstract)
abstract
This paper examines the transformative potential of Counterfactual World Simulation Models (CWSMs). CWSMs use pieces of multi-modal evidence, such as the CCTV footage or sound recordings of a road accident, to build a high-fidelity 3D reconstruction of the scene. They can also answer causal questions, such as whether the accident happened because the driver was speeding, by simulating what would have happened in relevant counterfactual situations. CWSMs will enhance our capacity to envision alternate realities and investigate the outcomes of counterfactual alterations to how events unfold. This also, however, raises questions about what alternative scenarios we should be considering and what to do with that knowledge. We present a normative and ethical framework that guides and constrains the simulation of counterfactuals. We address the challenge of ensuring fidelity in reconstructions while simultaneously preventing stereotype perpetuation during counterfactual simulations. We anticipate different modes of how users will interact with CWSMs and discuss how their outputs may be presented. Finally, we address the prospective applications of CWSMs in the legal domain, recognizing both their potential to revolutionize legal proceedings as well as the ethical concerns they engender. Anticipating a new type of AI, this paper seeks to illuminate a path forward for responsible and effective use of CWSMs.
Lara Kirfel, Rob MacCoun, Thomas Icard, Tobias Gerstenberg
AIES (1)4
2024 Without his cookies, he's just a monster: a counterfactual simulation model of social explanation
Erik Brockbank, Justin Yang, Mishika Govil, Judith E. Fan, Tobias Gerstenberg
CogSci5
2024 Procedural Dilemma Generation for Moral Reasoning in Humans and Language Models
Jan-Philipp Fränken, Kanishk Gandhi, Tori Qiu, Ayesha Khawaja, Noah D. Goodman, Tobias Gerstenberg
CogSci6
2024 Chain Versus Common Cause: Biased Causal Strength Judgments in Humans and Large Language Models
Anita Keshmirian, Moritz Willig, Babak Hemmatian, Kristian Kersting, Ulrike Hahn, Tobias Gerstenberg
CogSci6
2024 Do as I explain: Explanations communicate optimal interventions
Lara Kirfel, Jacqueline Harding, Jeong Yeon Shin, Cindy Xin, Thomas Icard, Tobias Gerstenberg
CogSci6
2024 Who is responsible for collective action?
Casey Lewry, Tania Lombrozo, Shannon Wing, Sydney Levine, Josh Tenenbaum, Lionel Wong, Sofia Bonicalzi, Tobias Gerstenberg
CogSci8
2024 Towards a computational model of responsibility judgments in sequential human-AI collaboration
Stratis Tsirtsis, Manuel Gomez-Rodriguez, Tobias Gerstenberg
CogSci3
2024 Whodunnit? Inferring what happened from multimodal evidence
Sarah A. Wu, Erik Brockbank, Hannah Cha, Jan-Philipp Fränken, Emily Jin, Zhuoyi Huang, Jiajun Wu 0001, Tobias Gerstenberg
CogSci10
2024 Resource-rational moral judgment
Sarah A. Wu, Xiang Ren 0001, Tobias Gerstenberg, Yejin Choi 0001, Sydney Levine
CogSci3
2024 Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
abstract
When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, technically challenging, and generally requires human preference labels or examples. We introduce SAMI, an iterative algorithm that finetunes a pretrained language model (without requiring preference labels or demonstrations) to increase the conditional mutual information between constitutions and self-generated responses given queries from a dataset. On single-turn dialogue and summarization, a SAMI-trained mistral-7b outperforms the initial pretrained model, with win rates between 66% and 77%. Strikingly, it also surpasses an instruction-finetuned baseline (mistral-7b-instruct) with win rates between 55% and 57% on single-turn dialogue. SAMI requires a model that writes the principles. To avoid dependence on strong models for writing principles, we align a strong pretrained model (mixtral-8x7b) using constitutions written by a weak instruction-finetuned model (mistral-7b-instruct), achieving a 65% win rate on summarization. Finally, we investigate whether SAMI generalizes to diverse summarization principles (e.g., "summaries should be scientific") and scales to stronger models (llama3-70b), finding that it achieves win rates of up to 68% for learned and 67% for held-out principles compared to the base model. Our results show that a pretrained LM can learn to follow constitutions without using preference labels, demonstrations, or human oversight.
Jan-Philipp Fränken, Eric Zelikman, Rafael Rafailov, Kanishk Gandhi, Tobias Gerstenberg, Noah D. Goodman
NeurIPS5
2024 MARPLE: A Benchmark for Long-Horizon Inference
abstract
Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchmark for evaluating long-horizon inference capabilities using multi-modal evidence. Our benchmark features agents interacting with simulated households, supporting vision, language, and auditory stimuli, as well as procedurally generated environments and agent behaviors. Inspired by classic ``whodunit'' stories, we ask AI models and human participants to infer which agent caused a change in the environment based on a step-by-step replay of what actually happened. The goal is to correctly identify the culprit as early as possible. Our findings show that human participants outperform both traditional Monte Carlo simulation methods and an LLM baseline (GPT-4) on this task. Compared to humans, traditional inference models are less robust and performant, while GPT-4 has difficulty comprehending environmental changes. We analyze factors influencing inference performance and ablate different modes of evidence, finding that all modes are valuable for performance. Overall, our experiments demonstrate that the long-horizon, multimodal inference tasks in our benchmark present a challenge to current models. Project website: https://marple-benchmark.github.io/.
Emily Jin, Zhuoyi Huang, Jan-Philipp Fränken, Hannah Cha, Erik Brockbank, Sarah A. Wu, Jiajun Wu 0001, Tobias Gerstenberg
NeurIPS10
2023 A Semantics for Causing, Enabling, and Preventing Verbs Using Structural Causal Models
Angela Cao, Atticus Geiger, Elisa Kreiss, Thomas Icard, Tobias Gerstenberg
CogSci5
2023 Causal Reasoning Across Agents and Objects
Bryan Gonzalez, Tobias Gerstenberg, Jonathan Phillips
CogSci2
2023 Father, don't forgive them, for they could have known what they're doing
Lara Kirfel, Xenia Bunk, Ro'i Zultan, Tobias Gerstenberg
CogSci4
2023 Show and tell: Learning causal structures from observations and explanations
Andrew Nam, Christopher Hughes, Thomas Icard, Tobias Gerstenberg
CogSci4
2023 Teleology and generics
David Rose, Siying Zhang, Tobias Gerstenberg
CogSci4
2023 Learning what matters: Causal abstraction in human inference
Steven M. Shin, Tobias Gerstenberg
CogSci2
2023 A computational model of responsibility judgments from counterfactual simulations and intention inferences
Sarah A. Wu, Shruti Sridhar, Tobias Gerstenberg
CogSci3
2023 You are what you're for: Essentialist categorization in large language models
Siying Zhang, Jingyuan Selena She, Tobias Gerstenberg, David Rose
CogSci3
2023 Understanding Social Reasoning in Language Models with Language Models
abstract
As Large Language Models (LLMs) become increasingly integrated into our everyday lives, understanding their ability to comprehend human mental states becomes critical for ensuring effective interactions. However, despite the recent attempts to assess the Theory-of-Mind (ToM) reasoning capabilities of LLMs, the degree to which these models can align with human ToM remains a nuanced topic of exploration. This is primarily due to two distinct challenges: (1) the presence of inconsistent results from previous evaluations, and (2) concerns surrounding the validity of existing evaluation methodologies. To address these challenges, we present a novel framework for procedurally generating evaluations with LLMs by populating causal templates. Using our framework, we create a new social reasoning benchmark (BigToM) for LLMs which consists of 25 controls and 5,000 model-written evaluations. We find that human participants rate the quality of our benchmark higher than previous crowd-sourced evaluations and comparable to expert-written evaluations. Using BigToM, we evaluate the social reasoning capabilities of a variety of LLMs and compare model performances with human performance. Our results suggest that GPT4 has ToM capabilities that mirror human inference patterns, though less reliable, while other LLMs struggle.
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, Noah D. Goodman
NeurIPS3
2023 MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment Tasks
abstract
Human commonsense understanding of the physical and social world is organized around intuitive theories. These theories support making causal and moral judgments. When something bad happens, we naturally ask: who did what, and why? A rich literature in cognitive science has studied people's causal and moral intuitions. This work has revealed a number of factors that systematically influence people's judgments, such as the violation of norms and whether the harm is avoidable or inevitable. We collected a dataset of stories from 24 cognitive science papers and developed a system to annotate each story with the factors they investigated. Using this dataset, we test whether large language models (LLMs) make causal and moral judgments about text-based scenarios that align with those of human participants. On the aggregate level, alignment has improved with more recent LLMs. However, using statistical analyses, we find that LLMs weigh the different factors quite differently from human participants. These results show how curated, challenge datasets combined with insights from cognitive science can help us go beyond comparisons based merely on aggregate metrics: we uncover LLMs implicit tendencies and show to what extent these align with human intuitions.
Allen Nie, Atharva Amdekar, Chris Piech, Tatsunori B. Hashimoto, Tobias Gerstenberg
NeurIPS6
2023 Explanations Can Reduce Overreliance on AI Systems During Decision-Making
abstract
Prior work has identified a resilient phenomenon that threatens the performance of human-AI decision-making teams: overreliance, when people agree with an AI, even when it is incorrect. Surprisingly, overreliance does not reduce when the AI produces explanations for its predictions, compared to only providing predictions. Some have argued that overreliance results from cognitive biases or uncalibrated trust, attributing overreliance to an inevitability of human cognition. By contrast, our paper argues that people strategically choose whether or not to engage with an AI explanation, demonstrating empirically that there are scenarios where AI explanations reduce overreliance. To achieve this, we formalize this strategic choice in a cost-benefit framework, where the costs and benefits of engaging with the task are weighed against the costs and benefits of relying on the AI. We manipulate the costs and benefits in a maze task, where participants collaborate with a simulated AI to find the exit of a maze. Through 5 studies (N = 731), we find that costs such as task difficulty (Study 1), explanation difficulty (Study 2, 3), and benefits such as monetary compensation (Study 4) affect overreliance. Finally, Study 5 adapts the Cognitive Effort Discounting paradigm to quantify the utility of different explanations, providing further support for our framework. Our results suggest that some of the null effects found in literature could be due in part to the explanation not sufficiently reducing the costs of verifying the AI's prediction.
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, Ranjay Krishna
Proc. ACM Hum. Comput. Interact.4
2022 Do Humans Trust Advice More if it Comes from AI?: An Analysis of Human-AI Interactions
abstract
In decision support applications of AI, the AI algorithm's output is framed as a suggestion to a human user. The user may ignore this advice or take it into consideration to modify their decision. With the increasing prevalence of such human-AI interactions, it is important to understand how users react to AI advice. In this paper, we recruited over 1100 crowdworkers to characterize how humans use AI suggestions relative to equivalent suggestions from a group of peer humans across several experimental settings. We find that participants' beliefs about how human versus AI performance on a given task affects whether they heed the advice. When participants do heed the advice, they use it similarly for human and AI suggestions. Based on these results, we propose a two-stage, "activation-integration" model for human behavior and use it to characterize the factors that affect human-AI interactions.
Kailas Vodrahalli, Roxana Daneshjou, Tobias Gerstenberg, James Zou 0001
AIES3
2022 Inferences from Disagreement
Jamie Amemiya, Gail D. Heyman, Tobias Gerstenberg
CogSci3
2022 Looking into the past: Eye-tracking mental simulation in physical inference
Ari Beller, Yingchen Xu, Scott W. Linderman, Tobias Gerstenberg
CogSci4
2022 Stop, children what's that sound? Multi-modal inference through mental simulation
Joseph Outa, Xijia Zhou, Hyowon Gweon, Tobias Gerstenberg
CogSci4
2022 That was close! A counterfactual simulation model of causal judgments about decisions
Sarah A. Wu, Shruti Sridhar, Tobias Gerstenberg
CogSci3
2022 Uncalibrated Models Can Improve Human-AI Collaboration
abstract
In many practical applications of AI, an AI model is used as a decision aid for human users. The AI provides advice that a human (sometimes) incorporates into their decision-making process. The AI advice is often presented with some measure of "confidence" that the human can use to calibrate how much they depend on or trust the advice. In this paper, we present an initial exploration that suggests showing AI models as more confident than they actually are, even when the original AI is well-calibrated, can improve human-AI performance (measured as the accuracy and confidence of the human's final prediction after seeing the AI advice). We first train a model to predict human incorporation of AI advice using data from thousands of human-AI interactions. This enables us to explicitly estimate how to transform the AI's prediction confidence, making the AI uncalibrated, in order to improve the final human prediction. We empirically validate our results across four different tasks---dealing with images, text and tabular data---involving hundreds of human participants. We further support our findings with simulation analysis. Our findings suggest the importance of jointly optimizing the human-AI system as opposed to the standard paradigm of optimizing the AI model alone.
Kailas Vodrahalli, Tobias Gerstenberg, James Zou 0001
NeurIPS2
2021 Eye-Tracking Multi-Modal Inference
Ari Beller, Yingchen Xu, Tobias Gerstenberg
CogSci3
2021 In Touch with Causation: Understanding the Impact of Kinesthetic Haptics on Causality
Elyse D. Z. Chase, Phillip Wolff, Tobias Gerstenberg, Sean Follmer
CogSci3
2021 Who went fishing? Inferences from social evaluations
Kelsey R. Allen, Tobias Gerstenberg
CogSci3
2021 The role of counterfactual reasoning in responsibility judgments
Sarah A. Wu, Tobias Gerstenberg
CogSci2
2020 The language of causation
Ari Beller, Erin D. Bennett, Tobias Gerstenberg
CogSci3
2020 Whom will Granny thank? Thinking about what could have been informs children's inferences about relative helpfulness
Sophie Bridgers, Chuyi Yang, Tobias Gerstenberg, Hyowon Gweon
CogSci3
2020 Learning from explanations
Lara Kirfel, Thomas Icard, Tobias Gerstenberg
CogSci3
2019 When circumstances change, update your pronouns
Joshua K. Hartshorne, Mariela Jennings, Tobias Gerstenberg, Josh Tenenbaum
CogSci3
2019 The trajectory of counterfactual simulation in development
Jonathan F. Kominsky, Tobias Gerstenberg, Madeline Pelz, Henrik Singmann, Mark Sheskin, Frank C. Keil
CogSci2
2019 Explaining intuitive difficulty judgments by modeling physical effort and risk
Ilker Yildirim, Basil Saeed, Grace Bennett-Pierre, Tobias Gerstenberg, Josh Tenenbaum, Hyowon Gweon
CogSci4
2018 Tiptoeing around it: Inference from absence in potentially offensive speech
Monica A. Gates, Tess L. Veuthey, Michael Henry Tessler, Kevin A. Smith 0001, Tobias Gerstenberg, Laurie Bayet, Josh Tenenbaum
CogSci5
2018 What happened? Reconstructing the past through vision and sound
Tobias Gerstenberg, Max H. Siegel, Josh Tenenbaum
CogSci1
2018 Moral Dynamics: A Computational Model of Moral Judgment
Felix Sosa, Tomer D. Ullman, Samuel Gershman, Josh Tenenbaum, Tobias Gerstenberg
CogSci5
2017 Causal learning from interventions and dynamics in continuous time
Neil Bramley, Ralf Mayrhofer, Tobias Gerstenberg, David A. Lagnado
CogSci3
2017 Faulty Towers: A hypothetical simulation model of physical support
Tobias Gerstenberg, Kevin A. Smith 0001, Josh Tenenbaum
CogSci1
2017 Eye movement-based probabilistic models for physical scene understanding
Eghbal Hosseini, Eli Pollock, Tobias Gerstenberg
CogSci3
2017 Marbles in Inaction: Counterfactual Simulation and Causation by Omission
Simon Stephan, Pascale Willemsen, Tobias Gerstenberg
CogSci3
2017 Physical problem solving: Joint planning with symbolic, geometric, and dynamic constraints
Ilker Yildirim, Tobias Gerstenberg, Basil Saeed, Marc Toussaint, Josh Tenenbaum
CogSci2
2016 Natural science: Active learning in dynamic physical microworlds
Neil Bramley, Tobias Gerstenberg, Josh Tenenbaum
CogSci2
2016 Understanding "almost": Empirical and computational studies of near misses
Tobias Gerstenberg, Josh Tenenbaum
CogSci1
2016 Implicit measurement of motivated causal attribution
Laura Niemi, Joshua K. Hartshorne, Tobias Gerstenberg, Liane Young
CogSci3
2015 Go fishing! Responsibility judgments when cooperation breaks down
Kelsey R. Allen, Julian Jara-Ettinger, Tobias Gerstenberg, Max Kleiman-Weiner, Josh Tenenbaum
CogSci3
2015 How, whether, why: Causal judgments as counterfactual contrasts
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum
CogSci1
2015 Responsibility judgments in voting scenarios
Tobias Gerstenberg, Joseph Y. Halpern, Josh Tenenbaum
CogSci1
2015 Inference of Intention and Permissibility in Moral Decision Making
Max Kleiman-Weiner, Tobias Gerstenberg, Sydney Levine, Josh Tenenbaum
CogSci2
2014 The order of things: Inferring causal structure from temporal patterns
Neil Bramley, Tobias Gerstenberg, David A. Lagnado
CogSci2
2014 From counterfactual simulation to causal judgment
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum
CogSci1
2014 Wins above replacement: Responsibility attributions as counterfactual replacements
Tobias Gerstenberg, Tomer D. Ullman, Max Kleiman-Weiner, David A. Lagnado, Josh Tenenbaum
CogSci1
2014 Causal Supersession
Jonathan F. Kominsky, Jonathan Phillips, Tobias Gerstenberg, David A. Lagnado, Joshua Knobe
CogSci3
2013 Back on track: Backtracking in counterfactual reasoning
Tobias Gerstenberg, Christos Bechlivanidis, David A. Lagnado
CogSci1
2012 Computational Models of Intuitive Physics
Peter W. Battaglia, Tomer D. Ullman, Josh Tenenbaum, Adam Sanborn, Kenneth D. Forbus, Tobias Gerstenberg, David A. Lagnado
CogSci6
2012 Ping Pong in Church: Productive use of concepts in human probabilistic inference
Tobias Gerstenberg, Noah D. Goodman
CogSci1
2012 Noisy Newtons: Unifying process and dependency accounts of causal attribution
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum
CogSci1
2012 Probabilistic generative models for counterfactual reasoning and blame attribution
John McCoy, Tomer D. Ullman, Andreas Stuhlmüller, Tobias Gerstenberg, Josh Tenenbaum
CogSci4
2011 Blame the Skilled
Tobias Gerstenberg, Anastasia Ejova, David A. Lagnado
CogSci1
2011 Rational Order Effects in Responsibility Attributions
Tobias Gerstenberg, David A. Lagnado, Maarten Speekenbrink, Catherine Cheung
CogSci1
2011 Beyond Outcomes: The Influence of Intentions and Deception
Simeon Schaechtele, Tobias Gerstenberg, David A. Lagnado
CogSci2