Sarah A. Wu

dblp:261/9255 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0003-0001-5440ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Spot the ball: Inferring Hidden Information from Human Behavioral Cues
Neha Balamurugan, Sarah A. Wu, Cristóbal Eyzaguirre, Adam Chun, Tobias Gerstenberg
CogSci2
2025 Leave a trace: Recursive reasoning about deceptive behavior
Verona Teo, Sarah A. Wu, Erik Brockbank, Tobias Gerstenberg
CogSci2
2025 Causal-PIK: Causality-based Physical Reasoning with a Physics-Informed Kernel
abstract
Tasks that involve complex interactions between objects with unknown dynamics make planning before execution difficult. These tasks require agents to iteratively improve their actions after actively exploring causes and effects in the environment. For these type of tasks, we propose Causal-PIK, a method that leverages Bayesian optimization to reason about causal interactions via a Physics-Informed Kernel to help guide efficient search for the best next action. Experimental results on Virtual Tools and PHYRE physical reasoning benchmarks show that Causal-PIK outperforms state-of-the-art results, requiring fewer actions to reach the goal. We also compare Causal-PIK to human studies, including results from a new user study we conducted on the PHYRE benchmark. We find that Causal-PIK remains competitive on tasks that are very challenging, even for human problem-solvers.
Carlota Parés-Morlans, Michelle Yi, Sarah A. Wu, Rika Antonova, Tobias Gerstenberg, Jeannette Bohg
ICML4
2024 Whodunnit? Inferring what happened from multimodal evidence
Sarah A. Wu, Erik Brockbank, Hannah Cha, Jan-Philipp Fränken, Emily Jin, Zhuoyi Huang, Jiajun Wu 0001, Tobias Gerstenberg
CogSci1
2024 Resource-rational moral judgment
Sarah A. Wu, Xiang Ren 0001, Tobias Gerstenberg, Yejin Choi 0001, Sydney Levine
CogSci1
2024 MARPLE: A Benchmark for Long-Horizon Inference
abstract
Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchmark for evaluating long-horizon inference capabilities using multi-modal evidence. Our benchmark features agents interacting with simulated households, supporting vision, language, and auditory stimuli, as well as procedurally generated environments and agent behaviors. Inspired by classic ``whodunit'' stories, we ask AI models and human participants to infer which agent caused a change in the environment based on a step-by-step replay of what actually happened. The goal is to correctly identify the culprit as early as possible. Our findings show that human participants outperform both traditional Monte Carlo simulation methods and an LLM baseline (GPT-4) on this task. Compared to humans, traditional inference models are less robust and performant, while GPT-4 has difficulty comprehending environmental changes. We analyze factors influencing inference performance and ablate different modes of evidence, finding that all modes are valuable for performance. Overall, our experiments demonstrate that the long-horizon, multimodal inference tasks in our benchmark present a challenge to current models. Project website: https://marple-benchmark.github.io/.
Emily Jin, Zhuoyi Huang, Jan-Philipp Fränken, Hannah Cha, Erik Brockbank, Sarah A. Wu, Jiajun Wu 0001, Tobias Gerstenberg
NeurIPS7
2023 A computational model of responsibility judgments from counterfactual simulations and intention inferences
Sarah A. Wu, Shruti Sridhar, Tobias Gerstenberg
CogSci1
2022 That was close! A counterfactual simulation model of causal judgments about decisions
Sarah A. Wu, Shruti Sridhar, Tobias Gerstenberg
CogSci1
2021 The role of counterfactual reasoning in responsibility judgments
Sarah A. Wu, Tobias Gerstenberg
CogSci1
2020 Too many cooks: Coordinating multi-agent collaboration through inverse planning
Sarah A. Wu, Rose E. Wang, James A. Evans, Josh Tenenbaum, David C. Parkes, Max Kleiman-Weiner
CogSci1