Erik Brockbank

dblp:249/6761 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0001-8702-239XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2025 How do we get to know someone? Diagnostic questions for inferring personal traits
Erik Brockbank, Tobias Gerstenberg, Judith E. Fan, Robert D. Hawkins
CogSci1
2025 Leave a trace: Recursive reasoning about deceptive behavior
Verona Teo, Sarah A. Wu, Erik Brockbank, Tobias Gerstenberg
CogSci3
2025 Linking student psychological orientation, engagement, and learning in college-level introductory data science
Kristine Zheng, Erik Brockbank, Shawn T. Schwartz, David Yeager, Chris Bryan, Carol S. Dweck, Judith E. Fan
CogSci2
2024 Without his cookies, he's just a monster: a counterfactual simulation model of social explanation
Erik Brockbank, Justin Yang, Mishika Govil, Judith E. Fan, Tobias Gerstenberg
CogSci1
2024 Whodunnit? Inferring what happened from multimodal evidence
Sarah A. Wu, Erik Brockbank, Hannah Cha, Jan-Philipp Fränken, Emily Jin, Zhuoyi Huang, Jiajun Wu 0001, Tobias Gerstenberg
CogSci2
2024 MARPLE: A Benchmark for Long-Horizon Inference
abstract
Reconstructing past events requires reasoning across long time horizons. To figure out what happened, humans draw on prior knowledge about the world and human behavior and integrate insights from various sources of evidence including visual, language, and auditory cues. We introduce MARPLE, a benchmark for evaluating long-horizon inference capabilities using multi-modal evidence. Our benchmark features agents interacting with simulated households, supporting vision, language, and auditory stimuli, as well as procedurally generated environments and agent behaviors. Inspired by classic ``whodunit'' stories, we ask AI models and human participants to infer which agent caused a change in the environment based on a step-by-step replay of what actually happened. The goal is to correctly identify the culprit as early as possible. Our findings show that human participants outperform both traditional Monte Carlo simulation methods and an LLM baseline (GPT-4) on this task. Compared to humans, traditional inference models are less robust and performant, while GPT-4 has difficulty comprehending environmental changes. We analyze factors influencing inference performance and ablate different modes of evidence, finding that all modes are valuable for performance. Overall, our experiments demonstrate that the long-horizon, multimodal inference tasks in our benchmark present a challenge to current models. Project website: https://marple-benchmark.github.io/.
Emily Jin, Zhuoyi Huang, Jan-Philipp Fränken, Hannah Cha, Erik Brockbank, Sarah A. Wu, Jiajun Wu 0001, Tobias Gerstenberg
NeurIPS6
2023 What is graph comprehension and how do you measure it?
Hannah Lloyd, Holly Huey, Erik Brockbank, Lace M. K. Padilla, Judith E. Fan
CogSci3
2022 How do people incorporate advice from artificial agents when making physical judgments?
Erik Brockbank, Justin Yang, Suvir Mirchandani, Erdem Biyik, Dorsa Sadigh, Judith E. Fan
CogSci1
2021 Humans fail to outwit adaptive rock, paper, scissors opponents
Erik Brockbank, Ed Vul
CogSci1
2020 Recursive Adversarial Reasoning in the Rock, Paper, Scissors Game
Erik Brockbank, Ed Vul
CogSci1
2020 Explanation Supports Hypothesis Generation in Learning
Erik Brockbank, Caren M. Walker
CogSci1
2020 Formalizing Interdisciplinary Collaboration in the CogSci Community
Lauren Oey, Isabella Destefano, Erik Brockbank, Ed Vul
CogSci3
2019 Mapping visual features onto numbers
Erik Brockbank, Ed Vul
CogSci1