EDBT 2026 Demo / reviewers in the wild / expert
Nick Haber
dblp:179/4983 · also Nicholas Haber
· DBLP profile ↗
40ranked-venue papers
3as first author
34since 2021 · last 2026
0000-0001-8804-7804ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 2 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-GENERALIST: Vision-Language-Action Models for Crafting 3D WorldsabstractCreating 3D graphics content for immersive and interactive worlds remains labor-intensive, limiting our ability to create large-scale synthetic data for training foundation models. Recent methods aim to alleviate this, but they often focus on a single aspect (e.g., layout) and do not improve generation quality by simply scaling computational resources. We recast 3D environment generation as a sequential decision-making problem, using Vision-Language Models (VLMs) as policies that output actions to jointly craft a 3D environment's layout, materials, lighting, and assets. Our framework, 3D-Generalist, trains VLMs to generate more prompt-aligned 3D environments via self-improvement fine-tuning. We demonstrate the effectiveness of 3D-Generalist and our training strategy in generating simulation-ready 3D environments. We also demonstrate its quality and scalability for synthetic data generation by pretraining a vision foundation model on the generated data. After fine-tuning on downstream tasks, we show that it surpasses models pre-trained on meticulously human-crafted synthetic data and approaches results achieved when training with orders of magnitude larger real data. Fan-Yun Sun, Shengguang Wu, Christian Jacobsen, Thomas Yim, Haoming Zou, Alexander Zook, Shangru Li, Yu-Hsin Chou, Ethem Can, Xunlei Wu, Clemens Eppner, Valts Blukis, Jonathan Tremblay, Jiajun Wu 0001, Stanley T. Birchfield, Nick Haber |
3DV | 16 |
| 2025 | Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive ImagesabstractRecent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations.We hypothesize that this issue arises because existing VLMs are not explicitly trained to generate texts that are accurately grounded in fine-grained image details.To enhance visual feedback during VLM training, we propose S-VCO (Symmetrical Visual Contrastive Optimization), a novel finetuning objective that steers the model toward capturing important visual details and aligning them with corresponding text tokens.To further facilitate this detailed alignment, we introduce MVC, a paired image-text dataset built by automatically filtering and augmenting visual counterfactual data to challenge the model with hard contrastive cases involving Minimal Visual Contrasts.Experiments show that our method consistently improves VLM performance across diverse benchmarks covering various abilities and domains, achieving up to a 22% reduction in hallucinations, and significant gains in vision-centric and general tasks.Notably, these improvements become increasingly pronounced in benchmarks with higher visual dependency.In short, S-VCO offers a significant enhancement of VLM's visuallydependent task performance while retaining or even improving the model's general abilities. Shengguang Wu, Fan-Yun Sun, Kaiyue Wen, Nick Haber |
ACL (1) | 4 |
| 2025 | CogGen: A Learner-Centered Generative AI Architecture for Intelligent Tutoring with Programming Videos
Wengxi Li, Roy D. Pea, Nick Haber, Hariharan Subramonyam |
AIED (5) | 3 |
| 2025 | Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions
Logan Matthew Cross, Nick Haber, Dan Yamins |
CogSci | 2 |
| 2025 | Do Large Language Models Have a Planning Theory of Mind? Evidence from MindGames: a Multi-Step Persuasion Task
Jared Moore, Rasmus Overmark, Ned Cooper, Beba Cibralic, Nick Haber, Cameron R. Jones |
CogSci | 5 |
| 2025 | Simulating variation in infant-caregiver attachment using reinforcement learning
Xijia Zhou, Chris Doyle, Logan Matthew Cross, Michael C. Frank, Nick Haber |
CogSci | 5 |
| 2025 | LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language ModelsabstractSpatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still struggle with 3D reasoning tasks like arranging objects in space according to open-ended language instructions, particularly in dense and physically constrained environments. We introduce LayoutVLM, a framework and scene layout representation that exploits the semantic knowledge of Vision-Language Models (VLMs) and supports differentiable optimization to ensure physical plausibility. LayoutVLM employs VLMs to generate two mutually reinforcing representations from visually marked images, and a self-consistent decoding process to improve VLMs spatial planning. Our experiments show that LayoutVLM addresses the limitations of existing LLM and constraint-based approaches, producing physically plausible 3D layouts better aligned with the semantic intent of input language instructions. We also demonstrate that fine-tuning VLMs with the proposed scene layout representation extracted from existing scene datasets can improve their reasoning performance. Fan-Yun Sun, Siyi Gu, Dylan Lim, Goutam Bhat, Federico Tombari, Manling Li, Nick Haber, Jiajun Wu 0001 |
CVPR | 8 |
| 2025 | Scaffold or Crutch? Examining College Students' Use and Views of Generative AI Tools for STEM EducationabstractDeveloping problem-solving competency is central to Science, Technology, Engineering, and Mathematics (STEM) education, yet translating this priority into effective approaches to problem-solving instruction and assessment has been a significant challenge. The recent proliferation of generative artificial intelligence (genAI) tools like ChatGPT in higher education introduces new considerations: how to define problem-solving competency in a genAI era, and how these tools can help or hinder students' development of STEM problem-solving competency. Our research takes steps in examining these considerations by studying how and why college students are currently using genAI tools in their STEM coursework, with a specific focus on how they employ these tools to support their problem-solving. We conducted an online survey of 40 STEM college students from diverse institutions across the US. In addition, we surveyed 28 STEM faculty to understand instructor views on effective and ineffective genAI tool use in STEM courses and their guidance for students. Our findings reveal high adoption rates and diverse applications of genAI tools among STEM students. The most common use cases of genAI tools in STEM coursework include finding explanations, exploring related topics, summarizing readings, and helping with problem-set questions. The primary motivation for using genAI tools in STEM coursework was to save time. Moreover, we found that over half of the student participants reported simply inputting a problem for AI to generate solutions, potentially bypassing their own problem-solving processes. These findings indicate that despite high adoption rates, students' current approaches to utilizing genAI tools often fall short in enhancing their own STEM problem-solving competencies. The study also explored students' and STEM instructors' perceptions of the benefits and risks associated with using genAI tools in STEM education. Our findings provide insights into how to guide students on appropriate genAI use in STEM courses and how to design genAI-based tools to foster students' problem-solving competency. Karen D. Wang, Zhangyang Wu, L'Nard Tufts II, Carl E. Wieman, Shima Salehi, Nick Haber |
EDUCON | 6 |
| 2025 | The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech PathologyabstractAccording to the U.S. National Institutes of Health, more than 3.4 million children experience speech disorders that require clinical intervention.The number of speech-language pathologists (SLPs) is roughly 20 times fewer than the number of affected children, highlighting a significant gap in children's care and a pressing need for technological support that improves the productivity of SLPs.State-ofthe-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored largely due to a limited understanding of their performance in highstakes clinical settings.To address this gap, we collaborate with domain experts to develop a taxonomy of real-world use cases of MLMs in speech-language pathologies.Building on this taxonomy, we introduce the first comprehensive benchmark for evaluating MLM across five core use cases, each containing 1,000 manually annotated data points.This benchmark includes robustness and sensitivity tests under various settings, including background noise, speaker gender, and accent.Our evaluation of 15 state-of-the-art MLMs reveals that no single model consistently outperforms others across all tasks.Notably, we find systematic disparities, with models performing better on male speakers, and observe that chain-of-thought prompting can degrade performance on classification tasks with large label spaces and narrow decision boundaries.Furthermore, we study fine-tuning MLMs on domain-specific data, achieving improvements of over 30% compared to base models.These findings highlight both the potential and limitations of current MLMs for speech-language pathology applications, underscoring the need for further research and targeted development 1 . Fagun Patel, Duc Q. Nguyen, Sang T. Truong, Jody Vaynshtok, Oluwasanmi Koyejo, Nick Haber |
EMNLP | 6 |
| 2025 | Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language ModelsabstractMulti-agent reinforcement learning (MARL) methods struggle with the non-stationarity of multi-agent systems and fail to adaptively learn online when tested with novel agents. Here, we leverage large language models (LLMs) to create an autonomous agent that can handle these challenges. Our agent, Hypothetical Minds, consists of a cognitively-inspired architecture, featuring modular components for perception, memory, and hierarchical planning over two levels of abstraction. We introduce the Theory of Mind module that scaffolds the high-level planning process by generating hypotheses about other agents' strategies in natural language. It then evaluates and iteratively refines these hypotheses by reinforcing hypotheses that make correct predictions about the other agents' behavior. Hypothetical Minds significantly improves performance over previous LLM-agent and RL baselines on a range of competitive, mixed motive, and collaborative domains in the Melting Pot benchmark, including both dyadic and population-based environments. Additionally, comparisons against LLM-agent baselines and ablations reveal the importance of hypothesis evaluation and refinement for succeeding on complex scenarios. Logan Matthew Cross, Violet Xiang, Agam Bhatia, Dan Yamins, Nick Haber |
ICLR | 5 |
| 2025 | ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research CodeabstractLarge language models (LLMs) have shown promise in transforming machine learning research, yet their capability to faithfully implement genuinely novel ideas from recent research papers—ideas unseen during pretraining—remains unclear. We introduce ResearchCodeBench, a benchmark that evaluates LLMs’ ability to translate cutting-edge ML contributions from top 2024-2025 research papers into executable code. We assessed 30+ proprietary and open-source LLMs, finding that even the best models correctly implement less than 40% of the code. We present empirical findings on performance comparison, contamination, and error patterns. By providing a rigorous evaluation platform, ResearchCodeBench enables continuous understanding and advancement of LLM-driven innovation in research code generation. Tianyu Hua, Harper Hua, Violet Xiang, Benjamin Klieger, Sang T. Truong, Weixin Liang, Fan-Yun Sun, Nick Haber |
NeurIPS | 8 |
| 2025 | When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI CollaborationabstractAs large language models (LLMs) increasingly serve as close collaborators for humans, it is crucial that they express their reasoning in ways that humans can understand and learn from. However, this capability remains relatively less understood and under-evaluated. To address this, we introduce a conceptual framework for such Human-AI knowledge transfer capabilities and conduct the first large-scale user study (N=118) explicitly designed to measure it. In our two-phase setup, humans first ideate with an LLM on problem-solving strategies, then independently implement solutions, isolating the influence of model reasoning on human understanding. Our findings reveal that while model benchmark performance correlates with collaborative outcomes, this relationship is notably inconsistent with significant outliers, highlighting that knowledge transfer is a distinct capability requiring dedicated optimization. Our analysis uncovers behavioral and strategic factors that mediate successful knowledge transfer, and we release our code, dataset, and evaluation framework to support future work on communicatively aligned models. Carlos E. Jimenez, Shunyu Yao 0006, Nick Haber, Diyi Yang, Karthik Narasimhan |
NeurIPS | 4 |
| 2025 | Fantastic Bugs and Where to Find Them in AI BenchmarksabstractBenchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we introduce a framework for systematic benchmark revision that leverages statistical analysis of response patterns to flag potentially invalid questions for further expert review. Our approach builds on a core assumption commonly used in AI evaluations that the mean score sufficiently summarizes model performance. This implies a unidimensional latent construct underlying the measurement experiment, yielding expected ranges for various statistics for each item. When empirically estimated values for these statistics fall outside the expected range for an item, the item is more likely to be problematic. Across nine widely used benchmarks, our method guides expert review to identify problematic questions with up to 84\% precision. In addition, we introduce an LLM‑judge first pass to review questions, further reducing human effort. Together, these components provide an efficient and scalable framework for systematic benchmark revision. Sang T. Truong, Yuheng Tu 0001, Michael Hardy, Anka Reuel, Zeyu Tang 0002, Jirayu Burapacheep, Jonathan Perera, Chibuike Uwakwe, Benjamin W. Domingue, Nick Haber, Oluwasanmi Koyejo |
NeurIPS | 10 |
| 2025 | From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer ReviewabstractThe advent of large language models (LLMs) offers unprecedented opportunities to reimagine peer review beyond the constraints of traditional workflows.
Despite these opportunities, prior efforts have largely focused on replicating traditional review workflows with LLMs serving as direct substitutes for human reviewers, while limited attention has been given to exploring new paradigms that fundamentally rethink how LLMs can participate in the academic review process.
In this paper, we introduce and explore a novel mechanism that employs LLM agents to perform pairwise comparisons among manuscripts instead of individual scoring. By aggregating outcomes from substantial pairwise evaluations, this approach enables a more accurate and robust measure of relative manuscript quality.
Our experiments demonstrate that this comparative approach significantly outperforms traditional rating-based methods in identifying high-impact papers. However, our analysis also reveals emergent biases in the selection process, notably a reduced novelty in research topics and an increased institutional imbalance. These findings highlight both the transformative potential of rethinking peer review with LLMs and critical challenges that future systems must address to ensure equity and diversity. Haijing Zhang, Wenlong Ji, Tianyu Hua, Nick Haber, Hancheng Cao, Weixin Liang |
NeurIPS | 5 |
| 2024 | Partial-View Object View Synthesis via Filtering InversionabstractWe propose Filtering Inversion (FINV), a learning framework and optimization process that predicts a renderable 3D object representation from one or few partial views. FINV addresses the challenge of synthesizing novel views of objects from partial observations, spanning cases where the object is not entirely in view, is partially occluded, or is only observed from similar views. To achieve this, FINV learns shape priors by training a 3D generative model. At inference, given one or more views of a novel real-world object, FINV first finds a set of latent codes for the object by inverting the generative model from multiple initial seeds. Maintaining the set of latent codes, FINV filters and resamples them after receiving each new observation, akin to particle filtering. The generator is then finetuned for each latent code on the available views in order to adapt to novel objects. We show that FINV successfully synthesizes novel views of real-world objects (e.g., chairs, tables, and cars), even if the generative prior is trained only on synthetic objects. The ability to address the sim-to-real problem allows FINV to be used for object categories without real-world datasets. FINV achieves state-of-the-art performance on multiple real-world datasets, recovers object shape and texture from partial and sparse views, is robust to occlusion, and is able to incrementally improves its representation with more observations. Fan-Yun Sun, Jonathan Tremblay, Valts Blukis, Danfei Xu, Boris Ivanovic, Péter Karkus, Stanley T. Birchfield, Dieter Fox, Yunzhu Li, Jiajun Wu 0001, Marco Pavone 0001, Nick Haber |
3DV | 14 |
| 2024 | Animate Agent World Modeling Benchmark
Logan Matthew Cross, Violet Xiang, Nick Haber, Dan Yamins |
CogSci | 3 |
| 2024 | Modeling Social Learning Through Demonstration in Multi-Armed Bandits
Julio Martinez, Michael C. Frank, Nick Haber |
CogSci | 3 |
| 2024 | Simulating Infants' Attachment: Behavioral Patterns of Caregiver Proximity Seeking and Environment Exploration Using Reinforcement Learning Models
Xijia Zhou, Chris Doyle, Michael C. Frank, Nick Haber |
CogSci | 4 |
| 2024 | Holodeck: Language Guided Generation of 3D Embodied AI Environmentsabstract3D simulated environments play a critical role in Embodied AI, but their creation requires expertise and extensive manual effort, restricting their diversity and scope. To miti-gate this limitation, we present Holodeck, a system that generates 3D environments to match a user-supplied prompt fullyautomatedly. Holodeck can generate diverse scenes, e.g., arcades, spas, and museums, adjust the designs for styles, and can capture the semantics of complex queries such as “apartment for a researcher with a cat” and “office of a professor who is a fan of Star Wars”. Holodeck leverages a large language model (i.e., GPT-4) for common sense knowledge about what the scene might look like and uses a large collection of 3D assets from Objaverse to populate the scene with diverse objects. To address the challenge of positioning objects correctly, we prompt GPT-4 to generate spatial relational constraints between objects and then optimize the layout to satisfy those constraints. Our large-scale human evaluation shows that annotators prefer Holodeck over manually designed procedural baselines in residential scenes and that Holodeck can produce high-quality outputs for diverse scene types. We also demonstrate an exciting application of Holodeck in Embodied AI, training agents to navigate in novel scenes like music rooms and daycares without human-constructed data, which is a significant step forward in developing general-purpose embodied agents. Yue Yang 0006, Fan-Yun Sun, Luca Weihs, Eli VanderBilt, Alvaro Herrasti, Winson Han, Jiajun Wu 0001, Nick Haber, Ranjay Krishna, Lingjie Liu, Chris Callison-Burch, Mark Yatskar, Aniruddha Kembhavi |
CVPR | 8 |
| 2024 | ContextRef: Evaluating Referenceless Metrics for Image Description GenerationabstractReferenceless metrics (e.g., CLIPScore) use pretrained vision--language models to assess image descriptions directly without costly ground-truth reference texts. Such methods can facilitate rapid progress, but only if they truly align with human preference judgments. In this paper, we introduce ContextRef, a benchmark for assessing referenceless metrics for such alignment. ContextRef has two components: human ratings along a variety of established quality dimensions, and ten diverse robustness checks designed to uncover fundamental weaknesses. A crucial aspect of ContextRef is that images and descriptions are presented in context, reflecting prior work showing that context is important for description quality. Using ContextRef, we assess a variety of pretrained models, scoring functions, and techniques for incorporating context. None of the methods is successful with ContextRef, but we show that careful fine-tuning yields substantial improvements. ContextRef remains a challenging benchmark though, in large part due to the challenge of context dependence. Elisa Kreiss, Eric Zelikman, Christopher Potts, Nick Haber |
ICLR | 4 |
| 2024 | Hypothesis Search: Inductive Reasoning with Language ModelsabstractInductive reasoning is a core problem-solving capacity: humans can identify underlying principles from a few examples, which can then be robustly generalized to novel scenarios. Recent work has evaluated large language models (LLMs) on inductive reasoning tasks by directly prompting them yielding "in context learning." This can work well for straightforward inductive tasks, but performs very poorly on more complex tasks such as the Abstraction and Reasoning Corpus (ARC). In this work, we propose to improve the inductive reasoning ability of LLMs by generating explicit hypotheses at multiple levels of abstraction: we prompt the LLM to propose multiple abstract hypotheses about the problem, in natural language, then implement the natural language hypotheses as concrete Python programs. These programs can be directly verified by running on the observed examples and generalized to novel inputs. To reduce the hypothesis search space, we explore steps to filter the set of hypotheses to be implemented as programs: we either ask the LLM to summarize them into a smaller set of hypotheses, or ask human annotators to select a subset. We verify our pipeline's effectiveness on the ARC visual inductive reasoning benchmark, its variant 1D-ARC, and string transformation dataset SyGuS. On a random 40-problem subset of ARC, our automated pipeline using LLM summaries achieves 27.5% accuracy, significantly outperforming the direct prompting baseline (accuracy of 12.5%). With the minimal human input of selecting from LLM-generated candidates, the performance is boosted to 37.5%. Our ablation studies show that abstract hypothesis generation and concrete program representations are both beneficial for LLMs to perform inductive reasoning tasks. Ruocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu, Nick Haber, Noah D. Goodman |
ICLR | 5 |
| 2024 | Discovering Players' Problem-Solving Behavioral Characteristics in a Puzzle Game through Sequence MiningabstractDigital games offer promising platforms for assessing student higher-order competencies such as problem-solving. However, processing and analyzing the large volume of interaction log data generated in these platforms to uncover meaningful behavioral patterns remain a complex research challenge. In this study, we employ sequence mining and clustering techniques to examine students’ log data in an interactive puzzle game that requires player to change rules to win the game. Our goal is to identify behavioral characteristics associated with the problem-solving practices adopted by individual students. The findings indicate that the most effective problem solvers made fewer rule changes and took longer time to make those changes across both an introductory and a more advanced level of the game. Conversely, rapid rule change actions were linked to ineffective problem-solving. This research underscores the potential of sequence mining and cluster analysis as generalizable methods for understanding student higher-order competencies through log data in digital gaming and learning environments. It also suggests future directions on how to provide just-in-time, in-game feedback to enhance student problem-solving competences. Karen D. Wang, David DeLiema, Nick Haber, Shima Salehi |
LAK | 4 |
| 2024 | Policy-shaped prediction: avoiding distractions in model-based reinforcement learningabstractModel-based reinforcement learning (MBRL) is a promising route to sample-efficient policy optimization. However, a known vulnerability of reconstruction-based MBRL consists of scenarios in which detailed aspects of the world are highly predictable, but irrelevant to learning a good policy. Such scenarios can lead the model to exhaust its capacity on meaningless content, at the cost of neglecting important environment dynamics. While existing approaches attempt to solve this problem, we highlight its continuing impact on leading MBRL methods ---including DreamerV3 and DreamerPro--- with a novel environment where background distractions are intricate, predictable, and useless for planning future actions. To address this challenge we develop a method for focusing the capacity of the world model through a synergy of a pretrained segmentation model, a task-aware reconstruction loss, and adversarial learning. Our method outperforms a variety of other approaches designed to reduce the impact of distractors, and is an advance towards robust model-based reinforcement learning. Miles Hutson, Isaac Kauvar, Nick Haber |
NeurIPS | 3 |
| 2024 | Learning Formal Mathematics From Intrinsic MotivationabstractHow did humanity coax mathematics from the aether? We explore the Platonic view that mathematics can be discovered from its axioms---a game of conjecture and proof. We describe an agent that jointly learns to pose challenging problems for itself (conjecturing) and solve them (theorem proving). Given a mathematical domain axiomatized in dependent type theory, we first combine methods for constrained decoding and type-directed synthesis to sample valid conjectures from a language model. Our method guarantees well-formed conjectures by construction, even as we start with a randomly initialized model. We use the same model to represent a policy and value function for guiding proof search. Our agent targets generating hard but provable conjectures --- a moving target, since its own theorem proving ability also improves as it trains. We propose novel methods for hindsight relabeling on proof search trees to significantly improve the agent's sample efficiency in both tasks. Experiments on 3 axiomatic domains (propositional logic, arithmetic and group theory) demonstrate that our agent can bootstrap from only the axioms, self-improving in generating true and challenging conjectures and in finding proofs. Gabriel Poesia, David Broman, Nick Haber, Noah D. Goodman |
NeurIPS | 3 |
| 2024 | FactorSim: Generative Simulation via Factorized RepresentationabstractGenerating simulations to train intelligent agents in game-playing and robotics from natural language input, user input, or task documentation remains an open-ended challenge. Existing approaches focus on parts of this challenge, such as generating reward functions or task hyperparameters. Unlike previous work, we introduce FACTORSIM that generates full simulations in code from language input that can be used to train agents. Exploiting the structural modularity specific to coded simulations, we propose to use a factored partially observable Markov decision process representation that allows us to reduce context dependence during each step of the generation. For evaluation, we introduce a generative simulation benchmark that assesses the generated simulation code’s accuracy and effectiveness in facilitating zero-shot transfers in reinforcement learning settings. We show that FACTORSIM outperforms existing methods in generating simulations regarding prompt alignment (i.e., accuracy), zero-shot transfer abilities, and human evaluation. We also demonstrate its effectiveness in generating robotic tasks. Fan-Yun Sun, S. I. Harini, Angela Yi, Alexander Zook, Jonathan Tremblay, Logan Matthew Cross, Jiajun Wu 0001, Nick Haber |
NeurIPS | 9 |
| 2024 | Math IDE: A Platform for Creating with MathabstractTo inspire student engagement in middle school math, we explore the possibility of using generative AI to enhance the creativity of math learning. We present the Math IDE, a math education environment in which students learn about math concepts by building artifacts. We aimed to create a platform in which students can engage with mathematical concepts, create an artifact that embodies the math that they are learning about, and practice their high-level specification skills. In the current iteration of the Math IDE, students can create custom web pages by describing and demonstrating understanding of the math that is involved in the web page. In this short overview, we describe our process and discuss several open questions regarding the design and application of this novel method of math education. Sierra Wang, John C. Mitchell, Nick Haber, Chris Piech |
SIGCSE (2) | 3 |
| 2023 | Developmental Curiosity and Social Interaction in Virtual Agents
Chris Doyle, Sarah Shader, Michelle Lau, Megumi Sano, Dan Yamins, Nick Haber |
CogSci | 6 |
| 2023 | Measuring and Modeling Physical Intrinsic Motivation
Julio Martinez, Felix J. Binder, Nick Haber, Judith E. Fan, Dan Yamins |
CogSci | 4 |
| 2023 | Characterizing Learning Progress of Problem-Solvers Using Puzzle-Solving Log Data
Fan-Yun Sun, Frieda Rong, Kumiko Nakajima, Nick Haber, Shima Salehi |
EDM | 5 |
| 2023 | Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading EfficiencyabstractDeveloping an educational test can be expensive and time-consuming, as each item must be written by experts and then evaluated by collecting hundreds of student responses.Moreover, many tests require multiple distinct sets of questions administered throughout the school year to closely monitor students' progress, known as parallel tests.In this study, we focus on tests of silent sentence reading efficiency, used to assess students' reading ability over time.To generate high-quality parallel tests, we propose to fine-tune large language models (LLMs) to simulate how previous students would have responded to unseen items.With these simulated responses, we can estimate each item's difficulty and ambiguity.We first use GPT-4 to generate new test items following a list of expert-developed rules and then apply a fine-tuned LLM to filter the items based on criteria from psychological measurements.We also propose an optimal-transport-inspired technique for generating parallel tests and show the generated tests closely correspond to the original test's difficulty and reliability based on crowdworker responses.Our evaluation of a generated test with 234 students from grades 2 to 8 produces test scores highly correlated (r=0.93) to those of a standard test form written by human experts and evaluated across thousands of K-12 students. Eric Zelikman, Wanjing Anya Ma, Jasmine E. Tran, Diyi Yang, Jason D. Yeatman, Nick Haber |
EMNLP | 6 |
| 2023 | Curious Replay for Model-based AdaptationabstractAgents must be able to adapt quickly as an environment changes. We find that existing model-based reinforcement learning agents are unable to do this well, in part because of how they use past experiences to train their world model. Here, we present Curious Replay—a form of prioritized experience replay tailored to model-based agents through use of a curiosity-based priority signal. Agents using Curious Replay exhibit improved performance in an exploration paradigm inspired by animal behavior and on the Crafter benchmark. DreamerV3 with Curious Replay surpasses state-of-the-art performance on Crafter, achieving a mean score of 19.4 that substantially improves on the previous high score of 14.5 by DreamerV3 with uniform replay, while also maintaining similar performance on the Deepmind Control Suite. Code for Curious Replay is available at github.com/AutonomousAgentsLab/curiousreplay. Isaac Kauvar, Chris Doyle, Linqi Zhou, Nick Haber |
ICML | 4 |
| 2023 | Parsel🦆: Algorithmic Reasoning with Language Models by Composing DecompositionsabstractDespite recent success in large language model (LLM) reasoning, LLMs struggle with hierarchical multi-step reasoning tasks like generating complex programs. For these tasks, humans often start with a high-level algorithmic design and implement each part gradually. We introduce Parsel, a framework enabling automatic implementation and validation of complex algorithms with code LLMs. With Parsel, we automatically decompose algorithmic tasks into hierarchical natural language function descriptions and then search over combinations of possible function implementations using tests. We show that Parsel can be used across domains requiring hierarchical reasoning, including program synthesis and robotic planning. We find that, using Parsel, LLMs solve more competition-level problems in the APPS dataset, resulting in pass rates over 75\% higher than prior results from directly sampling AlphaCode and Codex, while often using a smaller sample budget. Moreover, with automatically generated tests, we find that Parsel can improve the state-of-the-art pass@1 performance on HumanEval from 67\% to 85\%. We also find that LLM-generated robotic plans using Parsel are more than twice as likely to be considered accurate than directly generated plans. Lastly, we explore how Parsel addresses LLM limitations and discuss how Parsel may be useful for human programmers. We release our code at https://github.com/ezelikman/parsel. Eric Zelikman, Qian Huang 0006, Gabriel Poesia, Noah D. Goodman, Nick Haber |
NeurIPS | 5 |
| 2022 | Measuring social curiosity-driven attentional differences in children with autism using an augmented reality-based phone app
Samaher Radwan, Aaron Kline, Alejandro Galindo, Michael C. Frank, Dennis P. Wall, Nick Haber |
CogSci | 6 |
| 2022 | Interaction Modeling with Multiplex AttentionabstractModeling multi-agent systems requires understanding how agents interact. Such systems are often difficult to model because they can involve a variety of types of interactions that layer together to drive rich social behavioral dynamics. Here we introduce a method for accurately modeling multi-agent systems. We present Interaction Modeling with Multiplex Attention (IMMA), a forward prediction model that uses a multiplex latent graph to represent multiple independent types of interactions and attention to account for relations of different strengths. We also introduce Progressive Layer Training, a training strategy for this architecture. We show that our approach outperforms state-of-the-art models in trajectory forecasting and relation inference, spanning three multi-agent scenarios: social navigation, cooperative task achievement, and team sports. We further demonstrate that our approach can improve zero-shot generalization and allows us to probe how different interactions impact agent behavior. Fan-Yun Sun, Isaac Kauvar, Jiachen Li 0001, Mykel J. Kochenderfer, Jiajun Wu 0001, Nick Haber |
NeurIPS | 7 |
| 2020 | Learning in Social Environments with Curious Neural Agents
Megumi Sano, Julian De Freitas, Nick Haber, Dan Yamins |
CogSci | 3 |
| 2020 | Active World Model Learning with Progress CuriosityabstractWorld models are self-supervised predictive models of how the world evolves. Humans learn world models by curiously exploring their environment, in the process acquiring compact abstractions of high bandwidth sensory inputs, the ability to plan across long temporal horizons, and an understanding of the behavioral patterns of other agents. In this work, we study how to design such a curiosity-driven Active World Model Learning (AWML) system. To do so, we construct a curious agent building world models while visually exploring a 3D physical environment rich with distillations of representative real-world agents. We propose an AWML system driven by $\gamma$-Progress: a scalable and effective learning progress-based curiosity signal and show that $\gamma$-Progress naturally gives rise to an exploration policy that directs attention to complex but learnable dynamics in a balanced manner, as a result overcoming the “white noise problem”. As a result, our $\gamma$-Progress-driven controller achieves significantly higher AWML performance than baseline controllers equipped with state-of-the-art exploration strategies such as Random Network Distillation and Model Disagreement. Kuno Kim, Megumi Sano, Julian De Freitas, Nick Haber, Dan Yamins |
ICML | 4 |
| 2018 | Emergence of Structured Behaviors from Curiosity-Based Intrinsic Motivation
Nick Haber, Damian Mrowca, Li Fei-Fei 0001, Dan Yamins |
CogSci | 1 |
| 2018 | Learning to Play With Intrinsically-Motivated, Self-Aware AgentsabstractInfants are experts at playing, with an amazing ability to generate novel structured behaviors in unstructured environments that lack clear extrinsic reward signals. We seek to mathematically formalize these abilities using a neural network that implements curiosity-driven intrinsic motivation. Using a simple but ecologically naturalistic simulated environment in which an agent can move and interact with objects it sees, we propose a "world-model" network that learns to predict the dynamic consequences of the agent's actions. Simultaneously, we train a separate explicit "self-model" that allows the agent to track the error map of its world-model. It then uses the self-model to adversarially challenge the developing world-model. We demonstrate that this policy causes the agent to explore novel and informative interactions with its environment, leading to the generation of a spectrum of complex behaviors, including ego-motion prediction, object attention, and object gathering. Moreover, the world-model that the agent learns supports improved performance on object dynamics prediction, detection, localization and recognition tasks. Taken together, our results are initial steps toward creating flexible autonomous agents that self-supervise in realistic physical environments. Nick Haber, Damian Mrowca, Stephanie Wang, Li Fei-Fei 0001, Dan Yamins |
NeurIPS | 1 |
| 2018 | Flexible neural representation for physics predictionabstractHumans have a remarkable capacity to understand the physical dynamics of objects in their environment, flexibly capturing complex structures and interactions at multiple levels of detail. Inspired by this ability, we propose a hierarchical particle-based object representation that covers a wide variety of types of three-dimensional objects, including both arbitrary rigid geometrical shapes and deformable materials. We then describe the Hierarchical Relation Network (HRN), an end-to-end differentiable neural network based on hierarchical graph convolution, that learns to predict physical dynamics in this representation. Compared to other neural network baselines, the HRN accurately handles complex collisions and nonrigid deformations, generating plausible dynamics predictions at long time scales in novel settings, and scaling to large scene configurations. These results demonstrate an architecture with the potential to form the basis of next-generation physics predictors for use in computer vision, robotics, and quantitative cognitive science. Damian Mrowca, Chengxu Zhuang, Elias Wang, Nick Haber, Li Fei-Fei 0001, Josh Tenenbaum, Dan Yamins |
NeurIPS | 4 |
| 2016 | A practical approach to real-time neutral feature subtraction for facial expression recognitionabstractMethods for automated facial expression recognition - identifying faces as happy, sad, angry, etc. - typically rely on the classification of features extracted from images. These features, designed to encode shape and texture information, depend on both (1) the expression an individual is making, and (2) the individual's physical characteristics and lighting conditions of the image. To reduce the effect of (2), a common strategy is to establish a "baseline" for an individual and subtract out this individual's baseline neutral feature. This extra neutral feature information often is not available - in particular for in-the-wild, real-time classification of a previously unseen subject. Thus, in order to implement "neutral subtraction," one must estimate the individual's neutral feature. Existing methods to do this are susceptible to class imbalance at test time (e.g., averaging over all facial features), require a more complex model specific to the individual to be trained, or are restricted to features computed entirely from tracked landmark points (taking advantage of a subset of "stable points" which move little as an individual emotes). We extend neutral subtraction to different computer vision feature spaces as a method to correct for inter-face and lighting variance. We further propose a simple, real-time method which is robust to class imbalance and in principal works over a wide class of feature choices. We test this method on feature extraction techniques that lead to high baseline accuracy without neutral subtraction (97% on the Extended Cohn-Kanade Dataset). We find that on difficult classification tasks our method recovers almost 2/3 of the ~ 8% gain shown by a "cheating" neutral-subtracted feature classifier, which uses examples that have been labeled as neutral, validating with both HOG and SIFT features. Nick Haber, Catalin Voss, Azar Fazel, Terry Winograd, Dennis P. Wall |
WACV | 1 |