Noah D. Goodman

dblp:96/1216 · DBLP profile ↗
← Back
160ranked-venue papers
5as first author
60since 2021 · last 2025
0000-0002-9176-8802ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 151 · 4 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 80 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Language and Experience: A Computational Model of Social Learning in Complex Novel Tasks
Cédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler, Noah D. Goodman, Jacob Andreas, Josh Tenenbaum
CogSci5
2025 Thinking fast, slow, and everywhere in between in humans and language models
Ben Prystawski, Noah D. Goodman
CogSci2
2025 Non-literal Understanding of Number Words by Language Models
Polina Tsvilodub, Kanishk Gandhi, Jan-Philipp Fränken, Michael Franke, Noah D. Goodman
CogSci6
2025 Scaling up the think-aloud method
Daniel Wurgaft, Ben Prystawski, Kanishk Gandhi, Cedegao E. Zhang, Josh Tenenbaum, Noah D. Goodman
CogSci6
2025 Value Profiles for Encoding Human Variation
abstract
Taylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah Goodman, Verena Rieser. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Taylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah D. Goodman, Verena Rieser
EMNLP8
2025 What Makes a Maze Look Like a Maze?
abstract
A unique aspect of human visual understanding is the ability to flexibly interpret abstract concepts: acquiring lifted rules explaining what they symbolize, grounding them across familiar and unfamiliar contexts, and making predictions or reasoning about them. While off-the-shelf vision-language models excel at making literal interpretations of images (e.g., recognizing object categories such as tree branches), they still struggle to make sense of such visual abstractions (e.g., how an arrangement of tree branches may form the walls of a maze). To address this challenge, we introduce Deep Schema Grounding (DSG), a framework that leverages explicit structured representations of visual abstractions for grounding and reasoning. At the core of DSG are schemas—dependency graph descriptions of abstract concepts that decompose them into more primitive-level symbols. DSG uses large language models to extract schemas, then hierarchically grounds concrete to abstract components of the schema onto images with vision-language models. The grounded schema is used to augment visual abstraction understanding. We systematically evaluate DSG and different methods in reasoning on our new Visual Abstractions Benchmark, which consists of diverse, real-world images of abstract concepts and corresponding question-answer pairs labeled by humans. We show that DSG significantly improves the abstract visual reasoning performance of vision-language models, and is a step toward human-aligned understanding of visual abstractions.
Joy Hsu, Jiayuan Mao, Josh Tenenbaum, Noah D. Goodman, Jiajun Wu 0001
ICLR4
2025 Eliciting Human Preferences with Language Models
abstract
Language models (LMs) can be directed to perform user- and context-dependent tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts can be challenging---especially in tasks that require users to precisely articulate nebulous preferences or reason about complex edge cases. For such tasks, we introduce **Generative Active Task Elicitation (GATE)**, a method for using *LMs themselves* to guide the task specification process. GATE is a learning framework in which models elicit and infer human preferences through free-form, language-based interaction with users. We identify prototypical challenges that users face when specifying preferences, and design three preference modeling tasks to study these challenges: content recommendation, moral reasoning, and email validation. In preregistered experiments, we show that LMs that learn to perform these tasks using GATE (by interactively querying users with open-ended questions) obtain preference specifications that are more informative than user-written prompts or examples. GATE matches existing task specification methods in the moral reasoning task, and significantly outperforms them in the content recommendation and email validation tasks. Users additionally report that interactive task elicitation requires less effort than prompting or example labeling and surfaces considerations that they did not anticipate on their own. Our findings suggest that LM-driven elicitation can be a powerful tool for aligning models to complex human preferences and values.
Belinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob Andreas
ICLR3
2025 In-Context Learning Strategies Emerge Rationally
abstract
Recent work analyzing in-context learning (ICL) has identified a broad set of strategies that describe model behavior in different experimental conditions. We aim to unify these findings by asking why a model learns these disparate strategies in the first place. Specifically, we start with the observation that when trained to learn a mixture of tasks, as is popular in the literature, the strategies learned by a model for performing ICL can be captured by a family of Bayesian predictors: a memorizing predictor, which assumes a discrete prior on the set of seen tasks, and a generalizing predictor, where the prior matches the underlying task distribution. Adopting the normative lens of rational analysis, where a learner’s behavior is explained as an optimal adaptation to data given computational constraints, we develop a hierarchical Bayesian framework that almost perfectly predicts Transformer next- token predictions throughout training—without assuming access to its weights. Under this framework, pretraining is viewed as a process of updating the posterior probability of different strategies, and inference-time behavior as a posterior- weighted average over these strategies’ predictions. Our framework draws on common assumptions about neural network learning dynamics, which make explicit a tradeoff between loss and complexity among candidate strategies: beyond how well it explains the data, a model’s preference towards implementing a strategy is dictated by its complexity. This helps explain well-known ICL phenomena, while offering novel predictions: e.g., we show a superlinear trend in the timescale for transitioning from generalization to memorization as task diversity increases. Overall, our work advances an explanatory and predictive account of ICL grounded in tradeoffs between strategy loss and complexity.
Daniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka, Gautam Reddy, Noah D. Goodman
NeurIPS6
2025 Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
abstract
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering.
Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang 0014, Aryaman Arora, Zhengxuan Wu, Noah D. Goodman, Christopher Potts, Thomas Icard
J. Mach. Learn. Res.9
2025 Automated Discovery of Tactic Libraries for Interactive Theorem Proving
abstract
Enabling more concise and modular proofs is essential for advancing formal reasoning using interactive theorem provers (ITPs). Since many ITPs, such as Rocq and Lean, use tactic-style proofs, learning higher-level custom tactics is crucial for proof modularity and automation. This paper presents a novel approach to tactic discovery, which leverages Tactic Dependence Graphs (TDGs) to identify reusable proof strategies across multiple proofs. TDGs capture logical dependencies between tactic applications while abstracting away irrelevant syntactic details, allowing for both the discovery of new tactics and the refactoring of existing proofs into more modular forms. We have implemented this technique in a tool called TacMiner and compare it against an anti-unification-based approach ( Peano ) to tactic discovery. Our evaluation demonstrates that TacMiner can learn 3× as many tactics as Peano and reduces the size of proofs by 26% across all benchmarks. Furthermore, our evaluation demonstrates the benefits of learning custom tactics for proof automation, allowing a state-of-the-art proof automation tool to achieve a relative increase of 172% in terms of success rate.
Yutong Xin, Jimmy Xin, Gabriel Poesia, Noah D. Goodman, Qiaochu Chen, Isil Dillig
Proc. ACM Program. Lang.4
2024 Naturalistic Transmission of Causal Knowledge between Machines and Humans
Cédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler, Noah D. Goodman, Jacob Andreas, Josh Tenenbaum
CogSci5
2024 Procedural Dilemma Generation for Moral Reasoning in Humans and Language Models
Jan-Philipp Fränken, Kanishk Gandhi, Tori Qiu, Ayesha Khawaja, Noah D. Goodman, Tobias Gerstenberg
CogSci5
2024 Symbolic Variables in Distributed Networks that Count
Satchel Grant, Zhengxuan Wu, James L. McClelland, Noah D. Goodman
CogSci4
2024 Evaluating and Optimizing Educational Content with Large Language Model Judgments
Joy He-Yueya, Noah D. Goodman, Emma Brunskill
EDM2
2024 Is Child-Directed Speech Effective Training Data for Language Models?
abstract
While high-performing language models are typically trained on hundreds of billions of words, human children become fluent language users with a much smaller amount of data.What are the features of the data they receive, and how do these features support language modeling objectives?To investigate this question, we train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset (TinyDialogues), comparing to OpenSubtitles, Wikipedia, and a heterogeneous blend of datasets from the BabyLM challenge.We evaluate the syntactic and semantic knowledge of these models using developmentallyinspired evaluations.Through pretraining experiments, we test whether the global developmental ordering or the local discourse ordering of children's training data supports high performance relative to other datasets.The local properties of the data affect model results, but surprisingly, global properties do not.Further, child language input is not uniquely valuable for training language models.These findings support the hypothesis that, rather than proceeding from better data, the child's learning algorithm is substantially more data-efficient than current language modeling techniques.
Steven Y. Feng, Noah D. Goodman, Michael Frank 0005
EMNLP2
2024 Hypothesis Search: Inductive Reasoning with Language Models
abstract
Inductive reasoning is a core problem-solving capacity: humans can identify underlying principles from a few examples, which can then be robustly generalized to novel scenarios. Recent work has evaluated large language models (LLMs) on inductive reasoning tasks by directly prompting them yielding "in context learning." This can work well for straightforward inductive tasks, but performs very poorly on more complex tasks such as the Abstraction and Reasoning Corpus (ARC). In this work, we propose to improve the inductive reasoning ability of LLMs by generating explicit hypotheses at multiple levels of abstraction: we prompt the LLM to propose multiple abstract hypotheses about the problem, in natural language, then implement the natural language hypotheses as concrete Python programs. These programs can be directly verified by running on the observed examples and generalized to novel inputs. To reduce the hypothesis search space, we explore steps to filter the set of hypotheses to be implemented as programs: we either ask the LLM to summarize them into a smaller set of hypotheses, or ask human annotators to select a subset. We verify our pipeline's effectiveness on the ARC visual inductive reasoning benchmark, its variant 1D-ARC, and string transformation dataset SyGuS. On a random 40-problem subset of ARC, our automated pipeline using LLM summaries achieves 27.5% accuracy, significantly outperforming the direct prompting baseline (accuracy of 12.5%). With the minimal human input of selecting from LLM-generated candidates, the performance is boosted to 37.5%. Our ablation studies show that abstract hypothesis generation and concrete program representations are both beneficial for LLMs to perform inductive reasoning tasks.
Ruocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu, Nick Haber, Noah D. Goodman
ICLR6
2024 Automated Statistical Model Discovery with Language Models
abstract
Statistical model discovery is a challenging search over a vast space of models subject to domain-specific constraints. Efficiently searching over this space requires expertise in modeling and the problem domain. Motivated by the domain knowledge and programming capabilities of large language models (LMs), we introduce a method for language model driven automated statistical model discovery. We cast our automated procedure within the principled framework of Box’s Loop: the LM iterates between proposing statistical models represented as probabilistic programs, acting as a modeler, and critiquing those models, acting as a domain expert. By leveraging LMs, we do not have to define a domain-specific language of models or design a handcrafted search procedure, which are key restrictions of previous systems. We evaluate our method in three settings in probabilistic modeling: searching within a restricted space of models, searching over an open-ended space, and improving expert models under natural language constraints (e.g., this model should be interpretable to an ecologist). Our method identifies models on par with human expert designed models and extends classic models in interpretable ways. Our results highlight the promise of LM-driven model discovery.
Michael Y. Li, Emily B. Fox, Noah D. Goodman
ICML3
2024 Codebook Features: Sparse and Discrete Interpretability for Neural Networks
abstract
Understanding neural networks is challenging in part because of the dense, continuous nature of their hidden states. We explore whether we can train neural networks to have hidden states that are sparse, discrete, and more interpretable by quantizing their continuous features into what we call codebook features. Codebook features are produced by finetuning neural networks with vector quantization bottlenecks at each layer, producing a network whose hidden features are the sum of a small number of discrete vector codes chosen from a larger codebook. Surprisingly, we find that neural networks can operate under this extreme bottleneck with only modest degradation in performance. In addition, we can control a model’s behavior by finding codes that activate on a desired behavior, then activating those same codes during generation. We first validate codebook features on a finite state machine dataset with far more hidden states than neurons. In this setting, our approach overcomes the superposition problem by assigning states to distinct codes, and we find that we can make the neural network behave as if it is in a different state by activating the code for that state. We then train Transformer language models with up to 410M parameters on two natural language datasets. We identify codes in these models representing diverse, disentangled concepts (ranging from negative emotions to months of the year) and find that we can guide the model to generate different topics and pronoun genders by activating these codes during inference. Overall, codebook features appear to be a promising unit of analysis and control for neural networks and interpretability. Our codebase and models are open-sourced at this URL.
Alex Tamkin, Mohammad Taufeeque, Noah D. Goodman
ICML3
2024 Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
abstract
When prompting a language model (LM), users often expect the model to adhere to a set of behavioral principles across diverse tasks, such as producing insightful content while avoiding harmful or biased language. Instilling such principles (i.e., a constitution) into a model is resource-intensive, technically challenging, and generally requires human preference labels or examples. We introduce SAMI, an iterative algorithm that finetunes a pretrained language model (without requiring preference labels or demonstrations) to increase the conditional mutual information between constitutions and self-generated responses given queries from a dataset. On single-turn dialogue and summarization, a SAMI-trained mistral-7b outperforms the initial pretrained model, with win rates between 66% and 77%. Strikingly, it also surpasses an instruction-finetuned baseline (mistral-7b-instruct) with win rates between 55% and 57% on single-turn dialogue. SAMI requires a model that writes the principles. To avoid dependence on strong models for writing principles, we align a strong pretrained model (mixtral-8x7b) using constitutions written by a weak instruction-finetuned model (mistral-7b-instruct), achieving a 65% win rate on summarization. Finally, we investigate whether SAMI generalizes to diverse summarization principles (e.g., "summaries should be scientific") and scales to stronger models (llama3-70b), finding that it achieves win rates of up to 68% for learned and 67% for held-out principles compared to the base model. Our results show that a pretrained LM can learn to follow constitutions without using preference labels, demonstrations, or human oversight.
Jan-Philipp Fränken, Eric Zelikman, Rafael Rafailov, Kanishk Gandhi, Tobias Gerstenberg, Noah D. Goodman
NeurIPS6
2024 On scalable oversight with weak LLMs judging strong LLMs
abstract
Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, where a single AI tries to convince a judge that asks questions; and compare to a baseline of direct question-answering, where the judge just answers outright without the AI. We use large language models (LLMs) as both AI agents and as stand-ins for human judges, taking the judge models to be weaker than agent models. We benchmark on a diverse range of asymmetries between judges and agents, extending previous work on a single extractive QA task with information asymmetry, to also include mathematics, coding, logic and multimodal reasoning asymmetries. We find that debate outperforms consultancy across all tasks when the consultant is randomly assigned to argue for the correct/incorrect answer. Comparing debate to direct question answering, the results depend on the type of task: in extractive QA tasks with information asymmetry debate outperforms direct question answering, but in other tasks without information asymmetry the results are mixed. Previous work assigned debaters/consultants an answer to argue for. When we allow them to instead choose which answer to argue for, we find judges are less frequently convinced by the wrong answer in debate than in consultancy. Further, we find that stronger debater models increase judge accuracy, though more modestly than in previous studies.
Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, Rohin Shah
NeurIPS10
2024 Learning Formal Mathematics From Intrinsic Motivation
abstract
How did humanity coax mathematics from the aether? We explore the Platonic view that mathematics can be discovered from its axioms---a game of conjecture and proof. We describe an agent that jointly learns to pose challenging problems for itself (conjecturing) and solve them (theorem proving). Given a mathematical domain axiomatized in dependent type theory, we first combine methods for constrained decoding and type-directed synthesis to sample valid conjectures from a language model. Our method guarantees well-formed conjectures by construction, even as we start with a randomly initialized model. We use the same model to represent a policy and value function for guiding proof search. Our agent targets generating hard but provable conjectures --- a moving target, since its own theorem proving ability also improves as it trains. We propose novel methods for hindsight relabeling on proof search trees to significantly improve the agent's sample efficiency in both tasks. Experiments on 3 axiomatic domains (propositional logic, arithmetic and group theory) demonstrate that our agent can bootstrap from only the axioms, self-improving in generating true and challenging conjectures and in finding proofs.
Gabriel Poesia, David Broman, Nick Haber, Noah D. Goodman
NeurIPS4
2023 Cultural reinforcement learning: a framework for modeling cumulative culture on a limited channel
Ben Prystawski, Dilip Arumugam, Noah D. Goodman
CogSci3
2023 Psychologically-informed chain-of-thought prompts for metaphor understanding in large language models
Ben Prystawski, Paul H. Thibodeau, Christopher Potts, Noah D. Goodman
CogSci4
2023 Overinformative Question Answering by Humans and Machines
Polina Tsvilodub, Michael Franke, Robert D. Hawkins, Noah D. Goodman
CogSci4
2023 Characterizing tradeoffs between teaching via language and demonstrations in multi-agent systems
Dhara Yu, Noah D. Goodman, Jesse Mu
CogSci2
2023 Task Ambiguity in Humans and Language Models
Alex Tamkin, Kunal Handa, Avash Shrestha, Noah D. Goodman
ICLR4
2023 Generating Language Corrections for Teaching Physical Control Tasks
abstract
AI assistance continues to help advance applications in education, from language learning to intelligent tutoring systems, yet current methods for providing students feedback are still quite limited. Most automatic feedback systems either provide binary correctness feedback, which may not help a student understand how to improve, or require hand-coding feedback templates, which may not generalize to new domains. This can be particularly challenging for physical control tasks, where the rich diversity in student behavior and specialized domains make it challenging to leverage general-purpose assistive tools for providing feedback. We design and build CORGI, a model trained to generate language corrections for physical control tasks, such as learning to ride a bike. CORGI takes in as input a pair of student and expert trajectories, and then generates natural language corrections to help the student improve. We collect and train CORGI over data from three diverse physical control tasks (drawing, steering, and joint movement). Through both automatic and human evaluations, we show that CORGI can (i) generate valid feedback for novel student trajectories, (ii) outperform baselines on domains with novel control dynamics, and (iii) improve student learning in an interactive drawing task.
Megha Srivastava, Noah D. Goodman, Dorsa Sadigh
ICML2
2023 Understanding Social Reasoning in Language Models with Language Models
abstract
As Large Language Models (LLMs) become increasingly integrated into our everyday lives, understanding their ability to comprehend human mental states becomes critical for ensuring effective interactions. However, despite the recent attempts to assess the Theory-of-Mind (ToM) reasoning capabilities of LLMs, the degree to which these models can align with human ToM remains a nuanced topic of exploration. This is primarily due to two distinct challenges: (1) the presence of inconsistent results from previous evaluations, and (2) concerns surrounding the validity of existing evaluation methodologies. To address these challenges, we present a novel framework for procedurally generating evaluations with LLMs by populating causal templates. Using our framework, we create a new social reasoning benchmark (BigToM) for LLMs which consists of 25 controls and 5,000 model-written evaluations. We find that human participants rate the quality of our benchmark higher than previous crowd-sourced evaluations and comparable to expert-written evaluations. Using BigToM, we evaluate the social reasoning capabilities of a variety of LLMs and compare model performances with human performance. Our results suggest that GPT4 has ToM capabilities that mirror human inference patterns, though less reliable, while other LLMs struggle.
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, Noah D. Goodman
NeurIPS4
2023 Learning to Compress Prompts with Gist Tokens
abstract
Prompting is the primary way to utilize the multitask capabilities of language models (LMs), but prompts occupy valuable space in the input context window, and repeatedly encoding the same prompt is computationally inefficient. Finetuning and distillation methods allow for specialization of LMs without prompting, but require retraining the model for each task. To avoid this trade-off entirely, we present gisting, which trains an LM to compress prompts into smaller sets of "gist" tokens which can be cached and reused for compute efficiency. Gist models can be trained with no additional cost over standard instruction finetuning by simply modifying Transformer attention masks to encourage prompt compression. On decoder (LLaMA-7B) and encoder-decoder (FLAN-T5-XXL) LMs, gisting enables up to 26x compression of prompts, resulting in up to 40% FLOPs reductions, 4.2% wall time speedups, and storage savings, all with minimal loss in output quality.
Jesse Mu, Xiang Li 0063, Noah D. Goodman
NeurIPS3
2023 Why think step by step? Reasoning emerges from the locality of experience
abstract
Humans have a powerful and mysterious capacity to reason. Working through a set of mental steps enables us to make inferences we would not be capable of making directly even though we get no additional data from the world. Similarly, when large language models generate intermediate steps (a chain of thought) before answering a question, they often produce better answers than they would directly. We investigate why and how chain-of-thought reasoning is useful in language models, testing the hypothesis that reasoning is effective when training data consists of overlapping local clusters of variables that influence each other strongly. These training conditions enable the chaining of accurate local inferences to estimate relationships between variables that were not seen together in training. We prove that there will exist a "reasoning gap", where reasoning through intermediate variables reduces bias, for the simple case of an autoregressive density estimator trained on local samples from a chain-structured probabilistic model. We then test our hypothesis experimentally in more complex models, training an autoregressive language model on samples from Bayes nets but only including a subset of variables in each sample. We test language models’ ability to match conditional probabilities with and without intermediate reasoning steps, finding that intermediate steps are only helpful when the training data is locally structured with respect to dependencies between variables. The combination of locally structured observations and reasoning is much more data-efficient than training on all variables. Our results illustrate how the effectiveness of reasoning step by step is rooted in the local statistical structure of the training data.
Ben Prystawski, Michael Li, Noah D. Goodman
NeurIPS3
2023 Feature Dropout: Revisiting the Role of Augmentations in Contrastive Learning
abstract
What role do augmentations play in contrastive learning? Recent work suggests that good augmentations are label-preserving with respect to a specific downstream task. We complicate this picture by showing that label-destroying augmentations can be useful in the foundation model setting, where the goal is to learn diverse, general-purpose representations for multiple downstream tasks. We perform contrastive learning experiments on a range of image and audio datasets with multiple downstream tasks (e.g. for digits superimposed on photographs, predicting the class of one vs. the other). We find that Viewmaker Networks, a recently proposed model for learning augmentations for contrastive learning, produce label-destroying augmentations that stochastically destroy features needed for different downstream tasks. These augmentations are interpretable (e.g. altering shapes, digits, or letters added to images) and surprisingly often result in better performance compared to expert-designed augmentations, despite not preserving label information. To support our empirical results, we theoretically analyze a simple contrastive learning setting with a linear model. In this setting, label-destroying augmentations are crucial for preventing one set of features from suppressing the learning of features useful for another downstream task. Our results highlight the need for analyzing the interaction between multiple downstream tasks when trying to explain the success of foundation models.
Alex Tamkin, Margalit Glasgow, Xiluo He, Noah D. Goodman
NeurIPS4
2023 Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
abstract
Obtaining human-interpretable explanations of large, general-purpose language models is an urgent goal for AI safety. However, it is just as important that our interpretability methods are faithful to the causal dynamics underlying model behavior and able to robustly generalize to unseen inputs. Distributed Alignment Search (DAS) is a powerful gradient descent method grounded in a theory of causal abstraction that uncovered perfect alignments between interpretable symbolic algorithms and small deep learning models fine-tuned for specific tasks. In the present paper, we scale DAS significantly by replacing the remaining brute-force search steps with learned parameters -- an approach we call Boundless DAS. This enables us to efficiently search for interpretable causal structure in large language models while they follow instructions. We apply Boundless DAS to the Alpaca model (7B parameters), which, off the shelf, solves a simple numerical reasoning problem. With Boundless DAS, we discover that Alpaca does this by implementing a causal model with two interpretable boolean variables. Furthermore, we find that the alignment of neural representations with these variables is robust to changes in inputs and instructions. These findings mark a first step toward deeply understanding the inner-workings of our largest and most widely deployed language models.
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, Noah D. Goodman
NeurIPS5
2023 Parsel🦆: Algorithmic Reasoning with Language Models by Composing Decompositions
abstract
Despite recent success in large language model (LLM) reasoning, LLMs struggle with hierarchical multi-step reasoning tasks like generating complex programs. For these tasks, humans often start with a high-level algorithmic design and implement each part gradually. We introduce Parsel, a framework enabling automatic implementation and validation of complex algorithms with code LLMs. With Parsel, we automatically decompose algorithmic tasks into hierarchical natural language function descriptions and then search over combinations of possible function implementations using tests. We show that Parsel can be used across domains requiring hierarchical reasoning, including program synthesis and robotic planning. We find that, using Parsel, LLMs solve more competition-level problems in the APPS dataset, resulting in pass rates over 75\% higher than prior results from directly sampling AlphaCode and Codex, while often using a smaller sample budget. Moreover, with automatically generated tests, we find that Parsel can improve the state-of-the-art pass@1 performance on HumanEval from 67\% to 85\%. We also find that LLM-generated robotic plans using Parsel are more than twice as likely to be considered accurate than directly generated plans. Lastly, we explore how Parsel addresses LLM limitations and discuss how Parsel may be useful for human programmers. We release our code at https://github.com/ezelikman/parsel.
Eric Zelikman, Qian Huang 0006, Gabriel Poesia, Noah D. Goodman, Nick Haber
NeurIPS4
2022 Automated generation of sentence reading fluency test items
Julia White 0001, Amy Burkhardt, Jason D. Yeatman, Noah D. Goodman
CogSci4
2022 Color Overmodification Emerges from Data-Driven Learning and Pragmatic Reasoning
Fei Fang 0005, Kunal Sinha, Noah D. Goodman, Christopher Potts, Elisa Kreiss
CogSci3
2022 Two's company but six is a crowd: emergence of conventions in multiparty communication games
Veronica Boyce, Robert D. Hawkins, Noah D. Goodman, Michael C. Frank
CogSci3
2022 Left to the Reader: Abstracting Solutions in Mathematical Reasoning
Gabriel Poesia, Noah D. Goodman
CogSci2
2022 Mixed-effects transformers for hierarchical adaptation
abstract
Language differs dramatically from context to context.To some degree, large language models like GPT-3 account for such variation by conditioning on strings of initial input text, or prompts.However, prompting can be ineffective when contexts are sparse, out-of-sample, or extra-textual.In this paper, we introduce the mixed-effects transformer (MET), a novel approach for learning hierarchically-structured prefixes-lightweight modules prepended to an input sequence-to account for structured variation in language use.Specifically, we show how the popular class of mixedeffects regression models may be extended to transformer-based architectures using a regularized prefix-tuning procedure with dropout.We evaluate this approach on several domainadaptation benchmarks, finding that it learns contextual variation from minimal data while generalizing well to unseen contexts.
Julia White 0001, Noah D. Goodman, Robert D. Hawkins
EMNLP2
2022 Concadia: Towards Image-Based Text Generation with a Purpose
abstract
Current deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice.We argue that to close this gap, it is vital to distinguish descriptions from captions based on their distinct communicative roles.Descriptions focus on visual features and are meant to replace an image (often to increase accessibility), whereas captions appear alongside an image to supply additional information.To motivate this distinction and help people put it into practice, we introduce the publicly available Wikipedia-based dataset Concadia consisting of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.Using insights from Concadia, models trained on it, and a preregistered human-subjects experiment with human-and model-generated texts, we characterize the commonalities and differences between descriptions and captions.In addition, we show that, for generating both descriptions and captions, it is useful to augment image-totext models with representations of the textual context in which the image appeared.split datapoints unique articles avg length (words) avg word length vocab size train 77,534 31,240 caption: 12.79
Elisa Kreiss, Fei Fang 0005, Noah D. Goodman, Christopher Potts
EMNLP3
2022 Language modeling via stochastic processes
Rose E. Wang, Esin Durmus, Noah D. Goodman, Tatsunori B. Hashimoto
ICLR3
2022 Inducing Causal Structure for Interpretable Neural Networks
abstract
In many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashion. To achieve this, we present the new method of interchange intervention training (IIT). In IIT, we (1) align variables in a causal model (e.g., a deterministic program or Bayesian network) with representations in a neural model and (2) train the neural model to match the counterfactual behavior of the causal model on a base input when aligned representations in both models are set to be the value they would be for a source input. IIT is fully differentiable, flexibly combines with other objectives, and guarantees that the target causal model is a causal abstraction of the neural model when its loss is zero. We evaluate IIT on a structural vision task (MNIST-PVR), a navigational language task (ReaSCAN), and a natural language inference task (MQNLI). We compare IIT against multi-task training objectives and data augmentation. In all our experiments, IIT achieves the best results and produces neural models that are more interpretable in the sense that they more successfully realize the target causal model.
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, Christopher Potts
ICML7
2022 Causal Distillation for Language Models
abstract
Zhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah Goodman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhengxuan Wu, Atticus Geiger, Josh Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah D. Goodman
NAACL-HLT8
2022 Geoclidean: Few-Shot Generalization in Euclidean Geometry
abstract
Euclidean geometry is among the earliest forms of mathematical thinking. While the geometric primitives underlying its constructions, such as perfect lines and circles, do not often occur in the natural world, humans rarely struggle to perceive and reason with them. Will computer vision models trained on natural images show the same sensitivity to Euclidean geometry? Here we explore these questions by studying few-shot generalization in the universe of Euclidean geometry constructions. We introduce Geoclidean, a domain-specific language for Euclidean geometry, and use it to generate two datasets of geometric concept learning tasks for benchmarking generalization judgements of humans and machines. We find that humans are indeed sensitive to Euclidean geometry and generalize strongly from a few visual examples of a geometric concept. In contrast, low-level and high-level visual features from standard computer vision models pretrained on natural images do not support correct generalization. Thus Geoclidean represents a novel few-shot generalization benchmark for geometric concept learning, where the performance of humans and of AI models diverge. The Geoclidean framework and dataset are publicly available for download.
Joy Hsu, Jiajun Wu 0001, Noah D. Goodman
NeurIPS3
2022 CLEVRER-Humans: Describing Physical and Causal Events the Human Way
abstract
Building machines that can reason about physical events and their causal relationships is crucial for flexible interaction with the physical world. However, most existing physical and causal reasoning benchmarks are exclusively based on synthetically generated events and synthetic natural language descriptions of the causal relationships. This design brings up two issues. First, there is a lack of diversity in both event types and natural language descriptions; second, causal relationships based on manually-defined heuristics are different from human judgments. To address both shortcomings, we present the CLEVRER-Humans benchmark, a video reasoning dataset for causal judgment of physical events with human labels. We employ two techniques to improve data collection efficiency: first, a novel iterative event cloze task to elicit a new representation of events in videos, which we term Causal Event Graphs (CEGs); second, a data augmentation technique based on neural language generative models. We convert the collected CEGs into questions and answers to be consistent with prior work. Finally, we study a collection of baseline approaches for CLEVRER-Humans question-answering, highlighting great challenges set forth by our benchmark.
Jiayuan Mao, Xuelin Yang, Xikun Zhang 0001, Noah D. Goodman, Jiajun Wu 0001
NeurIPS4
2022 Improving Intrinsic Exploration with Language Abstractions
abstract
Reinforcement learning (RL) agents are particularly hard to train when rewards are sparse. One common solution is to use intrinsic rewards to encourage agents to explore their environment. However, recent intrinsic exploration methods often use state-based novelty measures which reward low-level exploration and may not scale to domains requiring more abstract skills. Instead, we explore natural language as a general medium for highlighting relevant abstractions in an environment. Unlike previous work, we evaluate whether language can improve over existing exploration methods by directly extending (and comparing to) competitive intrinsic exploration baselines: AMIGo (Campero et al., 2021) and NovelD (Zhang et al., 2021). These language-based variants outperform their non-linguistic forms by 47-85% across 13 challenging tasks from the MiniGrid and MiniHack environment suites.
Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah D. Goodman, Tim Rocktäschel, Edward Grefenstette
NeurIPS5
2022 Assistive Teaching of Motor Control Tasks to Humans
abstract
Recent works on shared autonomy and assistive-AI technologies, such as assistive robotic teleoperation, seek to model and help human users with limited ability in a fixed task. However, these approaches often fail to account for humans' ability to adapt and eventually learn how to execute a control task themselves. Furthermore, in applications where it may be desirable for a human to intervene, these methods may have inhibited their ability to learn how to succeed with full self-control. In this paper, we focus on the problem of assistive teaching of motor control tasks such as parking a car or landing an aircraft. Despite their ubiquitous role in humans' daily activities and occupations, motor tasks are rarely taught in a uniform way due to their high complexity and variance. We propose an AI-assisted teaching algorithm that leverages skill discovery methods from reinforcement learning (RL) literature to (i) break down any motor control task into teachable skills, (ii) construct novel drill sequences, and (iii) individualize curricula to students with different capabilities. Through an extensive mix of synthetic and user studies on two motor control tasks - parking a car with a joystick and writing characters from the Balinese alphabet - we show that assisted teaching with skills improve student performance by around 40% compared to practicing full trajectories without skills, and practicing with individualized drills can result in up to 25% further improvement.
Megha Srivastava, Erdem Biyik, Suvir Mirchandani, Noah D. Goodman, Dorsa Sadigh
NeurIPS4
2022 DABS 2.0: Improved Datasets and Algorithms for Universal Self-Supervision
abstract
Universal self-supervised (SSL) algorithms hold enormous promise for making machine learning accessible to high-impact domains such as protein biology, manufacturing, and genomics. We present DABS 2.0: a set of improved datasets and algorithms for advancing research on universal SSL. We extend the recently-introduced DABS benchmark with the addition of five real-world science and engineering domains: protein biology, bacterial genomics, multispectral satellite imagery, semiconductor wafers, and particle physics, bringing the total number of domains in the benchmark to twelve. We also propose a new universal SSL algorithm, Capri, and a generalized version of masked autoencoding, and apply both on all twelve domains---the most wide-ranging exploration of SSL yet. We find that multiple algorithms show gains across domains, outperforming previous baselines. In addition, we demonstrate the usefulness of DABS for scientific study of SSL by investigating the optimal corruption rate for each algorithm, showing that the best setting varies based on the domain. Code will be released at http://github.com/alextamkin/dabs}{http://github.com/alextamkin/dabs
Alex Tamkin, Gaurab Banerjee, Mohamed Owda, Shashank Rammoorthy, Noah D. Goodman
NeurIPS6
2022 Active Learning Helps Pretrained Models Learn the Intended Task
abstract
Models can fail in unpredictable ways during deployment due to task ambiguity, when multiple behaviors are consistent with the provided training data. An example is an object classifier trained on red squares and blue circles: when encountering blue squares, the intended behavior is undefined. We investigate whether pretrained models are better active learners, capable of disambiguating between the possible tasks a user may be trying to specify. Intriguingly, we find that better active learning is an emergent property of the pretraining process: pretrained models require up to 5 times fewer labels when using uncertainty-based active learning, while non-pretrained models see no or even negative benefit. We find these gains come from an ability to select examples with attributes that disambiguate the intended behavior, such as rare product categories or atypical backgrounds. These attributes are far more linearly separable in pretrained model's representation spaces vs non-pretrained models, suggesting a possible mechanism for this behavior.
Alex Tamkin, Salil Deshpande, Jesse Mu, Noah D. Goodman
NeurIPS5
2022 Foundation Posteriors for Approximate Probabilistic Inference
abstract
Probabilistic programs provide an expressive representation language for generative models. Given a probabilistic program, we are interested in the task of posterior inference: estimating a latent variable given a set of observed variables. Existing techniques for inference in probabilistic programs often require choosing many hyper-parameters, are computationally expensive, and/or only work for restricted classes of programs. Here we formulate inference as masked language modeling: given a program, we generate a supervised dataset of variables and assignments, and randomly mask a subset of the assignments. We then train a neural network to unmask the random values, defining an approximate posterior distribution. By optimizing a single neural network across a range of programs we amortize the cost of training, yielding a "foundation" posterior able to do zero-shot inference for new programs. The foundation posterior can also be fine-tuned for a particular program and dataset by optimizing a variational inference objective. We show the efficacy of the approach, zero-shot and fine-tuned, on a benchmark of STAN programs.
Mike Wu, Noah D. Goodman
NeurIPS2
2022 STaR: Bootstrapping Reasoning With Reasoning
abstract
Generating step-by-step "chain-of-thought" rationales improves language model performance on complex reasoning tasks like mathematics or commonsense question-answering. However, inducing language model rationale generation currently requires either constructing massive rationale datasets or sacrificing accuracy by using only few-shot inference. We propose a technique to iteratively leverage a small number of rationale examples and a large dataset without rationales, to bootstrap the ability to perform successively more complex reasoning. This technique, the "Self-Taught Reasoner" (STaR), relies on a simple loop: generate rationales to answer many questions, prompted with a few rationale examples; if the generated answers are wrong, try again to generate a rationale given the correct answer; fine-tune on all the rationales that ultimately yielded correct answers; repeat. We show that STaR significantly improves performance on multiple datasets compared to a model fine-tuned to directly predict final answers, and performs comparably to fine-tuning a 30$\times$ larger state-of-the-art language model on CommensenseQA. Thus, STaR lets a model improve itself by learning from its own generated reasoning.
Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman
NeurIPS4
2021 Pragmatic Code Autocomplete
abstract
Human language is ambiguous, with intended meanings recovered via pragmatic reasoning in context. Such reliance on context is essential for the efficiency of human communication. Programming languages, in stark contrast, are defined by unambiguous grammars. In this work, we aim to make programming languages more concise by allowing programmers to utilize a controlled level of ambiguity. Specifically, we allow single-character abbreviations for common keywords and identifiers. Our system first proposes a set of strings that can be abbreviated by the user. Using only 100 abbreviations, we observe that a large dataset of Python code can be compressed by 15%, a number that can be improved even further by specializing the abbreviations to a particular code base. We then use a contextualized sequence-to-sequence model to rank potential expansions of inputs that include abbreviations. In an offline reconstruction task our model achieves accuracies ranging from 93% to 99%, depending on the programming language and user settings. The model is small enough to run on a commodity CPU in real-time. We evaluate the usability of our system in a user study, integrating it in Microsoft VSCode, a popular code text editor. We observe that our system performs well and is complementary to traditional autocomplete features.
Gabriel Poesia, Noah D. Goodman
AAAI2
2021 Generative Grading: Near Human-level Accuracy for Automated Feedback on Richly Structured Problems
Ali Malik, Mike Wu, Vrinda Vasavada, Jinpeng Song, Madison Coots, Noah D. Goodman, Chris Piech
EDM7
2021 Open-domain clarification question generation without question examples
abstract
An overarching goal of natural language processing is to enable machines to communicate seamlessly with humans.However, natural language can be ambiguous or unclear.In cases of uncertainty, humans engage in an interactive process known as repair: asking questions and seeking clarification until their uncertainty is resolved.We propose a framework for building a visually grounded questionasking model capable of producing polar (yesno) clarification questions to resolve misunderstandings in dialogue.Our model uses an expected information gain objective to derive informative questions from an off-the-shelf image captioner without requiring any supervised question-answer data.We demonstrate our model's ability to pose questions that improve communicative success in a goal-oriented 20 questions game with synthetic and human answerers.
Julia White 0001, Gabriel Poesia, Robert D. Hawkins, Dorsa Sadigh, Noah D. Goodman
EMNLP (1)5
2021 Viewmaker Networks: Learning Views for Unsupervised Representation Learning
Alex Tamkin, Mike Wu, Noah D. Goodman
ICLR3
2021 Conditional Negative Sampling for Contrastive Learning of Visual Representations
Mike Wu, Milan Mossé, Chengxu Zhuang, Dan Yamins, Noah D. Goodman
ICLR5
2021 Emergent Communication of Generalizations
abstract
To build agents that can collaborate effectively with others, recent research has trained artificial agents to communicate with each other in Lewis-style referential games. However, this often leads to successful but uninterpretable communication. We argue that this is due to the game objective: communicating about a single object in a shared visual context is prone to overfitting and does not encourage language useful beyond concrete reference. In contrast, human language conveys a rich variety of abstract ideas. To promote such skills, we propose games that require communicating generalizations over sets of objects representing abstract visual concepts, optionally with separate contexts for each agent. We find that these games greatly improve systematicity and interpretability of the learned languages, according to several metrics in the literature. Finally, we propose a method for identifying logical operations embedded in the emergent languages by learning an approximate compositional reconstruction of the language.
Jesse Mu, Noah D. Goodman
NeurIPS2
2021 Contrastive Reinforcement Learning of Symbolic Reasoning Domains
abstract
Abstract symbolic reasoning, as required in domains such as mathematics and logic, is a key component of human intelligence. Solvers for these domains have important applications, especially to computer-assisted education. But learning to solve symbolic problems is challenging for machine learning algorithms. Existing models either learn from human solutions or use hand-engineered features, making them expensive to apply in new domains. In this paper, we instead consider symbolic domains as simple environments where states and actions are given as unstructured text, and binary rewards indicate whether a problem is solved. This flexible setup makes it easy to specify new domains, but search and planning become challenging. We introduce five environments inspired by the Mathematics Common Core Curriculum, and observe that existing Reinforcement Learning baselines perform poorly. We then present a novel learning algorithm, Contrastive Policy Learning (ConPoLe) that explicitly optimizes the InfoNCE loss, which lower bounds the mutual information between the current state and next states that continue on a path to the solution. ConPoLe successfully solves all four domains. Moreover, problem representations learned by ConPoLe enable accurate prediction of the categories of problems in a real mathematics curriculum. Our results suggest new directions for reinforcement learning in symbolic domains, as well as applications to mathematics education.
Gabriel Poesia, Wenxin Dong, Noah D. Goodman
NeurIPS3
2021 Improving Compositionality of Neural Networks by Decoding Representations to Inputs
abstract
In traditional software programs, it is easy to trace program logic from variables back to input, apply assertion statements to block erroneous behavior, and compose programs together. Although deep learning programs have demonstrated strong performance on novel applications, they sacrifice many of the functionalities of traditional software programs. With this as motivation, we take a modest first step towards improving deep learning programs by jointly training a generative model to constrain neural network activations to "decode" back to inputs. We call this design a Decodable Neural Network, or DecNN. Doing so enables a form of compositionality in neural networks, where one can recursively compose DecNN with itself to create an ensemble-like model with uncertainty. In our experiments, we demonstrate applications of this uncertainty to out-of-distribution detection, adversarial example detection, and calibration --- while matching standard neural networks in accuracy. We further explore this compositionality by combining DecNN with pretrained models, where we show promising results that neural networks can be regularized from using protected features.
Mike Wu, Noah D. Goodman, Stefano Ermon
NeurIPS2
2021 Neural Event Semantics for Grounded Language Understanding
abstract
Abstract We present a new conjunctivist framework, neural event semantics (NES), for compositional grounded language understanding. Our approach treats all words as classifiers that compose to form a sentence meaning by multiplying output scores. These classifiers apply to spatial regions (events) and NES derives its semantic structure from language by routing events to different classifier argument inputs via soft attention. NES is trainable end-to-end by gradient descent with minimal supervision. We evaluate our method on compositional grounded language tasks in controlled synthetic and real-world settings. NES offers stronger generalization capability than standard function-based compositional frameworks, while improving accuracy over state-of-the-art neural methods on real-world language tasks.
Shyamal Buch, Li Fei-Fei 0001, Noah D. Goodman
Trans. Assoc. Comput. Linguistics3
2021 Applying Probabilistic Programming to Affective Computing
abstract
Affective Computing is a rapidly growing field spurred by advancements in artificial intelligence, but often, held back by the inability to translate psychological theories of emotion into tractable computational models. To address this, we propose a probabilistic programming approach to affective computing, which models psychological-grounded theories as generative models of emotion, and implements them as stochastic, executable computer programs. We first review probabilistic approaches that integrate reasoning about emotions with reasoning about other latent mental states (e.g., beliefs, desires) in context. Recently-developed probabilistic programming languages offer several key desidarata over previous approaches, such as: (i) flexibility in representing emotions and emotional processes; (ii) modularity and compositionality; (iii) integration with deep learning libraries that facilitate efficient inference and learning from large, naturalistic data; and (iv) ease of adoption. Furthermore, using a probabilistic programming framework allows a standardized platform for theory-building and experimentation: Competing theories (e.g., of appraisal or other emotional processes) can be easily compared via modular substitution of code followed by model comparison. To jumpstart adoption, we illustrate our points with executable code that researchers can easily modify for their own models. We end with a discussion of applications and future directions of the probabilistic programming approach.
Desmond C. Ong, Harold Soh, Jamil Zaki, Noah D. Goodman
IEEE Trans. Affect. Comput.4
2020 Meta-Amortized Variational Inference and Learning
abstract
Despite the recent success in probabilistic modeling and their applications, generative models trained using traditional inference techniques struggle to adapt to new distributions, even when the target distribution may be closely related to the ones seen during training. In this work, we present a doubly-amortized variational inference procedure as a way to address this challenge. By sharing computation across not only a set of query inputs, but also a set of different, related probabilistic models, we learn transferable latent representations that generalize across several related distributions. In particular, given a set of distributions over images, we find the learned representations to transfer to different data transformations. We empirically demonstrate the effectiveness of our method by introducing the MetaVAE, and show that it significantly outperforms baselines on downstream image classification tasks on MNIST (10-50%) and NORB (10-35%).
Mike Wu, Kristy Choi, Noah D. Goodman, Stefano Ermon
AAAI3
2020 Shaping Visual Representations with Language for Few-Shot Classification
abstract
By describing the features and abstractions of our world, language is a crucial tool for human learning and a promising source of supervision for machine learning models.We use language to improve few-shot visual classification in the underexplored scenario where natural language task descriptions are available during training, but unavailable for novel tasks at test time.Existing models for this setting sample new descriptions at test time and use those to classify images.Instead, we propose language-shaped learning (LSL), an end-toend model that regularizes visual representations to predict language.LSL is conceptually simpler, more data efficient, and outperforms baselines in two challenging few-shot domains.
Jesse Mu, Percy Liang, Noah D. Goodman
ACL3
2020 Generalizing meanings from partners to populations: Hierarchical inference supports convention formation on networks
Robert D. Hawkins, Noah D. Goodman, Adele Goldberg 0002, Thomas L. Griffiths 0001
CogSci2
2020 Learning to refer informatively by amortizing pragmatic reasoning
Julia White 0001, Jesse Mu, Noah D. Goodman
CogSci3
2020 Continual Adaptation for Efficient Machine Communication
abstract
To communicate with new partners in new contexts, humans rapidly form new linguistic conventions.Recent neural language models are able to comprehend and produce the existing conventions present in their training data, but are not able to flexibly and interactively adapt those conventions on the fly as humans do.We introduce an interactive repeated reference task as a benchmark for models of adaptation in communication and propose a regularized continual learning framework that allows an artificial agent initialized with a generic language model to more accurately and efficiently communicate with a partner over time.We evaluate this framework through simulations on COCO and in real-time reference game experiments with human partners.
Robert D. Hawkins, Minae Kwon, Dorsa Sadigh, Noah D. Goodman
CoNLL4
2020 Variational Item Response Theory: Fast, Accurate, and Expressive
Mike Wu, Richard Lee Davis, Benjamin W. Domingue, Chris Piech, Noah D. Goodman
EDM5
2020 Language Through a Prism: A Spectral Approach for Multiscale Language Representations
abstract
Language exhibits structure at a wide range of scales, from subwords to words, sentences, paragraphs, and documents. We propose building models that isolate scale-specific information in deep representations, and develop methods for encouraging models during training to learn more about particular scales of interest. Our method for creating scale-specific neurons in deep NLP models constrains how the activation of a neuron can change across the tokens of an input by interpreting those activations as a digital signal and filtering out parts of its frequency spectrum. This technique enables us to extract scale-specific information from BERT representations: by filtering out different frequencies we can produce new representations that perform well on part of speech tagging (word-level), dialog speech acts classification (utterance-level), or topic classification (document-level), while performing poorly on the other tasks. We also present a prism layer for use during training, which constrains different neurons of a BERT model to different parts of the frequency spectrum. Our proposed BERT + Prism model is better able to predict masked tokens using long-range context, and produces individual multiscale representations that perform with comparable or improved performance across all three tasks. Our methods are general and readily applicable to other domains besides language, such as images, audio, and video.
Alex Tamkin, Daniel Jurafsky, Noah D. Goodman
NeurIPS3
2019 Zero Shot Learning for Code Education: Rubric Sampling with Deep Learning Inference
abstract
In modern computer science education, massive open online courses (MOOCs) log thousands of hours of data about how students solve coding challenges. Being so rich in data, these platforms have garnered the interest of the machine learning community, with many new algorithms attempting to autonomously provide feedback to help future students learn. But what about those first hundred thousand students? In most educational contexts (i.e. classrooms), assignments do not have enough historical data for supervised learning. In this paper, we introduce a human-in-the-loop “rubric sampling” approach to tackle the “zero shot” feedback challenge. We are able to provide autonomous feedback for the first students working on an introductory programming assignment with accuracy that substantially outperforms data-hungry algorithms and approaches human level fidelity. Rubric sampling requires minimal teacher effort, can associate feedback with specific parts of a student’s solution and can articulate a student’s misconceptions in the language of the instructor. Deep learning inference enables rubric sampling to further improve as more assignment specific student data is acquired. We demonstrate our results on a novel dataset from Code.org, the world’s largest programming education platform.
Mike Wu, Milan Mossé, Noah D. Goodman, Chris Piech
AAAI3
2019 Learning from Omission
abstract
Pragmatic reasoning allows humans to go beyond the literal meaning when interpreting language in context.Previous work has shown that such reasoning can improve the performance of already-trained language understanding systems.Here, we explore whether pragmatic reasoning during training can improve the quality of learned meanings.Our experiments on reference game data show that end-to-end pragmatic training produces more accurate utterance interpretation models, especially when data is sparse and language is complex.
Bill McDowell, Noah D. Goodman
ACL (1)2
2019 DisSent: Learning Sentence Representations from Explicit Discourse Relations
abstract
Learning effective representations of sentences is one of the core missions of natural language understanding. Existing models either train on a vast amount of text, or require costly, manually curated sentence relation datasets. We show that with dependency parsing and rule-based rubrics, we can curate a high quality sentence relation task by leveraging explicit discourse relations. We show that our curated dataset provides an excellent signal for learning vector representations of sentence meaning, representing relations that can only be determined when the meanings of two sentences are combined. We demonstrate that the automatically curated corpus allows a bidirectional LSTM sentence encoder to yield high quality sentence embeddings and can serve as a supervised fine-tuning dataset for larger models such as BERT. Our fixed sentence embeddings achieve high performance on a variety of transfer tasks, including SentEval, and we achieve state-of-the-art results on Penn Discourse Treebank's implicit relation prediction task.
Allen Nie, Erin D. Bennett, Noah D. Goodman
ACL (1)3
2019 Differentiable Antithetic Sampling for Variance Reduction in Stochastic Variational Inference
abstract
Stochastic optimization techniques are standard in variational inference algorithms. These methods estimate gradients by approximating expectations with independent Monte Carlo samples. In this paper, we explore a technique that uses correlated, but more representative, samples to reduce estimator variance. Specifically, we show how to generate antithetic samples that match sample moments with the true moments of an underlying importance distribution. Combining a differentiable antithetic sampler with modern stochastic variational inference, we showcase the effectiveness of this approach for learning a deep generative model. An implementation is available at https://github.com/mhw32/antithetic-vae-public.
Mike Wu, Noah D. Goodman, Stefano Ermon
AISTATS2
2019 The first crank of the cultural ratchet: Learning and transmitting concepts through language
Sahil Chopra, Michael Henry Tessler, Noah D. Goodman
CogSci3
2019 Disentangling contributions of visual information and interaction history in the formation of graphical conventions
Robert D. Hawkins, Megumi Sano, Noah D. Goodman, Judith W. Fan
CogSci3
2019 The interactions of rational, pragmatic agents lead to efficient language structure and use
Benjamin N. Peloquin, Noah D. Goodman, Michael C. Frank
CogSci2
2019 Extending Rationality
Emmanuel M. Pothos, Jerome R. Busemeyer, Timothy J. Pleskac, James M. Yearsley, Josh Tenenbaum, Noah D. Goodman, Michael Henry Tessler, Thomas L. Griffiths 0001, Falk Lieder, Ralph Hertwig, Thorsten Pachur, Christina Leuker, Richard M. Shiffrin
CogSci6
2019 Shapeglot: Learning Language for Shape Differentiation
abstract
In this work we explore how fine-grained differences between the shapes of common objects are expressed in language, grounded on 2D and/or 3D object representations. We first build a large scale, carefully controlled dataset of human utterances each of which refers to a 2D rendering of a 3D CAD model so as to distinguish it from a set of shape-wise similar alternatives. Using this dataset, we develop neural language understanding (listening) and production (speaking) models that vary in their grounding (pure 3D forms via point-clouds vs. rendered 2D images), the degree of pragmatic reasoning captured (e.g. speakers that reason about a listener or not), and the neural architecture (e.g. with or without attention). We find models that perform well with both synthetic and human partners, and with held out utterances and objects. We also find that these models are capable of zero-shot transfer learning to novel object classes (e.g. transfer from training on chairs to testing on lamps), as well as to real-world images drawn from furniture catalogs. Lesion studies indicate that the neural listeners depend heavily on part-related words and associate these words correctly with visual parts of objects (without any explicit supervision on such parts), and that transfer to novel classes is most successful when known part-related words are available. This work illustrates a practical approach to language grounding, and provides a novel case study in the relationship between object shape and linguistic structure when it comes to object differentiation.
Panos Achlioptas, Leonidas J. Guibas, Noah D. Goodman, Judy Fan, Robert D. Hawkins
ICCV3
2019 Tensor Variable Elimination for Plated Factor Graphs
abstract
A wide class of machine learning algorithms can be reduced to variable elimination on factor graphs. While factor graphs provide a unifying notation for these algorithms, they do not provide a compact way to express repeated structure when compared to plate diagrams for directed graphical models. To exploit efficient tensor algebra in graphs with plates of variables, we generalize undirected factor graphs to plated factor graphs and variable elimination to a tensor variable elimination algorithm that operates directly on plated factor graphs. Moreover, we generalize complexity bounds based on treewidth and characterize the class of plated factor graphs for which inference is tractable. As an application, we integrate tensor variable elimination into the Pyro probabilistic programming language to enable exact inference in discrete latent variable models with repeated structure. We validate our methods with experiments on both directed and undirected graphical models, including applications to polyphonic music modeling, animal movement modeling, and latent sentiment analysis.
Fritz Obermeyer, Eli Bingham, Martin Jankowiak, Neeraj Pradhan, Justin T. Chiu, Alexander M. Rush, Noah D. Goodman
ICML7
2019 Variational Bayesian Optimal Experimental Design
abstract
Bayesian optimal experimental design (BOED) is a principled framework for making efficient use of limited experimental resources. Unfortunately, its applicability is hampered by the difficulty of obtaining accurate estimates of the expected information gain (EIG) of an experiment. To address this, we introduce several classes of fast EIG estimators by building on ideas from amortized variational inference. We show theoretically and empirically that these estimators can provide significant gains in speed and accuracy over previous approaches. We further demonstrate the practicality of our approach on a number of end-to-end experiments.
Adam Foster 0001, Martin Jankowiak, Eli Bingham, Paul Horsfall, Yee Whye Teh, Tom Rainforth, Noah D. Goodman
NeurIPS7
2019 Pyro: Deep Universal Probabilistic Programming
abstract
Pyro is a probabilistic programming language built on Python as a platform for developing advanced probabilistic models in AI research. To scale to large data sets and high-dimensional models, Pyro uses stochastic variational inference algorithms and probability distributions built on top of PyTorch, a modern GPU-accelerated deep learning framework. To accommodate complex or model-specific algorithmic behavior, Pyro leverages Poutine, a library of composable building blocks for modifying the behavior of probabilistic programs.
Eli Bingham, Jonathan P. Chen, Martin Jankowiak, Fritz Obermeyer, Neeraj Pradhan, Theofanis Karaletsos, Paul A. Szerlip, Paul Horsfall, Noah D. Goodman
J. Mach. Learn. Res.10
2018 An Information-Theoretic Explanation of Adjective Ordering Preferences
Michael Hahn 0001, Judith Degen, Noah D. Goodman, Daniel Jurafsky, Richard Futrell
CogSci3
2018 Evaluating Compositionality in Sentence Embeddings
Ishita Dasgupta 0001, Demi Guo, Andreas Stuhlmüller, Samuel Gershman, Noah D. Goodman
CogSci5
2018 Emerging abstractions: Lexical conventions are shaped by communicative context
Robert D. Hawkins, Michael Franke, Kenny Smith, Noah D. Goodman
CogSci4
2018 webppl-oed: A practical optimal experiment design system
Long Ouyang, Michael Henry Tessler, Daniel Ly, Noah D. Goodman
CogSci4
2018 Deriving uniform information density behavior in pragmatic agents
Benjamin N. Peloquin, Noah D. Goodman, Michael C. Frank
CogSci2
2018 Statistics as Pottery: Bayesian Data Analysis using Probabilistic Programs
Michael Henry Tessler, Noah D. Goodman
CogSci2
2018 Generalizations, from representation to transmission
Michael Henry Tessler, Noah D. Goodman, David Danks, Emily Foster-Hanson, Marjorie Rhodes, Greg Carlson
CogSci2
2018 Multimodal Generative Models for Scalable Weakly-Supervised Learning
abstract
Multiple modalities often co-occur when describing natural phenomena. Learning a joint representation of these modalities should yield deeper and more useful representations.Previous generative approaches to multi-modal input either do not learn a joint distribution or require additional computation to handle missing data. Here, we introduce a multimodal variational autoencoder (MVAE) that uses a product-of-experts inference network and a sub-sampled training paradigm to solve the multi-modal inference problem. Notably, our model shares parameters to efficiently learn under any combination of missing modalities. We apply the MVAE on four datasets and match state-of-the-art performance using many fewer parameters. In addition, we show that the MVAE is directly applicable to weakly-supervised learning, and is robust to incomplete supervision. We then consider two case studies, one of learning image transformations---edge detection, colorization, segmentation---as a set of modalities, followed by one of machine translation between two languages. We find appealing results across this range of tasks.
Mike Wu, Noah D. Goodman
NeurIPS2
2018 Bias and Generalization in Deep Generative Models: An Empirical Study
abstract
In high dimensional settings, density estimation algorithms rely crucially on their inductive bias. Despite recent empirical success, the inductive bias of deep generative models is not well understood. In this paper we propose a framework to systematically investigate bias and generalization in deep generative models of images by probing the learning algorithm with carefully designed training datasets. By measuring properties of the learned distribution, we are able to find interesting patterns of generalization. We verify that these patterns are consistent across datasets, common models and architectures.
Shengjia Zhao, Hongyu Ren, Arianna Yuan, Jiaming Song, Noah D. Goodman, Stefano Ermon
NeurIPS5
2018 Planning, Inference, and Pragmatics in Sequential Language Games
abstract
We study sequential language games in which two players, each with private information, communicate to achieve a common goal. In such games, a successful player must (i) infer the partner’s private information from the partner’s messages, (ii) generate messages that are most likely to help with the goal, and (iii) reason pragmatically about the partner’s strategy. We propose a model that captures all three characteristics and demonstrate their importance in capturing human behavior on a new goal-oriented dataset we collected using crowdsourcing.
Fereshte Khani, Noah D. Goodman, Percy Liang
Trans. Assoc. Comput. Linguistics2
2017 Amortized Hypothesis Generation
Ishita Dasgupta 0001, Eric Schulz, Noah D. Goodman, Samuel Gershman
CogSci3
2017 Convention-formation in iterated reference games
Robert D. Hawkins, Mike Frank, Noah D. Goodman
CogSci3
2017 Mentioning atypical properties of objects is communicatively efficient
Elisa Kreiss, Robert D. Hawkins, Judith Degen, Noah D. Goodman
CogSci4
2017 Warm (for winter): Comparison class understanding in vague language
Michael Henry Tessler, Michael Lopez-Brau, Noah D. Goodman
CogSci3
2017 "I won't lie, it wasn't amazing": Modeling polite indirect speech
Erica J. Yoon, Michael Henry Tessler, Noah D. Goodman, Michael C. Frank
CogSci3
2017 Learning Disentangled Representations with Semi-Supervised Deep Generative Models
abstract
Variational autoencoders (VAEs) learn representations of data by jointly training a probabilistic encoder and decoder network. Typically these models encode all features of the data into a single variable. Here we are interested in learning disentangled representations that encode distinct aspects of the data into separate variables. We propose to learn such representations using model architectures that generalise from standard VAEs, employing a general graphical model structure in the encoder and decoder. This allows us to train partially-specified models that make relatively strong assumptions about a subset of interpretable variables and rely on the flexibility of neural networks to learn representations for the remaining variables. We further define a general objective for semi-supervised learning in this model class, which can be approximated using an importance sampling procedure. We evaluate our framework's ability to learn disentangled representations, both by qualitative exploration of its generative capacity, and quantitative evaluation of its discriminative ability on a variety of models and datasets.
N. Siddharth 0001, Brooks Paige, Jan-Willem van de Meent, Alban Desmaison, Noah D. Goodman, Pushmeet Kohli, Frank D. Wood, Philip Torr 0001
NIPS5
2017 Colors in Context: A Pragmatic Neural Model for Grounded Language Understanding
abstract
We present a model of pragmatic referring expression interpretation in a grounded communication task (identifying colors from descriptions) that draws upon predictions from two recurrent neural network classifiers, a speaker and a listener, unified by a recursive pragmatic reasoning framework. Experiments show that this combined pragmatic model interprets color descriptions more accurately than the classifiers from which it is built, and that much of this improvement results from combining the speaker and listener perspectives. We observe that pragmatic reasoning helps primarily in the hardest cases: when the model must distinguish very similar colors, or when few utterances adequately express the target color. Our findings make use of a newly-collected corpus of human utterances in color reference games, which exhibit a variety of pragmatic behaviors. We also show that the embedded speaker model reproduces many of these pragmatic behaviors.
Will Monroe, Robert D. Hawkins, Noah D. Goodman, Christopher Potts
Trans. Assoc. Comput. Linguistics3
2016 Learning the Preferences of Ignorant, Inconsistent Agents
abstract
An important use of machine learning is to learn what people value. What posts or photos should a user be shown? Which jobs or activities would a person find rewarding? In each case, observations of people's past choices can inform our inferences about their likes and preferences. If we assume that choices are approximately optimal according to some utility function, we can treat preference inference as Bayesian inverse planning. That is, given a prior on utility functions and some observed choices, we invert an optimal decision-making process to infer a posterior distribution on utility functions. However, people often deviate from approximate optimality. They have false beliefs, their planning is sub-optimal, and their choices may be temporally inconsistent due to hyperbolic discounting and other biases. We demonstrate how to incorporate these deviations into algorithms for preference inference by constructing generative models of planning for agents who are subject to false beliefs and time inconsistency. We explore the inferences these models make about preferences, beliefs, and biases. We present a behavioral experiment in which human subjects perform preference inference given the same observations of choices as our model. Results show that human subjects (like our model) explain choices in terms of systematic deviations from optimal behavior and suggest that they take such deviations into account when inferring preferences.
Owain Evans, Andreas Stuhlmüller, Noah D. Goodman
AAAI3
2016 C3: Lightweight Incrementalized MCMC for Probabilistic Programs using Continuations and Callsite Caching
abstract
Lightweight, source-to-source transformation approaches to implementing MCMC for probabilistic programming languages are popular for their simplicity, support of existing deterministic code, and ability to execute on existing fast runtimes. However, they are also inefficient, requiring a complete re-execution of the program on every Metropolis Hastings proposal. We present a new extension to the lightweight approach, C3, which enables efficient, incrementalized re-execution of MH proposals. C3 is based on two core ideas: transforming probabilistic programs into continuation passing style (CPS), and caching the results of function calls. It is particularly effective at speeding up recursive programs with many local latent variables. We show that on several common models, C3 reduces proposal runtime by 20-100x, in some cases reducing runtime complexity from linear in model size to constant. We also demonstrate nearly an order of magnitude speedup on a complex inverse procedural modeling application.
Daniel Ritchie 0001, Andreas Stuhlmüller, Noah D. Goodman
AISTATS3
2016 What does the crowd believe? A hierarchical approach to estimating subjective beliefs from empirical data
Michael Franke, Fabian Dablander, Anthea Schöller, Erin D. Bennett, Judith Degen, Michael Henry Tessler, Justine T. Kao, Noah D. Goodman
CogSci8
2016 Animal, dog, or dalmatian? Level of abstraction in nominal referring expressions
Caroline Graf, Judith Degen, Robert D. Hawkins, Noah D. Goodman
CogSci4
2016 Conversational expectations account for apparent limits on theory of mind use
Robert D. Hawkins, Noah D. Goodman
CogSci2
2016 The Emergence of Conventions
Robert D. Hawkins, Noah D. Goodman, Olga Feher, Kenny Smith, Robert L. Goldstone, Thomas L. Griffiths 0001
CogSci2
2016 Empirical and Computational Approaches to Metaphor and Figurative Meaning
Justine T. Kao, Noah D. Goodman
CogSci2
2016 Emotions in lay explanations of behavior
Desmond C. Ong, Jamil Zaki, Noah D. Goodman
CogSci3
2016 A rational speech-act model of projective content
Ciyang Qing, Noah D. Goodman, Daniel Lassiter
CogSci2
2016 Communicating generalizations about events
Michael Henry Tessler, Noah D. Goodman
CogSci2
2016 The Pragmatics of Spatial Language
Tomer D. Ullman, Yang Xu 0023, Noah D. Goodman
CogSci3
2016 Talking with tact: Polite language as a balance between informativity and kindness
Erica J. Yoon, Michael Henry Tessler, Noah D. Goodman, Michael C. Frank
CogSci3
2016 Learning to Generate Compositional Color Descriptions
abstract
The production of color language is essential for grounded language generation.Color descriptions have many challenging properties: they can be vague, compositionally complex, and denotationally rich.We present an effective approach to generating color descriptions using recurrent neural networks and a Fouriertransformed color representation.Our model outperforms previous work on a conditional language modeling task over a large corpus of naturalistic color descriptions.In addition, probing the model's output reveals that it can accurately produce not only basic color terms but also descriptors with non-convex denotations ("greenish"), bare modifiers ("bright", "dull"), and compositional phrases ("faded teal") not seen in training.
Will Monroe, Noah D. Goodman, Christopher Potts
EMNLP2
2016 Neurally-Guided Procedural Models: Amortized Inference for Procedural Graphics Programs using Neural Networks
abstract
Probabilistic inference algorithms such as Sequential Monte Carlo (SMC) provide powerful tools for constraining procedural models in computer graphics, but they require many samples to produce desirable results. In this paper, we show how to create procedural models which learn how to satisfy constraints. We augment procedural models with neural networks which control how the model makes random choices based on the output it has generated thus far. We call such models neurally-guided procedural models. As a pre-computation, we train these models to maximize the likelihood of example outputs generated via SMC. They are then used as efficient SMC importance samplers, generating high-quality results with very few samples. We evaluate our method on L-system-like models with image-based constraints. Given a desired quality threshold, neurally-guided models can generate satisfactory results up to 10x faster than unguided models.
Daniel Ritchie 0001, Anna Thomas, Pat Hanrahan, Noah D. Goodman
NIPS4
2015 Not by number alone: The effect of teachers' knowledge and its value in evaluating "sins of omission"
Ilona Bass, Daniel Hawthorne-Madell, Noah D. Goodman, Hyowon Gweon
CogSci3
2015 Extremely costly intensifiers are stronger than quite costly ones
Erin D. Bennett, Noah D. Goodman
CogSci2
2015 Wonky worlds: Listeners revise world knowledge when utterances are odd
Judith Degen, Michael Henry Tessler, Noah D. Goodman
CogSci3
2015 How, whether, why: Causal judgments as counterfactual contrasts
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum
CogSci2
2015 Why do you ask? Good questions provoke informative answers
Robert D. Hawkins, Andreas Stuhlmüller, Judith Degen, Noah D. Goodman
CogSci4
2015 So good it has to be true: Wishful thinking in theory of mind
Daniel Hawthorne-Madell, Noah D. Goodman
CogSci2
2015 A Resource-Rational Approach to the Causal Frame Problem
Thomas Icard, Noah D. Goodman
CogSci2
2015 Let's talk (ironically) about the weather: Modeling verbal irony
Justine T. Kao, Noah D. Goodman
CogSci2
2015 Emergent Collective Sensing in Human Groups
P. M. Krafft, Robert D. Hawkins, Alex Pentland, Noah D. Goodman, Josh Tenenbaum
CogSci4
2015 Near-misses sting even when they are uncontrollable
Desmond C. Ong, Noah D. Goodman, Jamil Zaki
CogSci2
2015 Toddlers Always Get the Last Word: Recency biases in early verbal behavior
Emily S. Sumner, Erika DeAngelis, Mara Hyatt, Noah D. Goodman, Celeste Kidd
CogSci4
2015 Generating Design Suggestions under Tight Constraints with Gradient-based Probabilistic Programming
abstract
Abstract We present a system for generating suggestions from highly‐constrained, continuous design spaces. We formulate suggestion as sampling from a probability distribution; constraints are represented as factors that concentrate probability mass around sub‐manifolds of the design space. These sampling problems are intractable using typical random walk MCMC techniques, so we adopt Hamiltonian Monte Carlo (HMC), a gradient‐based MCMC method. We implement HMC in a high‐performance probabilistic programming language, and we evaluate its ability to efficiently generate suggestions for two different, highly‐constrained example applications: vector art coloring and designing stable stacking structures.
Daniel Ritchie 0001, Sharon Lin, Noah D. Goodman, Pat Hanrahan
Comput. Graph. Forum3
2015 Controlling procedural modeling programs with stochastically-ordered sequential Monte Carlo
abstract
We present a method for controlling the output of procedural modeling programs using Sequential Monte Carlo (SMC). Previous probabilistic methods for controlling procedural models use Markov Chain Monte Carlo (MCMC), which receives control feedback only for completely-generated models. In contrast, SMC receives feedback incrementally on incomplete models, allowing it to reallocate computational resources and converge quickly. To handle the many possible sequentializations of a structured, recursive procedural modeling program, we develop and prove the correctness of a new SMC variant, Stochastically-Ordered Sequential Monte Carlo (SOSMC). We implement SOSMC for general-purpose programs using a new programming primitive: the stochastic future. Finally, we show that SOSMC reliably generates high-quality outputs for a variety of programs and control scoring functions. For small computational budgets, SOSMC's outputs often score nearly twice as high as those of MCMC or normal SMC.
Daniel Ritchie 0001, Ben Mildenhall, Noah D. Goodman, Pat Hanrahan
ACM Trans. Graph.3
2014 Generating Efficient MCMC Kernels from Probabilistic Programs
abstract
Universal probabilistic programming languages (such as Church) trade performance for abstraction: any model can be represented compactly as an arbitrary stochastic computation, but costly online analyses are required for inference. We present a technique that recovers hand-coded levels of performance from a universal probabilistic language, for the Metropolis-Hastings (MH) MCMC inference algorithm. It takes a Church program as input and traces its execution to remove computation overhead. It then analyzes the trace for each proposal, using slicing, to identify the minimal computation needed to evaluate the MH acceptance probability. Generated incremental code is much faster than a baseline implementation (up to 600x) and usually as fast as hand-coded MH kernels.
Lingfeng Yang, Pat Hanrahan, Noah D. Goodman
AISTATS3
2014 The strategic use of noise in pragmatic reasoning
Leon Bergen, Noah D. Goodman
CogSci2
2014 Lost your marbles? The puzzle of dependent measures in experimental pragmatics
Judith Degen, Noah D. Goodman
CogSci2
2014 Symposium: The Role of Alternatives in Pragmatic Inference
Judith Degen, Noah D. Goodman, Roni Katzir, David Barner, Albert Gatt
CogSci2
2014 Amortized Inference in Probabilistic Reasoning
Samuel Gershman, Noah D. Goodman
CogSci2
2014 From counterfactual simulation to causal judgment
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum
CogSci2
2014 Probability, programs, and the mind: Building structured Bayesian models of cognition
Noah D. Goodman, Josh Tenenbaum
CogSci1
2014 Formalizing the Pragmatics of Metaphor Understanding
Justine T. Kao, Leon Bergen, Noah D. Goodman
CogSci3
2014 Understanding Affective Cognition: Frontiers in modeling reasoning about others' emotions
Desmond C. Ong, Jamil Zaki, Noah D. Goodman
CogSci3
2014 Some arguments are probably valid: Syllogistic reasoning as communication
Michael Henry Tessler, Noah D. Goodman
CogSci2
2014 Learning physical theories from dynamical scenes
Tomer D. Ullman, Andreas Stuhlmüller, Noah D. Goodman, Josh Tenenbaum
CogSci3
2013 The Funny Thing About Incongruity: A Computational Model of Humor in Puns
Justine T. Kao, Roger Levy, Noah D. Goodman
CogSci3
2013 Learned helplessness and generalization
Falk Lieder, Noah D. Goodman, Quentin J. M. Huys
CogSci2
2013 Minimal Nativism: How does cognitive development get off the ground?
Tomer D. Ullman, Josh Tenenbaum, Noah D. Goodman, Shimon Ullman, Elizabeth S. Spelke
CogSci3
2013 Learning and using language via recursive pragmatic reasoning about other agents
abstract
Language users are remarkably good at making inferences about speakers' intentions in context, and children learning their native language also display substantial skill in acquiring the meanings of unknown words. These two cases are deeply related: Language users invent new terms in conversation, and language learners learn the literal meanings of words based on their pragmatic inferences about how those words are used. While pragmatic inference and word learning have both been independently characterized in probabilistic terms, no current work unifies these two. We describe a model in which language learners assume that they jointly approximate a shared, external lexicon and reason recursively about the goals of others in using this lexicon. This model captures phenomena in word learning and pragmatic inference; it additionally leads to insights about the emergence of communicative systems in conversation and the mechanisms by which pragmatic inferences become incorporated into word meanings.
Nathaniel J. Smith, Noah D. Goodman, Michael C. Frank
NIPS2
2013 Learning Stochastic Inverses
abstract
We describe a class of algorithms for amortized inference in Bayesian networks. In this setting, we invest computation upfront to support rapid online inference for a wide range of queries. Our approach is based on learning an inverse factorization of a model's joint distribution: a factorization that turns observations into root nodes. Our algorithms accumulate information to estimate the local conditional distributions that constitute such a factorization. These stochastic inverses can be used to invert each of the computation steps leading to an observation, sampling backwards in order to quickly find a likely explanation. We show that estimated inverses converge asymptotically in number of (prior or posterior) training samples. To make use of inverses before convergence, we describe the Inverse MCMC algorithm, which uses stochastic inverses to make block proposals for a Metropolis-Hastings sampler. We explore the efficiency of this sampler for a variety of parameter regimes and Bayes nets.
Andreas Stuhlmüller, Jessica Taylor, Noah D. Goodman
NIPS3
2013 The principles and practice of probabilistic programming
abstract
models, probabilistic programs Probabilities describe degrees of belief, and probabilistic inference describes rational reasoning under uncertainty. It is no wonder, then, that probabilistic models have exploded onto the scene of modern artificial intelligence, cognitive science, and applied statistics: these are all sciences of inference under uncertainty. But as probabilistic models have become more sophisticated, the tools to formally describe them and to perform probabilistic inference have wrestled with new complexity. Just as programming beyond the simplest algorithms requires tools for abstraction and composition, complex probabilistic modeling requires new progress in model representation—probabilistic programming languages. These languages provide compositional means for describing complex probability distributions; implementations of these languages provide generic inference engines: tools for performing efficient probabilistic
Noah D. Goodman
POPL1
2012 That's what she (could have) said: How alternative utterances affect language use
Leon Bergen, Noah D. Goodman, Roger Levy
CogSci2
2012 Ping Pong in Church: Productive use of concepts in human probabilistic inference
Tobias Gerstenberg, Noah D. Goodman
CogSci2
2012 Noisy Newtons: Unifying process and dependency accounts of causal attribution
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum
CogSci2
2012 Knowledge and implicature: Modeling language understanding as social cognition
Noah D. Goodman, Andreas Stuhlmüller
CogSci1
2012 Probability, programs, and the mind: Building structured Bayesian models of cognition
Noah D. Goodman, Josh Tenenbaum
CogSci1
2012 How many kinds of reasoning? Inference, probability, and natural language semantics
Daniel Lassiter, Noah D. Goodman
CogSci2
2012 "Burn-in, bias, and the rationality of anchoring"
abstract
Bayesian inference provides a unifying framework for addressing problems in machine learning, artificial intelligence, and robotics, as well as the problems facing the human mind. Unfortunately, exact Bayesian inference is intractable in all but the simplest models. Therefore minds and machines have to approximate Bayesian inference. Approximate inference algorithms can achieve a wide range of time-accuracy tradeoffs, but what is the optimal tradeoff? We investigate time-accuracy tradeoffs using the Metropolis-Hastings algorithm as a metaphor for the mind's inference algorithm(s). We find that reasonably accurate decisions are possible long before the Markov chain has converged to the posterior distribution, i.e. during the period known as burn-in. Therefore the strategy that is optimal subject to the mind's bounded processing speed and opportunity costs may perform so few iterations that the resulting samples are biased towards the initial value. The resulting cognitive process model provides a rational basis for the anchoring-and-adjustment heuristic. The model's quantitative predictions are tested against published data on anchoring in numerical estimation tasks. Our theoretical and empirical results suggest that the anchoring bias is consistent with approximate Bayesian inference.
Falk Lieder, Thomas L. Griffiths 0001, Noah D. Goodman
NIPS3
2012 Learning design patterns with bayesian grammar induction
abstract
Design patterns have proven useful in many creative fields, providing content creators with archetypal, reusable guidelines to leverage in projects. Creating such patterns, however, is a time-consuming, manual process, typically relegated to a few experts in any given domain. In this paper, we describe an algorithmic method for learning design patterns directly from data using techniques from natural language processing and structured concept learning. Given a set of labeled, hierarchical designs as input, we induce a probabilistic formal grammar over these exemplars. Once learned, this grammar encodes a set of generative rules for the class of designs, which can be sampled to synthesize novel artifacts. We demonstrate the method on geometric models and Web pages, and discuss how the learned patterns can drive new interaction mechanisms for content creators.
Jerry O. Talton, Lingfeng Yang, Ranjitha Kumar, Maxine Lim, Noah D. Goodman, Radomír Mech
UIST5
2012 Synthesizing open worlds with constraints using locally annealed reversible jump MCMC
abstract
We present a novel Markov chain Monte Carlo (MCMC) algorithm that generates samples from transdimensional distributions encoding complex constraints. We use factor graphs, a type of graphical model, to encode constraints as factors. Our proposed MCMC method, called locally annealed reversible jump MCMC, exploits knowledge of how dimension changes affect the structure of the factor graph. We employ a sequence of annealed distributions during the sampling process, allowing us to explore the state space across different dimensionalities more freely. This approach is motivated by the application of layout synthesis where relationships between objects are characterized as constraints. In particular, our method addresses the challenge of synthesizing open world layouts where the number of objects are not fixed and optimal configurations for different numbers of objects may be drastically different. We demonstrate the applicability of our approach on two open world layout synthesis problems: coffee shops and golf courses.
Lingfeng Yang, Noah D. Goodman, Pat Hanrahan
ACM Trans. Graph.4
2011 Productivity and Reuse in Language
Timothy J. O'Donnell, Jesse Snedeker, Josh Tenenbaum, Noah D. Goodman
CogSci4
2011 Productivity and Reuse in Language: a Developmental Study
Timothy J. O'Donnell, Jesse Snedeker, Josh Tenenbaum, Noah D. Goodman
CogSci4
2011 Ad-hoc scalar implicature in adults and children
Alex Stiller, Noah D. Goodman, Michael C. Frank
CogSci2
2011 Forward Physics: How people learn and generalize novel dynamical models
Tomer D. Ullman, Noah D. Goodman, Josh Tenenbaum
CogSci2
2011 Bayesian Policy Search with Policy Priors
David Wingate, Noah D. Goodman, Daniel M. Roy 0001, Leslie Pack Kaelbling, Josh Tenenbaum
IJCAI2
2011 Nonstandard Interpretations of Probabilistic Programs for Efficient Inference
abstract
Probabilistic programming languages allow modelers to specify a stochastic process using syntax that resembles modern programming languages. Because the program is in machine-readable format, a variety of techniques from compiler design and program analysis can be used to examine the structure of the distribution represented by the probabilistic program. We show how nonstandard interpretations of probabilistic programs can be used to craft efficient inference algorithms: information about the structure of a distribution (such as gradients or dependencies) is generated as a monad-like side computation while executing the program. These interpretations can be easily coded using special-purpose objects and operator overloading. We implement two examples of nonstandard interpretations in two different languages, and use them as building blocks to construct inference algorithms: automatic differentiation, which enables gradient based methods, and provenance tracking, which enables efficient construction of global proposals.
David Wingate, Noah D. Goodman, Andreas Stuhlmüller, Jeffrey Mark Siskind
NIPS2
2009 Help or Hinder: Bayesian Models of Social Goal Inference
abstract
Everyday social interactions are heavily influenced by our snap judgments about others goals. Even young infants can infer the goals of intentional agents from observing how they interact with objects and other agents in their environment: e.g., that one agent is helping orhindering anothers attempt to get up a hill or open a box. We propose a model for how people can infer these social goals from actions, based on inverse planning in multiagent Markov decision problems (MDPs). The model infers the goal most likely to be driving an agents behavior by assuming the agent acts approximately rationally given environmental constraints and its model of other agents present. We also present behavioral evidence in support of this model over a simpler, perceptual cue-based alternative.
Tomer D. Ullman, Chris L. Baker, Owen Macindoe, Owain Evans, Noah D. Goodman, Josh Tenenbaum
NIPS5
2009 The Infinite Latent Events Model
David Wingate, Noah D. Goodman, Daniel M. Roy 0001, Josh Tenenbaum
UAI2
2008 Church: a language for generative models
Noah D. Goodman, Vikash Mansinghka 0001, Daniel M. Roy 0001, Kallista A. Bonawitz, Josh Tenenbaum
UAI1
2007 A Bayesian Framework for Cross-Situational Word-Learning
abstract
For infants, early word learning is a chicken-and-egg problem. One way to learn a word is to observe that it co-occurs with a particular referent across different situations. Another way is to use the social context of an utterance to infer the in- tended referent of a word. Here we present a Bayesian model of cross-situational word learning, and an extension of this model that also learns which social cues are relevant to determining reference. We test our model on a small corpus of mother-infant interaction and find it performs better than competing models. Fi- nally, we show that our model accounts for experimental phenomena including mutual exclusivity, fast-mapping, and generalization from social cues. To understand the difficulty of an infant word-learner, imagine walking down the street with a friend who suddenly says “dax blicket philbin na fivy!” while at the same time wagging her elbow. If you knew any of these words you might infer from the syntax of her sentence that blicket is a novel noun, and hence the name of a novel object. At the same time, if you knew that this friend indicated her attention by wagging her elbow at objects, you might infer that she intends to refer to an object in a nearby show window. On the other hand if you already knew that “blicket” meant the object in the window, you might be able to infer these elements of syntax and social cues. Thus, the problem of early word-learning is a classic chicken-and-egg puzzle: in order to learn word meanings, learners must use their knowledge of the rest of language (including rules of syntax, parts of speech, and other word meanings) as well as their knowledge of social situations. But in order to learn about the facts of their language they must first learn some words, and in order to determine which cues matter for establishing reference (for instance, pointing and looking at an object but normally not waggling your elbow) they must first have a way to know the intended referent in some situations. For theories of language acquisition, there are two common ways out of this dilemma. The first involves positing a wide range of innate structures which determine the syntax and categories of a language and which social cues are informative. (Though even when all of these elements are innately determined using them to learn a language from evidence may not be trivial [1].) The other alternative involves bootstrapping: learning some words, then using those words to learn how to learn more. This paper gives a proposal for the second alternative. We first present a Bayesian model of how learners could use a statistical strategy—cross-situational word-learning—to learn how words map to objects, independent of syntactic and social cues. We then extend this model to a true bootstrapping situation: using social cues to learn words while using words to learn social cues. Finally, we examine several important phenomena in word learning: mutual exclusivity (the tendency to assign novel words to novel referents), fast-mapping (the ability to assign a novel word in a linguistic context to a novel referent after only a single use), and social generalization (the ability to use social context to learn the referent of a novel word). Without adding additional specialized machinery, we show how these can be explained within our model as the result of domain-general probabilistic inference mechanisms operating over the linguistic domain. 1 Figure 1: Graphical model de- scribing the generation of words (Ws) from an intention (Is) and lexicon ((cid:96)), and intention from the objects present in a situa- tion (Os). The plate indicates multiple copies of the model for different situation/utterance pairs (s). Dotted portions indicate ad- ditions to include the generation of social cues Ss from intentions.
Michael C. Frank, Noah D. Goodman, Josh Tenenbaum
NIPS2
2007 Learning and using relational theories
abstract
Much of human knowledge is organized into sophisticated systems that are often called intuitive theories. We propose that intuitive theories are mentally repre- sented in a logical language, and that the subjective complexity of a theory is determined by the length of its representation in this language. This complexity measure helps to explain how theories are learned from relational data, and how they support inductive inferences about unobserved relations. We describe two experiments that test our approach, and show that it provides a better account of human learning and reasoning than an approach developed by Goodman [1]. What is a theory, and what makes one theory better than another? Questions like these are of obvious interest to philosophers of science but are also discussed by psychologists, who have argued that everyday knowledge is organized into rich and complex systems that are similar in many respects to scientific theories. Even young children, for instance, have systematic beliefs about domains including folk physics, folk biology, and folk psychology [2]. Intuitive theories like these play many of the same roles as scientific theories: in particular, both kinds of theories are used to explain and encode observations of the world, and to predict future observations. This paper explores the nature, use and acquisition of simple theories. Consider, for instance, an anthropologist who has just begun to study the social structure of a remote tribe, and observes that certain words are used to indicate relationships between selected pairs of individuals. Suppose that term T1(·, ·) can be glossed as ancestor(·, ·), and that T2(·, ·) can be glossed as friend(·, ·). The anthropologist might discover that the first term is transitive, and that the second term is symmetric with a few exceptions. Suppose that term T3(·, ·) can be glossed as defers to(·, ·), and that the tribe divides into two castes such that members of the second caste defer to members of the first caste. In this case the anthropologist might discover two latent concepts (caste 1(·) and caste 2(·)) along with the relationship between these concepts. As these examples suggest, a theory can be defined as a system of laws and concepts that specify the relationships between the elements in some domain [2]. We will consider how these theories are learned, how they are used to encode relational data, and how they support predictions about unob- served relations. Our approach to all three problems relies on the notion of subjective complexity. We propose that theory learners prefer simple theories, that people remember relational data in terms of the simplest underlying theory, and that people extend a partially observed data set according to the simplest theory that is consistent with their observations. There is no guarantee that a single measure of subjective complexity can do all of the work that we require [3]. This paper, however, explores the strong hypothesis that a single measure will suffice. Our formal treatment of subjective complexity begins with the question of how theories are mentally represented. We suggest that theories are represented in some logical language, and propose a spe- cific first-order language that serves as a hypothesis about the “language of thought.” We then pursue the idea that the subjective complexity of a theory corresponds to the length of its representation in this language. Our approach therefore builds on the work of Feldman [4], and is related to other psychological applications of the notion of Kolmogorov complexity [5]. The complexity measure we describe can be used to define a probability distribution over a space of theories, and we develop a model of theory acquisition by using this distribution as the prior for a Bayesian learner. We also
Charles Kemp, Noah D. Goodman, Josh Tenenbaum
NIPS2