Brenden M. Lake

dblp:47/9567 · DBLP profile ↗
← Back
60ranked-venue papers
10as first author
31since 2021 · last 2025
0000-0001-8959-3401ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 10 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 43 · 7 first-author · 23 since 2021
YearPublicationVenuePosition
2025 Goal Inference using Reward-Producing Programs in a Novel Physics Environment
Guy Davidson, Graham Todd, Cédric Colas, Junyi Chu, Julian Togelius, Josh Tenenbaum, Todd M. Gureckis, Brenden M. Lake
CogSci8
2025 Do Large Language Models Reason Causally Like Us? Even Better?
Hanna M. Dettki, Brenden M. Lake, Charley M. Wu, Bob Rehder
CogSci2
2025 A Neurosymbolic Model of Human Reasoning on the Abstraction and Reasoning Corpus
Solim LeGris, Brenden M. Lake, Todd M. Gureckis
CogSci2
2025 Rapid Word Learning Through Meta In-Context Learning
abstract
Humans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts.Yet the abilities of current language models for fewshot word learning, and methods for improving these abilities, are underexplored.In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow).This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word.This training is repeated on many new words to develop a general word-learning ability.We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pretrained on orders of magnitude more data.Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples.These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks.
Guangyuan Jiang, Tal Linzen, Brenden M. Lake
EMNLP4
2025 SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
abstract
Do LLMs robustly generalize critical safety facts to novel situations? Lacking this ability is dangerous when users ask naive questions—for instance, ``I'm considering packing melon balls for my 10-month-old's lunch. What other foods would be good to include?'' Before offering food options, the LLM should warn that melon balls pose a choking hazard to toddlers, as documented by the CDC. Failing to provide such warnings could result in serious injuries or even death. To evaluate this, we introduce SAGE-Eval, SAfety-fact systematic GEneralization evaluation, the first benchmark that tests whether LLMs properly apply well‑established safety facts to naive user queries. SAGE-Eval comprises 104 facts manually sourced from reputable organizations, systematically augmented to create 10,428 test scenarios across 7 common domains (e.g., Outdoor Activities, Medicine). We find that the top model, Claude-3.7-sonnet, passes only 58% of all the safety facts tested. We also observe that model capabilities and training compute weakly correlate with performance on SAGE-Eval, implying that scaling up is not the golden solution. Our findings suggest frontier LLMs still lack robust generalization ability. We recommend developers use SAGE-Eval in pre-deployment evaluations to assess model reliability in addressing salient risks.
Yueh-Han Chen, Guy Davidson, Brenden M. Lake
NeurIPS3
2025 Do different prompting methods yield a common task representation in language models?
abstract
Demonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks. Do identical tasks elicited in different ways result in similar representations of the task? An improved understanding of task representation mechanisms would offer interpretability insights and may aid in steering models. We study this through function vectors (FVs), recently proposed as a mechanism to extract few-shot ICL task representations. We generalize FVs to alternative task presentations, focusing on short textual instruction prompts, and successfully extract instruction function vectors that promote zero-shot task accuracy. We find evidence that demonstration- and instruction-based function vectors leverage different model components, and offer several controls to dissociate their contributions to task performance. Our results suggest that different task prompting forms do not induce a common task representation through FVs but elicit different, partly overlapping mechanisms. Our findings offer principled support to the practice of combining instructions and task demonstrations, imply challenges in universally monitoring task inference across presentation forms, and encourage further examinations of LLM task inference mechanisms.
Guy Davidson, Todd M. Gureckis, Brenden M. Lake, Adina Williams
NeurIPS3
2024 Comparing Abstraction in Humans and Machines Using Multimodal Serial Reproduction
Sreejan Kumar, Raja Marjieh, Byron Zhang, Declan Campbell, Michael Y. Hu, Umang Bhatt, Brenden M. Lake, Thomas L. Griffiths 0001
CogSci7
2024 Predicting Insight during Physical Reasoning
Solim LeGris, Brenden M. Lake, Todd M. Gureckis
CogSci2
2024 Prompting invokes expert-like downward shifts in GPT-4V's conceptual hierarchies
Cara Su-Yi Leong, Brenden M. Lake
CogSci2
2024 An Infant-Cognition Inspired Machine Benchmark for Identifying Agency, Affiliation, Belief, and Intention
Shannon Yasuda, Moira R. Dillon, Brenden M. Lake
CogSci4
2024 Finding Unsupervised Alignment of Conceptual Systems in Image-Word Representations
Kexin Luo, Yajie Xiao, Brenden M. Lake
CogSci4
2024 Self-supervised learning of video representations from a child's perspective
A. Emin Orhan, Alex N. Wang, Mengye Ren, Brenden M. Lake
CogSci5
2024 A systematic investigation of learnability from single child linguistic input
Yulu Qin, Brenden M. Lake
CogSci3
2024 Is Deep Learning the Answer for Understanding Human Cognitive Dynamics?
John P. Spencer, Brenden M. Lake, Raul Grieben, Gregor Schöner, Mariya Toneva, Gina R. Kuperberg
CogSci2
2024 Compositional learning of functions in humans and machines
Yanli Zhou, Brenden M. Lake, Adina Williams
CogSci2
2024 Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
abstract
Though vision transformers (ViTs) have achieved state-of-the-art performance in a variety of settings, they exhibit surprising failures when performing tasks involving visual relations. This begs the question: how do ViTs attempt to perform tasks that require computing visual relations between objects? Prior efforts to interpret ViTs tend to focus on characterizing relevant low-level visual features. In contrast, we adopt methods from mechanistic interpretability to study the higher-level visual algorithms that ViTs use to perform abstract visual reasoning. We present a case study of a fundamental, yet surprisingly difficult, relational reasoning task: judging whether two visual entities are the same or different. We find that pretrained ViTs fine-tuned on this task often exhibit two qualitatively different stages of processing despite having no obvious inductive biases to do so: 1) a perceptual stage wherein local object features are extracted and stored in a disentangled representation, and 2) a relational stage wherein object representations are compared. In the second stage, we find evidence that ViTs can learn to represent somewhat abstract visual relations, a capability that has long been considered out of reach for artificial neural networks. Finally, we demonstrate that failures at either stage can prevent a model from learning a generalizable solution to our fairly simple tasks. By understanding ViTs in terms of discrete processing stages, one can more precisely diagnose and rectify shortcomings of existing and future models.
Michael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre, Brenden M. Lake, Ellie Pavlick
NeurIPS5
2022 Creativity, Compositionality, and Common Sense in Human Goal Generation
Guy Davidson, Todd M. Gureckis, Brenden M. Lake
CogSci3
2022 Evaluating locality in NMT models
Itay Itzhak, Koustuv Sinha, Brenden M. Lake, Adina Williams, Dieuwke Hupkes
CogSci3
2022 Improving Systematic Generalization Through Modularity and Augmentation
Laura Ruis, Brenden M. Lake
CogSci2
2022 A Developmentally-Inspired Examination of Shape versus Texture Bias in Machines
Alexa R. Tartaglini, Wai Keen Vong, Brenden M. Lake
CogSci3
2022 Categorising images by generating natural language rules
Wai Keen Vong, Brenden M. Lake
CogSci2
2021 Examining Infant Relation Categorization Through Deep Neural Networks
Guy Davidson, Brenden M. Lake
CogSci2
2021 Fast and Flexible: Human program induction in abstract reasoning tasks
Aysja Johnson, Wai Keen Vong, Brenden M. Lake, Todd M. Gureckis
CogSci3
2021 The Omniglot Jr. challenge; Can a model achieve child-level character generation and classification?
Eliza Kosoy, Masha Belyi, Charlie Snell, Brenden M. Lake, Josh Tenenbaum, Alison Gopnik
CogSci4
2021 Evaluating infants' reasoning about agents using the Baby Intuitions Benchmark (BIB)
Gala Stojnic, Kanishk Gandhi, Brenden M. Lake, Moira R. Dillon
CogSci3
2021 Modeling artificial category learning from pixels: Revisiting Shepard, Hovland, and Jenkins (1961) with deep neural networks
Alexa R. Tartaglini, Wai Keen Vong, Brenden M. Lake
CogSci3
2021 Modeling Question Asking Using Neural Program Generation
Brenden M. Lake
CogSci2
2021 Learning Task-General Representations with Generative Neuro-Symbolic Modeling
Reuben Feinman, Brenden M. Lake
ICLR2
2021 CURI: A Benchmark for Productive Concept Learning Under Uncertainty
abstract
Humans can learn and reason under substantial uncertainty in a space of infinitely many compositional, productive concepts. For example, if a scene with two blue spheres qualifies as “daxy,” one can reason that the underlying concept may require scenes to have “only blue spheres” or “only spheres” or “only two objects.” In contrast, standard benchmarks for compositional reasoning do not explicitly capture a notion of reasoning under uncertainty or evaluate compositional concept acquisition. We introduce a new benchmark, Compositional Reasoning Under Uncertainty (CURI) that instantiates a series of few-shot, meta-learning tasks in a productive concept space to evaluate different aspects of systematic generalization under uncertainty, including splits that test abstract understandings of disentangling, productive generalization, learning boolean operations, variable binding, etc. Importantly, we also contribute a model-independent “compositionality gap” to evaluate the difficulty of generalizing out-of-distribution along each of these axes, allowing objective comparison of the difficulty of each compositional split. Evaluations across a range of modeling choices and splits reveal substantial room for improvement on the proposed benchmark.
Ramakrishna Vedantam, Arthur Szlam, Maximilian Nickel, Ari S. Morcos, Brenden M. Lake
ICML5
2021 Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of others
abstract
To achieve human-like common sense about everyday life, machine learning systems must understand and reason about the goals, preferences, and actions of other agents in the environment. By the end of their first year of life, human infants intuitively achieve such common sense, and these cognitive achievements lay the foundation for humans' rich and complex understanding of the mental states of others. Can machines achieve generalizable, commonsense reasoning about other agents like human infants? The Baby Intuitions Benchmark (BIB) challenges machines to predict the plausibility of an agent's behavior based on the underlying causes of its actions. Because BIB's content and paradigm are adopted from developmental cognitive science, BIB allows for direct comparison between human and machine performance. Nevertheless, recently proposed, deep-learning-based agency reasoning models fail to show infant-like reasoning, leaving BIB an open challenge.
Kanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. Dillon
NeurIPS3
2021 Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic Reasoning
abstract
Human reasoning can be understood as an interplay between two systems: the intuitive and associative ("System 1") and the deliberative and logical ("System 2"). Neural sequence models---which have been increasingly successful at performing complex, structured tasks---exhibit the advantages and failure modes of System 1: they are fast and learn patterns from data, but are often inconsistent and incoherent. In this work, we seek a lightweight, training-free means of improving existing System 1-like sequence models by adding System 2-inspired logical reasoning. We explore several variations on this theme in which candidate generations from a neural sequence model are examined for logical consistency by a symbolic reasoning module, which can either accept or reject the generations. Our approach uses neural inference to mediate between the neural System 1 and the logical System 2. Results in robust story generation and grounded instruction-following show that this approach can increase the coherence and accuracy of neurally-based generations.
Maxwell I. Nye, Michael Henry Tessler, Josh Tenenbaum, Brenden M. Lake
NeurIPS4
2020 Investigating Simple Object Representations in Model-Free Deep Reinforcement Learning
Guy Davidson, Brenden M. Lake
CogSci2
2020 Generating new concepts with hybrid neuro-symbolic models
Reuben Feinman, Brenden M. Lake
CogSci2
2020 Extending the Rogers and McClelland Model of Semantic Cognition (2003) to work with Raw Pixel Information
Arihant Jain, Brenden M. Lake, Todd M. Gureckis
CogSci2
2020 Learning word-referent mappings and concepts from raw inputs
Wai Keen Vong, Brenden M. Lake
CogSci2
2020 Mutual exclusivity as a challenge for deep neural networks
abstract
Strong inductive biases allow children to learn in fast and adaptable ways. Children use the mutual exclusivity (ME) bias to help disambiguate how words map to referents, assuming that if an object has one label then it does not need another. In this paper, we investigate whether or not vanilla neural architectures have an ME bias, demonstrating that they lack this learning assumption. Moreover, we show that their inductive biases are poorly matched to lifelong learning formulations of classification and translation. We demonstrate that there is a compelling case for designing task-general neural networks that learn through mutual exclusivity, which remains an open challenge.
Kanishk Gandhi, Brenden M. Lake
NeurIPS2
2020 Learning Compositional Rules via Neural Program Synthesis
abstract
Many aspects of human reasoning, including language, require learning rules from very little data. Humans can do this, often learning systematic rules from very few examples, and combining these rules to form compositional rule-based systems. Current neural architectures, on the other hand, often fail to generalize in a compositional manner, especially when evaluated in ways that vary systematically from training. In this work, we present a neuro-symbolic model which learns entire rule systems from a small set of examples. Instead of directly predicting outputs from inputs, we train our model to induce the explicit system of rules governing a set of previously seen examples, drawing upon techniques from the neural program synthesis literature. Our rule-synthesis approach outperforms neural meta-learning techniques in three domains: an artificial instruction-learning domain used to evaluate human learning, the SCAN challenge datasets, and learning rule-based translations of number words into integers for a wide range of human languages.
Maxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. Lake
NeurIPS4
2020 Self-supervised learning through the eyes of a child
abstract
Within months of birth, children develop meaningful expectations about the world around them. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and recent progress in deep learning. In this paper, our goal is precisely to achieve such progress by utilizing modern self-supervised deep learning methods and a recent longitudinal, egocentric video dataset recorded from the perspective of three young children (Sullivan et al., 2020). Our results demonstrate the emergence of powerful, high-level visual representations from developmentally realistic natural videos using generic self-supervised learning objectives.
A. Emin Orhan, Vaibhav V. Gupta, Brenden M. Lake
NeurIPS3
2020 A Benchmark for Systematic Generalization in Grounded Language Understanding
abstract
Humans easily interpret expressions that describe unfamiliar situations composed from familiar parts ("greet the pink brontosaurus by the ferris wheel"). Modern neural networks, by contrast, struggle to interpret novel compositions. In this paper, we introduce a new benchmark, gSCAN, for evaluating compositional generalization in situated language understanding. Going beyond a related benchmark that focused on syntactic aspects of generalization, gSCAN defines a language grounded in the states of a grid world, facilitating novel evaluations of acquiring linguistically motivated rules. For example, agents must understand how adjectives such as 'small' are interpreted relative to the current world state or how adverbs such as 'cautiously' combine with new verbs. We test a strong multi-modal baseline model and a state-of-the-art compositional method finding that, in most cases, they fail dramatically when generalization requires systematic compositional rules.
Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt, Brenden M. Lake
NeurIPS5
2019 Learning a smooth kernel regularizer for convolutional neural networks
Reuben Feinman, Brenden M. Lake
CogSci2
2019 Human few-shot learning of compositional instructions
Brenden M. Lake, Tal Linzen, Marco Baroni
CogSci1
2019 Asking goal-oriented questions and learning from answers
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis
CogSci2
2019 Compositional generalization through meta sequence-to-sequence learning
abstract
People can learn a new concept and use it compositionally, understanding how to "blicket twice" after learning how to "blicket." In contrast, powerful sequence-to-sequence (seq2seq) neural networks fail such tests of compositionality, especially when composing new concepts together with existing concepts. In this paper, I show how memory-augmented neural networks can be trained to generalize compositionally through meta seq2seq learning. In this approach, models train on a series of seq2seq problems to acquire the compositional skills needed to solve new seq2seq problems. Meta se2seq learning solves several of the SCAN tests for compositional learning and can learn to apply implicit rules to variables.
Brenden M. Lake
NeurIPS1
2018 Learning Inductive Biases with Simple Neural Networks
Reuben Feinman, Brenden M. Lake
CogSci2
2018 Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks
abstract
Humans can understand and produce new utterances effortlessly, thanks to their compositional skills. Once a person learns the meaning of a new verb "dax," he or she can immediately understand the meaning of "dax twice" or "sing and dax." In this paper, we introduce the SCAN domain, consisting of a set of simple compositional navigation commands paired with the corresponding action sequences. We then test the zero-shot generalization capabilities of a variety of recurrent neural networks (RNNs) trained on SCAN with sequence-to-sequence methods. We find that RNNs can make successful zero-shot generalizations when the differences between training and test commands are small, so that they can apply "mix-and-match" strategies to solve the task. However, when generalization requires systematic compositional skills (as in the "dax" example above), RNNs fail spectacularly. We conclude with a proof-of-concept experiment in neural machine translation, suggesting that lack of systematicity might be partially responsible for neural networks’ notorious training data thirst.
Brenden M. Lake, Marco Baroni
ICML1
2017 One-shot Learning and Classification in Children
Eliza Kosoy, Brenden M. Lake, Josh Tenenbaum
CogSci2
2017 Progress in building a machine that can ask interesting and informative questions
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis
CogSci2
2017 Question Asking as Program Generation
abstract
A hallmark of human intelligence is the ability to ask rich, creative, and revealing questions. Here we introduce a cognitive model capable of constructing human-like questions. Our approach treats questions as formal programs that, when executed on the state of the world, output an answer. The model specifies a probability distribution over a complex, compositional space of programs, favoring concise programs that help the agent learn in the current context. We evaluate our approach by modeling the types of open-ended questions generated by humans who were attempting to learn about an ambiguous situation in a game. We find that our model predicts what questions people will ask, and can creatively produce novel questions that were not present in the training set. In addition, we compare a number of model variants, finding that both question informativeness and complexity are important for producing human-like questions.
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis
NIPS2
2016 Searching large hypothesis spaces by asking questions
Alexander Cohen, Brenden M. Lake
CogSci2
2016 Asking and evaluating natural language questions
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis
CogSci2
2015 Deep Neural Networks Predict Category Typicality Ratings for Images
Brenden M. Lake, Wojciech Zaremba, Rob Fergus, Todd M. Gureckis
CogSci1
2015 Asking useful questions: Active learning with rich queries
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis
CogSci2
2015 Softstar: Heuristic-Guided Probabilistic Inference
abstract
Recent machine learning methods for sequential behavior prediction estimate the motives of behavior rather than the behavior itself. This higher-level abstraction improves generalization in different prediction settings, but computing predictions often becomes intractable in large decision spaces. We propose the Softstar algorithm, a softened heuristic-guided search technique for the maximum entropy inverse optimal control model of sequential behavior. This approach supports probabilistic search with bounded approximation error at a significantly reduced computational cost when compared to sampling based methods. We present the algorithm, analyze approximation guarantees, and compare performance with simulation-based inference on two distinct complex decision tasks.
Mathew Monfort, Brenden M. Lake, Brian D. Ziebart, Patrick Lucey, Josh Tenenbaum
NIPS2
2014 Adaptive teaching: Improving the efficiency of learning through hypothesis-dependent selection of training data
Patricia Angie Chan, Douglas Markant, Brenden M. Lake, Todd M. Gureckis
CogSci3
2014 One-shot learning of generative speech concepts
Brenden M. Lake, Chia-ying Lee, James R. Glass, Josh Tenenbaum
CogSci1
2014 Computational Creativity: Generating new objects with a hierarchical Bayesian model
Brenden M. Lake, Josh Tenenbaum
CogSci1
2013 One-shot learning by inverting a compositional causal process
abstract
People can learn a new visual class from just one example, yet machine learning algorithms typically require hundreds or thousands of examples to tackle the same problems. Here we present a Hierarchical Bayesian model based on compositionality and causality that can learn a wide range of natural (although simple) visual concepts, generalizing in human-like ways from just one image. We evaluated performance on a challenging one-shot classification task, where our model achieved a human-level error rate while substantially outperforming two deep learning models. We also used a visual Turing test" to show that our model produces human-like performance on other conceptual tasks, including generating new examples and parsing."
Brenden M. Lake, Ruslan Salakhutdinov, Josh Tenenbaum
NIPS1
2012 Concept learning as motor program induction: A large-scale empirical study
Brenden M. Lake, Ruslan Salakhutdinov, Josh Tenenbaum
CogSci1
2011 Estimating the strength of unlabeled information during semi-supervised learning
Brenden M. Lake, James L. McClelland
CogSci1
2011 One shot learning of simple visual concepts
Brenden M. Lake, Ruslan Salakhutdinov, Jason Gross, Josh Tenenbaum
CogSci1