VLDB 2026 Research / reviewers in the wild / expert
Brenden M. Lake
dblp:47/9567
· DBLP profile ↗
60ranked-venue papers
10as first author
31since 2021 · last 2025
0000-0001-8959-3401ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 10 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 43 · 7 first-author · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Goal Inference using Reward-Producing Programs in a Novel Physics Environment
Guy Davidson, Graham Todd, Cédric Colas, Junyi Chu, Julian Togelius, Josh Tenenbaum, Todd M. Gureckis, Brenden M. Lake |
CogSci | 8 |
| 2025 | Do Large Language Models Reason Causally Like Us? Even Better?
Hanna M. Dettki, Brenden M. Lake, Charley M. Wu, Bob Rehder |
CogSci | 2 |
| 2025 | A Neurosymbolic Model of Human Reasoning on the Abstraction and Reasoning Corpus
Solim LeGris, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2025 | Rapid Word Learning Through Meta In-Context LearningabstractHumans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts.Yet the abilities of current language models for fewshot word learning, and methods for improving these abilities, are underexplored.In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow).This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word.This training is repeated on many new words to develop a general word-learning ability.We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pretrained on orders of magnitude more data.Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples.These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks. Guangyuan Jiang, Tal Linzen, Brenden M. Lake |
EMNLP | 4 |
| 2025 | SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety FactsabstractDo LLMs robustly generalize critical safety facts to novel situations? Lacking this ability is dangerous when users ask naive questions—for instance, ``I'm considering packing melon balls for my 10-month-old's lunch. What other foods would be good to include?'' Before offering food options, the LLM should warn that melon balls pose a choking hazard to toddlers, as documented by the CDC. Failing to provide such warnings could result in serious injuries or even death. To evaluate this, we introduce SAGE-Eval, SAfety-fact systematic GEneralization evaluation, the first benchmark that tests whether LLMs properly apply well‑established safety facts to naive user queries. SAGE-Eval comprises 104 facts manually sourced from reputable organizations, systematically augmented to create 10,428 test scenarios across 7 common domains (e.g., Outdoor Activities, Medicine). We find that the top model, Claude-3.7-sonnet, passes only 58% of all the safety facts tested. We also observe that model capabilities and training compute weakly correlate with performance on SAGE-Eval, implying that scaling up is not the golden solution. Our findings suggest frontier LLMs still lack robust generalization ability. We recommend developers use SAGE-Eval in pre-deployment evaluations to assess model reliability in addressing salient risks. Yueh-Han Chen, Guy Davidson, Brenden M. Lake |
NeurIPS | 3 |
| 2025 | Do different prompting methods yield a common task representation in language models?abstractDemonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks.
Do identical tasks elicited in different ways result in similar representations of the task? An improved understanding of task representation mechanisms would offer interpretability insights and may aid in steering models. We study this through function vectors (FVs), recently proposed as a mechanism to extract few-shot ICL task representations. We generalize FVs to alternative task presentations, focusing on short textual instruction prompts, and successfully extract instruction function vectors that promote zero-shot task accuracy. We find evidence that demonstration- and instruction-based function vectors leverage different model components, and offer several controls to dissociate their contributions to task performance. Our results suggest that different task prompting forms do not induce a common task representation through FVs but elicit different, partly overlapping mechanisms. Our findings offer principled support to the practice of combining instructions and task demonstrations, imply challenges in universally monitoring task inference across presentation forms, and encourage further examinations of LLM task inference mechanisms. Guy Davidson, Todd M. Gureckis, Brenden M. Lake, Adina Williams |
NeurIPS | 3 |
| 2024 | Comparing Abstraction in Humans and Machines Using Multimodal Serial Reproduction
Sreejan Kumar, Raja Marjieh, Byron Zhang, Declan Campbell, Michael Y. Hu, Umang Bhatt, Brenden M. Lake, Thomas L. Griffiths 0001 |
CogSci | 7 |
| 2024 | Predicting Insight during Physical Reasoning
Solim LeGris, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2024 | Prompting invokes expert-like downward shifts in GPT-4V's conceptual hierarchies
Cara Su-Yi Leong, Brenden M. Lake |
CogSci | 2 |
| 2024 | An Infant-Cognition Inspired Machine Benchmark for Identifying Agency, Affiliation, Belief, and Intention
Shannon Yasuda, Moira R. Dillon, Brenden M. Lake |
CogSci | 4 |
| 2024 | Finding Unsupervised Alignment of Conceptual Systems in Image-Word Representations
Kexin Luo, Yajie Xiao, Brenden M. Lake |
CogSci | 4 |
| 2024 | Self-supervised learning of video representations from a child's perspective
A. Emin Orhan, Alex N. Wang, Mengye Ren, Brenden M. Lake |
CogSci | 5 |
| 2024 | A systematic investigation of learnability from single child linguistic input
Yulu Qin, Brenden M. Lake |
CogSci | 3 |
| 2024 | Is Deep Learning the Answer for Understanding Human Cognitive Dynamics?
John P. Spencer, Brenden M. Lake, Raul Grieben, Gregor Schöner, Mariya Toneva, Gina R. Kuperberg |
CogSci | 2 |
| 2024 | Compositional learning of functions in humans and machines
Yanli Zhou, Brenden M. Lake, Adina Williams |
CogSci | 2 |
| 2024 | Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsabstractThough vision transformers (ViTs) have achieved state-of-the-art performance in a variety of settings, they exhibit surprising failures when performing tasks involving visual relations. This begs the question: how do ViTs attempt to perform tasks that require computing visual relations between objects? Prior efforts to interpret ViTs tend to focus on characterizing relevant low-level visual features. In contrast, we adopt methods from mechanistic interpretability to study the higher-level visual algorithms that ViTs use to perform abstract visual reasoning. We present a case study of a fundamental, yet surprisingly difficult, relational reasoning task: judging whether two visual entities are the same or different. We find that pretrained ViTs fine-tuned on this task often exhibit two qualitatively different stages of processing despite having no obvious inductive biases to do so: 1) a perceptual stage wherein local object features are extracted and stored in a disentangled representation, and 2) a relational stage wherein object representations are compared. In the second stage, we find evidence that ViTs can learn to represent somewhat abstract visual relations, a capability that has long been considered out of reach for artificial neural networks. Finally, we demonstrate that failures at either stage can prevent a model from learning a generalizable solution to our fairly simple tasks. By understanding ViTs in terms of discrete processing stages, one can more precisely diagnose and rectify shortcomings of existing and future models. Michael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre, Brenden M. Lake, Ellie Pavlick |
NeurIPS | 5 |
| 2022 | Creativity, Compositionality, and Common Sense in Human Goal Generation
Guy Davidson, Todd M. Gureckis, Brenden M. Lake |
CogSci | 3 |
| 2022 | Evaluating locality in NMT models
Itay Itzhak, Koustuv Sinha, Brenden M. Lake, Adina Williams, Dieuwke Hupkes |
CogSci | 3 |
| 2022 | Improving Systematic Generalization Through Modularity and Augmentation
Laura Ruis, Brenden M. Lake |
CogSci | 2 |
| 2022 | A Developmentally-Inspired Examination of Shape versus Texture Bias in Machines
Alexa R. Tartaglini, Wai Keen Vong, Brenden M. Lake |
CogSci | 3 |
| 2022 | Categorising images by generating natural language rules
Wai Keen Vong, Brenden M. Lake |
CogSci | 2 |
| 2021 | Examining Infant Relation Categorization Through Deep Neural Networks
Guy Davidson, Brenden M. Lake |
CogSci | 2 |
| 2021 | Fast and Flexible: Human program induction in abstract reasoning tasks
Aysja Johnson, Wai Keen Vong, Brenden M. Lake, Todd M. Gureckis |
CogSci | 3 |
| 2021 | The Omniglot Jr. challenge; Can a model achieve child-level character generation and classification?
Eliza Kosoy, Masha Belyi, Charlie Snell, Brenden M. Lake, Josh Tenenbaum, Alison Gopnik |
CogSci | 4 |
| 2021 | Evaluating infants' reasoning about agents using the Baby Intuitions Benchmark (BIB)
Gala Stojnic, Kanishk Gandhi, Brenden M. Lake, Moira R. Dillon |
CogSci | 3 |
| 2021 | Modeling artificial category learning from pixels: Revisiting Shepard, Hovland, and Jenkins (1961) with deep neural networks
Alexa R. Tartaglini, Wai Keen Vong, Brenden M. Lake |
CogSci | 3 |
| 2021 | Modeling Question Asking Using Neural Program Generation
Brenden M. Lake |
CogSci | 2 |
| 2021 | Learning Task-General Representations with Generative Neuro-Symbolic Modeling
Reuben Feinman, Brenden M. Lake |
ICLR | 2 |
| 2021 | CURI: A Benchmark for Productive Concept Learning Under UncertaintyabstractHumans can learn and reason under substantial uncertainty in a space of infinitely many compositional, productive concepts. For example, if a scene with two blue spheres qualifies as “daxy,” one can reason that the underlying concept may require scenes to have “only blue spheres” or “only spheres” or “only two objects.” In contrast, standard benchmarks for compositional reasoning do not explicitly capture a notion of reasoning under uncertainty or evaluate compositional concept acquisition. We introduce a new benchmark, Compositional Reasoning Under Uncertainty (CURI) that instantiates a series of few-shot, meta-learning tasks in a productive concept space to evaluate different aspects of systematic generalization under uncertainty, including splits that test abstract understandings of disentangling, productive generalization, learning boolean operations, variable binding, etc. Importantly, we also contribute a model-independent “compositionality gap” to evaluate the difficulty of generalizing out-of-distribution along each of these axes, allowing objective comparison of the difficulty of each compositional split. Evaluations across a range of modeling choices and splits reveal substantial room for improvement on the proposed benchmark. Ramakrishna Vedantam, Arthur Szlam, Maximilian Nickel, Ari S. Morcos, Brenden M. Lake |
ICML | 5 |
| 2021 | Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of othersabstractTo achieve human-like common sense about everyday life, machine learning systems must understand and reason about the goals, preferences, and actions of other agents in the environment. By the end of their first year of life, human infants intuitively achieve such common sense, and these cognitive achievements lay the foundation for humans' rich and complex understanding of the mental states of others. Can machines achieve generalizable, commonsense reasoning about other agents like human infants? The Baby Intuitions Benchmark (BIB) challenges machines to predict the plausibility of an agent's behavior based on the underlying causes of its actions. Because BIB's content and paradigm are adopted from developmental cognitive science, BIB allows for direct comparison between human and machine performance. Nevertheless, recently proposed, deep-learning-based agency reasoning models fail to show infant-like reasoning, leaving BIB an open challenge. Kanishk Gandhi, Gala Stojnic, Brenden M. Lake, Moira R. Dillon |
NeurIPS | 3 |
| 2021 | Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic ReasoningabstractHuman reasoning can be understood as an interplay between two systems: the intuitive and associative ("System 1") and the deliberative and logical ("System 2"). Neural sequence models---which have been increasingly successful at performing complex, structured tasks---exhibit the advantages and failure modes of System 1: they are fast and learn patterns from data, but are often inconsistent and incoherent. In this work, we seek a lightweight, training-free means of improving existing System 1-like sequence models by adding System 2-inspired logical reasoning. We explore several variations on this theme in which candidate generations from a neural sequence model are examined for logical consistency by a symbolic reasoning module, which can either accept or reject the generations. Our approach uses neural inference to mediate between the neural System 1 and the logical System 2. Results in robust story generation and grounded instruction-following show that this approach can increase the coherence and accuracy of neurally-based generations. Maxwell I. Nye, Michael Henry Tessler, Josh Tenenbaum, Brenden M. Lake |
NeurIPS | 4 |
| 2020 | Investigating Simple Object Representations in Model-Free Deep Reinforcement Learning
Guy Davidson, Brenden M. Lake |
CogSci | 2 |
| 2020 | Generating new concepts with hybrid neuro-symbolic models
Reuben Feinman, Brenden M. Lake |
CogSci | 2 |
| 2020 | Extending the Rogers and McClelland Model of Semantic Cognition (2003) to work with Raw Pixel Information
Arihant Jain, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2020 | Learning word-referent mappings and concepts from raw inputs
Wai Keen Vong, Brenden M. Lake |
CogSci | 2 |
| 2020 | Mutual exclusivity as a challenge for deep neural networksabstractStrong inductive biases allow children to learn in fast and adaptable ways. Children use the mutual exclusivity (ME) bias to help disambiguate how words map to referents, assuming that if an object has one label then it does not need another. In this paper, we investigate whether or not vanilla neural architectures have an ME bias, demonstrating that they lack this learning assumption. Moreover, we show that their inductive biases are poorly matched to lifelong learning formulations of classification and translation. We demonstrate that there is a compelling case for designing task-general neural networks that learn through mutual exclusivity, which remains an open challenge. Kanishk Gandhi, Brenden M. Lake |
NeurIPS | 2 |
| 2020 | Learning Compositional Rules via Neural Program SynthesisabstractMany aspects of human reasoning, including language, require learning rules from very little data. Humans can do this, often learning systematic rules from very few examples, and combining these rules to form compositional rule-based systems. Current neural architectures, on the other hand, often fail to generalize in a compositional manner, especially when evaluated in ways that vary systematically from training. In this work, we present a neuro-symbolic model which learns entire rule systems from a small set of examples. Instead of directly predicting outputs from inputs, we train our model to induce the explicit system of rules governing a set of previously seen examples, drawing upon techniques from the neural program synthesis literature. Our rule-synthesis approach outperforms neural meta-learning techniques in three domains: an artificial instruction-learning domain used to evaluate human learning, the SCAN challenge datasets, and learning rule-based translations of number words into integers for a wide range of human languages. Maxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. Lake |
NeurIPS | 4 |
| 2020 | Self-supervised learning through the eyes of a childabstractWithin months of birth, children develop meaningful expectations about the world around them. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and recent progress in deep learning. In this paper, our goal is precisely to achieve such progress by utilizing modern self-supervised deep learning methods and a recent longitudinal, egocentric video dataset recorded from the perspective of three young children (Sullivan et al., 2020). Our results demonstrate the emergence of powerful, high-level visual representations from developmentally realistic natural videos using generic self-supervised learning objectives. A. Emin Orhan, Vaibhav V. Gupta, Brenden M. Lake |
NeurIPS | 3 |
| 2020 | A Benchmark for Systematic Generalization in Grounded Language UnderstandingabstractHumans easily interpret expressions that describe unfamiliar situations composed from familiar parts ("greet the pink brontosaurus by the ferris wheel"). Modern neural networks, by contrast, struggle to interpret novel compositions. In this paper, we introduce a new benchmark, gSCAN, for evaluating compositional generalization in situated language understanding. Going beyond a related benchmark that focused on syntactic aspects of generalization, gSCAN defines a language grounded in the states of a grid world, facilitating novel evaluations of acquiring linguistically motivated rules. For example, agents must understand how adjectives such as 'small' are interpreted relative to the current world state or how adverbs such as 'cautiously' combine with new verbs. We test a strong multi-modal baseline model and a state-of-the-art compositional method finding that, in most cases, they fail dramatically when generalization requires systematic compositional rules. Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt, Brenden M. Lake |
NeurIPS | 5 |
| 2019 | Learning a smooth kernel regularizer for convolutional neural networks
Reuben Feinman, Brenden M. Lake |
CogSci | 2 |
| 2019 | Human few-shot learning of compositional instructions
Brenden M. Lake, Tal Linzen, Marco Baroni |
CogSci | 1 |
| 2019 | Asking goal-oriented questions and learning from answers
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2019 | Compositional generalization through meta sequence-to-sequence learningabstractPeople can learn a new concept and use it compositionally, understanding how to "blicket twice" after learning how to "blicket." In contrast, powerful sequence-to-sequence (seq2seq) neural networks fail such tests of compositionality, especially when composing new concepts together with existing concepts. In this paper, I show how memory-augmented neural networks can be trained to generalize compositionally through meta seq2seq learning. In this approach, models train on a series of seq2seq problems to acquire the compositional skills needed to solve new seq2seq problems. Meta se2seq learning solves several of the SCAN tests for compositional learning and can learn to apply implicit rules to variables. Brenden M. Lake |
NeurIPS | 1 |
| 2018 | Learning Inductive Biases with Simple Neural Networks
Reuben Feinman, Brenden M. Lake |
CogSci | 2 |
| 2018 | Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent NetworksabstractHumans can understand and produce new utterances effortlessly, thanks to their compositional skills. Once a person learns the meaning of a new verb "dax," he or she can immediately understand the meaning of "dax twice" or "sing and dax." In this paper, we introduce the SCAN domain, consisting of a set of simple compositional navigation commands paired with the corresponding action sequences. We then test the zero-shot generalization capabilities of a variety of recurrent neural networks (RNNs) trained on SCAN with sequence-to-sequence methods. We find that RNNs can make successful zero-shot generalizations when the differences between training and test commands are small, so that they can apply "mix-and-match" strategies to solve the task. However, when generalization requires systematic compositional skills (as in the "dax" example above), RNNs fail spectacularly. We conclude with a proof-of-concept experiment in neural machine translation, suggesting that lack of systematicity might be partially responsible for neural networks’ notorious training data thirst. Brenden M. Lake, Marco Baroni |
ICML | 1 |
| 2017 | One-shot Learning and Classification in Children
Eliza Kosoy, Brenden M. Lake, Josh Tenenbaum |
CogSci | 2 |
| 2017 | Progress in building a machine that can ask interesting and informative questions
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2017 | Question Asking as Program GenerationabstractA hallmark of human intelligence is the ability to ask rich, creative, and revealing questions. Here we introduce a cognitive model capable of constructing human-like questions. Our approach treats questions as formal programs that, when executed on the state of the world, output an answer. The model specifies a probability distribution over a complex, compositional space of programs, favoring concise programs that help the agent learn in the current context. We evaluate our approach by modeling the types of open-ended questions generated by humans who were attempting to learn about an ambiguous situation in a game. We find that our model predicts what questions people will ask, and can creatively produce novel questions that were not present in the training set. In addition, we compare a number of model variants, finding that both question informativeness and complexity are important for producing human-like questions. Anselm Rothe, Brenden M. Lake, Todd M. Gureckis |
NIPS | 2 |
| 2016 | Searching large hypothesis spaces by asking questions
Alexander Cohen, Brenden M. Lake |
CogSci | 2 |
| 2016 | Asking and evaluating natural language questions
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2015 | Deep Neural Networks Predict Category Typicality Ratings for Images
Brenden M. Lake, Wojciech Zaremba, Rob Fergus, Todd M. Gureckis |
CogSci | 1 |
| 2015 | Asking useful questions: Active learning with rich queries
Anselm Rothe, Brenden M. Lake, Todd M. Gureckis |
CogSci | 2 |
| 2015 | Softstar: Heuristic-Guided Probabilistic InferenceabstractRecent machine learning methods for sequential behavior prediction estimate the motives of behavior rather than the behavior itself. This higher-level abstraction improves generalization in different prediction settings, but computing predictions often becomes intractable in large decision spaces. We propose the Softstar algorithm, a softened heuristic-guided search technique for the maximum entropy inverse optimal control model of sequential behavior. This approach supports probabilistic search with bounded approximation error at a significantly reduced computational cost when compared to sampling based methods. We present the algorithm, analyze approximation guarantees, and compare performance with simulation-based inference on two distinct complex decision tasks. Mathew Monfort, Brenden M. Lake, Brian D. Ziebart, Patrick Lucey, Josh Tenenbaum |
NIPS | 2 |
| 2014 | Adaptive teaching: Improving the efficiency of learning through hypothesis-dependent selection of training data
Patricia Angie Chan, Douglas Markant, Brenden M. Lake, Todd M. Gureckis |
CogSci | 3 |
| 2014 | One-shot learning of generative speech concepts
Brenden M. Lake, Chia-ying Lee, James R. Glass, Josh Tenenbaum |
CogSci | 1 |
| 2014 | Computational Creativity: Generating new objects with a hierarchical Bayesian model
Brenden M. Lake, Josh Tenenbaum |
CogSci | 1 |
| 2013 | One-shot learning by inverting a compositional causal processabstractPeople can learn a new visual class from just one example, yet machine learning algorithms typically require hundreds or thousands of examples to tackle the same problems. Here we present a Hierarchical Bayesian model based on compositionality and causality that can learn a wide range of natural (although simple) visual concepts, generalizing in human-like ways from just one image. We evaluated performance on a challenging one-shot classification task, where our model achieved a human-level error rate while substantially outperforming two deep learning models. We also used a visual Turing test" to show that our model produces human-like performance on other conceptual tasks, including generating new examples and parsing." Brenden M. Lake, Ruslan Salakhutdinov, Josh Tenenbaum |
NIPS | 1 |
| 2012 | Concept learning as motor program induction: A large-scale empirical study
Brenden M. Lake, Ruslan Salakhutdinov, Josh Tenenbaum |
CogSci | 1 |
| 2011 | Estimating the strength of unlabeled information during semi-supervised learning
Brenden M. Lake, James L. McClelland |
CogSci | 1 |
| 2011 | One shot learning of simple visual concepts
Brenden M. Lake, Ruslan Salakhutdinov, Jason Gross, Josh Tenenbaum |
CogSci | 1 |