VLDB 2026 Research / reviewers in the wild / expert
Josh Tenenbaum
dblp:t/JoshuaBTenenbaum · also Joshua B. Tenenbaum, Joshua Tenenbaum
· DBLP profile ↗
497ranked-venue papers
10as first author
226since 2021 · last 2025
0000-0002-1925-2035ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 487 · 10 first-author · 218 since 2021Applied, interdisciplinary, general and emerging computing · 207 · 79 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 25 since 2021Systems, architecture and hardware · 27 · 18 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Studying Mathematical Reasoning through the Gadget Game
Jonas Bayer, Jacob Loader, Katie Collins, Simon Frieder, Adrian Weller, Josh Tenenbaum, Timothy Gowers |
CogSci | 6 |
| 2025 | Surprise isn't symmetrical: Adults' looking suggests non-perceptual considerations during dishabituation
Qiong Cao, Anjie Cao, Gal Raz, Josh Tenenbaum, Shari Liu |
CogSci | 4 |
| 2025 | A Computational Model of Human Vocal Imitation
Matthew Caren, Kartik Chandra, Josh Tenenbaum, Jonathan Ragan-Kelley, Karima Ma |
CogSci | 3 |
| 2025 | Theories of Mind as Languages of Thought for Thought about Thought
Kartik Chandra, Jonathan Ragan-Kelley, Josh Tenenbaum |
CogSci | 3 |
| 2025 | Decompose, Deduce, and Dispose: A Memory-Limited Metacognitive Model of Human Problem Solving
Samuel J. Cheyette, Tony Chen 0003, Matthias Hofer 0002, Frederick Callaway, Neil Bramley, Josh Tenenbaum |
CogSci | 6 |
| 2025 | Language and Experience: A Computational Model of Social Learning in Complex Novel Tasks
Cédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler, Noah D. Goodman, Jacob Andreas, Josh Tenenbaum |
CogSci | 7 |
| 2025 | Empathy in Explanation
Katie Collins, Kartik Chandra, Jonathan Ragan-Kelley, Adrian Weller, Josh Tenenbaum |
CogSci | 5 |
| 2025 | Generation and Evaluation in the Human Invention Process through the Lens of Game Design
Katie Collins, Graham Todd, Cedegao E. Zhang, Adrian Weller, Julian Togelius, Junyi Chu, Lionel Wong, Thomas L. Griffiths 0001, Josh Tenenbaum |
CogSci | 9 |
| 2025 | Seeing through Occlusion: Uncertainty-aware Joint Physical Tracking and Prediction
Arijit Dasgupta, Andrew D. Bolton, Vikash Mansinghka 0001, Josh Tenenbaum, Kevin A. Smith 0001 |
CogSci | 4 |
| 2025 | Goal Inference using Reward-Producing Programs in a Novel Physics Environment
Guy Davidson, Graham Todd, Cédric Colas, Junyi Chu, Julian Togelius, Josh Tenenbaum, Todd M. Gureckis, Brenden M. Lake |
CogSci | 6 |
| 2025 | Tracking Uncertainty During Uncertain Tracking
Yoni Friedman, Matin Ghavamizadeh, Maddy Bowers, Andrew D. Bolton, Max H. Siegel, Vikash Mansinghka 0001, Josh Tenenbaum |
CogSci | 7 |
| 2025 | Finding structure in logographic writing with library learning II: Grapheme, sound, and meaning systematicity
Guangyuan Jiang, Matthias Hofer 0002, Jiayuan Mao, Lionel Wong, Josh Tenenbaum, Roger Levy |
CogSci | 5 |
| 2025 | Calculating probabilities from imagined possibilities: Limitations in 4-year-olds
Brian Leahy, Vicente Vivanco, Samuel J. Cheyette, Kevin A. Smith 0001, Lucy J. White, Roman Feiman, Laura Schulz, Josh Tenenbaum |
CogSci | 8 |
| 2025 | Flexible Physical Problem Solving with Strategy Acquisition and Composition
Jung-Chun Liu, Jiayuan Mao, Joyce Y. Chai, Josh Tenenbaum |
CogSci | 4 |
| 2025 | Minding the Politeness Gap in Cross-cultural Communication
Yuka Machino, Max H. Siegel, Matthias Hofer 0002, Josh Tenenbaum, Robert D. Hawkins |
CogSci | 4 |
| 2025 | Strategy selection in complex tasks through adaptive integration of learned and online metareasoning
Tracey Mills, Samuel Gershman, Josh Tenenbaum |
CogSci | 3 |
| 2025 | Reverse-Engineering an Intuitive Psychology of Power
Junior Chinomso Okoroafor, Rebecca Saxe, Josh Tenenbaum, Max Kleiman-Weiner |
CogSci | 3 |
| 2025 | The trade-off between rule-based thinking and mutual benefit in tacit coordination
Arthur Le Pargneux, Sydney Levine, Josh Tenenbaum, Fiery Cushman |
CogSci | 3 |
| 2025 | Ensemble Physics: Perceiving the Mass of Groups of Objects is More Than the Sum of Its Parts
Vicente Vivanco, Josh Tenenbaum, Vivian C. Paulun, Kevin A. Smith 0001 |
CogSci | 2 |
| 2025 | From pixels to physics: an image-computable model of physical predictions
Nishad Gothoskar, Josh Tenenbaum, Kevin A. Smith 0001 |
CogSci | 3 |
| 2025 | Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
Lionel Wong, Katie Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy J. O'Donnell, Alexander K. Lew, Jacob Andreas, Tyler Brooke-Wilson, Josh Tenenbaum |
CogSci | 11 |
| 2025 | Scaling up the think-aloud method
Daniel Wurgaft, Ben Prystawski, Kanishk Gandhi, Cedegao E. Zhang, Josh Tenenbaum, Noah D. Goodman |
CogSci | 5 |
| 2025 | Belief Attribution as Mental Explanation: The Role of Accuracy, Informativity, and Causality
Lance Ying, Almog Hilel, Ryan Truong, Vikash Mansinghka 0001, Josh Tenenbaum, Tan Zhi-Xuan |
CogSci | 5 |
| 2025 | Adaptive Social Learning using Theory of Mind
Lance Ying, Ryan Truong, Josh Tenenbaum, Samuel Gershman |
CogSci | 3 |
| 2025 | What's in the Box? Reasoning about Unseen Objects from Multimodal Cues
Lance Ying, Daniel Xu, Alicia Zhang, Katie Collins, Max H. Siegel, Josh Tenenbaum |
CogSci | 6 |
| 2025 | On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad ConceptsabstractLanguage use is shaped by pragmatics-i.e., reasoning about communicative goals and norms in context.As language models (LMs) are increasingly used as conversational agents, it becomes ever more important to understand their pragmatic reasoning abilities.We propose an evaluation framework derived from Wavelength, a popular communication game where a speaker and a listener communicate about a broad range of concepts in a granular manner.We study a range of LMs on both language comprehension and language production using direct and Chain-of-Thought (CoT) prompting, and further explore a Rational Speech Act (RSA) approach to incorporating Bayesian pragmatic reasoning into LM inference.We find that state-of-the-art LMs, but not smaller ones, achieve strong performance on language comprehension, obtaining similar-to-human accuracy and exhibiting high correlations with human judgments even without CoT prompting or RSA.On language production, CoT can outperform direct prompting, and using RSA provides significant improvements over both approaches.Our study helps identify the strengths and limitations in LMs' pragmatic reasoning abilities and demonstrates the potential for improving them with RSA, opening up future avenues for understanding conceptual representation, language understanding, and social reasoning in LMs and humans. 1 Left Concept (0) Target Value Right Concept (100) Human-written Clues Chosen Clue Human Mean Deep thought 10 Shallow thought Evolution, Solving complex problems, Chess, Einstein, Meditation, Quantum mechanics Solv.complex prob. Linlu Qiu, Cedegao E. Zhang, Josh Tenenbaum, Roger Levy |
EMNLP | 3 |
| 2025 | What Makes a Maze Look Like a Maze?abstractA unique aspect of human visual understanding is the ability to flexibly interpret abstract concepts: acquiring lifted rules explaining what they symbolize, grounding them across familiar and unfamiliar contexts, and making predictions or reasoning about them. While off-the-shelf vision-language models excel at making literal interpretations of images (e.g., recognizing object categories such as tree branches), they still struggle to make sense of such visual abstractions (e.g., how an arrangement of tree branches may form the walls of a maze). To address this challenge, we introduce Deep Schema Grounding (DSG), a framework that leverages explicit structured representations of visual abstractions for grounding and reasoning. At the core of DSG are schemas—dependency graph descriptions of abstract concepts that decompose them into more primitive-level symbols. DSG uses large language models to extract schemas, then hierarchically grounds concrete to abstract components of the schema onto images with vision-language models. The grounded schema is used to augment visual abstraction understanding. We systematically evaluate DSG and different methods in reasoning on our new Visual Abstractions Benchmark, which consists of diverse, real-world images of abstract concepts and corresponding question-answer pairs labeled by humans. We show that DSG significantly improves the abstract visual reasoning performance of vision-language models, and is a step toward human-aligned understanding of visual abstractions. Joy Hsu, Jiayuan Mao, Josh Tenenbaum, Noah D. Goodman, Jiajun Wu 0001 |
ICLR | 3 |
| 2025 | VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot PlanningabstractBroadly intelligent agents should form task-specific abstractions that selectively expose the essential elements of a task, while abstracting away the complexity of the raw sensorimotor space. In this work, we present Neuro-Symbolic Predicates, a first-order abstraction language that combines the strengths of symbolic and neural knowledge representations. We outline an online algorithm for inventing such predicates and learning abstract world models. We compare our approach to hierarchical reinforcement learning, vision-language model planning, and symbolic predicate invention approaches, on both in- and out-of-distribution tasks across five simulated robotic domains. Results show that our approach offers better sample complexity, stronger out-of-distribution generalization, and improved interpretability. Yichao Liang, Nishanth Kumar, Hao Tang 0008, Adrian Weller, Josh Tenenbaum, Tom Silver, João F. Henriques, Kevin Ellis |
ICLR | 5 |
| 2025 | Can Large Language Models Understand Symbolic Graphics Programs?abstractAgainst the backdrop of enthusiasm for large language models (LLMs), there is a growing need to scientifically assess their capabilities and shortcomings. This is nontrivial in part because it is difficult to find tasks which the models have not encountered during training. Utilizing symbolic graphics programs, we propose a domain well-suited to test multiple spatial-semantic reasoning skills of LLMs. Popular in computer graphics, these programs procedurally generate visual data. While LLMs exhibit impressive skills in general program synthesis and analysis, symbolic graphics programs offer a new layer of evaluation: they allow us to test an LLM's ability to answer semantic questions about the images or 3D geometries without a vision encoder. To semantically understand the symbolic programs, LLMs would need to possess the ability to "imagine" and reason how the corresponding graphics content would look with only the symbolic description of the local curvatures and strokes. We use this task to evaluate LLMs by creating a large benchmark for the semantic visual understanding of symbolic graphics programs, built procedurally with minimal human effort. Particular emphasis is placed on transformations of images that leave the image level semantics invariant while introducing significant changes to the underlying program. We evaluate commercial and open-source LLMs on our benchmark to assess their ability to reason about visual output of programs, finding that LLMs considered stronger at reasoning generally perform better. Lastly, we introduce a novel method to improve this ability -- Symbolic Instruction Tuning (SIT), in which the LLM is finetuned with pre-collected instruction data on symbolic graphics programs. Interestingly, we find that SIT not only improves LLM's understanding on symbolic programs, but it also improves general reasoning ability on various other benchmarks. Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu 0019, Tim Z. Xiao, Katie Collins, Josh Tenenbaum, Adrian Weller, Michael J. Black, Bernhard Schölkopf |
ICLR | 7 |
| 2025 | Multiagent Finetuning: Self Improvement with Diverse Reasoning ChainsabstractLarge language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. However, successive steps of self-improvement can reach a point of diminishing returns. In this work, we propose a complementary approach towards self-improvement where finetuning is applied to a multiagent society of language models. A group of language models, all starting from the same base model, are independently specialized by updating each one using data generated through multiagent interactions among the models. By training each model on independent sets of data, we illustrate how this approach enables specialization across models and diversification over the set of models. As a result, our overall system is able to preserve diverse reasoning chains and autonomously improve over many more rounds of fine-tuning than single-agent self-improvement methods. We quantitatively illustrate the efficacy of the approach across a wide suite of reasoning tasks. Vighnesh Subramaniam, Yilun Du, Josh Tenenbaum, Antonio Torralba 0001, Shuang Li 0013, Igor Mordatch |
ICLR | 3 |
| 2025 | Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representationsabstractStudies of the functional role of the primate ventral visual stream have traditionally focused on object categorization, often ignoring -- despite much prior evidence -- its role in estimating "spatial" latents such as object position and pose. Most leading ventral stream models are derived by optimizing networks for object categorization, which seems to imply that the ventral stream is also derived under such an objective. Here, we explore an alternative hypothesis: Might the ventral stream be optimized for estimating spatial latents? And a closely related question: How different -- if at all -- are representations learned from spatial latent estimation compared to categorization? To ask these questions, we leveraged synthetic image datasets generated by a 3D graphic engine and trained convolutional neural networks (CNNs) to estimate different combinations of spatial and category latents. We found that models trained to estimate just a few spatial latents achieve neural alignment scores comparable to those trained on hundreds of categories, and the spatial latent performance of models strongly correlates with their neural alignment. Spatial latent and category-trained models have very similar -- but not identical -- internal representations, especially in their early and middle layers. We provide evidence that this convergence is partly driven by non-target latent variability in the training data, which facilitates the implicit learning of representations of those non-target latents. Taken together, these results suggest that many training objectives, such as spatial latents, can lead to similar models aligned neurally with the ventral stream. Thus, one should not assume that the ventral stream is optimized for object categorization only. As a field, we need to continue to sharpen our measures of comparing models to brains to better understand the functional roles of the ventral stream. Yudi Xie, Weichen Huang, Esther Alter, Jeremy Schwartz, Josh Tenenbaum, James J. DiCarlo |
ICLR | 5 |
| 2025 | Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language ModelsabstractPre-trained vision language models still fall short of human visual cognition. In an effort to improve visual cognition and align models with human behavior, we introduce visual stimuli and human judgments on visual cognition tasks, allowing us to systematically evaluate performance across cognitive domains under a consistent environment. We fine-tune models on ground truth data for intuitive physics and causal reasoning and find that this improves model performance in the respective fine-tuning domain. Furthermore, it can improve model alignment with human behavior. However, we find that task-specific fine-tuning does not contribute to robust human-like generalization to data with other visual characteristics or to tasks in other cognitive domains. Luca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata, Matthias Bethge, Josh Tenenbaum, Eric Schulz |
ICML | 5 |
| 2025 | KALM: Keypoint Abstraction Using Large Models for Object-Relative Imitation LearningabstractGeneralization to novel object configurations and instances across diverse tasks and environments is a critical challenge in robotics. Keypoint-based representations have been proven effective as a succinct representation for capturing essential object features, and for establishing a reference frame in action prediction, enabling data-efficient learning of robot skills. However, their manual design nature and reliance on additional human labels limit their scalability. In this paper, we propose KALM, a framework that leverages large pre-trained vision-language models (LMs) to automatically generate taskrelevant and cross-instance consistent keypoints. KALM distills robust and consistent keypoints across views and objects by generating proposals using LMs and verifies them against a small set of robot demonstration data. Based on the generated keypoints, we can train keypoint-conditioned policy models that predict actions in keypoint-centric frames, enabling robots to generalize effectively across varying object poses, camera views, and object instances with similar functional shapes. Our method demonstrates strong performance in the real world, adapting to different tasks and environments from only a handful of demonstrations while requiring no additional labels. Videos can be found at https://kalm-il.github.io/. Xiaolin Fang 0002, Bo-Ruei Huang, Jiayuan Mao, Jasmine Shone, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICRA | 5 |
| 2025 | One-Shot Manipulation Strategy Learning by Making Contact AnalogiesabstractWe present a novel approach, MAGIC (manipulation analogies for generalizable intelligent contacts), for one-shot learning of manipulation strategies with fast and extensive generalization to novel objects. By leveraging a reference action trajectory, MAGIC effectively identifies similar contact points and sequences of actions on novel objects to replicate a demonstrated strategy, such as using different hooks to retrieve distant objects of different shapes and sizes. Our method is based on a twostage contact-point matching process that combines global shape matching using pretrained neural features with local curvature analysis to ensure precise and physically plausible contact points. We experiment with three tasks including scooping, hanging, and hooking objects. MAGIC demonstrates superior performance over existing methods, achieving significant improvements in runtime speed and generalization to different object categories. Website: https://magic-2024.github.io/. Yuyao Liu, Jiayuan Mao, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICRA | 3 |
| 2025 | When Is It Acceptable to Break the Rules? Knowledge Representation of Moral Judgements Based on Empirical Data (Extended Abstract)
Edmond Awad, Sydney Levine, Andrea Loreggia, Nicholas Mattei, Iyad Rahwan, Francesca Rossi 0001, Kartik Talamadupula, Josh Tenenbaum, Max Kleiman-Weiner |
AAMAS | 8 |
| 2025 | Learning Linear Attention in Polynomial TimeabstractPrevious research has explored the expressivity of Transformer models in simulating Boolean circuits or Turing machines. However, the efficient learnability of Transformers from data has remained an open question. Our study addresses this gap by providing the first polynomial-time learnability results (specifically strong, agnostic PAC learning) for single-layer Transformers with linear attention. We show that learning the optimal multi head linear attention can be recast as finding the optimal kernel predictor in a suitably defined RKHS. Moving to generalization, we construct an algorithm that, given a dataset, checks in polynomial time whether the set of best fit multi head linear attention networks on this data all perform an identical computation--a powerful notion for out of distribution generalization. We empirically validate our theoretical findings on several canonical tasks: learning random linear attention networks, key--value associations, and learning to execute finite automata. Our findings bridge a critical gap between theoretical expressivity and learnability of Transformer models. Morris Yau, Ekin Akyürek, Jiayuan Mao, Josh Tenenbaum, Stefanie Jegelka, Jacob Andreas |
NeurIPS | 4 |
| 2025 | Stochastic Lazy Knowledge Compilation for Inference in Discrete Probabilistic ProgramsabstractWe present new techniques for exact and approximate inference in discrete probabilistic programs, based on two new ways of exploiting lazy evaluation. First, we show how knowledge compilation, a state-of-the art technique for exact inference in discrete probabilistic programs, can be made lazy, enabling asymptotic speed-ups. Second, we show how a probabilistic program’s lazy semantics naturally give rise to a division of its random choices into subproblems, which can be solved in sequence by sequential Monte Carlo with locallyoptimal proposals automatically computed via lazy knowledge compilation. We implement our approach in a new tool, Pluck , and evaluate its performance against state-of-the-art approaches to inference in discrete probabilistic languages. We !nd that on a suite of inference benchmarks, lazy knowledge compilation can be faster than state-of-the-art approaches, sometimes by orders of magnitude. Maddy Bowers, Alexander K. Lew, Josh Tenenbaum, Armando Solar-Lezama, Vikash Mansinghka 0001 |
Proc. ACM Program. Lang. | 3 |
| 2025 | A Domain-Specific Probabilistic Programming Language for Reasoning about Reasoning (Or: A Memo on memo)abstractThe human ability to think about thinking (“theory of mind”) is a fundamental object of study in many disciplines. In recent decades, researchers across these disciplines have converged on a rich computational paradigm for modeling theory of mind, grounded in recursive probabilistic reasoning. However, practitioners often find programming in this paradigm challenging: first, because thinking-about-thinking is confusing for programmers, and second, because models are slow to run. This paper presents memo , a new domain-specific probabilistic programming language that overcomes these challenges: first, by providing specialized syntax and semantics for theory of mind, and second, by taking a unique approach to inference that scales well on modern hardware via array programming. memo enables practitioners to write dramatically faster models with much less code, and has already been adopted by several research groups. Kartik Chandra, Tony Chen 0003, Josh Tenenbaum, Jonathan Ragan-Kelley |
Proc. ACM Program. Lang. | 3 |
| 2025 | Compositional Physical Reasoning of Objects and Events From VideosabstractUnderstanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and shapes can be directly observed, others, such as mass and electric charge, are hidden from the objects' visual appearance. This paper addresses the unique challenge of inferring these hidden physical properties from objects' motion and interactions and predicting corresponding dynamics based on the inferred physical properties. We first introduce the Compositional Physical Reasoning (ComPhy) dataset. For a given set of objects, ComPhy includes limited videos of them moving and interacting under different initial conditions. The model is evaluated based on its capability to unravel the compositional hidden properties, such as mass and charge, and use this knowledge to answer a set of questions. Besides the synthetic videos from simulators, we also collect a real-world dataset to show further test physical reasoning abilities of different models. We evaluate state-of-the-art video reasoning models on ComPhy and reveal their limited ability to capture these hidden properties, which leads to inferior performance. We also propose a novel neuro-symbolic framework, Physical Concept Reasoner (PCR), that learns and reasons about both visible and hidden physical properties from question answering. Leveraging an object-centric representation, PCR utilizes videos and the associated natural language to infer objects' physical properties without dense object annotations. Furthermore, It incorporates property-aware graph networks to approximate the dynamic interactions among objects. PCR also employs a semantic parser to convert questions into semantic programs, and a program executor to execute the programs based on the learned physical properties and dynamics. After training, PCR demonstrates remarkable capabilities. It can detect and associate objects across frames, ground visible and hidden physical properties, make future and counterfactual predictions, and utilize these extracted representations to answer challenging questions. We hope the proposed ComPhy dataset and the PCR model present a promising step towards more comprehensive physical reasoning in AI systems. Zhenfang Chen, Shilong Dong, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba 0001, Josh Tenenbaum, Chuang Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas |
Trans. Assoc. Comput. Linguistics | 19 |
| 2025 | Understanding Epistemic Language with a Language-augmented Bayesian Theory of MindabstractAbstract How do people understand and evaluate claims about others’ beliefs, even though these beliefs cannot be directly observed? In this paper, we introduce a cognitive model of epistemic language interpretation, grounded in Bayesian inferences about other agents’ goals, beliefs, and intentions: a language-augmented Bayesian theory-of-mind (LaBToM). By translating natural language into an epistemic “language-of-thought” with grammar-constrained LLM decoding, then evaluating these translations against the inferences produced by inverting a generative model of rational action and perception, LaBToM captures graded plausibility judgments of epistemic claims. We validate our model in an experiment where participants watch an agent navigate a maze to find keys hidden in boxes needed to reach their goal, then rate sentences about the agent’s beliefs. In contrast with multimodal LLMs (GPT-4o, Gemini Pro) and ablated models, our model correlates highly with human judgments for a wide range of expressions, including modal language, uncertainty expressions, knowledge claims, likelihood comparisons, and attributions of false belief. Lance Ying, Tan Zhi-Xuan, Lionel Wong, Vikash Mansinghka 0001, Josh Tenenbaum |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | Meschers: Geometry Processing of Impossible ObjectsabstractImpossible objects, geometric constructions that humans can perceive but that cannot exist in real life, have been a topic of intrigue in visual arts, perception, and graphics, yet no satisfying computer representation of such objects exists. Previous work embeds impossible objects in 3D, cutting them or twisting/bending them in the depth axis. Cutting an impossible object changes its local geometry at the cut, which can hamper downstream graphics applications, such as smoothing, while bending makes it difficult to relight the object. Both of these can invalidate geometry operations, such as distance computation. As an alternative, we introduce Meschers, meshes capable of representing impossible constructions akin to those found in M.C. Escher's woodcuts. Our representation has a theoretical foundation in discrete exterior calculus and supports the use-cases above, as we demonstrate in a number of example applications. Moreover, because we can do discrete geometry processing on our representation, we can inverse-render impossible objects. We also compare our representation to cut and bend representations of impossible objects. Ana Dodik, Isabella Yu, Kartik Chandra, Jonathan Ragan-Kelley, Josh Tenenbaum, Vincent Sitzmann, Justin Solomon 0001 |
ACM Trans. Graph. | 5 |
| 2024 | Neural Amortized Inference for Nested Multi-Agent ReasoningabstractMulti-agent interactions, such as communication, teaching, and bluffing, often rely on higher-order social inference, i.e., understanding how others infer oneself. Such intricate reasoning can be effectively modeled through nested multi-agent reasoning. Nonetheless, the computational complexity escalates exponentially with each level of reasoning, posing a significant challenge. However, humans effortlessly perform complex social inferences as part of their daily lives. To bridge the gap between human-like inference capabilities and computational limitations, we propose a novel approach: leveraging neural networks to amortize high-order social inference, thereby expediting nested multi-agent reasoning. We evaluate our method in two challenging multi-agent interaction domains. The experimental results demonstrate that our method is computationally efficient while exhibiting minimal degradation in accuracy. Kunal Jha, Tuan Anh Le 0001, Chuanyang Jin, Yen-Ling Kuo, Josh Tenenbaum, Tianmin Shu |
AAAI | 5 |
| 2024 | Generalized Planning in PDDL Domains with Pretrained Large Language ModelsabstractRecent work has considered whether large language models (LLMs) can function as planners: given a task, generate a plan. We investigate whether LLMs can serve as generalized planners: given a domain and training tasks, generate a program that efficiently produces plans for other tasks in the domain. In particular, we consider PDDL domains and use GPT-4 to synthesize Python programs. We also consider (1) Chain-of-Thought (CoT) summarization, where the LLM is prompted to summarize the domain and propose a strategy in words before synthesizing the program; and (2) automated debugging, where the program is validated with respect to the training tasks, and in case of errors, the LLM is re-prompted with four types of feedback. We evaluate this approach in seven PDDL domains and compare it to four ablations and four baselines. Overall, we find that GPT-4 is a surprisingly powerful generalized planner. We also conclude that automated debugging is very important, that CoT summarization has non-uniform impact, that GPT-4 is far superior to GPT-3.5, and that just two training tasks are often sufficient for strong generalization. Tom Silver, Soham Dan, Kavitha Srinivas, Josh Tenenbaum, Leslie Pack Kaelbling, Michael Katz 0001 |
AAAI | 4 |
| 2024 | MMToM-QA: Multimodal Theory of Mind Question AnsweringabstractChuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, Tianmin Shu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chuanyang Jin, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum, Tianmin Shu |
ACL (1) | 9 |
| 2024 | Intervening on Emotions by Planning Over a Theory of Mind
Tony Chen 0003, Sean Dae Houlihan, Kartik Chandra, Josh Tenenbaum, Rebecca Saxe |
CogSci | 4 |
| 2024 | Concept Learning as Coarse-to-Fine Probabilistic Program Induction
Maddy Bowers, Alexander K. Lew, Wenhao Qi, Joshua S. Rule, Vikash Mansinghka 0001, Josh Tenenbaum, Armando Solar-Lezama |
CogSci | 6 |
| 2024 | Cooperative Explanation as Rational Communication
Kartik Chandra, Tony Chen 0003, Tzu-Mao Li, Jonathan Ragan-Kelley, Josh Tenenbaum |
CogSci | 5 |
| 2024 | Naturalistic Transmission of Causal Knowledge between Machines and Humans
Cédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler, Noah D. Goodman, Jacob Andreas, Josh Tenenbaum |
CogSci | 7 |
| 2024 | Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
Gabriel Grand, Valerio Pepe, Jacob Andreas, Josh Tenenbaum |
CogSci | 4 |
| 2024 | Finding structure in logographic writing with library learning
Guangyuan Jiang, Matthias Hofer 0002, Jiayuan Mao, Lionel Wong, Josh Tenenbaum, Roger Levy |
CogSci | 5 |
| 2024 | Neuro-Symbolic Models of Human Moral Judgment
Joseph Kwon, Josh Tenenbaum, Sydney Levine |
CogSci | 2 |
| 2024 | Who is responsible for collective action?
Casey Lewry, Tania Lombrozo, Shannon Wing, Sydney Levine, Josh Tenenbaum, Lionel Wong, Sofia Bonicalzi, Tobias Gerstenberg |
CogSci | 5 |
| 2024 | Listener Knowledge Structures Commonsense Explanation
Yuka Machino, Ron Shprints, Max H. Siegel, Lionel Wong, Josh Tenenbaum |
CogSci | 5 |
| 2024 | Connecting the dots: a comparative and developmental analysis of spatiotemporal pattern learning
Tracey Mills, Nicole Coates, Alessandra Acadia Silva, Stephen Ferrigno, Laura Schulz, Josh Tenenbaum, Samuel J. Cheyette |
CogSci | 6 |
| 2024 | Knowing What Counts for Counting
Karla E. Perez, Max H. Siegel, Nicole Coates, Josh Tenenbaum, Laura Schulz |
CogSci | 4 |
| 2024 | Probabilistic simulation supports generalizable intuitive physics
Khaled Jedoui, Rahul M. V., Felix J. Binder, Josh Tenenbaum, Judith E. Fan, Dan Yamins, Kevin A. Smith 0001 |
CogSci | 5 |
| 2024 | Moral flexibility in applying queuing norms can be explained by contractualist principles and game-theoretic considerations
Joshua P. White, Rahul Bhui, Fiery Cushman, Josh Tenenbaum, Sydney Levine |
CogSci | 4 |
| 2024 | Grounding Language about Belief in a Bayesian Theory-of-Mind
Lance Ying, Tan Zhi-Xuan, Lionel Wong, Vikash Mansinghka 0001, Josh Tenenbaum |
CogSci | 5 |
| 2024 | People use fast, goal-directed simulation to reason about novel games
Cedegao E. Zhang, Katie Collins, Lionel Wong, Adrian Weller, Josh Tenenbaum |
CogSci | 5 |
| 2024 | Infinite Ends from Finite Samples: Open-Ended Goal Inference as Top-Down Bayesian Filtering of Bottom-Up Proposals
Tan Zhi-Xuan, Gloria Kang, Vikash Mansinghka 0001, Josh Tenenbaum |
CogSci | 4 |
| 2024 | Video Language PlanningabstractWe are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present video language planning (VLP), an algorithm that consists of a tree search procedure, where we train (i) vision-language models to serve as both policies and value functions, and (ii) text-to-video models as dynamics models. VLP takes as input a long-horizon task instruction and current image observation, and outputs a long video plan that provides detailed multimodal (video and language) specifications that describe how to complete the final task. VLP scales with increasing computation budget where more computation time results in improved video plans, and is able to synthesize long-horizon video plans across different robotics domains -- from multi-object rearrangement, to multi-camera bi-arm dexterous manipulation. Generated video plans can be translated into real robot actions via goal-conditioned policies, conditioned on each intermediate frame of the generated video. Experiments show that VLP substantially improves long-horizon task success rates compared to prior methods on both simulated and real robots (across 3 hardware platforms). Yilun Du, Sherry Yang 0001, Peter R. Florence, Fei Xia 0002, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Josh Tenenbaum, Leslie Pack Kaelbling, Andy Zeng 0001, Jonathan Tompson |
ICLR | 10 |
| 2024 | LILO: Learning Interpretable Libraries by Compressing and Documenting CodeabstractWhile large language models (LLMs) now excel at code generation, a key aspect of software development is the art of refactoring: consolidating code into libraries of reusable and readable programs. In this paper, we introduce LILO, a neurosymbolic framework that iteratively synthesizes, compresses, and documents code to build libraries tailored to particular problem domains. LILO combines LLM-guided program synthesis with recent algorithmic advances in automated refactoring from Stitch: a symbolic compression system that efficiently identifies optimal lambda abstractions across large code corpora. To make these abstractions interpretable, we introduce an auto-documentation (AutoDoc) procedure that infers natural language names and docstrings based on contextual examples of usage. In addition to improving human readability, we find that AutoDoc boosts performance by helping LILO's synthesizer to interpret and deploy learned abstractions. We evaluate LILO on three inductive program synthesis benchmarks for string editing, scene reasoning, and graphics composition. Compared to existing neural and symbolic methods—including the state-of-the-art library learning algorithm DreamCoder—LILO solves more complex tasks and learns richer libraries that are grounded in linguistic knowledge. Gabriel Grand, Lionel Wong, Matthew Bowers, Theo X. Olausson, Muxin Liu, Josh Tenenbaum, Jacob Andreas |
ICLR | 6 |
| 2024 | Learning to Act from Actionless Videos through Dense CorrespondencesabstractIn this work, we present an approach to construct a video-based robot policy capable of reliably executing diverse tasks across different robots and environments from few video demonstrations without using any action annotations. Our method leverages images as a task-agnostic representation, encoding both the state and action information, and text as a general representation for specifying robot goals. By synthesizing videos that "hallucinate" robot executing actions and in combination with dense correspondences between frames, our approach can infer the closed-formed action to execute to an environment without the need of any explicit action labels. This unique capability allows us to train the policy solely based on RGB videos and deploy learned policies to various robotic tasks. We demonstrate the efficacy of our approach in learning policies on table-top manipulation and navigation tasks. Additionally, we contribute an open-source framework for efficient video modeling, enabling the training of high-fidelity policy models with four GPUs within a single day. Po-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun, Josh Tenenbaum |
ICLR | 5 |
| 2024 | Learning to Jointly Understand Visual and Tactile SignalsabstractModeling and analyzing object and shape has been well studied in the past. However, manipulation of these complex tools and articulated objects remains difficult for autonomous agents. Our human hands, however, are dexterous and adaptive. We can easily adapt a manipulation skill on one object to all objects in the class and to other similar classes. Our intuition comes from that there is a close connection between manipulations and topology and articulation of objects. The possible articulation of objects indicates the types of manipulation necessary to operate the object. In this work, we aim to take a manipulation perspective to understand everyday objects and tools. We collect a multi-modal visual-tactile dataset that contains paired full-hand force pressure maps and manipulation videos. We also propose a novel method to learn a cross-modal latent manifold that allow for cross-modal prediction and discovery of latent structure in different data modalities. We conduct extensive experiments to demonstrate the effectiveness of our method. Yichen Li 0004, Yilun Du, Chao Liu 0021, Chao Liu 0064, Francis Williams, Michael Foshey, Benjamin Eckart, Jan Kautz, Josh Tenenbaum, Antonio Torralba 0001, Wojciech Matusik |
ICLR | 9 |
| 2024 | Learning Grounded Action Abstractions from LanguageabstractEffective planning in the real world requires not only world knowledge, but the ability to leverage that knowledge to build the right representation of the task at hand. Decades of hierarchical planning techniques have used domain-specific temporal action abstractions to support efficient and accurate planning, almost always relying on human priors and domain knowledge to decompose hard tasks into smaller subproblems appropriate for a goal or set of goals. This paper describes Ada (Action Domain Acquisition), a framework for automatically constructing task-specific planning representations using task-general background knowledge from language models (LMs). Starting with a general-purpose hierarchical planner and a low-level goal-conditioned policy, Ada interactively learns a library of planner-compatible high-level action abstractions and low-level controllers adapted to a particular domain of planning tasks. On two language-guided interactive planning benchmarks (Mini Minecraft and ALFRED Household Tasks), Ada strongly outperforms other approaches that use LMs for sequential decision-making, offering more accurate plans and better generalization to complex tasks. Lionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary S. Siegel, Jiahai Feng, Noa Korneev, Josh Tenenbaum, Jacob Andreas |
ICLR | 7 |
| 2024 | Probabilistic Adaptation of Black-Box Text-to-Video ModelsabstractLarge text-to-video models trained on internet-scale data have demonstrated exceptional capabilities in generating high-fidelity videos from arbitrary textual descriptions. However, similar to proprietary language models, large text-to-video models are often black boxes whose weight parameters are not publicly available, posing a significant challenge to adapting these models to specific domains such as robotics, animation, and personalized stylization. Inspired by how a large language model can be prompted to perform new tasks without access to the model weights, we investigate how to adapt a black-box pretrained text-to-video model to a variety of downstream domains without weight access to the pretrained model. In answering this question, we propose \emph{\methodname}, which leverages the score function of a large pretrained video diffusion model as a probabilistic prior to guide the generation of a task-specific small video model. Our experiments show that, by incorporating broad knowledge and fidelity of the pretrained model probabilistically, a small model with as few as 1.25% parameters of the pretrained model can generate high-quality yet domain-specific videos for a variety of downstream domains such as animation, egocentric modeling, and modeling of simulated and real-world robotics data. As large text-to-video models starting to become available as a service similar to large language models, we advocate for private institutions to expose scores of video diffusion models as outputs in addition to generated videos to allow flexible adaptation of large pretrained text-to-video models by the general public. Sherry Yang 0001, Yilun Du, Bo Dai 0001, Dale Schuurmans, Josh Tenenbaum, Pieter Abbeel |
ICLR | 5 |
| 2024 | Building Cooperative Embodied Agents Modularly with Large Language ModelsabstractIn this work, we address challenging multi-agent cooperation problems with decentralized control, raw sensory observations, costly communication, and multi-objective tasks instantiated in various embodied environments. While previous research either presupposes a cost-free communication channel or relies on a centralized controller with shared observations, we harness the commonsense knowledge, reasoning ability, language comprehension, and text generation prowess of LLMs and seamlessly incorporate them into a cognitive-inspired modular framework that integrates with perception, memory, and execution. Thus building a Cooperative Embodied Language Agent CoELA, who can plan, communicate, and cooperate with others to accomplish long-horizon tasks efficiently. Our experiments on C-WAH and TDW-MAT demonstrate that CoELA driven by GPT-4 can surpass strong planning-based methods and exhibit emergent effective communication. Though current Open LMs like LLAMA-2 still underperform, we fine-tune a CoELA with data collected with our agents and show how they can achieve promising performance. We also conducted a user study for human-agent interaction and discovered that CoELA communicating in natural language can earn more trust and cooperate more effectively with humans. Our research underscores the potential of LLMs for future research in multi-agent cooperation. Videos can be found on the project website https://vis-www.cs.umass.edu/Co-LLM-Agents/. Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Josh Tenenbaum, Tianmin Shu, Chuang Gan 0001 |
ICLR | 6 |
| 2024 | HAZARD Challenge: Embodied Decision Making in Dynamically Changing EnvironmentsabstractRecent advances in high-fidelity virtual environments serve as one of the major driving forces for building intelligent embodied agents to perceive, reason and interact with the physical world. Typically, these environments remain unchanged unless agents interact with them. However, in real-world scenarios, agents might also face dynamically changing environments characterized by unexpected events and need to rapidly take action accordingly. To remedy this gap, we propose a new simulated embodied benchmark, called HAZARD, specifically designed to assess the decision-making abilities of embodied agents in dynamic situations. HAZARD consists of three unexpected disaster scenarios, including fire, flood, and wind, and specifically supports the utilization of large language models (LLMs) to assist common sense reasoning and decision-making. This benchmark enables us to evaluate autonomous agents' decision-making capabilities across various pipelines, including reinforcement learning (RL), rule-based, and search-based methods in dynamically changing environments. As a first step toward addressing this challenge using large language models, we further develop an LLM-based agent and perform an in-depth analysis of its promise and challenge of solving these challenging tasks. HAZARD is available at https://vis-www.cs.umass.edu/hazard/. Qinhong Zhou, Sunli Chen, Haozhe Xu, Weihua Du, Yilun Du, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 8 |
| 2024 | Improving Factuality and Reasoning in Language Models through Multiagent DebateabstractLarge language models (LLMs) have demonstrated remarkable capabilities in language generation, understanding, and few-shot learning in recent years. An extensive body of work has explored how their performance may be further improved through the tools of prompting, ranging from verification, self-consistency, or intermediate scratchpads. In this paper, we present a complementary approach to improve language responses where multiple language model instances propose and debate their individual responses and reasoning processes over multiple rounds to arrive at a common final answer. Our findings indicate that this approach significantly enhances mathematical and strategic reasoning across a number of tasks. We also demonstrate that our approach improves the factual validity of generated content, reducing fallacious answers and hallucinations that contemporary models are prone to. Our approach may be directly applied to existing black-box models and uses identical procedure and prompts for all tasks we investigate. Overall, our findings suggest that such "society of minds" approach has the potential to significantly advance the capabilities of LLMs and pave the way for further breakthroughs in language generation and understanding. Yilun Du, Shuang Li 0013, Antonio Torralba 0001, Josh Tenenbaum, Igor Mordatch |
ICML | 4 |
| 2024 | Learning Iterative Reasoning through Energy DiffusionabstractWe introduce iterative reasoning through energy diffusion (IRED), a novel framework for learning to reason for a variety of tasks by formulating reasoning and decision-making problems with energy-based optimization. IRED learns energy functions to represent the constraints between input conditions and desired outputs. After training, IRED adapts the number of optimization steps during inference based on problem difficulty, enabling it to solve problems outside its training distribution --- such as more complex Sudoku puzzles, matrix completion with large value magnitudes, and path finding in larger graphs. Key to our method’s success is two novel techniques: learning a sequence of annealed energy landscapes for easier inference and a combination of score function and energy landscape supervision for faster and more stable training. Our experiments show that IRED outperforms existing methods in continuous-space reasoning, discrete-space reasoning, and planning tasks, particularly in more challenging scenarios. Yilun Du, Jiayuan Mao, Josh Tenenbaum |
ICML | 3 |
| 2024 | Potential Based Diffusion Motion PlanningabstractEffective motion planning in high dimensional spaces is a long-standing open problem in robotics. One class of traditional motion planning algorithms corresponds to potential-based motion planning. An advantage of potential based motion planning is composability – different motion constraints can easily combined by adding corresponding potentials. However, constructing motion paths from potentials requires solving a global optimization across configuration space potential landscape, which is often prone to local minima. We propose a new approach towards learning potential based motion planning, where we train a neural network to capture and learn an easily optimizable potentials over motion planning trajectories. We illustrate the effectiveness of such approach, significantly outperforming both classical and recent learned motion planning approaches and avoiding issues with local minima. We further illustrate its inherent composability, enabling us to generalize to a multitude of different motion constraints. Project website at https://energy-based-model.github.io/potential-motion-plan. Yunhao Luo 0001, Chen Sun 0002, Josh Tenenbaum, Yilun Du |
ICML | 3 |
| 2024 | LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific DiscoveryabstractLarge Language Models have recently gained significant attention in scientific discovery for their extensive knowledge and advanced reasoning capabilities. However, they encounter challenges in effectively simulating observational feedback and grounding it with language to propel advancements in physical scientific discovery. Conversely, human scientists undertake scientific discovery by formulating hypotheses, conducting experiments, and revising theories through observational analysis. Inspired by this, we propose to enhance the knowledge-driven, abstract reasoning abilities of LLMs with the computational strength of simulations. We introduce Scientific Generative Agent (SGA), a bilevel optimization framework: LLMs act as knowledgeable and versatile thinkers, proposing scientific hypotheses and reason about discrete components, such as physics equations or molecule structures; meanwhile, simulations function as experimental platforms, providing observational feedback and optimizing via differentiability for continuous parts, such as physical parameters. We conduct extensive experiments to demonstrate our framework’s efficacy in constitutive law discovery and molecular design, unveiling novel solutions that differ from conventional human expectations yet remain coherent upon analysis. Pingchuan Ma 0002, Tsun-Hsuan Wang, Zhiqing Sun, Josh Tenenbaum, Daniela Rus, Chuang Gan 0001, Wojciech Matusik |
ICML | 5 |
| 2024 | Compositional Image Decomposition with Diffusion ModelsabstractGiven an image of a natural scene, we are able to quickly decompose it into a set of components such as objects, lighting, shadows, and foreground. We can then envision a scene where we combine certain components with those from other images, for instance a set of objects from our bedroom and animals from a zoo under the lighting conditions of a forest, even if we have never encountered such a scene before. In this paper, we present a method to decompose an image into such compositional components. Our approach, Decomp Diffusion, is an unsupervised method which, when given a single image, infers a set of different components in the image, each represented by a diffusion model. We demonstrate how components can capture different factors of the scene, ranging from global scene descriptors like shadows or facial expression to local scene descriptors like constituent objects. We further illustrate how inferred factors can be flexibly composed, even with factors inferred from other models, to generate a variety of scenes sharply different than those seen in training time. Code and visualizations are at https://energy-based-model.github.io/decomp-diffusion. Jocelin Su, Nan Liu 0010, Josh Tenenbaum, Yilun Du |
ICML | 4 |
| 2024 | ContPhy: Continuum Physical Concept Learning and Reasoning from VideosabstractWe introduce the Continuum Physical Dataset (ContPhy), a novel benchmark for assessing machine physical commonsense. ContPhy complements existing physical reasoning benchmarks by encompassing the inference of diverse physical properties, such as mass and density, across various scenarios and predicting corresponding dynamics. We evaluated a range of AI models and found that they still struggle to achieve satisfactory performance on ContPhy, which shows that current AI models still lack physical commonsense for the continuum, especially soft-bodies, and illustrates the value of the proposed dataset. We also introduce an oracle model (ContPRO) that marries the particle-based physical dynamic models with the recent large language models, which enjoy the advantages of both models, precise dynamic predictions, and interpretable reasoning. ContPhy aims to spur progress in perception and reasoning within diverse physical settings, narrowing the divide between human and machine intelligence in understanding the physical world. Zhicheng Zheng, Xin Yan 0008, Zhenfang Chen, Jingzhou Wang, Qin Zhi Eddie Lim, Josh Tenenbaum, Chuang Gan 0001 |
ICML | 6 |
| 2024 | ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and PlanningabstractFor robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D representations. However, these approaches tend to produce maps with per-point feature vectors, which do not scale well in larger environments, nor do they contain semantic spatial relationships between entities in the environment, which are useful for downstream planning. In this work, we propose ConceptGraphs, an open-vocabulary graph-structured representation for 3D scenes. ConceptGraphs is built by leveraging 2D foundation models and fusing their output to 3D by multi-view association. The resulting representations generalize to novel semantic classes, without the need to collect large 3D datasets or finetune models. We demonstrate the utility of this representation through a number of downstream planning tasks that are specified through abstract (language) prompts and require complex reasoning over spatial and semantic concepts. To explore the full scope of our experiments and results, we encourage readers to visit our project webpage. Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, Chuang Gan 0001, Celso de Melo, Josh Tenenbaum, Antonio Torralba 0001, Florian Shkurti, Liam Paull |
ICRA | 13 |
| 2024 | Tactile Estimation of Extrinsic Contact Patch for Stable PlacementabstractPrecise perception of contact interactions is essential for fine-grained manipulation skills for robots. In this paper, we present the design of feedback skills for robots that must learn to stack complex-shaped objects on top of each other (see Fig. 1). To design such a system, a robot should be able to reason about the stability of placement from very gentle contact interactions. Our results demonstrate that it is possible to infer the stability of object placement based on tactile readings during contact formation between the object and its environment. In particular, we estimate the contact patch between a grasped object and its environment using force and tactile observations to estimate the stability of the object during a contact formation. The contact patch could be used to estimate the stability of the object upon release of the grasp. The proposed method is demonstrated in various pairs of objects that are used in a very popular board game. Kei Ota, Devesh K. Jha, Krishna Murthy Jatavallabhula, Asako Kanezaki, Josh Tenenbaum |
ICRA | 5 |
| 2024 | PickScan: Object discovery and reconstruction from handheld interactionsabstractReconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong appearance priors for object discovery, therefore only working on those classes of objects on which the method has been trained, or do not allow for object manipulation, which is necessary to scan objects fully and to guide object discovery in challenging scenarios. We address these limitations with a novel interaction-guided and class-agnostic method based on object displacements that allows a user to move around a scene with an RGB-D camera, hold up objects, and finally outputs one 3D model per held-up object. Our main contribution to this end is a novel approach to detecting user-object interactions and extracting the masks of manipulated objects. On a custom-captured dataset, our pipeline discovers manipulated objects with 78.3% precision at 100% recall and reconstructs them with a mean chamfer distance of 0.90cm. Compared to Co-Fusion, the only comparable interaction-based and class-agnostic baseline, this corresponds to a reduction in chamfer distance of 73% while detecting 99% fewer false positives. Vincent van der Brugge, Marc Pollefeys, Josh Tenenbaum, Krishna Murthy Jatavallabhula, Ayush Tewari |
IROS | 3 |
| 2024 | GOMA: Proactive Embodied Cooperative Communication via Goal-Oriented Mental AlignmentabstractVerbal communication plays a crucial role in human cooperation, particularly when the partners only have incomplete information about the task, environment, and each other’s mental state. In this paper, we propose a novel cooperative communication framework, Goal-Oriented Mental Alignment (GOMA). GOMA formulates verbal communication as a planning problem that minimizes the misalignment between the parts of agents’ mental states that are relevant to the goals. This approach enables an embodied assistant to reason about when and how to proactively initialize communication with humans verbally using natural language to help achieve better cooperation. We evaluate our approach against strong baselines in two challenging environments, Overcooked (a multiplayer game) and VirtualHome (a household simulator). Our experimental results demonstrate that large language models struggle with generating meaningful communication that is grounded in the social and physical context. In contrast, our approach can successfully generate concise verbal communication for the embodied assistant to effectively boost the performance of the cooperation as well as human users’ perception of the assistant. Lance Ying, Kunal Jha, Shivam Aarya, Josh Tenenbaum, Antonio Torralba 0001, Tianmin Shu |
IROS | 4 |
| 2024 | Evaluating Multiview Object Consistency in Humans and Image ModelsabstractWe introduce a benchmark to directly evaluate the alignment between human observers and vision models on a 3D shape inference task. We leverage an experimental design from the cognitive sciences: given a set of images, participants identify which contain the same/different objects, despite considerable viewpoint variation. We draw from a diverse range of images that include common objects (e.g., chairs) as well as abstract shapes (i.e., procedurally generated 'nonsense' objects). After constructing over 2000 unique image sets, we administer these tasks to human participants, collecting 35K trials of behavioral data from over 500 participants. This includes explicit choice behaviors as well as intermediate measures, such as reaction time and gaze data. We then evaluate the performance of common vision models (e.g., DINOv2, MAE, CLIP). We find that humans outperform all models by a wide margin. Using a multi-scale evaluation approach, we identify underlying similarities and differences between models and humans: while human-model performance is correlated, humans allocate more time/processing on challenging trials. All images, data, and code can be accessed via our project page. Tyler Bonnen, Stephanie Fu, Yutong Bai, Thomas P. O'Connell, Yoni Friedman, Nancy Kanwisher, Josh Tenenbaum, Alexei A. Efros |
NeurIPS | 7 |
| 2024 | Evaluating Large Vision-and-Language Models on Children's Mathematical OlympiadsabstractRecent years have seen a significant progress in the general-purpose problem solving abilities of large vision and language models (LVLMs), such as ChatGPT, Gemini, etc.; some of these breakthroughs even seem to enable AI models to outperform human abilities in varied tasks that demand higher-order cognitive skills. Are the current large AI models indeed capable of generalized problem solving as humans do? A systematic analysis of AI capabilities for joint vision and text reasoning, however, is missing in the current scientific literature. In this paper, we make an effort towards filling this gap, by evaluating state-of-the-art LVLMs on their mathematical and algorithmic reasoning abilities using visuo-linguistic problems from children's Olympiads. Specifically, we consider problems from the Mathematical Kangaroo (MK) Olympiad, which is a popular international competition targeted at children from grades 1-12, that tests children's deeper mathematical abilities using puzzles that are appropriately gauged to their age and skills. Using the puzzles from MK, we created a dataset, dubbed SMART-840, consisting of 840 problems from years 2020-2024. With our dataset, we analyze LVLMs power on mathematical reasoning; their responses on our puzzles offer a direct way to compare against that of children. Our results show that modern LVLMs do demonstrate increasingly powerful reasoning skills in solving problems for higher grades, but lack the foundations to correctly answer problems designed for younger children. Further analysis shows that there is no significant correlation between the reasoning capabilities of AI models and that of young children, and their capabilities appear to be based on a different type of reasoning than the cumulative knowledge that underlies children's mathematical skills. Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Joanna Matthiesen, Kevin A. Smith 0001, Josh Tenenbaum |
NeurIPS | 6 |
| 2024 | Physically Compatible 3D Object Modeling from a Single ImageabstractWe present a computational framework that transforms single images into 3D physical objects. The visual geometry of a physical object in an image is determined by three orthogonal attributes: mechanical properties, external forces, and rest-shape geometry. Existing single-view 3D reconstruction methods often overlook this underlying composition, presuming rigidity or neglecting external forces. Consequently, the reconstructed objects fail to withstand real-world physical forces, resulting in instability or undesirable deformation -- diverging from their intended designs as depicted in the image. Our optimization framework addresses this by embedding physical compatibility into the reconstruction process. We explicitly decompose the three physical attributes and link them through static equilibrium, which serves as a hard constraint, ensuring that the optimized physical shapes exhibit desired physical behaviors. Evaluations on a dataset collected from Objaverse demonstrate that our framework consistently enhances the physical realism of 3D models over existing methods. The utility of our framework extends to practical applications in dynamic simulations and 3D printing, where adherence to physical compatibility is paramount. Pingchuan Ma 0002, Crystal Elaine Owens, Chuang Gan 0001, Josh Tenenbaum, Kaiming He, Wojciech Matusik |
NeurIPS | 7 |
| 2024 | Few-Shot Task Learning through Inverse Generative ModelingabstractLearning the intents of an agent, defined by its goals or motion style, is often extremely challenging from just a few examples. We refer to this problem as task concept learning and present our approach, Few-Shot Task Learning through Inverse Generative Modeling (FTL-IGM), which learns new task concepts by leveraging invertible neural generative models. The core idea is to pretrain a generative model on a set of basic concepts and their demonstrations. Then, given a few demonstrations of a new concept (such as a new goal or a new action), our method learns the underlying concepts through backpropagation without updating the model weights, thanks to the invertibility of the generative model. We evaluate our method in five domains -- object rearrangement, goal-oriented navigation, motion caption of human actions, autonomous driving, and real-world table-top manipulation. Our experimental results demonstrate that via the pretrained generative model, we successfully learn novel concepts and generate agent plans or motion corresponding to these concepts in (1) unseen environments and (2) in composition with training concepts. Aviv Netanyahu, Yilun Du, Antonia Bronars, Jyothish Pari, Josh Tenenbaum, Tianmin Shu, Pulkit Agrawal 0001 |
NeurIPS | 5 |
| 2024 | Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal ImitationabstractSA Conference Papers ’24, December 03–06, 2024, Tokyo, Japan Matthew Caren, Kartik Chandra, Josh Tenenbaum, Jonathan Ragan-Kelley, Karima Ma |
SIGGRAPH Asia | 3 |
| 2024 | When is it acceptable to break the rules? Knowledge representation of moral judgements based on empirical dataabstractAbstract Constraining the actions of AI systems is one promising way to ensure that these systems behave in a way that is morally acceptable to humans. But constraints alone come with drawbacks as in many AI systems, they are not flexible. If these constraints are too rigid, they can preclude actions that are actually acceptable in certain, contextual situations. Humans, on the other hand, can often decide when a simple and seemingly inflexible rule should actually be overridden based on the context. In this paper, we empirically investigate the way humans make these contextual moral judgements, with the goal of building AI systems that understand when to follow and when to override constraints. We propose a novel and general preference-based graphical model that captures a modification of standard dual process theories of moral judgment. We then detail the design, implementation, and results of a study of human participants who judge whether it is acceptable to break a well-established rule: no cutting in line. We then develop an instance of our model and compare its performance to that of standard machine learning approaches on the task of predicting the behavior of human participants in the study, showing that our preference-based approach more accurately captures the judgments of human decision-makers. It also provides a flexible method to model the relationship between variables for moral decision-making tasks that can be generalized to other settings. Edmond Awad, Sydney Levine, Andrea Loreggia, Nicholas Mattei, Iyad Rahwan, Francesca Rossi 0001, Kartik Talamadupula, Josh Tenenbaum, Max Kleiman-Weiner |
Auton. Agents Multi Agent Syst. | 8 |
| 2024 | Building 3D Generative Models from Minimal DataabstractWe propose a method for constructing generative models of 3D objects from a single 3D mesh and improving them through unsupervised low-shot learning from 2D images. Our method produces a 3D morphable model that represents shape and albedo in terms of Gaussian processes. Whereas previous approaches have typically built 3D morphable models from multiple high-quality 3D scans through principal component analysis, we build 3D morphable models from a single scan or template. As we demonstrate in the face domain, these models can be used to infer 3D reconstructions from 2D data (inverse graphics) or 3D data (registration). Specifically, we show that our approach can be used to perform face recognition using only a single 3D template (one scan total, not one per person). We extend our model to a preliminary unsupervised learning framework that enables the learning of the distribution of 3D faces using one 3D template and a small number of 2D images. Our approach is motivated as a potential model for the origins of face perception in human infants, who appear to start with an innate face template and subsequently develop a flexible system for perceiving the 3D structure of any novel face from experience with only 2D images of a relatively small number of familiar faces. Skylar Sutherland, Bernhard Egger 0001, Josh Tenenbaum |
Int. J. Comput. Vis. | 3 |
| 2024 | Approximate planning in spatial searchabstractHow people plan is an active area of research in cognitive science, neuroscience, and artificial intelligence. However, tasks traditionally used to study planning in the laboratory tend to be constrained to artificial environments, such as Chess and bandit problems. To date there is still no agreed-on model of how people plan in realistic contexts, such as navigation and search, where values intuitively derive from interactions between perception and cognition. To address this gap and move towards a more naturalistic study of planning, we present a novel spatial Maze Search Task (MST) where the costs and rewards are physically situated as distances and locations. We used this task in two behavioral experiments to evaluate and contrast multiple distinct computational models of planning, including optimal expected utility planning, several one-step heuristics inspired by studies of information search, and a family of planners that deviate from optimal planning, in which action values are estimated by the interactions between perception and cognition. We found that people's deviations from optimal expected utility are best explained by planners with a limited horizon, however our results do not exclude the possibility that in human planning action values may be also affected by cognitive mechanisms of numerosity and probability perception. This result makes a novel theoretical contribution in showing that limited planning horizon generalizes to spatial planning, and demonstrates the value of our multi-model approach for understanding cognition. Marta Kryven, Suhyoun Yu, Max Kleiman-Weiner, Tomer D. Ullman, Josh Tenenbaum |
PLoS Comput. Biol. | 5 |
| 2023 | Learning Rational Subgoals from Demonstrations and InstructionsabstractWe present a framework for learning useful subgoals that support efficient long-term planning to achieve novel goals. At the core of our framework is a collection of rational subgoals (RSGs), which are essentially binary classifiers over the environmental states. RSGs can be learned from weakly-annotated data, in the form of unsegmented demonstration trajectories, paired with abstract task descriptions, which are composed of terms initially unknown to the agent (e.g., collect-wood then craft-boat then go-across-river). Our framework also discovers dependencies between RSGs, e.g., the task collect-wood is a helpful subgoal for the task craft-boat. Given a goal description, the learned subgoals and the derived dependencies facilitate off-the-shelf planning algorithms, such as A* and RRT, by setting helpful subgoals as waypoints to the planner, which significantly improves performance-time efficiency. Project page: https://rsg.csail.mit.edu Zhezheng Luo, Jiayuan Mao, Jiajun Wu 0001, Tomás Lozano-Pérez, Josh Tenenbaum, Leslie Pack Kaelbling |
AAAI | 5 |
| 2023 | Predicate Invention for Bilevel PlanningabstractEfficient planning in continuous state and action spaces is fundamentally hard, even when the transition model is deterministic and known. One way to alleviate this challenge is to perform bilevel planning with abstractions, where a high-level search for abstract plans is used to guide planning in the original transition space. Previous work has shown that when state abstractions in the form of symbolic predicates are hand-designed, operators and samplers for bilevel planning can be learned from demonstrations. In this work, we propose an algorithm for learning predicates from demonstrations, eliminating the need for manually specified state abstractions. Our key idea is to learn predicates by optimizing a surrogate objective that is tractable but faithful to our real efficient-planning objective. We use this surrogate objective in a hill-climbing search over predicate sets drawn from a grammar. Experimentally, we show across four robotic planning environments that our learned abstractions are able to quickly solve held-out tasks, outperforming six baselines. Tom Silver, Rohan Chitnis, Nishanth Kumar, Willie McClinton, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Josh Tenenbaum |
AAAI | 7 |
| 2023 | Zero-Shot Linear Combinations of Grounded Social Interactions with Linear Social MDPsabstractHumans and animals engage in rich social interactions. It is often theorized that a relatively small number of basic social interactions give rise to the full range of behavior observed. But no computational theory explaining how social interactions combine together has been proposed before. We do so here. We take a model, the Social MDP, which is able to express a range of social interactions, and extend it to represent linear combinations of social interactions. Practically for robotics applications, such models are now able to not just express that an agent should help another agent, but to express goal-centric social interactions. Perhaps an agent is helping someone get dressed, but preventing them from falling, and is happy to exchange stories in the meantime. How an agent responds socially, should depend on what it thinks the other agent is doing at that point in time. To encode this notion, we take linear combinations of social interactions as defined in Social MDPs, and compute the weights on those combinations on the fly depending on the estimated goals of other agents. This new model, the Linear Social MDP, enables zero-shot reasoning about complex social interactions, provides a mathematical basis for the long-standing intuition that social interactions should compose, and leads to interesting new behaviors that we validate using human observers. Complex social interactions are part of the future of intelligent agents, and having principled mathematical models built on a foundation like MDPs will make it possible to bring social interactions to every robotic application. Ravi Tejwani, Yen-Ling Kuo, Tianmin Shu, Bennett Stankovits, Dan Gutfreund, Josh Tenenbaum, Boris Katz, Andrei Barbu |
AAAI | 6 |
| 2023 | Storytelling as Inverse Inverse Planning
Kartik Chandra, Tzu-Mao Li, Josh Tenenbaum, Jonathan Ragan-Kelley |
CogSci | 3 |
| 2023 | "Just In Time" Representations for Mental Simulation in Intuitive Physics
Tony Chen 0003, Kelsey R. Allen, Samuel J. Cheyette, Josh Tenenbaum, Kevin A. Smith 0001 |
CogSci | 4 |
| 2023 | People seek easily interpretable information
Samuel J. Cheyette, Frederick Callaway, Neil Bramley, Jonathan D. Nelson, Josh Tenenbaum |
CogSci | 5 |
| 2023 | Representations of Abstract Relations in Early Childhood
Nicole Coates, Max H. Siegel, Josh Tenenbaum, Laura Schulz |
CogSci | 3 |
| 2023 | Perception of Mooney Faces: Extreme Generalization through Inverse Rendering?
Shreya Kapoor, Maximilian Weiherer, Max H. Siegel, Amir Arsalan Soltani, Ilker Yildirim, Josh Tenenbaum, Bernhard Egger 0001 |
CogSci | 6 |
| 2023 | When it's not out of line to get out of line: Principles of universalizability, welfare, and harm
Joseph Kwon, Tan Zhi-Xuan, Josh Tenenbaum, Sydney Levine |
CogSci | 3 |
| 2023 | Evaluating statistical language models as pragmatic reasoners
Benjamin Lipkin, Lionel Wong, Gabriel Grand, Josh Tenenbaum |
CogSci | 4 |
| 2023 | Towards a model of confidence judgements in concept learning
Tracey Mills, Cedegao E. Zhang, Tony Chen 0003, Josh Tenenbaum |
CogSci | 4 |
| 2023 | Strategy choice for physical reasoning is (partially) sensitive to cognitive costs
Thomas Ngo, Samuel J. Cheyette, Josh Tenenbaum, Kevin A. Smith 0001 |
CogSci | 3 |
| 2023 | Grounded physical language understanding with probabilistic programs and simulated worlds
Cedegao E. Zhang, Lionel Wong, Gabriel Grand, Josh Tenenbaum |
CogSci | 4 |
| 2023 | Language Models as Informative Goal Priors in a Bayesian Theory of Mind
Tan Zhi-Xuan, Paul Stefan Lunis, Nathalie Fernandez Echeverri, Vikash Mansinghka 0001, Josh Tenenbaum |
CogSci | 5 |
| 2023 | Are Deep Neural Networks SMARTer Than Second Graders?abstractRecent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, question answering (e.g., ChatGPT), etc. Such a dramatic progress raises the question: how generalizable are neural networks in solving problems that demand broad skills? To answer this question, we propose SMART: a Simple Multimodal Algorithmic Reasoning Task and the associated SMART-101 dataset11The SMART-101 dataset is publicly available at: https://doi.org/10.5281/zenodo.7761800, for evaluating the abstraction, deduction, and generalization abilities of neural networks in solving visuo-linguistic puzzles designed specifically for children in the 6–8 age group. Our dataset consists of 101 unique puzzles; each puzzle comprises a picture and a question, and their solution needs a mix of several elementary skills, including arithmetic, algebra, and spatial reasoning, among others. To scale our dataset towards training deep neural networks, we programmatically generate entirely new instances for each puzzle while retaining their solution algorithm. To benchmark the performance on the SMART-101 dataset, we propose a vision-and-language meta-learning model that can incorporate varied state-of-the-art neural backbones. Our experiments reveal that while powerful deep models offer reasonable performances on puzzles in a supervised setting, they are not better than random accuracy when analyzed for generalization –filling this gap may demand new multimodal learning approaches. Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith 0001, Josh Tenenbaum |
CVPR | 5 |
| 2023 | Visual Dependency Transformers: Dependency Tree Emerges from Reversed AttentionabstractHumans possess a versatile mechanism for extracting structured representations of our visual world. When looking at an image, we can decompose the scene into entities and their parts as well as obtain the dependencies between them. To mimic such capability, we propose Visual Dependency Transformers (DependencyViT)11https://github.com/dingmyu/DependencyViT that can induce visual dependencies without any labels. We achieve that with a novel neural operator called reversed attention that can naturally capture long-range visual dependencies between image patches. Specifically, we formulate it as a dependency graph where a child token in reversed attention is trained to attend to its parent tokens and send information following a normalized probability distribution rather than gathering information in conventional self-attention. With such a design, hierarchies naturally emerge from reversed attention layers, and a dependency tree is progressively induced from leaf nodes to the root node unsupervisedly. DependencyViT offers several appealing benefits. (i) Entities and their parts in an image are represented by different subtrees, enabling part partitioning from dependencies; (ii) Dynamic visual pooling is made possible. The leaf nodes which rarely send messages can be pruned without hindering the model performance, based on which we propose the lightweight DependencyViT-Lite to reduce the computational and memory footprints; (iii) DependencyViT works well on both self- and weakly-supervised pretraining paradigms on ImageNet, and demonstrates its effectiveness on 8 datasets and 5 tasks, such as unsupervised part and saliency segmentation, recognition, and detection. Mingyu Ding, Yikang Shen, Lijie Fan, Zhenfang Chen, Zitian Chen, Ping Luo 0002, Josh Tenenbaum, Chuang Gan 0001 |
CVPR | 7 |
| 2023 | 3D Concept Learning and Reasoning from Multi-View ImagesabstractHumans are able to accurately reason in 3D by gathering multi-view observations of the surrounding world. Inspired by this insight, we introduce a new large-scale benchmark for 3D multi-view visual question answering (3DMV-VQA). This dataset is collected by an embodied agent actively moving and capturing RGB images in an environment using the Habitat simulator. In total, it consists of approximately 5k scenes, 600k images, paired with 50k questions. We evaluate various state-of-the-art models for visual reasoning on our benchmark and find that they all perform poorly. We suggest that a principled approach for 3D reasoning from multi-view images should be to infer a compact 3D representation of the world from the multi-view images, which is further grounded on open-vocabulary semantic concepts, and then to execute reasoning on these 3D representations. As the first step towards this approach, we propose a novel 3D concept learning and reasoning (3D-CLR) framework that seamlessly combines these components via neural fields, 2D pre-trained vision-language models, and neural reasoning operators. Experimental results suggest that our framework outperforms baseline models by a large margin, but the challenge remains largely unsolved. We further perform an in-depth analysis of the challenges and highlight potential future directions.. Yining Hong, Chunru Lin, Yilun Du, Zhenfang Chen, Josh Tenenbaum, Chuang Gan 0001 |
CVPR | 5 |
| 2023 | LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversabstractLogical reasoning, i.e., deductively inferring the truth value of a conclusion from a set of premises, is an important task for artificial intelligence with wide potential impacts on science, mathematics, and society.While many prompting-based strategies have been proposed to enable Large Language Models (LLMs) to do such reasoning more effectively, they still appear unsatisfactory, often failing in subtle and unpredictable ways.In this work, we investigate the validity of instead reformulating such tasks as modular neurosymbolic programming, which we call LINC: Logical Inference via Neurosymbolic Computation.In LINC, the LLM acts as a semantic parser, translating premises and conclusions from natural language to expressions in first-order logic.These expressions are then offloaded to an external theorem prover, which symbolically performs deductive inference.Leveraging this approach, we observe significant performance gains on FOLIO and a balanced subset of ProofWriter for three different models in nearly all experimental conditions we evaluate.On ProofWriter, augmenting the comparatively small open-source StarCoder+ (15.5B parameters) with LINC even outperforms GPT-3.5 and GPT-4 with Chain-of-Thought (CoT) prompting by an absolute 38% and 10%, respectively.When used with GPT-4, LINC scores 26% higher than CoT on ProofWriter while performing comparatively on FOLIO.Further analysis reveals that although both methods on average succeed roughly equally often on this dataset, they exhibit distinct and complementary failure modes.We thus provide promising evidence for how logical reasoning over natural language can be tackled through jointly leveraging LLMs alongside symbolic provers.All corresponding code is publicly available. Theo X. Olausson, Alex Gu, Benjamin Lipkin, Cedegao E. Zhang, Armando Solar-Lezama, Josh Tenenbaum, Roger Levy |
EMNLP | 6 |
| 2023 | Unsupervised Compositional Concepts Discovery with Text-to-Image Generative ModelsabstractText-to-image generative models have enabled high-resolution image synthesis across different domains, but require users to specify the content they wish to generate. In this paper, we consider the inverse problem – given a collection of different images, can we discover the generative concepts that represent each image? We present an unsupervised approach to discover generative concepts from a collection of images, disentangling different art styles in paintings, objects, and lighting from kitchen scenes, and discovering image classes given ImageNet images. We show how such generative concepts can accurately represent the content of images, be recombined and composed to generate new artistic and hybrid images, and be further used as a representation for downstream classification tasks. Nan Liu 0010, Yilun Du, Shuang Li 0013, Josh Tenenbaum, Antonio Torralba 0001 |
ICCV | 4 |
| 2023 | 3D Neural Embedding Likelihood: Probabilistic Inverse Graphics for Robust 6D Pose EstimationabstractThe ability to perceive and understand 3D scenes is crucial for many applications in computer vision and robotics. Inverse graphics is an appealing approach to 3D scene understanding that aims to infer the 3D scene structure from 2D images. In this paper, we introduce probabilistic modeling to the inverse graphics framework to quantify uncertainty and achieve robustness in 6D pose estimation tasks. Specifically, we propose 3D Neural Embedding Likelihood (3DNEL) as a unified probabilistic model over RGBD images, and develop efficient inference procedures on 3D scene descriptions. 3DNEL effectively combines learned neural embeddings from RGB with depth information to improve robustness in sim-to-real 6D object pose estimation from RGB-D images. Performance on the YCB-Video dataset is on par with state-of-the-art yet is much more robust in challenging regimes. In contrast to discriminative approaches, 3DNEL’s probabilistic generative formulation jointly models multiple objects in a scene, quantifies uncertainty in a principled way, and handles object pose tracking under heavy occlusion. Finally, 3DNEL provides a principled framework for incorporating prior knowledge about the scene and objects, which allows natural extension to additional tasks like camera pose tracking from video. Nishad Gothoskar, Lirui Wang, Josh Tenenbaum, Dan Gutfreund, Miguel Lázaro-Gredilla, Dileep George, Vikash Mansinghka 0001 |
ICCV | 4 |
| 2023 | Is Conditional Generative Modeling all you need for Decision Making?
Anurag Ajay, Yilun Du, Abhi Gupta, Josh Tenenbaum, Tommi S. Jaakkola, Pulkit Agrawal 0001 |
ICLR | 4 |
| 2023 | Planning with Sequence Models through Iterative Energy Minimization
Yilun Du, Yiye Chen, Josh Tenenbaum, Patricio A. Vela |
ICLR | 4 |
| 2023 | Composing Ensembles of Pre-trained Models via Iterative Consensus
Shuang Li 0013, Yilun Du, Josh Tenenbaum, Antonio Torralba 0001, Igor Mordatch |
ICLR | 3 |
| 2023 | DexDeform: Dexterous Deformable Object Manipulation with Human Demonstrations and Differentiable Physics
Zhiao Huang, Tao Chen 0046, Tao Du 0001, Hao Su 0001, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 6 |
| 2023 | Neural Groundplans: Persistent Neural Scene Representations from a Single Image
Prafull Sharma, Ayush Tewari, Yilun Du, Sergey Zakharov, Rares Ambrus, Adrien Gaidon, William T. Freeman, Frédo Durand, Josh Tenenbaum, Vincent Sitzmann |
ICLR | 9 |
| 2023 | SoftZoo: A Soft Robot Co-design Benchmark For Locomotion In Diverse Environments
Tsun-Hsuan Wang, Pingchuan Ma 0002, Andrew Spielberg, Zhou Xian, Josh Tenenbaum, Daniela Rus, Chuang Gan 0001 |
ICLR | 6 |
| 2023 | Planning with Large Language Models for Code Generation
Zhenfang Chen, Yikang Shen, Mingyu Ding, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 5 |
| 2023 | Learning Neural Constitutive Laws from Motion Observations for Generalizable PDE DynamicsabstractWe propose a hybrid neural network (NN) and PDE approach for learning generalizable PDE dynamics from motion observations. Many NN approaches learn an end-to-end model that implicitly models both the governing PDE and constitutive models (or material models). Without explicit PDE knowledge, these approaches cannot guarantee physical correctness and have limited generalizability. We argue that the governing PDEs are often well-known and should be explicitly enforced rather than learned. Instead, constitutive models are particularly suitable for learning due to their data-fitting nature. To this end, we introduce a new framework termed "Neural Constitutive Laws" (NCLaw), which utilizes a network architecture that strictly guarantees standard constitutive priors, including rotation equivariance and undeformed state equilibrium. We embed this network inside a differentiable simulation and train the model by minimizing a loss function based on the difference between the simulation and the motion observation. We validate NCLaw on various large-deformation dynamical systems, ranging from solids to fluids. After training on a single motion trajectory, our method generalizes to new geometries, initial/boundary conditions, temporal ranges, and even multi-physics systems. On these extremely out-of-distribution generalization tasks, NCLaw is orders-of-magnitude more accurate than previous NN approaches. Real-world experiments demonstrate our method’s ability to learn constitutive laws from videos. Pingchuan Ma 0002, Peter Yichen Chen, Bolei Deng, Josh Tenenbaum, Tao Du 0001, Chuang Gan 0001, Wojciech Matusik |
ICML | 4 |
| 2023 | Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCabstractSince their introduction, diffusion models have quickly become the prevailing approach to generative modeling in many domains. They can be interpreted as learning the gradients of a time-varying sequence of log-probability density functions. This interpretation has motivated classifier-based and classifier-free guidance as methods for post-hoc control of diffusion models. In this work, we build upon these ideas using the score-based interpretation of diffusion models, and explore alternative ways to condition, modify, and reuse diffusion models for tasks involving compositional generation and guidance. In particular, we investigate why certain types of composition fail using current techniques and present a number of solutions. We conclude that the sampler (not the model) is responsible for this failure and propose new samplers, inspired by MCMC, which enable successful compositional generation. Further, we propose an energy-based parameterization of diffusion models which enables the use of new compositional operators and more sophisticated, Metropolis-corrected samplers. Intriguingly we find these samplers lead to notable improvements in compositional generation across a wide variety of problems such as classifier-guided ImageNet modeling and compositional text-to-image generation. Yilun Du, Conor Durkan, Robin Strudel, Josh Tenenbaum, Sander Dieleman, Rob Fergus, Jascha Sohl-Dickstein, Arnaud Doucet, Will Grathwohl |
ICML | 4 |
| 2023 | Inferring Relational Potentials in Interacting SystemsabstractSystems consisting of interacting agents are prevalent in the world, ranging from dynamical systems in physics to complex biological networks. To build systems which can interact robustly in the real world, it is thus important to be able to infer the precise interactions governing such systems. Existing approaches typically discover such interactions by explicitly modeling the feed-forward dynamics of the trajectories. In this work, we propose Neural Interaction Inference with Potentials (NIIP) as an alternative approach to discover such interactions that enables greater flexibility in trajectory modeling: it discovers a set of relational potentials, represented as energy functions, which when minimized reconstruct the original trajectory. NIIP assigns low energy to the subset of trajectories which respect the relational constraints observed. We illustrate that with these representations NIIP displays unique capabilities in test-time. First, it allows trajectory manipulation, such as interchanging interaction types across separately trained models, as well as trajectory forecasting. Additionally, it allows adding external hand-crafted potentials at test-time. Finally, NIIP enables the detection of out-of-distribution samples and anomalies without explicit training. Armand Comas Massague, Yilun Du, Christian Fernandez Lopez, Sandesh Ghimire, Mario Sznaier, Josh Tenenbaum, Octavia I. Camps |
ICML | 6 |
| 2023 | On the Complexity of Bayesian GeneralizationabstractWe examine concept generalization at a large scale in the natural visual spectrum. Established computational modes (*i.e.*, rule-based or similarity-based) are primarily studied isolated, focusing on confined and abstract problem spaces. In this work, we study these two modes when the *problem space* scales up and when the *complexity* of concepts becomes diverse. At the **representational level**, we investigate how the complexity varies when a visual concept is mapped to the representation space. Prior literature has shown that two types of complexities (Griffiths & Tenenbaum, 2003) build an inverted-U relation (Donderi, 2006; Sun & Firestone, 2021). Leveraging *Representativeness of Attribute* (RoA), we computationally confirm: Models use attributes with high RoA to describe visual concepts, and the description length falls in an inverted-U relation with the increment in visual complexity. At the **computational level**, we examine how the complexity of representation affects the shift between the rule- and similarity-based generalization. We hypothesize that category-conditioned visual modeling estimates the co-occurrence frequency between visual and categorical attributes, thus potentially serving as the prior for the natural visual world. Experimental results show that representations with relatively high subjective complexity outperform those with relatively low subjective complexity in rule-based generalization, while the trend is the opposite in similarity-based generalization. Yu-Zhe Shi, Manjie Xu, John E. Hopcroft, Kun He 0001, Josh Tenenbaum, Song-Chun Zhu, Ying Nian Wu, Wenjuan Han, Yixin Zhu 0001 |
ICML | 5 |
| 2023 | H-SAUR: Hypothesize, Simulate, Act, Update, and Repeat for Understanding Object Articulations from InteractionsabstractThe world is filled with articulated objects that are difficult to determine how to use from vision alone, e.g., a door might open inwards or outwards. Humans handle these objects with strategic trial-and-error: first pushing a door then pulling if that doesn't work. We enable these capabilities in autonomous agents by proposing “Hypothesize, Simulate, Act, Update, and Repeat” (H-SAUR), a probabilistic generative framework that simultaneously generates a distribution of hypotheses about how objects articulate given input observations, captures certainty over hypotheses over time, and infer plausible actions for exploration and goal-conditioned manipulation. We compare our model with existing work in manipulating objects after a handful of exploration actions, on the PartNet-Mobility dataset. We further propose a novel PuzzleBoxes benchmark that contains locked boxes that require multiple steps to solve. We show that the proposed model significantly outperforms the current state-of-the-art articulated object manipulation framework, despite using zero training data. We further improve the test-time efficiency of H-SAUR by integrating a learned prior from learning-based vision models. Kei Ota, Hsiao-Yu Fish Tung, Kevin A. Smith 0001, Anoop Cherian, Tim K. Marks, Alan Sullivan, Asako Kanezaki, Josh Tenenbaum |
ICRA | 8 |
| 2023 | NOPA: Neurally-guided Online Probabilistic Assistance for Building Socially Intelligent Home AssistantsabstractIn this work, we study how to build socially intelligent robots to assist people in their homes. In particular, we focus on assistance with online goal inference, where robots must simultaneously infer humans' goals and how to help them achieve those goals. Prior assistance methods either lack the adaptivity to adjust helping strategies (i.e., when and how to help) in response to uncertainty about goals or the scalability to conduct fast inference in a large goal space. Our NOPA (Neurally-guided Online Probabilistic Assistance) method addresses both of these challenges. NOPA consists of (1) an online goal inference module combining neural goal proposals with inverse planning and particle filtering for robust inference under uncertainty, and (2) a helping planner that discovers valuable subgoals to help with and is aware of the uncertainty in goal inference. We compare NOPA against multiple baselines in a new embodied AI assistance challenge: Online Watch-And-Help, in which a helper agent needs to simultaneously watch a main agent's action, infer its goal, and help perform a common household task faster in realistic virtual home environments. Experiments show that our helper agent robustly updates its goal inference and adapts its helping plans to the changing level of uncertainty.11Code and a supplementary video are available at https://www.tshu.io/online_watch_and_help. Xavier Puig, Tianmin Shu, Josh Tenenbaum, Antonio Torralba 0001 |
ICRA | 3 |
| 2023 | Compositional Foundation Models for Hierarchical PlanningabstractTo make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with the devised plan through visual-motor control. We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model which leverages multiple expert foundation model trained on language, vision and action data individually jointly together to solve long-horizon tasks. We use a large language model to construct symbolic plans that are grounded in the environment through a large video diffusion model. Generated video plans are then grounded to visual-motor control, through an inverse dynamics model that infers actions from generated videos. To enable effective reasoning within this hierarchy, we enforce consistency between the models via iterative refinement. We illustrate the efficacy and adaptability of our approach in three different long-horizon table-top manipulation tasks. Anurag Ajay, Seungwook Han, Yilun Du, Shuang Li 0013, Abhi Gupta, Tommi S. Jaakkola, Josh Tenenbaum, Leslie Pack Kaelbling, Akash Srivastava, Pulkit Agrawal 0001 |
NeurIPS | 7 |
| 2023 | Inferring the Future by Imagining the PastabstractA single panel of a comic book can say a lot: it can depict not only where the characters currently are, but also their motions, their motivations, their emotions, and what they might do next. More generally, humans routinely infer complex sequences of past and future events from a *static snapshot* of a *dynamic scene*, even in situations they have never seen before.
In this paper, we model how humans make such rapid and flexible inferences. Building on a long line of work in cognitive science, we offer a Monte Carlo algorithm whose inferences correlate well with human intuitions in a wide variety of domains, while only using a small, cognitively-plausible number of samples. Our key technical insight is a surprising connection between our inference problem and Monte Carlo path tracing, which allows us to apply decades of ideas from the computer graphics community to this seemingly-unrelated theory of mind task. Kartik Chandra, Tony Chen 0003, Tzu-Mao Li, Jonathan Ragan-Kelley, Josh Tenenbaum |
NeurIPS | 5 |
| 2023 | Learning Universal Policies via Text-Guided Video GenerationabstractA goal of artificial intelligence is to construct an agent that can solve a wide variety of tasks. Recent progress in text-guided image synthesis has yielded models with an impressive ability to generate complex novel images, exhibiting combinatorial generalization across domains. Motivated by this success, we investigate whether such tools can be used to construct more general-purpose agents. Specifically, we cast the sequential decision making problem as a text-conditioned video generation problem, where, given a text-encoded specification of a desired goal, a planner synthesizes a set of future frames depicting its planned actions in the future, after which control actions are extracted from the generated video. By leveraging text as the underlying goal specification, we are able to naturally and combinatorially generalize to novel goals. The proposed policy-as-video formulation can further represent environments with different state and action spaces in a unified space of images, which, for example, enables learning and generalization across a variety of robot manipulation tasks. Finally, by leveraging pretrained language embeddings and widely available videos from the internet, the approach enables knowledge transfer through predicting highly realistic video plans for real robots. Yilun Du, Sherry Yang 0001, Bo Dai 0001, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, Pieter Abbeel |
NeurIPS | 6 |
| 2023 | What's Left? Concept Grounding with Logic-Enhanced Foundation ModelsabstractRecent works such as VisProg and ViperGPT have smartly composed foundation models for visual reasoning—using large language models (LLMs) to produce programs that can be executed by pre-trained vision-language models. However, they operate in limited domains, such as 2D images, not fully exploiting the generalization of language: abstract concepts like “*left*” can also be grounded in 3D, temporal, and action data, as in moving to your *left*. This limited generalization stems from these inference-only methods’ inability to learn or adapt pre-trained models to a new domain. We propose the **L**ogic-**E**nhanced **F**ounda**T**ion Model (**LEFT**), a unified framework that *learns* to ground and reason with concepts across domains with a differentiable, domain-independent, first-order logic-based program executor. LEFT has an LLM interpreter that outputs a program represented in a general, logic-based reasoning language, which is shared across all domains and tasks. LEFT’s executor then executes the program with trainable domain-specific grounding modules. We show that LEFT flexibly learns concepts in four domains: 2D images, 3D scenes, human motions, and robotic manipulation. It exhibits strong reasoning ability in a wide variety of tasks, including those that are complex and not seen during training, and can be easily applied to new domains. Joy Hsu, Jiayuan Mao, Josh Tenenbaum, Jiajun Wu 0001 |
NeurIPS | 3 |
| 2023 | What Planning Problems Can A Relational Neural Network Solve?abstractGoal-conditioned policies are generally understood to be "feed-forward" circuits, in the form of neural networks that map from the current state and the goal specification to the next action to take. However, under what circumstances such a policy can be learned and how efficient the policy will be are not well understood. In this paper, we present a circuit complexity analysis for relational neural networks (such as graph neural networks and transformers) representing policies for planning problems, by drawing connections with serialized goal regression search (S-GRS). We show that there are three general classes of planning problems, in terms of the growth of circuit width and depth as a function of the number of objects and planning horizon, providing constructive proofs. We also illustrate the utility of this analysis for designing neural networks for policy learning. Jiayuan Mao, Tomás Lozano-Pérez, Josh Tenenbaum, Leslie Pack Kaelbling |
NeurIPS | 3 |
| 2023 | Human spatiotemporal pattern learning as probabilistic program synthesisabstractPeople are adept at learning a wide variety of structured patterns from small amounts of data, presenting a conundrum from the standpoint of the bias-variance tradeoff: what kinds of representations and algorithms support the joint flexibility and data-paucity of human learning? One possibility is that people "learn by programming": inducing probabilistic models to fit observed data. Here, we experimentally test human learning in the domain of structured 2-dimensional patterns, using a task in which participants repeatedly predicted where a dot would move based on its previous trajectory. We evaluate human performance against standard parametric and non-parametric time-series models, as well as two Bayesian program synthesis models whose hypotheses vary in their degree of structure: a compositional Gaussian Process model and a structured "Language of Thought" (LoT) model. We find that signatures of human pattern learning are best explained by the LoT model, supporting the idea that the flexibility and data-efficiency of human structure learning can be understood as probabilistic inference over an expressive space of programs. Tracey Mills, Josh Tenenbaum, Samuel J. Cheyette |
NeurIPS | 2 |
| 2023 | Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct SupervisionabstractDenoising diffusion models are a powerful type of generative models used to capture complex distributions of real-world signals. However, their applicability is limited to scenarios where training samples are readily available, which is not always the case in real-world applications. For example, in inverse graphics, the goal is to generate samples from a distribution of 3D scenes that align with a given image, but ground-truth 3D scenes are unavailable and only 2D images are accessible. To address this limitation, we propose a novel class of denoising diffusion probabilistic models that learn to sample from distributions of signals that are never directly observed. Instead, these signals are measured indirectly through a known differentiable forward model, which produces partial observations of the unknown signal. Our approach involves integrating the forward model directly into the denoising process. A key contribution of our work is the integration of a differentiable forward model into the denoising process. This integration effectively connects the generative modeling of observations with the generative modeling of the underlying signals, allowing for end-to-end training of a conditional generative model over signals. During inference, our approach enables sampling from the distribution of underlying signals that are consistent with a given partial observation. We demonstrate the effectiveness of our method on three challenging computer vision tasks. For instance, in the context of inverse graphics, our model enables direct sampling from the distribution of 3D scenes that align with a single 2D input image. Ayush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov, Josh Tenenbaum, Frédo Durand, William T. Freeman, Vincent Sitzmann |
NeurIPS | 5 |
| 2023 | Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical PropertiesabstractGeneral physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g., mass or elasticity), and that those properties affect the outcome of physical events. While there has been great progress in physical and video prediction models in recent years, benchmarks to test their performance typically do not require an understanding that objects have individual physical properties, or at best test only those properties that are directly observable (e.g., size or color). This work proposes a novel dataset and benchmark, termed Physion++, that rigorously evaluates visual physical prediction in artificial systems under circumstances where those predictions rely on accurate estimates of the latent physical properties of objects in the scene. Specifically, we test scenarios where accurate prediction relies on estimates of properties such as mass, friction, elasticity, and deformability, and where the values of those properties can only be inferred by observing how objects move and interact with other objects or fluids. We evaluate the performance of a number of state-of-the-art prediction models that span a variety of levels of learning vs. built-in knowledge, and compare that performance to a set of human predictions. We find that models that have been trained using standard regimes and datasets do not spontaneously learn to make inferences about latent properties, but also that models that encode objectness and physical states tend to make better predictions. However, there is still a huge gap between all models and human performance, and all models' predictions correlate poorly with those made by humans, suggesting that no state-of-the-art model is learning to make physical predictions in a human-like way. These results show that current deep learning models that succeed in some settings nevertheless fail to achieve human-level physical prediction in other cases, especially those where latent property inference is required. Project page: https://dingmyu.github.io/physion_v2/ Hsiao-Yu Fish Tung, Mingyu Ding, Zhenfang Chen, Daniel Bear, Chuang Gan 0001, Josh Tenenbaum, Dan Yamins, Judith E. Fan, Kevin A. Smith 0001 |
NeurIPS | 6 |
| 2023 | DiffuseBot: Breeding Soft Robots With Physics-Augmented Generative Diffusion ModelsabstractNature evolves creatures with a high complexity of morphological and behavioral intelligence, meanwhile computational methods lag in approaching that diversity and efficacy. Co-optimization of artificial creatures' morphology and control in silico shows promise for applications in physical soft robotics and virtual character creation; such approaches, however, require developing new learning algorithms that can reason about function atop pure structure. In this paper, we present DiffuseBot, a physics-augmented diffusion model that generates soft robot morphologies capable of excelling in a wide spectrum of tasks. \name bridges the gap between virtually generated content and physical utility by (i) augmenting the diffusion process with a physical dynamical simulation which provides a certificate of performance, and (ii) introducing a co-design procedure that jointly optimizes physical design and control by leveraging information about physical sensitivities from differentiable simulation. We showcase a range of simulated and fabricated robots along with their capabilities. Check our website: https://diffusebot.github.io/ Tsun-Hsuan Wang, Juntian Zheng, Pingchuan Ma 0002, Yilun Du, Byungchul Kim, Andrew Spielberg, Josh Tenenbaum, Chuang Gan 0001, Daniela Rus |
NeurIPS | 7 |
| 2023 | 3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging ScenesabstractGiven a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on extensive trial and error. In this paper, we present a framework capable of learning 3D-grounded visual intuitive physics models from videos of complex scenes with fluids. Our method is composed of a conditional Neural Radiance Field (NeRF)-style visual frontend and a 3D point-based dynamics prediction backend, using which we can impose strong relational and structural inductive bias to capture the structure of the underlying environment. Unlike existing intuitive point-based dynamics works that rely on the supervision of dense point trajectory from simulators, we relax the requirements and only assume access to multi-view RGB images and (imperfect) instance masks acquired using color prior. This enables the proposed model to handle scenarios where accurate point estimation and tracking are hard or impossible. We generate datasets including three challenging scenarios involving fluid, granular materials, and rigid objects in the simulation. The datasets do not include any dense particle information so most previous 3D-based intuitive physics pipelines can barely deal with that. We show our model can make long-horizon future predictions by learning from raw images and significantly outperforms models that do not employ an explicit 3D representation space. We also show that once trained, our model can achieve strong generalization in complex scenarios under extrapolate settings. Antonio Torralba 0001, Josh Tenenbaum, Dan Yamins, Yunzhu Li, Hsiao-Yu Fish Tung |
NeurIPS | 3 |
| 2023 | Top-Down Synthesis for Library LearningabstractThis paper introduces corpus-guided top-down synthesis as a mechanism for synthesizing library functions that capture common functionality from a corpus of programs in a domain specific language (DSL). The algorithm builds abstractions directly from initial DSL primitives, using syntactic pattern matching of intermediate abstractions to intelligently prune the search space and guide the algorithm towards abstractions that maximally capture shared structures in the corpus. We present an implementation of the approach in a tool called Stitch and evaluate it against the state-of-the-art deductive library learning algorithm from DreamCoder. Our evaluation shows that Stitch is 3-4 orders of magnitude faster and uses 2 orders of magnitude less memory while maintaining comparable or better library quality (as measured by compressivity). We also demonstrate Stitch’s scalability on corpora containing hundreds of complex programs that are intractable with prior deductive approaches and show empirically that it is robust to terminating the search procedure early—further allowing it to scale to challenging datasets by means of early stopping. Matthew Bowers, Theo X. Olausson, Lionel Wong, Gabriel Grand, Josh Tenenbaum, Kevin Ellis, Armando Solar-Lezama |
Proc. ACM Program. Lang. | 5 |
| 2023 | Combining Functional and Automata Synthesis to Discover Causal Reactive ProgramsabstractWe present a new algorithm that synthesizes functional reactive programs from observation data. The key novelty is to iterate between a functional synthesis step, which attempts to generate a transition function over observed states, and an automata synthesis step, which adds any additional latent state necessary to fully account for the observations. We develop a functional reactive DSL called Autumn that can express a rich variety of causal dynamics in time-varying, Atari-style grid worlds, and apply our method to synthesize Autumn programs from data. We evaluate our algorithm on a benchmark suite of 30 Autumn programs as well as a third-party corpus of grid-world-style video games. We find that our algorithm synthesizes 27 out of 30 programs in our benchmark suite and 21 out of 27 programs from the third-party corpus, including several programs describing complex latent state transformations, and from input traces containing hundreds of observations. We expect that our approach will provide a template for how to integrate functional and automata synthesis in other induction domains. Ria Das, Josh Tenenbaum, Armando Solar-Lezama, Zenna Tavares |
Proc. ACM Program. Lang. | 2 |
| 2022 | Discovering State and Action Abstractions for Generalized Task and Motion PlanningabstractGeneralized planning accelerates classical planning by finding an algorithm-like policy that solves multiple instances of a task. A generalized plan can be learned from a few training examples and applied to an entire domain of problems. Generalized planning approaches perform well in discrete AI planning problems that involve large numbers of objects and extended action sequences to achieve the goal. In this paper, we propose an algorithm for learning features, abstractions, and generalized plans for continuous robotic task and motion planning (TAMP) and examine the unique difficulties that arise when forced to consider geometric and physical constraints as a part of the generalized plan. Additionally, we show that these simple generalized plans learned from only a handful of examples can be used to improve the search efficiency of TAMP solvers. Aidan Curtis, Tom Silver, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
AAAI | 3 |
| 2022 | MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image ManipulationabstractRecent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningful and disentangled way. Prior approaches aim to achieve such semantic control and disentanglement within the latent space of a previously trained GAN. In contrast, we propose a framework that a priori models physical attributes of the face such as 3D shape, albedo, pose, and lighting explicitly, thus providing disentanglement by design. Our method, MOST-GAN, integrates the expressive power and photorealism of style-based GANs with the physical disentanglement and flexibility of nonlinear 3D morphable models, which we couple with a state-of-the-art 2D hair manipulation network. MOST-GAN achieves photorealistic manipulation of portrait images with fully disentangled 3D control over their physical attributes, enabling extreme manipulation of lighting, facial expression, and pose variations up to full profile view. Safa C. Medin, Bernhard Egger 0001, Anoop Cherian, Ye Wang 0001, Josh Tenenbaum, Xiaoming Liu 0002, Tim K. Marks |
AAAI | 5 |
| 2022 | Preschoolers' sensitivity to abstract correlations in the properties of sets and functions
Nicole Coates, Max H. Siegel, Junyi Chu, Melissa Kline Struhl, Josh Tenenbaum, Laura Schulz |
CogSci | 5 |
| 2022 | Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks
Katie Collins, Catherine Wong, Jiahai Feng, Megan Wei, Josh Tenenbaum |
CogSci | 5 |
| 2022 | Benchmarking mid-level vision with texture-defined 3D objects
Yoni Friedman, Thomas P. O'Connell, Max H. Siegel, Daniel Bear, Tuan Anh Le 0001, Bernhard Egger 0001, Josh Tenenbaum |
CogSci | 7 |
| 2022 | Modeling risky food sharing as rational communication about relationships
Michelle Simona Hung, Ashley J. Thomas, Setayesh Radkani, Josh Tenenbaum, Rebecca Saxe |
CogSci | 4 |
| 2022 | Flexibility in Moral Cognition: When is it okay to break the rules?
Joseph Kwon, Josh Tenenbaum, Sydney Levine |
CogSci | 2 |
| 2022 | Joint online inference of material properties and object shape
Vivian C. Paulun, Max H. Siegel, Josh Tenenbaum |
CogSci | 3 |
| 2022 | Modeling punishment as a rational communicative social action
Setayesh Radkani, Josh Tenenbaum, Rebecca Saxe |
CogSci | 2 |
| 2022 | Learning as Programming: Modeling Efficient Search in Human Concept Learning
Joshua S. Rule, Steve Piantadosi, Josh Tenenbaum |
CogSci | 3 |
| 2022 | Efficient exploration of spatial environments through Map Induction using adaptable compositional map representations
Sugandha Sharma, Aidan Curtis, Marta Kryven, Josh Tenenbaum, Ila Fiete |
CogSci | 4 |
| 2022 | Identifying concept libraries from language about object structure
Catherine Wong, William P. McCarthy, Gabriel Grand, Yoni Friedman, Josh Tenenbaum, Jacob Andreas, Robert D. Hawkins, Judith E. Fan |
CogSci | 5 |
| 2022 | Finding Fallen Objects Via Asynchronous Audio-Visual IntegrationabstractThe way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in which to study multi-modal object localization in 3D virtual environments. An object is dropped somewhere in a room. An embodied robot agent, equipped with camera and microphone, must determine what object has been dropped - and where - by combining audio and visual signals with knowledge of the underlying physics. To study this problem, we have generated a large-scale dataset - the Fallen Objects dataset - that includes 8000 instances of 30 physical object categories in 64 rooms. The dataset uses the ThreeDWorld Platform that can simulate physics-based impact sounds and complex physical interactions between objects in a photorealistic setting. As a first step toward addressing this challenge, we develop a set of embodied agent baselines, based on imitation learning, reinforcement learning, and modular planning, and perform an in-depth analysis of the challenge of this new task. This dataset is publicly available11Project page: http://fallen-object.csail.mit.edu. Chuang Gan 0001, Yi Gu 0002, Jeremy Schwartz, Seth Alter, James Traer, Dan Gutfreund, Josh Tenenbaum, Josh H. McDermott, Antonio Torralba 0001 |
CVPR | 8 |
| 2022 | Fixing Malfunctional Objects With Learned Physical Simulation and Functional PredictionabstractThis paper studies the problem of fixing malfunctional 3D objects. While previous works focus on building passive perception models to learn the functionality from static 3D objects, we argue that functionality is reckoned with respect to the physical interactions between the object and the user. Given a malfunctional object, humans can perform mental simulations to reason about its functionality and figure out how to fix it. Inspired by this, we propose FixIt, a dataset that contains about 5k poorly-designed 3D physical objects paired with choices to fix them. To mimic humans' mental simulation process, we present FixNet, a novel framework that seamlessly incorporates perception and physical dynamics. Specifically, FixNet consists of a perception module to extract the structured representation from the 3D point cloud, a physical dynamics prediction module to simulate the results of interactions on 3D objects, and a functionality prediction module to evaluate the functionality and choose the correct fix. Experimental results show that our framework outperforms baseline models by a large margin, and can generalize well to objects with similar interaction types. Code and dataset are publicly available11http://fixing-malfunctional.csail.mit.edu. Yining Hong, Kaichun Mo, Li Yi 0001, Leonidas J. Guibas, Antonio Torralba 0001, Josh Tenenbaum, Chuang Gan 0001 |
CVPR | 6 |
| 2022 | Unsupervised Segmentation in Real-World Images via Spelke Object Inference
Rahul M. V., Yoni Friedman, Jiajun Wu 0001, Josh Tenenbaum, Dan Yamins, Daniel Bear |
ECCV (29) | 5 |
| 2022 | Compositional Visual Generation with Composable Diffusion Models
Nan Liu 0010, Shuang Li 0013, Yilun Du, Antonio Torralba 0001, Josh Tenenbaum |
ECCV (17) | 5 |
| 2022 | Hybrid Memoised Wake-Sleep: Approximate Inference at the Discrete-Continuous Interface
Tuan Anh Le 0001, Katie Collins, Luke Hewitt, Kevin Ellis, N. Siddharth 0001, Samuel Gershman, Josh Tenenbaum |
ICLR | 7 |
| 2022 | RISP: Rendering-Invariant State Predictor with Differentiable Simulation and Rendering for Cross-Domain Parameter Estimation
Pingchuan Ma 0002, Tao Du 0001, Josh Tenenbaum, Wojciech Matusik, Chuang Gan 0001 |
ICLR | 3 |
| 2022 | ComPhy: Compositional Physical Reasoning of Objects and Events from Videos
Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba 0001, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 6 |
| 2022 | Contact Points Discovery for Soft-Body Manipulations with Differentiable Physics
Zhiao Huang, Tao Du 0001, Hao Su 0001, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 5 |
| 2022 | DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools
Zhiao Huang, Yunzhu Li, Josh Tenenbaum, David Held, Chuang Gan 0001 |
ICLR | 4 |
| 2022 | FALCON: Fast Visual Concept Learning by Integrating Images, Linguistic descriptions, and Conceptual Relations
Lingjie Mei, Jiayuan Mao, Chuang Gan 0001, Josh Tenenbaum |
ICLR | 5 |
| 2022 | Map Induction: Compositional spatial submap learning for efficient exploration in novel environments
Sugandha Sharma, Aidan Curtis, Marta Kryven, Josh Tenenbaum, Ila Fiete |
ICLR | 4 |
| 2022 | Linking Emergent and Natural Languages via Corpus Transfer
Shunyu Yao 0006, Mo Yu, Yang Zhang 0001, Karthik Narasimhan, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 5 |
| 2022 | Learning Iterative Reasoning through Energy MinimizationabstractDeep learning has excelled on complex pattern recognition tasks such as image classification and object recognition. However, it struggles with tasks requiring nontrivial reasoning, such as algorithmic computation. Humans are able to solve such tasks through iterative reasoning – spending more time to think about harder tasks. Most existing neural networks, however, exhibit a fixed computational budget controlled by the neural network architecture, preventing additional computational processing on harder tasks. In this work, we present a new framework for iterative reasoning with neural networks. We train a neural network to parameterize an energy landscape over all outputs, and implement each step of the iterative reasoning as an energy minimization step to find a minimal energy solution. By formulating reasoning as an energy minimization problem, for harder problems that lead to more complex energy landscapes, we may then adjust our underlying computational budget by running a more complex optimization procedure. We empirically illustrate that our iterative reasoning approach can solve more accurate and generalizable algorithmic reasoning tasks in both graph and continuous domains. Finally, we illustrate that our approach can recursively solve algorithmic problems requiring nested reasoning. Yilun Du, Shuang Li 0013, Josh Tenenbaum, Igor Mordatch |
ICML | 3 |
| 2022 | Planning with Diffusion for Flexible Behavior SynthesisabstractModel-based reinforcement learning methods often use learning only for the purpose of recovering an approximate dynamics model, offloading the rest of the decision-making work to classical trajectory optimizers. While conceptually simple, this combination has a number of empirical shortcomings, suggesting that learned models may not be well-suited to standard trajectory optimization. In this paper, we consider what it would look like to fold as much of the trajectory optimization pipeline as possible into the modeling problem, such that sampling from the model and planning with it become nearly identical. The core of our technical approach lies in a diffusion probabilistic model that plans by iteratively denoising trajectories. We show how classifier-guided sampling and image inpainting can be reinterpreted as coherent planning strategies, explore the unusual and useful properties of diffusion-based planning methods, and demonstrate the effectiveness of our framework in control settings that emphasize long-horizon decision-making and test-time flexibility. Michael Janner, Yilun Du, Josh Tenenbaum, Sergey Levine |
ICML | 3 |
| 2022 | Discovering Generalizable Spatial Goal Representations via Graph-based Active Reward LearningabstractIn this work, we consider one-shot imitation learning for object rearrangement tasks, where an AI agent needs to watch a single expert demonstration and learn to perform the same task in different environments. To achieve a strong generalization, the AI agent must infer the spatial goal specification for the task. However, there can be multiple goal specifications that fit the given demonstration. To address this, we propose a reward learning approach, Graph-based Equivalence Mappings (GEM), that can discover spatial goal representations that are aligned with the intended goal specification, enabling successful generalization in unseen environments. Specifically, GEM represents a spatial goal specification by a reward function conditioned on i) a graph indicating important spatial relationships between objects and ii) state equivalence mappings for each edge in the graph indicating invariant properties of the corresponding relationship. GEM combines inverse reinforcement learning and active reward learning to efficiently improve the reward function by utilizing the graph structure and domain randomization enabled by the equivalence mappings. We conducted experiments with simulated oracles and with human subjects. The results show that GEM can drastically improve the generalizability of the learned goal representations over strong baselines. Aviv Netanyahu, Tianmin Shu, Josh Tenenbaum, Pulkit Agrawal 0001 |
ICML | 3 |
| 2022 | Prompting Decision Transformer for Few-Shot Policy GeneralizationabstractHuman can leverage prior experience and learn novel tasks from a handful of demonstrations. In contrast to offline meta-reinforcement learning, which aims to achieve quick adaptation through better algorithm design, we investigate the effect of architecture inductive bias on the few-shot learning capability. We propose a Prompt-based Decision Transformer (Prompt-DT), which leverages the sequential modeling ability of the Transformer architecture and the prompt framework to achieve few-shot adaptation in offline RL. We design the trajectory prompt, which contains segments of the few-shot demonstrations, and encodes task-specific information to guide policy generation. Our experiments in five MuJoCo control benchmarks show that Prompt-DT is a strong few-shot learner without any extra finetuning on unseen target tasks. Prompt-DT outperforms its variants and strong meta offline RL baselines by a large margin with a trajectory prompt containing only a few timesteps. Prompt-DT is also robust to prompt length changes and can generalize to out-of-distribution (OOD) environments. Project page: \href{https://mxu34.github.io/PromptDT/}{https://mxu34.github.io/PromptDT/}. Mengdi Xu, Yikang Shen, Ding Zhao, Josh Tenenbaum, Chuang Gan 0001 |
ICML | 6 |
| 2022 | The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark Towards Physically Realistic Embodied AIabstractWe introduce a visually-guided task-and-motion planning benchmark, which we call the ThreeDWorld Trans-port Challenge. In this challenge, an embodied agent is spawned randomly in a simulated physical home environment and required to transport a small set of objects scattered around the house with containers. We build this benchmark challenge using the ThreeDWorld simulation: a virtual 3D environment where all objects respond to physics, and a robot agent can be controlled using a fully physics-driven navigation and interaction API. We evaluate several existing agents on this benchmark. Experimental results suggest that: 1) a pure RL model struggles on this challenge; 2) state-of-the-art hierarchical planning-based agents can transport some objects but are still far from solving this task. We anticipate that this benchmark will empower researchers to develop more intelligent physics-aware robot learning algorithms. Chuang Gan 0001, Jeremy Schwartz, Seth Alter, Abhishek Bhandwaldar, Dan Gutfreund, Dan Yamins, James J. DiCarlo, Josh H. McDermott, Antonio Torralba 0001, Josh Tenenbaum |
ICRA | 11 |
| 2022 | DURableVS: Data-efficient Unsupervised Recalibrating Visual Servoing via online learning in a structured generative modelabstractVisual servoing enables robotic systems to perform accurate closed-loop control, which is required in many applications. However, existing methods require either precise calibration of the robot kinematic model and cameras or use neural architectures that require large amounts of data to train. In this work, we present a method for unsupervised learning of visual servoing that does not require any prior calibration and is extremely data-efficient. Our key insight is that visual servoing does not depend on identifying the veridical kinematic and camera parameters, but instead only on an accurate generative model of image feature observations from the joint positions of the robot. We demonstrate that with our model architecture and learning algorithm, we can consistently learn accurate models from less than 50 training samples (which amounts to less than 1 min of unsupervised data collection), and that such data-efficient learning is not possible with standard neural architectures. Further, we show that by using the generative model in the loop and learning online, we can enable a robotic system to recover from calibration errors and to detect and quickly adapt to possibly unexpected changes in the robot-camera system (e.g. bumped camera, new objects). Nishad Gothoskar, Miguel Lázaro-Gredilla, Yasemin Bekiroglu, Josh Tenenbaum, Vikash Mansinghka 0001, Dileep George |
ICRA | 5 |
| 2022 | Neural Descriptor Fields: SE(3)-Equivariant Object Representations for ManipulationabstractWe present Neural Descriptor Fields (NDFs), an object representation that encodes both points and relative poses between an object and a target (such as a robot gripper or a rack used for hanging) via category-level descriptors. We employ this representation for object manipulation, where given a task demonstration, we want to repeat the same task on a new object instance from the same category. We propose to achieve this objective by searching (via optimization) for the pose whose descriptor matches that observed in the demonstration. NDFs are conveniently trained in a self-supervised fashion via a 3D auto-encoding task that does not rely on expert-labeled keypoints. Further, NDFs are SE(3)-equivariant, guaranteeing performance that generalizes across all possible 3D object translations and rotations. We demonstrate learning of manipulation tasks from few (∼5-10) demonstrations both in simulation and on a real robot. Our performance generalizes across both object instances and 6-DoF object poses, and significantly outperforms a recent baseline that relies on 2D descriptors. Project website: https://yilundu.github.io/ndf/ Anthony Simeonov, Yilun Du, Andrea Tagliasacchi, Josh Tenenbaum, Alberto Rodriguez 0003, Pulkit Agrawal 0001, Vincent Sitzmann |
ICRA | 4 |
| 2022 | Incorporating Rich Social Interactions Into MDPsabstractMuch of what we do as humans is engage socially with other agents, a skill that robots must also eventually possess. We demonstrate that a rich theory of social interactions originating from microsociology can be formalized by extending a nested MDP where agents reason about arbitrary functions of each other's rewards. This extended Social MDP allows us to encode the five basic interactions that underlie microsociology: cooperation, conflict, coercion, competition, and exchange. The result is a robotic agent capable of executing social interactions in new environments with no interaction-specific training; like humans it can engage socially in novel ways even without a single example of that social interaction. Moreover, the estimations of these Social MDPs align closely with the judge-ments of humans when considering which social interaction is taking place in an environment. This method both sheds light on the nature of social interactions, by providing concrete mathematical definitions, and brings rich social interactions into a mathematical framework that has proven to be natural for robotics. Ravi Tejwani, Yen-Ling Kuo, Tianmin Shu, Bennett Stankovits, Dan Gutfreund, Josh Tenenbaum, Boris Katz, Andrei Barbu |
ICRA | 6 |
| 2022 | Learning Neuro-Symbolic Relational Transition Models for Bilevel PlanningabstractIn robotic domains, learning and planning are complicated by continuous state spaces, continuous action spaces, and long task horizons. In this work, we address these challenges with Neuro-Symbolic Relational Transition Models (NSRTs), a novel class of models that are data-efficient to learn, compatible with powerful robotic planning methods, and generalizable over objects. NSRTs have both symbolic and neural components, enabling a bilevel planning scheme where symbolic AI planning in an outer loop guides continuous planning with neural models in an inner loop. Experiments in four robotic planning domains show that NSRTs can be learned very data-efficiently, and then used for fast planning in new tasks that require up to 60 actions and involve many more objects than were seen during training. Rohan Chitnis, Tom Silver, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IROS | 3 |
| 2022 | Robust Change Detection Based on Neural Descriptor FieldsabstractThe ability to reason about changes in the environment is crucial for robots operating over extended periods of time. Agents are expected to capture changes during operation so that actions can be followed to ensure a smooth progression of the working session. However, varying viewing angles and accumulated localization errors make it easy for robots to falsely detect changes in the surrounding world due to low observation overlap and drifted object associations. In this paper, based on the recently proposed category-level Neural Descriptor Fields (NDFs), we develop an object-level online change detection approach that is robust to partially overlapping observations and noisy localization results. Utilizing the shape completion capability and SE(3)-equivariance of NDFs, we represent objects with compact shape codes encoding full object shapes from partial observations. The objects are then organized in a spatial tree structure based on object centers recovered from NDFs for fast queries of object neighborhoods. By associating objects via shape code similarity and comparing local object-neighbor spatial layout, our proposed approach demonstrates robustness to low observation overlap and localization noises. We conduct experiments on both synthetic and real-world sequences and achieve improved change detection results compared to multiple baseline methods. Project web-page: ?http://yilundu.github.io/ndf_change Jiahui Fu 0002, Yilun Du, Kurran Singh, Josh Tenenbaum, John J. Leonard |
IROS | 4 |
| 2022 | Noisy Agents: Self-supervised Exploration by Predicting Auditory EventsabstractHumans integrate multiple sensory modalities (e.g., visual and audio) to build a causal understanding of the physical world. In this work, we propose a novel type of intrinsic motivation for Reinforcement Learning (RL) that encourages the agent to understand the causal effect of its actions through auditory event prediction. First, we allow the agent to collect a small amount of acoustic data and use K-means to discover underlying auditory event clusters. We then train a neural network to predict the auditory events and use the prediction errors as intrinsic rewards to guide RL exploration. We first conduct proof-of-concept experiments using a set of Atari games for an in-depth analysis of our module. We then apply our model to embodied audio-visual exploration using the Habitat simulator and active exploration with a rolling robot using the ThreeDWorld (TDW) simulator. Experimental results demonstrate the advantages of using audio signals over vision-based models as intrinsic rewards to guide RL explorations. Chuang Gan 0001, Phillip Isola, Antonio Torralba 0001, Josh Tenenbaum |
IROS | 5 |
| 2022 | Communicating Natural Programs to Humans and MachinesabstractThe Abstraction and Reasoning Corpus (ARC) is a set of procedural tasks that tests an agent's ability to flexibly solve novel problems. While most ARC tasks are easy for humans, they are challenging for state-of-the-art AI. What makes building intelligent systems that can generalize to novel situations such as ARC difficult? We posit that the answer might be found by studying the difference of $\textit{language}$: While humans readily generate and interpret instructions in a general language, computer systems are shackled to a narrow domain-specific language that they can precisely execute. We present LARC, the $\textit{Language-complete ARC}$: a collection of natural language descriptions by a group of human participants who instruct each other on how to solve ARC tasks using language alone, which contains successful instructions for 88\% of the ARC tasks. We analyze the collected instructions as `natural programs', finding that while they resemble computer programs, they are distinct in two ways: First, they contain a wide range of primitives; Second, they frequently leverage communicative strategies beyond directly executable codes. We demonstrate that these two distinctions prevent current program synthesis techniques from leveraging LARC to its full potential, and give concrete suggestions on how to build the next-generation program synthesizers. Samuel Acquaviva, Yewen Pu, Marta Kryven, Theodoros Sechopoulos, Catherine Wong, Gabrielle E. Ecanow, Maxwell I. Nye, Michael Henry Tessler, Josh Tenenbaum |
NeurIPS | 9 |
| 2022 | Learning Physical Dynamics with Subequivariant Graph Neural NetworksabstractGraph Neural Networks (GNNs) have become a prevailing tool for learning physical dynamics. However, they still encounter several challenges: 1) Physical laws abide by symmetry, which is a vital inductive bias accounting for model generalization and should be incorporated into the model design. Existing simulators either consider insufficient symmetry, or enforce excessive equivariance in practice when symmetry is partially broken by gravity. 2) Objects in the physical world possess diverse shapes, sizes, and properties, which should be appropriately processed by the model. To tackle these difficulties, we propose a novel backbone, called Subequivariant Graph Neural Network, which 1) relaxes equivariance to subequivariance by considering external fields like gravity, where the universal approximation ability holds theoretically; 2) introduces a new subequivariant object-aware message passing for learning physical interactions between multiple objects of various shapes in particle-based representation; 3) operates in a hierarchical fashion, allowing for modeling long-range and complex interactions. Our model achieves on average over 3% enhancement in contact prediction accuracy across 8 scenarios on Physion and 2$\times$ lower rollout MSE on RigidFall compared with state-of-the-art GNN simulators, while exhibiting strong generalization and data efficiency. Jiaqi Han 0001, Wenbing Huang 0001, Hengbo Ma, Jiachen Li 0001, Josh Tenenbaum, Chuang Gan 0001 |
NeurIPS | 5 |
| 2022 | 3D Concept Grounding on Neural FieldsabstractIn this paper, we address the challenging problem of 3D concept grounding (i.e., segmenting and learning visual concepts) by looking at RGBD images and reasoning about paired questions and answers. Existing visual reasoning approaches typically utilize supervised methods to extract 2D segmentation masks on which concepts are grounded. In contrast, humans are capable of grounding concepts on the underlying 3D representation of images. However, traditionally inferred 3D representations (e.g., point clouds, voxelgrids and meshes) cannot capture continuous 3D features flexibly, thus making it challenging to ground concepts to 3D regions based on the language description of the object being referred to. To address both issues, we propose to leverage the continuous, differentiable nature of neural fields to segment and learn concepts. Specifically, each 3D coordinate in a scene is represented as a high dimensional descriptor. Concept grounding can then be performed by computing the similarity between the descriptor vector of a 3D coordinate and the vector embedding of a language concept, which enables segmentations and concept learning to be jointly learned on neural fields in a differentiable fashion. As a result, both 3D semantic and instance segmentations can emerge directly from question answering supervision using a set of defined neural operators on top of neural fields (e.g., filtering and counting). Experimental results show that our proposed framework outperforms unsupervised / language-mediated segmentation models on semantic and instance segmentation tasks, as well as outperforms existing models on the challenging 3D aware visual reasoning tasks. Furthermore, our framework can generalize well to unseen shape categories and real scans. Yining Hong, Yilun Du, Chunru Lin, Josh Tenenbaum, Chuang Gan 0001 |
NeurIPS | 4 |
| 2022 | When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentabstractAI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A central challenge for AI safety is capturing the flexibility of the human moral mind — the ability to determine when a rule should be broken, especially in novel or unusual situations. In this paper, we present a novel challenge set consisting of moral exception question answering (MoralExceptQA) of cases that involve potentially permissible moral exceptions – inspired by recent moral psychology studies. Using a state-of-the-art large language model (LLM) as a basis, we propose a novel moral chain of thought (MoralCoT) prompting strategy that combines the strengths of LLMs with theories of moral reasoning developed in cognitive science to predict human moral judgments. MoralCoT outperforms seven existing LLMs by 6.2% F1, suggesting that modeling human reasoning might be necessary to capture the flexibility of the human moral mind. We also conduct a detailed error analysis to suggest directions for future work to improve AI safety using MoralExceptQA. Our data is open-sourced at https://huggingface.co/datasets/feradauto/MoralExceptQA and code at https://github.com/feradauto/MoralCoT. Zhijing Jin 0001, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf |
NeurIPS | 8 |
| 2022 | Drawing out of Distribution with Neuro-Symbolic Generative ModelsabstractLearning general-purpose representations from perceptual inputs is a hallmark of human intelligence. For example, people can write out numbers or characters, or even draw doodles, by characterizing these tasks as different instantiations of the same generic underlying process---compositional arrangements of different forms of pen strokes. Crucially, learning to do one task, say writing, implies reasonable competence at another, say drawing, on account of this shared process. We present Drawing out of Distribution (DooD), a neuro-symbolic generative model of stroke-based drawing that can learn such general-purpose representations. In contrast to prior work, DooD operates directly on images, requires no supervision or expensive test-time inference, and performs unsupervised amortized inference with a symbolic stroke model that better enables both interpretability and generalization. We evaluate DooD on its ability to generalize across both data and tasks. We first perform zero-shot transfer from one dataset (e.g. MNIST) to another (e.g. Quickdraw), across five different datasets, and show that DooD clearly outperforms different baselines. An analysis of the learnt representations further highlights the benefits of adopting a symbolic stroke model. We then adopt a subset of the Omniglot challenge tasks, and evaluate its ability to generate new exemplars (both unconditionally and conditionally), and perform one-shot classification, showing that DooD matches the state of the art. Taken together, we demonstrate that DooD does indeed capture general-purpose representations across both data and task, and takes a further step towards building general and robust concept-learning systems. Yichao Liang, Josh Tenenbaum, Tuan Anh Le 0001, N. Siddharth 0001 |
NeurIPS | 2 |
| 2022 | Learning Neural Acoustic FieldsabstractOur environment is filled with rich and dynamic acoustic information. When we walk into a cathedral, the reverberations as much as appearance inform us of the sanctuary's wide open space. Similarly, as an object moves around us, we expect the sound emitted to also exhibit this movement. While recent advances in learned implicit functions have led to increasingly higher quality representations of the visual world, there have not been commensurate advances in learning spatial auditory representations. To address this gap, we introduce Neural Acoustic Fields (NAFs), an implicit representation that captures how sounds propagate in a physical scene. By modeling acoustic propagation in a scene as a linear time-invariant system, NAFs learn to continuously map all emitter and listener location pairs to a neural impulse response function that can then be applied to arbitrary sounds. We demonstrate NAFs on both synthetic and real data, and show that the continuous nature of NAFs enables us to render spatial acoustics for a listener at arbitrary locations. We further show that the representation learned by NAFs can help improve visual learning with sparse views. Finally we show that a representation informative of scene structure emerges during the learning of NAFs. Andrew Luo 0001, Yilun Du, Michael J. Tarr, Josh Tenenbaum, Antonio Torralba 0001, Chuang Gan 0001 |
NeurIPS | 4 |
| 2022 | PDSketch: Integrated Domain Programming, Learning, and PlanningabstractThis paper studies a model learning and online planning approach towards building flexible and general robots. Specifically, we investigate how to exploit the locality and sparsity structures in the underlying environmental transition model to improve model generalization, data-efficiency, and runtime-efficiency. We present a new domain definition language, named PDSketch. It allows users to flexibly define high-level structures in the transition models, such as object and feature dependencies, in a way similar to how programmers use TensorFlow or PyTorch to specify kernel sizes and hidden dimensions of a convolutional neural network. The details of the transition model will be filled in by trainable neural networks. Based on the defined structures and learned parameters, PDSketch automatically generates domain-independent planning heuristics without additional training. The derived heuristics accelerate the performance-time planning for novel goals. Jiayuan Mao, Tomás Lozano-Pérez, Josh Tenenbaum, Leslie Pack Kaelbling |
NeurIPS | 3 |
| 2022 | HandMeThat: Human-Robot Communication in Physical and Social EnvironmentsabstractWe introduce HandMeThat, a benchmark for a holistic evaluation of instruction understanding and following in physical and social environments. While previous datasets primarily focused on language grounding and planning, HandMeThat considers the resolution of human instructions with ambiguities based on the physical (object states and relations) and social (human actions and goals) information. HandMeThat contains 10,000 episodes of human-robot interactions. In each episode, the robot first observes a trajectory of human actions towards her internal goal. Next, the robot receives a human instruction and should take actions to accomplish the subgoal set through the instruction. In this paper, we present a textual interface for our benchmark, where the robot interacts with a virtual environment through textual commands. We evaluate several baseline models on HandMeThat, and show that both offline and online reinforcement learning algorithms perform poorly on HandMeThat, suggesting significant room for future work on physical and social human-robot communications and interactions. Yanming Wan, Jiayuan Mao, Josh Tenenbaum |
NeurIPS | 3 |
| 2022 | Neural Networks special issue on Artificial Intelligence and Brain Science
Kenji Doya, Karl J. Friston, Masashi Sugiyama, Josh Tenenbaum |
Neural Networks | 4 |
| 2021 | GLIB: Efficient Exploration for Relational Model-Based Reinforcement Learning via Goal-Literal BabblingabstractWe address the problem of efficient exploration for transition model learning in the relational model-based reinforcement learning setting without extrinsic goals or rewards. Inspired by human curiosity, we propose goal-literal babbling (GLIB), a simple and general method for exploration in such problems. GLIB samples relational conjunctive goals that can be understood as specific, targeted effects that the agent would like to achieve in the world, and plans to achieve these goals using the transition model being learned. We provide theoretical guarantees showing that exploration with GLIB will converge almost surely to the ground truth model. Experimentally, we find GLIB to strongly outperform existing methods in both prediction and planning on a range of tasks, encompassing standard PDDL and PPDDL planning benchmarks and a robotic manipulation task implemented in the PyBullet physics simulator. Video: https://youtu.be/F6lmrPT6TOY Code: https://git.io/JIsTB Rohan Chitnis, Tom Silver, Josh Tenenbaum, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 3 |
| 2021 | PHASE: PHysically-grounded Abstract Social Events for Machine Social PerceptionabstractThe ability to perceive and reason about social interactions in the context of physical environments is core to human social intelligence and human-machine cooperation. However, no prior dataset or benchmark has systematically evaluated physically grounded perception of complex social interactions that go beyond short actions, such as high-fiving, or simple group activities, such as gathering. In this work, we create a dataset of physically-grounded abstract social events, PHASE, that resemble a wide range of real-life social interactions by including social concepts such as helping another agent. PHASE consists of 2D animations of pairs of agents moving in a continuous space generated procedurally using a physics engine and a hierarchical planner. Agents have a limited field of view, and can interact with multiple objects, in an environment that has multiple landmarks and obstacles. Using PHASE, we design a social recognition task and a social prediction task. PHASE is validated with human experiments demonstrating that humans perceive rich interactions in the social events, and that the simulated agents behave similarly to humans. As a baseline model, we introduce a Bayesian inverse planning approach, SIMPLE (SIMulation, Planning and Local Estimation), which outperforms state-of-the-art feed-forward neural networks. We hope that PHASE can serve as a difficult new challenge for developing new models that can recognize complex social interactions. Aviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu, Josh Tenenbaum |
AAAI | 5 |
| 2021 | Planning with Learned Object Importance in Large Problem Instances using Graph Neural NetworksabstractReal-world planning problems often involve hundreds or even thousands of objects, straining the limits of modern planners. In this work, we address this challenge by learning to predict a small set of objects that, taken together, would be sufficient for finding a plan. We propose a graph neural network architecture for predicting object importance in a single inference pass, thus incurring little overhead while greatly reducing the number of objects that must be considered by the planner. Our approach treats the planner and transition model as black boxes, and can be used with any off-the-shelf planner. Empirically, across classical planning, probabilistic planning, and robotic task and motion planning, we find that our method results in planning that is significantly faster than several baselines, including other partial grounding strategies and lifted planners. We conclude that learning to predict a sufficient set of objects for a planning problem is a simple, powerful, and general mechanism for planning in large instances. Video: https://youtu.be/FWsVJc2fvCE Code: https://git.io/JIsqX Tom Silver, Rohan Chitnis, Aidan Curtis, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
AAAI | 4 |
| 2021 | Augmenting Policy Learning with Routines Discovered from a Single DemonstrationabstractHumans can abstract prior knowledge from very little data and use it to boost skill learning. In this paper, we propose routine-augmented policy learning (RAPL), which discovers routines composed of primitive actions from a single demonstration and uses discovered routines to augment policy learning. To discover routines from the demonstration, we first abstract routine candidates by identifying grammar over the demonstrated action trajectory. Then, the best routines measured by length and frequency are selected to form a routine library. We propose to learn policy simultaneously at primitive-level and routine-level with discovered routines, leveraging the temporal structure of routines. Our approach enables imitating expert behavior at multiple temporal scales for imitation learning and promotes reinforcement learning exploration. Extensive experiments on Atari games demonstrate that RAPL improves the state-of-the-art imitation learning method SQIL and reinforcement learning method A2C. Further, we show that discovered routines can generalize to unseen levels and difficulties on the CoinRun benchmark. Zelin Zhao 0001, Chuang Gan 0001, Jiajun Wu 0001, Josh Tenenbaum |
AAAI | 5 |
| 2021 | Learning Evolved Combinatorial Symbols with a Neuro-symbolic Generative Model
Matthias Hofer 0002, Tuan Anh Le 0001, Roger Levy, Josh Tenenbaum |
CogSci | 4 |
| 2021 | LARC: Language annotated Abstraction and Reasoning Corpus
Samuel Acquaviva, Yewen Pu, Maxwell I. Nye, Catherine Wong, Michael Henry Tessler, Josh Tenenbaum |
CogSci | 6 |
| 2021 | Modeling the Mistakes of Boundedly Rational Agents Within a Bayesian Theory of Mind
Arwa Alanqary, Gloria Z. Lin, Joie Le, Tan Zhi-Xuan, Vikash Mansinghka 0001, Josh Tenenbaum |
CogSci | 6 |
| 2021 | Meta-strategy learning in physical problem-solving: the effect of embodied experience
Kelsey R. Allen, Kevin A. Smith 0001, Laura-Ashleigh Bird, Josh Tenenbaum, Tamar R. Makin, Dorothy Cowie |
CogSci | 4 |
| 2021 | Using Games to Understand Intelligence
Franziska Brändle, Kelsey R. Allen, Josh Tenenbaum, Eric Schulz |
CogSci | 3 |
| 2021 | Explore, Exploit, Create: Inventing goals in play
Sophia Diggs-Galligan, Junyi Chu, Josh Tenenbaum, Laura Schulz |
CogSci | 3 |
| 2021 | Core knowledge objects in reasoning and language use for highly abstract inductive tasks
Gabrielle E. Ecanow, Catherine Wong, Samuel Acquaviva, Yewen Pu, Marta Kryven, Josh Tenenbaum |
CogSci | 6 |
| 2021 | Explaining the Gestalt principle of common fate as amortized inference
Yoni Friedman, Tuan Anh Le 0001, Bernhard Egger 0001, Max H. Siegel, Josh Tenenbaum |
CogSci | 5 |
| 2021 | The Omniglot Jr. challenge; Can a model achieve child-level character generation and classification?
Eliza Kosoy, Masha Belyi, Charlie Snell, Brenden M. Lake, Josh Tenenbaum, Alison Gopnik |
CogSci | 5 |
| 2021 | Engineering and reverse-engineering morality
Sydney Levine, Fiery Cushman, Iyad Rahwan, Josh Tenenbaum |
CogSci | 4 |
| 2021 | Combining rules and simulation to explain infant physical learning
João Loula, Kelsey R. Allen, Josh Tenenbaum |
CogSci | 3 |
| 2021 | Growing knowledge culturally across generations to solve novel, complex task
Michael Henry Tessler, Pedro Tsividis, Jason Madeano, Brin Harper, Josh Tenenbaum |
CogSci | 5 |
| 2021 | Language as a bootstrap for compositional visual reasoning
Catherine Wong, Yoni Friedman, Jacob Andreas, Josh Tenenbaum |
CogSci | 4 |
| 2021 | Modeling human planning in a life-like search-and-rescue mission
Zhutian Yang, Marta Kryven, Howard E. Shrobe, Josh Tenenbaum |
CogSci | 4 |
| 2021 | Seeing in the dark: Testing deep neural network and analysis-by-synthesis accounts of 3D shape perception with highly degraded images
Hakan Yilmaz, Gargi Singh, Bernhard Egger 0001, Josh Tenenbaum, Ilker Yildirim |
CogSci | 4 |
| 2021 | Unpacking the computations of human spatial search under uncertainty: noisy utility maximization, discounting, and probability warping
Suhyoun Yu, Marta Kryven, Josh Tenenbaum, Max Kleiman-Weiner |
CogSci | 3 |
| 2021 | Identity-Expression Ambiguity in 3D Morphable Face Modelsabstract3D Morphable Models are a class of generative models commonly used to model faces. They are typically applied to ill-posed problems such as 3D reconstruction from 2D data. Several ambiguities in this problem's image formation process have been studied explicitly. We demonstrate that nonorthogonality of the variation in identity and expression can cause identity-expression ambiguity in 3D Morphable Models, and that in practice expression and identity are far from orthogonal and can explain each other surprisingly well. Whilst previously reported ambiguities only arise in an inverse rendering setting, identity-expression ambiguity emerges in the 3D shape generation process itself. We demonstrate this effect with 3D shapes directly as well as through an inverse rendering task, and use two popular models built from high quality 3D scans as well as a model built from a large collection of 2D images and videos. We explore this issue's implications for inverse rendering and observe that it cannot be resolved by a purely statistical prior on identity and expression deformations. Bernhard Egger 0001, Skylar Sutherland, Safa C. Medin, Josh Tenenbaum |
FG | 4 |
| 2021 | Neural Radiance Flow for 4D View Synthesis and Video ProcessingabstractWe present a method, Neural Radiance Flow (NeRFlow), to learn a 4D spatial-temporal representation of a dynamic scene from a set of RGB images. Key to our approach is the use of a neural implicit representation that learns to capture the 3D occupancy, radiance, and dynamics of the scene. By enforcing consistency across different modalities, our representation enables multi-view rendering in diverse dynamic scenes, including water pouring, robotic interaction, and real images, outperforming state-of-the-art methods for spatial-temporal view synthesis. Our approach works even when being provided only a single monocular real video. We further demonstrate that the learned representation can serve as an implicit scene prior, enabling video processing tasks such as image super-resolution and de-noising without any additional supervision. Yilun Du, Hong-Xing Yu, Josh Tenenbaum, Jiajun Wu 0001 |
ICCV | 4 |
| 2021 | Learning with AMIGo: Adversarially Motivated Intrinsic Goals
Andres Campero, Roberta Raileanu, Heinrich Küttler, Josh Tenenbaum, Tim Rocktäschel, Edward Grefenstette |
ICLR | 4 |
| 2021 | Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning
Zhenfang Chen, Jiayuan Mao, Jiajun Wu 0001, Kwan-Yee Kenneth Wong, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 5 |
| 2021 | Unsupervised Discovery of 3D Physical Objects from Video
Yilun Du, Kevin A. Smith 0001, Tomer D. Ullman, Josh Tenenbaum, Jiajun Wu 0001 |
ICLR | 4 |
| 2021 | PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable Physics
Zhiao Huang, Yuanming Hu, Tao Du 0001, Hao Su 0001, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 6 |
| 2021 | Learning Task Decomposition with Ordered Memory Policy Network
Yikang Shen, Aaron C. Courville, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 5 |
| 2021 | Representing Partial Programs with Blended Abstract Semantics
Maxwell I. Nye, Yewen Pu, Matthew Bowers, Jacob Andreas, Josh Tenenbaum, Armando Solar-Lezama |
ICLR | 5 |
| 2021 | Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
Xavier Puig, Tianmin Shu, Shuang Li 0013, Yuan-Hong Liao, Josh Tenenbaum, Sanja Fidler, Antonio Torralba 0001 |
ICLR | 6 |
| 2021 | A large-scale benchmark for few-shot program induction and synthesisabstractA landmark challenge for AI is to learn flexible, powerful representations from small numbers of examples. On an important class of tasks, hypotheses in the form of programs provide extreme generalization capabilities from surprisingly few examples. However, whereas large natural few-shot learning image benchmarks have spurred progress in meta-learning for deep networks, there is no comparably big, natural program-synthesis dataset that can play a similar role. This is because, whereas images are relatively easy to label from internet meta-data or annotated by non-experts, generating meaningful input-output examples for program induction has proven hard to scale. In this work, we propose a new way of leveraging unit tests and natural inputs for small programs as meaningful input-output examples for each sub-program of the overall program. This allows us to create a large-scale naturalistic few-shot program-induction benchmark and propose new challenges in this domain. The evaluation of multiple program induction and synthesis algorithms points to shortcomings of current methods and suggests multiple avenues for future work. Ferran Alet, Javier Lopez-Contreras, James Koppel, Maxwell I. Nye, Armando Solar-Lezama, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Josh Tenenbaum |
ICML | 8 |
| 2021 | Improved Contrastive Divergence Training of Energy-Based ModelsabstractContrastive divergence is a popular method of training energy-based models, but is known to have difficulties with training stability. We propose an adaptation to improve contrastive divergence training by scrutinizing a gradient term that is difficult to calculate and is often left out for convenience. We show that this gradient term is numerically significant and in practice is important to avoid training instabilities, while being tractable to estimate. We further highlight how data augmentation and multi-scale processing can be used to improve model robustness and generation quality. Finally, we empirically evaluate stability of model architectures and show improved performance on a host of benchmarks and use cases, such as image generation, OOD detection, and compositional generation. Yilun Du, Shuang Li 0013, Josh Tenenbaum, Igor Mordatch |
ICML | 3 |
| 2021 | AGENT: A Benchmark for Core Psychological ReasoningabstractFor machine agents to successfully interact with humans in real-world settings, they will need to develop an understanding of human mental life. Intuitive psychology, the ability to reason about hidden mental variables that drive observable actions, comes naturally to people: even pre-verbal infants can tell agents from objects, expecting agents to act efficiently to achieve goals given constraints. Despite recent interest in machine agents that reason about other agents, it is not clear if such agents learn or hold the core psychology principles that drive human reasoning. Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four scenarios (goal preferences, action efficiency, unobserved constraints, and cost-reward trade-offs) that probe key concepts of core intuitive psychology. We validate AGENT with human-ratings, propose an evaluation protocol emphasizing generalization, and compare two strong baselines built on Bayesian inverse planning and a Theory of Mind neural network. Our results suggest that to pass the designed tests of core intuitive psychology at human levels, a model must acquire or have built-in representations of how agents plan, combining utility computations and core knowledge of objects and physics. Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan 0001, Kevin A. Smith 0001, Shari Liu, Dan Gutfreund, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
ICML | 8 |
| 2021 | Leveraging Language to Learn Program Abstractions and Search HeuristicsabstractInductive program synthesis, or inferring programs from examples of desired behavior, offers a general paradigm for building interpretable, robust, andgeneralizable machine learning systems. Effective program synthesis depends on two key ingredients: a strong library of functions from which to build programs, and an efficient search strategy for finding programs that solve a given task. We introduce LAPS (Language for Abstraction and Program Search), a technique for using natural language annotations to guide joint learning of libraries and neurally-guided search models for synthesis. When integrated into a state-of-the-art library learning system (DreamCoder), LAPS produces higher-quality libraries and improves search efficiency and generalization on three domains {–} string editing, image composition, and abstract reasoning about scenes {–} even when no natural language hints are available at test time. Catherine Wong, Kevin Ellis, Josh Tenenbaum, Jacob Andreas |
ICML | 3 |
| 2021 | Temporal and Object Quantification NetworksabstractWe present Temporal and Object Quantification Networks (TOQ-Nets), a new class of neuro-symbolic networks with a structural bias that enables them to learn to recognize complex relational-temporal events. This is done by including reasoning layers that implement finite-domain quantification over objects and time. The structure allows them to generalize directly to input instances with varying numbers of objects in temporal sequences of varying lengths. We evaluate TOQ-Nets on input domains that require recognizing event-types in terms of complex temporal relational patterns. We demonstrate that TOQ-Nets can generalize from small amounts of data to scenarios containing more objects than were present during training and to temporal warpings of input sequences. Jiayuan Mao, Zhezheng Luo, Chuang Gan 0001, Josh Tenenbaum, Jiajun Wu 0001, Leslie Pack Kaelbling, Tomer D. Ullman |
IJCAI | 4 |
| 2021 | OPEn: An Open-ended Physics Environment for Learning Without a TaskabstractHumans have mental models that allow them to plan, experiment, and reason in the physical world. How should an intelligent agent go about learning such models? In this paper, we will study if models of the world learned in an open-ended physics environment, without any specific tasks, can be reused for downstream physics reasoning tasks. To this end, we build a benchmark Open-ended Physics Environment (OPEn) and also design several tasks to test learning representations in this environment explicitly. This setting reflects the conditions in which real agents (i.e. rolling robots) find themselves, where they may be placed in a new kind of environment and must adapt without any teacher to tell them how this environment works. This setting is challenging because it requires solving an exploration problem in addition to a model building and representation learning problem. We test several existing RL-based exploration methods on this benchmark and find that an agent using unsupervised contrastive learning for representation learning, and impact-driven learning for exploration, achieved the best results. However, all models still fall short in sample efficiency when transferring to the downstream tasks. We expect that OPEn will encourage the development of novel rolling robot agents that can build reusable mental models of the world that facilitate many tasks. Chuang Gan 0001, Abhishek Bhandwaldar, Antonio Torralba 0001, Josh Tenenbaum, Phillip Isola |
IROS | 4 |
| 2021 | Learning Symbolic Operators for Task and Motion PlanningabstractRobotic planning problems in hybrid state and action spaces can be solved by integrated task and motion planners (TAMP) that handle the complex interaction between motion-level decisions and task-level plan feasibility. TAMP approaches rely on domain-specific symbolic operators to guide the task-level search, making planning efficient. In this work, we formalize and study the problem of operator learning for TAMP. Central to this study is the view that operators define a lossy abstraction of the transition model of a domain. We then propose a bottom-up relational learning method for operator learning and show how the learned operators can be used for planning in a TAMP system. Experimentally, we provide results in three domains, including long-horizon robotic planning tasks. We find our approach to substantially outperform several baselines, including three graph neural network-based model-free approaches from the recent literature. Video: https://youtu.be/iVfpX9BpBRo. Code: https://git.io/JCT0g Tom Silver, Rohan Chitnis, Josh Tenenbaum, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 3 |
| 2021 | Dynamic Modeling of Hand-Object Interactions via Tactile SensingabstractTactile sensing is critical for humans to perform everyday tasks. While significant progress has been made in analyzing object grasping from vision, it remains unclear how we can utilize tactile sensing to reason about and model the dynamics of hand-object interactions. In this work, we employ a high-resolution tactile glove to perform four different interactive activities on a diversified set of objects. We propose a framework aiming at predicting the 3d locations of both the hand and the object purely from the touch data by combining a predictive model and a contrastive learning module. This framework can reason about the interaction patterns from the tactile data, hallucinate the changes in the environment, esti-mate the uncertainty of the prediction, and generalize to unseen objects. We also provide detailed ablation studies regarding different system designs as well as visualizations of the predicted trajectories. This work takes a step on dynamics modeling in hand-object interactions from dense tactile sensing, which opens the door for future applications in activity learning, human-computer interactions, and imitation learning for robotics. Yunzhu Li, Yiyue Luo, Wan Shou, Michael Foshey, Junchi Yan, Josh Tenenbaum, Wojciech Matusik, Antonio Torralba 0001 |
IROS | 7 |
| 2021 | Noether Networks: meta-learning useful conserved quantitiesabstractProgress in machine learning (ML) stems from a combination of data availability, computational resources, and an appropriate encoding of inductive biases. Useful biases often exploit symmetries in the prediction problem, such as convolutional networks relying on translation equivariance. Automatically discovering these useful symmetries holds the potential to greatly improve the performance of ML systems, but still remains a challenge. In this work, we focus on sequential prediction problems and take inspiration from Noether's theorem to reduce the problem of finding inductive biases to meta-learning useful conserved quantities. We propose Noether Networks: a new type of architecture where a meta-learned conservation loss is optimized inside the prediction function. We show, theoretically and experimentally, that Noether Networks improve prediction quality, providing a general framework for discovering inductive biases in sequential problems. Ferran Alet, Dylan Doblar, Allan Zhou, Josh Tenenbaum, Kenji Kawaguchi, Chelsea Finn |
NeurIPS | 4 |
| 2021 | Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageabstractIn this work, we propose a unified framework, called Visual Reasoning with Differ-entiable Physics (VRDP), that can jointly learn visual concepts and infer physics models of objects and their interactions from videos and language. This is achieved by seamlessly integrating three components: a visual perception module, a concept learner, and a differentiable physics engine. The visual perception module parses each video frame into object-centric trajectories and represents them as latent scene representations. The concept learner grounds visual concepts (e.g., color, shape, and material) from these object-centric representations based on the language, thus providing prior knowledge for the physics engine. The differentiable physics model, implemented as an impulse-based differentiable rigid-body simulator, performs differentiable physical simulation based on the grounded concepts to infer physical properties, such as mass, restitution, and velocity, by fitting the simulated trajectories into the video observations. Consequently, these learned concepts and physical models can explain what we have seen and imagine what is about to happen in future and counterfactual scenarios. Integrating differentiable physics into the dynamic reasoning framework offers several appealing benefits. More accurate dynamics prediction in learned physics models enables state-of-the-art performance on both synthetic and real-world benchmarks while still maintaining high transparency and interpretability; most notably, VRDP improves the accuracy of predictive and counterfactual questions by 4.5% and 11.5% compared to its best counterpart. VRDP is also highly data-efficient: physical parameters can be optimized from very few videos, and even a single video can be sufficient. Finally, with all physical parameters inferred, VRDP can quickly learn new concepts from a few examples. Mingyu Ding, Zhenfang Chen, Tao Du 0001, Ping Luo 0002, Josh Tenenbaum, Chuang Gan 0001 |
NeurIPS | 5 |
| 2021 | Learning Signal-Agnostic Manifolds of Neural FieldsabstractDeep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are generally modality-dependent, requiring custom architectures and objectives to process different classes of signals. We leverage neural fields to capture the underlying structure in image, shape, audio and cross-modal audiovisual domains in a modality-independent manner. We cast our task as one of learning a manifold, where we aim to infer a low-dimensional, locally linear subspace in which our data resides. By enforcing coverage of the manifold, local linearity, and local isometry, our model -- dubbed GEM -- learns to capture the underlying structure of datasets across modalities. We can then travel along linear regions of our manifold to obtain perceptually consistent interpolations between samples, and can further use GEM to recover points on our manifold and glean not only diverse completions of input images, but cross-modal hallucinations of audio or image signals. Finally, we show that by walking across the underlying manifold of GEM, we may generate new samples in our signal domains. Yilun Du, Katie Collins, Josh Tenenbaum, Vincent Sitzmann |
NeurIPS | 3 |
| 2021 | Unsupervised Learning of Compositional Energy ConceptsabstractHumans are able to rapidly understand scenes by utilizing concepts extracted from prior experience. Such concepts are diverse, and include global scene descriptors, such as the weather or lighting, as well as local scene descriptors, such as the color or size of a particular object. So far, unsupervised discovery of concepts has focused on either modeling the global scene-level or the local object-level factors of variation, but not both. In this work, we propose COMET, which discovers and represents concepts as separate energy functions, enabling us to represent both global concepts as well as objects under a unified framework. COMET discovers energy functions through recomposing the input image, which we find captures independent factors without additional supervision. Sample generation in COMET is formulated as an optimization process on underlying energy functions, enabling us to generate images with permuted and composed concepts. Finally, discovered visual concepts in COMET generalize well, enabling us to compose concepts between separate modalities of images as well as with other concepts discovered by a separate instance of COMET trained on a different dataset. Code and data available at https://energy-based-model.github.io/comet/. Yilun Du, Shuang Li 0013, Yash Sharma 0001, Josh Tenenbaum, Igor Mordatch |
NeurIPS | 4 |
| 2021 | 3DP3: 3D Scene Perception via Probabilistic ProgrammingabstractWe present 3DP3, a framework for inverse graphics that uses inference in a structured generative model of objects, scenes, and images. 3DP3 uses (i) voxel models to represent the 3D shape of objects, (ii) hierarchical scene graphs to decompose scenes into objects and the contacts between them, and (iii) depth image likelihoods based on real-time graphics. Given an observed RGB-D image, 3DP3's inference algorithm infers the underlying latent 3D scene, including the object poses and a parsimonious joint parametrization of these poses, using fast bottom-up pose proposals, novel involutive MCMC updates of the scene graph structure, and, optionally, neural object detectors and pose estimators. We show that 3DP3 enables scene understanding that is aware of 3D shape, occlusion, and contact structure. Our results demonstrate that 3DP3 is more accurate at 6DoF object pose estimation from real images than deep learning baselines and shows better generalization to challenging scenes with novel viewpoints, contact, and partial observability. Nishad Gothoskar, Marco F. Cusumano-Towner, Ben Zinberg, Matin Ghavamizadeh, Falk Pollok, Austin Garrett, Josh Tenenbaum, Dan Gutfreund, Vikash Mansinghka 0001 |
NeurIPS | 7 |
| 2021 | PTR: A Benchmark for Part-based Conceptual, Relational, and Physical ReasoningabstractA critical aspect of human visual perception is the ability to parse visual scenes into individual objects and further into object parts, forming part-whole hierarchies. Such composite structures could induce a rich set of semantic concepts and relations, thus playing an important role in the interpretation and organization of visual signals as well as for the generalization of visual perception and reasoning. However, existing visual reasoning benchmarks mostly focus on objects rather than parts. Visual reasoning based on the full part-whole hierarchy is much more challenging than object-centric reasoning due to finer-grained concepts, richer geometry relations, and more complex physics. Therefore, to better serve for part-based conceptual, relational and physical reasoning, we introduce a new large-scale diagnostic visual reasoning dataset named PTR. PTR contains around 80k RGBD synthetic images with ground truth object and part level annotations regarding semantic instance segmentation, color attributes, spatial and geometric relationships, and certain physical properties such as stability. These images are paired with 800k machine-generated questions covering various types of reasoning types, making them a good testbed for visual reasoning models. We examine several state-of-the-art visual reasoning models on this dataset and observe that they still make many surprising mistakes in situations where humans can easily infer the correct answer. We believe this dataset will open up new opportunities for part-based reasoning. PTR dataset and baseline models are publicly available. Yining Hong, Li Yi 0001, Josh Tenenbaum, Antonio Torralba 0001, Chuang Gan 0001 |
NeurIPS | 3 |
| 2021 | Learning to Compose Visual RelationsabstractThe visual world around us can be described as a structured set of objects and their associated relations. An image of a room may be conjured given only the description of the underlying objects and their associated relations. While there has been significant work on designing deep neural networks which may compose individual objects together, less work has been done on composing the individual relations between objects. A principal difficulty is that while the placement of objects is mutually independent, their relations are entangled and dependent on each other. To circumvent this issue, existing works primarily compose relations by utilizing a holistic encoder, in the form of text or graphs. In this work, we instead propose to represent each relation as an unnormalized density (an energy-based model), enabling us to compose separate relations in a factorized manner. We show that such a factorized decomposition allows the model to both generate and edit scenes that have multiple sets of relations more faithfully. We further show that decomposition enables our model to effectively understand the underlying relational scene structure. Nan Liu 0010, Shuang Li 0013, Yilun Du, Josh Tenenbaum, Antonio Torralba 0001 |
NeurIPS | 4 |
| 2021 | Grammar-Based Grounded Lexicon LearningabstractWe present Grammar-Based Grounded Language Learning (G2L2), a lexicalist approach toward learning a compositional and grounded meaning representation of language from grounded data, such as paired images and texts. At the core of G2L2 is a collection of lexicon entries, which map each word to a tuple of a syntactic type and a neuro-symbolic semantic program. For example, the word shiny has a syntactic type of adjective; its neuro-symbolic semantic program has the symbolic form $\lambda x.\textit{filter}(x, \textbf{SHINY})$, where the concept SHINY is associated with a neural network embedding, which will be used to classify shiny objects. Given an input sentence, G2L2 first looks up the lexicon entries associated with each token. It then derives the meaning of the sentence as an executable neuro-symbolic program by composing lexical meanings based on syntax. The recovered meaning programs can be executed on grounded inputs. To facilitate learning in an exponentially-growing compositional space, we introduce a joint parsing and expected execution algorithm, which does local marginalization over derivations to reduce the training time. We evaluate G2L2 on two domains: visual reasoning and language-driven navigation. Results show that G2L2 can generalize from small amounts of data to novel compositions of words. Jiayuan Mao, Freda Shi, Jiajun Wu 0001, Roger Levy, Josh Tenenbaum |
NeurIPS | 5 |
| 2021 | Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic ReasoningabstractHuman reasoning can be understood as an interplay between two systems: the intuitive and associative ("System 1") and the deliberative and logical ("System 2"). Neural sequence models---which have been increasingly successful at performing complex, structured tasks---exhibit the advantages and failure modes of System 1: they are fast and learn patterns from data, but are often inconsistent and incoherent. In this work, we seek a lightweight, training-free means of improving existing System 1-like sequence models by adding System 2-inspired logical reasoning. We explore several variations on this theme in which candidate generations from a neural sequence model are examined for logical consistency by a symbolic reasoning module, which can either accept or reject the generations. Our approach uses neural inference to mediate between the neural System 1 and the logical System 2. Results in robust story generation and grounded instruction-following show that this approach can increase the coherence and accuracy of neurally-based generations. Maxwell I. Nye, Michael Henry Tessler, Josh Tenenbaum, Brenden M. Lake |
NeurIPS | 3 |
| 2021 | Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingabstractInferring representations of 3D scenes from 2D observations is a fundamental problem of computer graphics, computer vision, and artificial intelligence. Emerging 3D-structured neural scene representations are a promising approach to 3D scene understanding. In this work, we propose a novel neural scene representation, Light Field Networks or LFNs, which represent both geometry and appearance of the underlying 3D scene in a 360-degree, four-dimensional light field parameterized via a neural implicit representation. Rendering a ray from an LFN requires only a single network evaluation, as opposed to hundreds of evaluations per ray for ray-marching or volumetric based renderers in 3D-structured neural scene representations. In the setting of simple scenes, we leverage meta-learning to learn a prior over LFNs that enables multi-view consistent light field reconstruction from as little as a single image observation. This results in dramatic reductions in time and memory complexity, and enables real-time rendering. The cost of storing a 360-degree light field via an LFN is two orders of magnitude lower than conventional methods such as the Lumigraph. Utilizing the analytical differentiability of neural implicit representations and a novel parameterization of light space, we further demonstrate the extraction of sparse depth maps from LFNs. Vincent Sitzmann, Semon Rezchikov, William T. Freeman, Josh Tenenbaum, Frédo Durand |
NeurIPS | 4 |
| 2021 | A Bayesian-Symbolic Approach to Reasoning and Learning in Intuitive PhysicsabstractHumans can reason about intuitive physics in fully or partially observed environments even after being exposed to a very limited set of observations. This sample-efficient intuitive physical reasoning is considered a core domain of human common sense knowledge. One hypothesis to explain this remarkable capacity, posits that humans quickly learn approximations to the laws of physics that govern the dynamics of the environment. In this paper, we propose a Bayesian-symbolic framework (BSP) for physical reasoning and learning that is close to human-level sample-efficiency and accuracy. In BSP, the environment is represented by a top-down generative model of entities, which are assumed to interact with each other under unknown force laws over their latent and observed properties. BSP models each of these entities as random variables, and uses Bayesian inference to estimate their unknown properties. For learning the unknown forces, BSP leverages symbolic regression on a novel grammar of Newtonian physics in a bilevel optimization setup. These inference and regression steps are performed in an iterative manner using expectation-maximization, allowing BSP to simultaneously learn force laws while maintaining uncertainty over entity properties. We show that BSP is more sample-efficient compared to neural alternatives on controlled synthetic datasets, demonstrate BSP's applicability to real-world common sense scenes and study BSP's performance on tasks previously used to study human physical reasoning. Kai Xu 0016, Akash Srivastava, Dan Gutfreund, Felix Sosa, Tomer D. Ullman, Josh Tenenbaum, Charles Sutton |
NeurIPS | 6 |
| 2021 | DreamCoder: bootstrapping inductive program synthesis with wake-sleep library learningabstractWe present a system for inductive program synthesis called DreamCoder, which inputs a corpus of synthesis problems each specified by one or a few examples, and automatically derives a library of program components and a neural search policy that can be used to efficiently solve other similar synthesis problems. The library and search policy bootstrap each other iteratively through a variant of "wake-sleep" approximate Bayesian learning. A new refactoring algorithm based on E-graph matching identifies common sub-components across synthesized programs, building a progressively deepening library of abstractions capturing the structure of the input domain. We evaluate on eight domains including classic program synthesis areas and AI tasks such as planning, inverse graphics, and equation discovery. We show that jointly learning the library and neural search policy leads to solving more problems, and solving them more quickly. Kevin Ellis, Catherine Wong, Maxwell I. Nye, Mathias Sablé-Meyer, Lucas Morales, Luke B. Hewitt, Luc Cary, Armando Solar-Lezama, Josh Tenenbaum |
PLDI | 9 |
| 2021 | World model learning and inferenceabstractUnderstanding information processing in the brain-and creating general-purpose artificial intelligence-are long-standing aspirations of scientists and engineers worldwide. The distinctive features of human intelligence are high-level cognition and control in various interactions with the world including the self, which are not defined in advance and are vary over time. The challenge of building human-like intelligent machines, as well as progress in brain science and behavioural analyses, robotics, and their associated theoretical formalisations, speaks to the importance of the world-model learning and inference. In this article, after briefly surveying the history and challenges of internal model learning and probabilistic learning, we introduce the free energy principle, which provides a useful framework within which to consider neuronal computation and probabilistic world models. Next, we showcase examples of human behaviour and cognition explained under that principle. We then describe symbol emergence in the context of probabilistic modelling, as a topic at the frontiers of cognitive robotics. Lastly, we review recent progress in creating human-like intelligence by using novel probabilistic programming languages. The striking consensus that emerges from these studies is that probabilistic descriptions of learning and inference are powerful and effective ways to create human-like artificial intelligent machines and to understand intelligence in the context of how humans interact with their world. Karl J. Friston, Rosalyn J. Moran, Yukie Nagai, Tadahiro Taniguchi, Hiroaki Gomi, Josh Tenenbaum |
Neural Networks | 6 |
| 2020 | Few-Shot Bayesian Imitation Learning with Logical Program PoliciesabstractHumans can learn many novel tasks from a very small number (1–5) of demonstrations, in stark contrast to the data requirements of nearly tabula rasa deep learning methods. We propose an expressive class of policies, a strong but general prior, and a learning algorithm that, together, can learn interesting policies from very few examples. We represent policies as logical combinations of programs drawn from a domain-specific language (DSL), define a prior over policies with a probabilistic grammar, and derive an approximate Bayesian inference algorithm to learn policies from demonstrations. In experiments, we study six strategy games played on a 2D grid with one shared DSL. After a few demonstrations of each game, the inferred policies generalize to new game instances that differ substantially from the demonstrations. Our policy learning is 20–1,000x more data efficient than convolutional and fully convolutional policy learning and many orders of magnitude more computationally efficient than vanilla program induction. We argue that the proposed method is an apt choice for tasks that have scarce training data and feature significant, structured variation between task instances. Tom Silver, Kelsey R. Allen, Alex K. Lew, Leslie Pack Kaelbling, Josh Tenenbaum |
AAAI | 5 |
| 2020 | The Origins of Common Sense in Humans and Machines
Kevin A. Smith 0001, Eliza Kosoy, Alison Gopnik, Deepak Pathak, Alan Fern, Josh Tenenbaum, Tomer D. Ullman |
CogSci | 6 |
| 2020 | The fine structure of surprise in intuitive physics: when, why, and how much?
Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
CogSci | 6 |
| 2020 | Abstract strategy learning underlies flexible transfer in physical problem solving
Kelsey R. Allen, Kevin A. Smith 0001, Ulyana Piterbarg, Josh Tenenbaum |
CogSci | 5 |
| 2020 | Perceiving unseen objects
Katie Collins, Josh Tenenbaum, Kevin A. Smith 0001 |
CogSci | 2 |
| 2020 | Inverse Rendering Best Explains Face Perception Under Extreme Illuminations
Bernhard Egger 0001, Max H. Siegel, Riya Arora, Amir Arsalan Soltani, Ilker Yildirim, Josh Tenenbaum |
CogSci | 6 |
| 2020 | Jointly learning motion verbs and frame semantics from natural language and grounded scenes
Jon Gauthier, Jiayuan Mao, Tianmin Shu, Roger Levy, Josh Tenenbaum |
CogSci | 5 |
| 2020 | Leveraging Unstructured Statistical Knowledge in a Probabilistic Language of Thought
Alexander K. Lew, Michael Henry Tessler, Vikash Mansinghka 0001, Josh Tenenbaum |
CogSci | 4 |
| 2020 | A Task and Motion Approach to the Development of Planning
João Loula, Kelsey R. Allen, Josh Tenenbaum |
CogSci | 3 |
| 2020 | Humans measure algorithmic complexity to guide engagement with event sequences
Gal Raz, Setayesh Radkani, Josh Tenenbaum, Rebecca Saxe |
CogSci | 3 |
| 2020 | Learning sequential patterns from graphical programs
Anselm Rothe, Eric Schulz, Mathias Sablé-Meyer, Josh Tenenbaum, Azzurra Ruggeri |
CogSci | 4 |
| 2020 | Social Learning with Sparse Belief Samples
Rabih Salhab, Amir Ajorlou, Ali Jadbabaie, Josh Tenenbaum |
CogSci | 4 |
| 2020 | Adventures in Flatland: Perceiving Social Interactions Under Physical Dynamics
Tianmin Shu, Marta Kryven, Tomer D. Ullman, Josh Tenenbaum |
CogSci | 4 |
| 2020 | Learning a Generative Model of Human Faces Through Inverse Rendering
Skylar Sutherland, Bernhard Egger 0001, Josh Tenenbaum |
CogSci | 3 |
| 2020 | How many observations is one generic worth?
Michael Henry Tessler, Sophie Bridgers, Josh Tenenbaum |
CogSci | 3 |
| 2020 | Too many cooks: Coordinating multi-agent collaboration through inverse planning
Sarah A. Wu, Rose E. Wang, James A. Evans, Josh Tenenbaum, David C. Parkes, Max Kleiman-Weiner |
CogSci | 4 |
| 2020 | Music Gesture for Visual Sound SeparationabstractRecent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited abilities to find the correlations between audio signals and visual points, especially when separating multiple instruments of the same types, such as multiple violins in a scene. To address this, we propose ``Music Gesture," a keypoint-based structured representation to explicitly model the body and finger movements of musicians when they perform music. We first adopt a context-aware graph network to integrate visual semantic context with body dynamics and then apply an audio-visual fusion model to associate body movements with the corresponding audio signals. Experimental results on three music performance datasets show: 1) strong improvements upon benchmark metrics for hetero-musical separation tasks (i.e. different instruments); 2) new ability for effective homo-musical separation for piano, flute, and trumpet duets, which to our best knowledge has never been achieved with alternative methods. Chuang Gan 0001, Deng Huang, Hang Zhao 0021, Josh Tenenbaum, Antonio Torralba 0001 |
CVPR | 4 |
| 2020 | Perspective Plane Program Induction From a Single ImageabstractWe study the inverse graphics problem of inferring a holistic representation for natural images. Given an input image, our goal is to induce a neuro-symbolic, program-like representation that jointly models camera poses, object locations, and global scene structures. Such high-level, holistic scene representations further facilitate low-level image manipulation tasks such as inpainting. We formulate this problem as jointly finding the camera pose and scene structure that best describe the input image. The benefits of such joint inference are two-fold: scene regularity serves as a new cue for perspective correction, and in turn, correct perspective correction leads to a simplified scene structure, similar to how the correct shape leads to the most regular texture in shape from texture. Our proposed framework, Perspective Plane Program Induction (P3I), combines search-based and gradient-based algorithms to efficiently solve the problem. P3I outperforms a set of baselines on a collection of Internet images, across tasks including camera pose estimation, global structure inference, and down-stream image manipulation tasks. Jiayuan Mao, Xiuming Zhang, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
CVPR | 5 |
| 2020 | End-to-End Optimization of Scene LayoutabstractWe propose an end-to-end variational generative model for scene layout synthesis conditioned on scene graphs. Unlike unconditional scene layout generation, we use scene graphs as an abstract but general representation to guide the synthesis of diverse scene layouts that satisfy relationships included in the scene graph. This gives rise to more flexible control over the synthesis process, allowing various forms of inputs such as scene layouts extracted from sentences or inferred from a single color image. Using our conditional layout synthesizer, we can generate various layouts that share the same structure of the input example. In addition to this conditional generation design, we also integrate a differentiable rendering module that enables layout refinement using only 2D projections of the scene. Given a depth and a semantics map, the differentiable rendering module enables optimizing over the synthesized layout to fit the given input in an analysis-by-synthesis fashion. Experiments suggest that our model achieves higher accuracy and diversity in conditional scene synthesis and allows exemplar-based scene generation from various input forms. Andrew Luo 0001, Zhoutong Zhang, Jiajun Wu 0001, Josh Tenenbaum |
CVPR | 4 |
| 2020 | A Morphable Face Albedo ModelabstractIn this paper, we bring together two divergent strands of research: photometric face capture and statistical 3D face appearance modelling. We propose a novel lightstage capture and processing pipeline for acquiring ear-to-ear, truly intrinsic diffuse and specular albedo maps that fully factor out the effects of illumination, camera and geometry. Using this pipeline, we capture a dataset of 50 scans and combine them with the only existing publicly available albedo dataset (3DRFE) of 23 scans. This allows us to build the first morphable face albedo model. We believe this is the first statistical analysis of the variability of facial specular albedo maps. This model can be used as a plug in replacement for the texture model of the Basel Face Model and we make our new albedo model publicly available. We ensure careful spectral calibration such that our model is built in a linear sRGB space, suitable for inverse rendering of images taken by typical cameras. We demonstrate our model in a state of the art analysis-by-synthesis 3DMM fitting pipeline, are the first to integrate specular map estimation and outperform the Basel Face Model in albedo reconstruction. William A. P. Smith, Alassane Seck, Hannah M. Dee, Bernard Tiddeman, Josh Tenenbaum, Bernhard Egger 0001 |
CVPR | 5 |
| 2020 | Probabilistic Video Prediction From Noisy Data With a Posterior ConfidenceabstractWe study a new research problem of probabilistic future frames prediction from a sequence of noisy inputs, which is useful because it is difficult to guarantee the quality of input frames in practical spatiotemporal prediction applications. It is also challenging because it involves two levels of uncertainty: the perceptual uncertainty from noisy observations and the dynamics uncertainty in forward modeling. In this paper, we propose to tackle this problem with an end-to-end trainable model named Bayesian Predictive Network (BP-Net). Unlike previous work in stochastic video prediction that assumes spatiotemporal coherence and therefore fails to deal with perceptual uncertainty, BP-Net models both levels of uncertainty in an integrated framework. Furthermore, unlike previous work that can only provide unsorted estimations of future frames, BP-Net leverages a differentiable sequential importance sampling (SIS) approach to make future predictions based on the inference of underlying physical states, thereby providing sorted prediction candidates in accordance with the SIS importance weights, i.e., the confidences. Our experiment results demonstrate that BP-Net remarkably outperforms existing approaches on predicting future frames from noisy data. Yunbo Wang, Jiajun Wu 0001, Mingsheng Long, Josh Tenenbaum |
CVPR | 4 |
| 2020 | Foley Music: Learning to Generate Music from Videos
Chuang Gan 0001, Deng Huang, Peihao Chen, Josh Tenenbaum, Antonio Torralba 0001 |
ECCV (11) | 4 |
| 2020 | Rethinking Few-Shot Image Classification: A Good Embedding is All You Need?
Yonglong Tian, Yue Wang 0041, Dilip Krishnan, Josh Tenenbaum, Phillip Isola |
ECCV (14) | 4 |
| 2020 | CLEVRER: Collision Events for Video Representation and Reasoning
Kexin Yi, Chuang Gan 0001, Yunzhu Li, Pushmeet Kohli, Jiajun Wu 0001, Antonio Torralba 0001, Josh Tenenbaum |
ICLR | 7 |
| 2020 | Deep Audio Priors Emerge From Harmonic Convolutional Networks
Zhoutong Zhang, Chuang Gan 0001, Jiajun Wu 0001, Josh Tenenbaum, Antonio Torralba 0001, William T. Freeman |
ICLR | 5 |
| 2020 | Visual Grounding of Learned Physical ModelsabstractHumans intuitively recognize objects’ physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environments, while intrinsic to humans, remain challenging to state-of-the-art computational models. In this work, we present a neural model that simultaneously reasons about physics and makes future predictions based on visual and dynamics priors. The visual prior predicts a particle-based representation of the system from visual observations. An inference module operates on those particles, predicting and refining estimates of particle locations, object states, and physical parameters, subject to the constraints imposed by the dynamics prior, which we refer to as visual grounding. We demonstrate the effectiveness of our method in environments involving rigid objects, deformable materials, and fluids. Experiments show that our model can infer the physical properties within a few observations, which allows the model to quickly adapt to unseen scenarios and make accurate predictions into the future. Yunzhu Li, Toru Lin, Kexin Yi, Daniel Bear, Dan Yamins, Jiajun Wu 0001, Josh Tenenbaum, Antonio Torralba 0001 |
ICML | 7 |
| 2020 | Look, Listen, and Act: Towards Audio-Visual Embodied NavigationabstractA crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the problem of Audio-Visual Embodied Navigation, the task of planning the shortest path from a random starting location in a scene to the sound source in an indoor environment, given only raw egocentric visual and audio sensory data. To accomplish this task, the agent is required to learn from various modalities, i.e., relating the audio signal to the visual environment. Here we describe an approach to audio-visual embodied navigation that takes advantage of both visual and audio pieces of evidence. Our solution is based on three key ideas: a visual perception mapper module that constructs its spatial memory of the environment, a sound perception module that infers the relative location of the sound source from the agent, and a dynamic path planner that plans a sequence of actions based on the audio-visual observations and the spatial memory of the environment to navigate toward the goal. Experimental results on a newly collected Visual-Audio-Room dataset using the simulated multi-modal environment demonstrate the effectiveness of our approach over several competitive baselines. Chuang Gan 0001, Yiwei Zhang 0010, Jiajun Wu 0001, Boqing Gong, Josh Tenenbaum |
ICRA | 5 |
| 2020 | Accurate Vision-based Manipulation through Contact ReasoningabstractPlanning contact interactions is one of the core challenges of many robotic tasks. Optimizing contact locations while taking dynamics into account is computationally costly and, in environments that are only partially observable, executing contact-based tasks often suffers from low accuracy. We present an approach that addresses these two challenges for the problem of vision-based manipulation. First, we propose to disentangle contact from motion optimization. Thereby, we improve planning efficiency by focusing computation on promising contact locations. Second, we use a hybrid approach for perception and state estimation that combines neural networks with a physically meaningful state representation. In simulation and real-world experiments on the task of planar pushing, we show that our method is more efficient and achieves a higher manipulation accuracy than previous vision-based approaches. Alina Kloss, Maria Bauzá 0001, Jiajun Wu 0001, Josh Tenenbaum, Alberto Rodriguez 0003, Jeannette Bohg |
ICRA | 4 |
| 2020 | DualSMC: Tunneling Differentiable Filtering and Planning under Continuous POMDPsabstractA major difficulty of solving continuous POMDPs is to infer the multi-modal distribution of the unobserved true states and to make the planning algorithm dependent on the perceived uncertainty. We cast POMDP filtering and planning problems as two closely related Sequential Monte Carlo (SMC) processes, one over the real states and the other over the future optimal trajectories, and combine the merits of these two parts in a new model named the DualSMC network. In particular, we first introduce an adversarial particle filter that leverages the adversarial relationship between its internal components. Based on the filtering results, we then propose a planning algorithm that extends the previous SMC planning approach [Piche et al., 2018] to continuous POMDPs with an uncertainty-dependent policy. Crucially, not only can DualSMC handle complex observations such as image input but also it remains highly interpretable. It is shown to be effective in three continuous POMDP domains: the floor positioning domain, the 3D light-dark navigation domain, and a modified Reacher domain. Yunbo Wang, Bo Liu 0042, Jiajun Wu 0001, Yuke Zhu, Simon S. Du, Li Fei-Fei 0001, Josh Tenenbaum |
IJCAI | 7 |
| 2020 | Learning constraint-based planning models from demonstrationsabstractHow can we learn representations for planning that are both efficient and flexible? Task and motion planning models are a good candidate, having been very successful in long-horizon planning tasks-however, they've proved challenging for learning, relying mostly on hand-coded representations. We present a framework for learning constraint-based task and motion planning models using gradient descent. Our model observes expert demonstrations of a task and decomposes them into modes-segments which specify a set of constraints on a trajectory optimization problem. We show that our model learns these modes from few demonstrations, that modes can be used to plan flexibly in different environments and to achieve different types of goals, and that the model can recombine these modes in novel ways. João Loula, Kelsey R. Allen, Tom Silver, Josh Tenenbaum |
IROS | 4 |
| 2020 | Learning Physical Graph Representations from Visual ScenesabstractConvolutional Neural Networks (CNNs) have proved exceptional at learning representations for visual object categorization. However, CNNs do not explicitly encode objects, parts, and their physical properties, which has limited CNNs' success on tasks that require structured understanding of visual scenes. To overcome these limitations, we introduce the idea of ``Physical Scene Graphs'' (PSGs), which represent scenes as hierarchical graphs, with nodes in the hierarchy corresponding intuitively to object parts at different scales, and edges to physical connections between parts. Bound to each node is a vector of latent attributes that intuitively represent object properties such as surface shape and texture. We also describe PSGNet, a network architecture that learns to extract PSGs by reconstructing scenes through a PSG-structured bottleneck. PSGNet augments standard CNNs by including: recurrent feedback connections to combine low and high-level image information; graph pooling and vectorization operations that convert spatially-uniform feature maps into object-centric graph structures; and perceptual grouping principles to encourage the identification of meaningful scene elements. We show that PSGNet outperforms alternative self-supervised scene representation algorithms at scene segmentation tasks, especially on complex real-world images, and generalizes well to unseen object types and scene arrangements. PSGNet is also able learn from physical motion, enhancing scene estimates even for static images. We present a series of ablation studies illustrating the importance of each component of the PSGNet architecture, analyses showing that learned latent attributes capture intuitive scene properties, and illustrate the use of PSGs for compositional scene inference. Daniel Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li Fei-Fei 0001, Jiajun Wu 0001, Josh Tenenbaum, Dan Yamins |
NeurIPS | 10 |
| 2020 | Multi-Plane Program Induction with 3D Box PriorsabstractWe consider two important aspects in understanding and editing images: modeling regular, program-like texture or patterns in 2D planes, and 3D posing of these planes in the scene. Unlike prior work on image-based program synthesis, which assumes the image contains a single visible 2D plane, we present Box Program Induction (BPI), which infers a program-like scene representation that simultaneously models repeated structure on multiple 2D planes, the 3D position and orientation of the planes, and camera parameters, all from a single image. Our model assumes a box prior, i.e., that the image captures either an inner view or an outer view of a box in 3D. It uses neural networks to infer visual cues such as vanishing points, wireframe lines to guide a search-based algorithm to find the program that best explains the image. Such a holistic, structured scene representation enables 3D-aware interactive image editing operations such as inpainting missing pixels, changing camera parameters, and extrapolate the image contents. Jiayuan Mao, Xiuming Zhang, William T. Freeman, Josh Tenenbaum, Noah Snavely, Jiajun Wu 0001 |
NeurIPS | 5 |
| 2020 | Learning Compositional Rules via Neural Program SynthesisabstractMany aspects of human reasoning, including language, require learning rules from very little data. Humans can do this, often learning systematic rules from very few examples, and combining these rules to form compositional rule-based systems. Current neural architectures, on the other hand, often fail to generalize in a compositional manner, especially when evaluated in ways that vary systematically from training. In this work, we present a neuro-symbolic model which learns entire rule systems from a small set of examples. Instead of directly predicting outputs from inputs, we train our model to induce the explicit system of rules governing a set of previously seen examples, drawing upon techniques from the neural program synthesis literature. Our rule-synthesis approach outperforms neural meta-learning techniques in three domains: an artificial instruction-learning domain used to evaluate human learning, the SCAN challenge datasets, and learning rule-based translations of number words into integers for a wide range of human languages. Maxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. Lake |
NeurIPS | 3 |
| 2020 | Program Synthesis with Pragmatic CommunicationabstractProgram synthesis techniques construct or infer programs from user-provided specifications, such as input-output examples. Yet most specifications, especially those given by end-users, leave the synthesis problem radically ill-posed, because many programs may simultaneously satisfy the specification. Prior work resolves this ambiguity by using various inductive biases, such as a preference for simpler programs. This work introduces a new inductive bias derived by modeling the program synthesis task as rational communication, drawing insights from recursive reasoning models of pragmatics. Given a specification, we score a candidate program both on its consistency with the specification, and also whether a rational speaker would chose this particular specification to communicate that program. We develop efficient algorithms for such an approach when learning from input-output examples, and build a pragmatic program synthesizer over a simple grid-like layout domain. A user study finds that end-user participants communicate more effectively with the pragmatic program synthesizer over a non-pragmatic one. Yewen Pu, Kevin Ellis, Marta Kryven, Josh Tenenbaum, Armando Solar-Lezama |
NeurIPS | 4 |
| 2020 | Learning abstract structure for drawing by efficient motor program inductionabstractHumans flexibly solve new problems that differ from those previously practiced. This ability to flexibly generalize is supported by learned concepts that represent useful structure common across different problems. Here we develop a naturalistic drawing task to study how humans rapidly acquire structured prior knowledge. The task requires drawing visual figures that share underlying structure, based on a set of composable geometric rules and simple objects. We show that people spontaneously learn abstract drawing procedures that support generalization, and propose a model of how learners can discover these reusable drawing procedures. Trained in the same setting as humans, and constrained to produce efficient motor actions, this model discovers new drawing program subroutines that generalize to test figures and resemble learned features of human behavior. These results suggest that two principles guiding motor program induction in the model - abstraction (programs can reflect high-level structure that ignores figure-specific details) and compositionality (new programs are discovered by recombining previously learned programs) - are key for explaining how humans learn structured internal representations that guide flexible reasoning and learning. Lucas Yanan Tian, Kevin Ellis, Marta Kryven, Josh Tenenbaum |
NeurIPS | 4 |
| 2020 | Online Bayesian Goal Inference for Boundedly Rational Planning AgentsabstractPeople routinely infer the goals of others by observing their actions over time. Remarkably, we can do so even when those actions lead to failure, enabling us to assist others when we detect that they might not achieve their goals. How might we endow machines with similar capabilities? Here we present an architecture capable of inferring an agent’s goals online from both optimal and non-optimal sequences of actions. Our architecture models agents as boundedly-rational planners that interleave search with execution by replanning, thereby accounting for sub-optimal behavior. These models are specified as probabilistic programs, allowing us to represent and perform efficient Bayesian inference over an agent's goals and internal planning processes. To perform such inference, we develop Sequential Inverse Plan Search (SIPS), a sequential Monte Carlo algorithm that exploits the online replanning assumption of these models, limiting computation by incrementally extending inferred plans as new actions are observed. We present experiments showing that this modeling and inference architecture outperforms Bayesian inverse reinforcement learning baselines, accurately inferring goals from both optimal and non-optimal trajectories involving failure and back-tracking, while generalizing across domains with compositional structure and sparse rewards. Tan Zhi-Xuan, Jordyn L. Mann, Tom Silver, Josh Tenenbaum, Vikash Mansinghka 0001 |
NeurIPS | 4 |
| 2020 | Learning to learn generative programs with Memoised Wake-SleepabstractWe study a class of neuro-symbolic generative models in which neural networks are used both for inference and as priors over symbolic, data-generating programs. As generative models, these programs capture compositional structures in a naturally explainable form. To tackle the challenge of performing program induction as an ‘inner-loop’ to learning, we propose the Memoised Wake-Sleep (MWS) algorithm, which extends Wake Sleep by explicitly storing and reusing the best programs discovered by the inference network throughout training. We use MWS to learn accurate, explainable models in three challenging domains: stroke-based character modelling, cellular automata, and few-shot learning in a novel dataset of real-world string concepts. Luke B. Hewitt, Tuan Anh Le 0001, Josh Tenenbaum |
UAI | 3 |
| 2019 | Theory of Minds: Understanding Behavior in Groups through Inverse PlanningabstractHuman social behavior is structured by relationships. We form teams, groups, tribes, and alliances at all scales of human life. These structures guide multi-agent cooperation and competition, but when we observe others these underlying relationships are typically unobservable and hence must be inferred. Humans make these inferences intuitively and flexibly, often making rapid generalizations about the latent relationships that underlie behavior from just sparse and noisy observations. Rapid and accurate inferences are important for determining who to cooperate with, who to compete with, and how to cooperate in order to compete. Towards the goal of building machine-learning algorithms with human-like social intelligence, we develop a generative model of multiagent action understanding based on a novel representation for these latent relationships called Composable Team Hierarchies (CTH). This representation is grounded in the formalism of stochastic games and multi-agent reinforcement learning. We use CTH as a target for Bayesian inference yielding a new algorithm for understanding behavior in groups that can both infer hidden relationships as well as predict future actions for multiple agents interacting together. Our algorithm rapidly recovers an underlying causal model of how agents relate in spatial stochastic games from just a few observations. The patterns of inference made by this algorithm closely correspond with human judgments and the algorithm makes the same rapid generalizations that people do. Michael Shum, Max Kleiman-Weiner, Michael L. Littman, Josh Tenenbaum |
AAAI | 4 |
| 2019 | Rapid Trial-and-Error Learning in Physical Problem Solving
Kelsey R. Allen, Kevin A. Smith 0001, Josh Tenenbaum |
CogSci | 3 |
| 2019 | Simplicity and Probability in Human Judgment
Tyler Brooke-Wilson, Jonathan S. Rosenfeld, Matthias Hofer 0002, Junyi Chu, Josh Tenenbaum |
CogSci | 5 |
| 2019 | Query-guided visual search
Junyi Chu, Jon Gauthier, Roger Levy, Josh Tenenbaum, Laura Schulz |
CogSci | 4 |
| 2019 | Heuristics, hacks, and habits: Boundedly optimal approaches to learning, reasoning and decision making
Ishita Dasgupta 0001, Eric Schulz, Jessica B. Hamrick, Josh Tenenbaum |
CogSci | 4 |
| 2019 | A rational model of syntactic bootstrapping
Jon Gauthier, Roger Levy, Josh Tenenbaum |
CogSci | 3 |
| 2019 | When circumstances change, update your pronouns
Joshua K. Hartshorne, Mariela Jennings, Tobias Gerstenberg, Josh Tenenbaum |
CogSci | 4 |
| 2019 | Emotion attributions echo the structure of people's intuitive theory of psychology
Sean Dae Houlihan, Max Kleiman-Weiner, Josh Tenenbaum, Rebecca Saxe |
CogSci | 3 |
| 2019 | Look out, it's going to fall!: Does physical instability capture attention and lead to distraction?
Marta Kryven, Sholei Croom, Brian J. Scholl, Josh Tenenbaum |
CogSci | 4 |
| 2019 | Choosing the unimaginable: Social psychological factors in seeking transformative experiences
Marta Kryven, Laura Niemi, Laurie Paul, Josh Tenenbaum |
CogSci | 4 |
| 2019 | What if everybody did that?: Universalization as a mechanism of moral decision-making
Sydney Levine, Max Kleiman-Weiner, Laura Schulz, Josh Tenenbaum, Fiery Cushman |
CogSci | 4 |
| 2019 | Discovering a symbolic planning language from continuous experience
João Loula, Tom Silver, Kelsey R. Allen, Josh Tenenbaum |
CogSci | 4 |
| 2019 | Extending Rationality
Emmanuel M. Pothos, Jerome R. Busemeyer, Timothy J. Pleskac, James M. Yearsley, Josh Tenenbaum, Noah D. Goodman, Michael Henry Tessler, Thomas L. Griffiths 0001, Falk Lieder, Ralph Hertwig, Thorsten Pachur, Christina Leuker, Richard M. Shiffrin |
CogSci | 5 |
| 2019 | Inferring Structured Visual Concepts from Minimal Data
Luke B. Hewitt, Josh Tenenbaum, Roger Levy |
CogSci | 3 |
| 2019 | Learning a novel rule-based conceptual system
Joshua S. Rule, Josh Tenenbaum, Steve Piantadosi |
CogSci | 2 |
| 2019 | Real-time inference of physical properties in dynamic scenes
Kevin A. Smith 0001, Mario Belledonne, Ilker Yildirim, Jiajun Wu 0001, Josh Tenenbaum |
CogSci | 5 |
| 2019 | Draping an Elephant: Uncovering Children's Reasoning About Cloth-Covered Objects
Tomer D. Ullman, Eliza Kosoy, Ilker Yildirim, Amir Arsalan Soltani, Max H. Siegel, Josh Tenenbaum, Elizabeth S. Spelke |
CogSci | 6 |
| 2019 | Modeling Expertise with Neurally-Guided Bayesian Program Induction
Catherine Wong, Kevin Ellis, Mathias Sablé-Meyer, Josh Tenenbaum |
CogSci | 4 |
| 2019 | The effects of object motion observations on physical prediction
Moyuru Yamada, Kevin A. Smith 0001, Josh Tenenbaum |
CogSci | 3 |
| 2019 | Explaining intuitive difficulty judgments by modeling physical effort and risk
Ilker Yildirim, Basil Saeed, Grace Bennett-Pierre, Tobias Gerstenberg, Josh Tenenbaum, Hyowon Gweon |
CogSci | 5 |
| 2019 | Program-Guided Image ManipulatorsabstractHumans are capable of building holistic representations for images at various levels, from local objects, to pairwise relations, to global structures. The interpretation of structures involves reasoning over repetition and symmetry of the objects in the image. In this paper, we present the Program-Guided Image Manipulator (PG-IM), inducing neuro-symbolic program-like representations to represent and manipulate images. Given an image, PG-IM detects repeated patterns, induces symbolic programs, and manipulates the image using a neural network that is guided by the program. PG-IM learns from a single image, exploiting its internal statistics. Despite trained only on image inpainting, PG-IM is directly capable of extrapolation and regularity editing in a unified framework. Extensive experiments show that PG-IM achieves superior performance on all the tasks. Xiuming Zhang, Jiayuan Mao, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
ICCV | 5 |
| 2019 | GAN Dissection: Visualizing and Understanding Generative Adversarial Networks
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Josh Tenenbaum, William T. Freeman, Antonio Torralba 0001 |
ICLR (Poster) | 5 |
| 2019 | Reasoning About Physical Interactions with Object-Oriented Prediction and Planning
Michael Janner, Sergey Levine, William T. Freeman, Josh Tenenbaum, Chelsea Finn, Jiajun Wu 0001 |
ICLR (Poster) | 4 |
| 2019 | Learning Particle Dynamics for Manipulating Rigid Bodies, Deformable Objects, and Fluids
Yunzhu Li, Jiajun Wu 0001, Russ Tedrake, Josh Tenenbaum, Antonio Torralba 0001 |
ICLR (Poster) | 4 |
| 2019 | Learning to Describe Scenes with Programs
Daniel Ritchie 0001, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
ICLR (Poster) | 5 |
| 2019 | The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Jiayuan Mao, Chuang Gan 0001, Pushmeet Kohli, Josh Tenenbaum, Jiajun Wu 0001 |
ICLR | 4 |
| 2019 | Stochastic Prediction of Multi-Agent Interactions from Partial Observations
Chen Sun 0002, Per Karlsson, Jiajun Wu 0001, Josh Tenenbaum, Kevin Murphy 0002 |
ICLR (Poster) | 4 |
| 2019 | Learning to Infer and Execute 3D Shape Programs
Yonglong Tian, Andrew Luo 0001, Xingyuan Sun, Kevin Ellis, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
ICLR (Poster) | 6 |
| 2019 | Unsupervised Discovery of Parts, Structure, and Dynamics
Zhenjia Xu, Chen Sun 0002, Kevin Murphy 0002, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
ICLR (Poster) | 6 |
| 2019 | Infinite Mixture Prototypes for Few-shot LearningabstractWe propose infinite mixture prototypes to adaptively represent both simple and complex data distributions for few-shot learning. Infinite mixture prototypes combine deep representation learning with Bayesian nonparametrics, representing each class by a set of clusters, unlike existing prototypical methods that represent each class by a single cluster. By inferring the number of clusters, infinite mixture prototypes interpolate between nearest neighbor and prototypical representations in a learned feature space, which improves accuracy and robustness in the few-shot regime. We show the importance of adaptive capacity for capturing complex data distributions such as super-classes (like alphabets in character recognition), with 10-25% absolute accuracy improvements over prototypical networks, while still maintaining or improving accuracy on standard few-shot learning benchmarks. By clustering labeled and unlabeled data with the same rule, infinite mixture prototypes achieve state-of-the-art semi-supervised accuracy, and can perform purely unsupervised clustering, unlike existing fully- and semi-supervised prototypical methods. Kelsey R. Allen, Evan Shelhamer, Hanul Shin, Josh Tenenbaum |
ICML | 4 |
| 2019 | Neurally-Guided Structure InferenceabstractMost structure inference methods either rely on exhaustive search or are purely data-driven. Exhaustive search robustly infers the structure of arbitrarily complex data, but it is slow. Data-driven methods allow efficient inference, but do not generalize when test data have more complex structures than training data. In this paper, we propose a hybrid inference algorithm, the Neurally-Guided Structure Inference (NG-SI), keeping the advantages of both search-based and data-driven methods. The key idea of NG-SI is to use a neural network to guide the hierarchical, layer-wise search over the compositional space of structures. We evaluate our algorithm on two representative structure inference tasks: probabilistic matrix decomposition and symbolic program parsing. It outperforms data-driven and search-based alternatives on both tasks. Sidi Lu, Jiayuan Mao, Josh Tenenbaum, Jiajun Wu 0001 |
ICML | 3 |
| 2019 | Learning to Infer Program SketchesabstractOur goal is to build systems which write code automatically from the kinds of specifications humans can most easily provide, such as examples and natural language instruction. The key idea of this work is that a flexible combination of pattern recognition and explicit reasoning can be used to solve these complex programming problems. We propose a method for dynamically integrating these types of information. Our novel intermediate representation and training algorithm allow a program synthesis system to learn, without direct supervision, when to rely on pattern recognition and when to perform symbolic search. Our model matches the memorization and generalization performance of neural synthesis and symbolic search, respectively, and achieves state-of-the-art performance on a dataset of simple English description-to-code programming problems. Maxwell I. Nye, Luke B. Hewitt, Josh Tenenbaum, Armando Solar-Lezama |
ICML | 3 |
| 2019 | Combining Physical Simulators and Object-Based Networks for ControlabstractPhysics engines play an important role in robot planning and control; however, many real-world control problems involve complex contact dynamics that cannot be characterized analytically. Most physics engines therefore employ approximations that lead to a loss in precision. In this paper, we propose a hybrid dynamics model, simulator-augmented interaction networks (SAIN), combining a physics engine with an object-based neural network for dynamics modeling. Compared with existing models that are purely analytical or purely data-driven, our hybrid model captures the dynamics of interacting objects in a more accurate and data-efficient manner. Experiments both in simulation and on a real robot suggest that it also leads to better performance when used in complex control tasks. Finally, we show that our model generalizes to novel environments with varying object shapes and materials. Anurag Ajay, Maria Bauzá 0001, Jiajun Wu 0001, Nima Fazeli, Josh Tenenbaum, Alberto Rodriguez 0003, Leslie Pack Kaelbling |
ICRA | 5 |
| 2019 | ChainQueen: A Real-Time Differentiable Physical Simulator for Soft RoboticsabstractPhysical simulators have been widely used in robot planning and control. Among them, differentiable simulators are particularly favored, as they can be incorporated into gradient-based optimization algorithms that are efficient in solving inverse problems such as optimal control and motion planning. Therefore, rigid body simulators and recently their differentiable variants are studied extensively. Simulating deformable objects is, however, more challenging compared to rigid body dynamics. The underlying physical laws of deformable objects are more complex, and the resulting systems have orders of magnitude more degrees of freedom and there-fore they are significantly more computationally expensive to simulate. Computing gradients with respect to physical design or controller parameters is typically even more computationally challenging. In this paper, we propose a real-time, differentiable hybrid Lagrangian-Eulerian physical simulator for deformable objects, ChainQueen, based on the Moving Least Squares Material Point Method (MLS-MPM). MLS-MPM can simulate deformable objects with collisions and can be seamlessly incorporated into soft robotic systems. We demonstrate that our simulator achieves high precision in both forward simulation and backward gradient computation. We have successfully employed it in a diverse set of inference, control and co-design tasks for soft robotics. Yuanming Hu, Jiancheng Liu, Andrew Spielberg, Josh Tenenbaum, William T. Freeman, Jiajun Wu 0001, Daniela Rus, Wojciech Matusik |
ICRA | 4 |
| 2019 | Propagation Networks for Model-Based Control Under Partial ObservationabstractThere has been an increasing interest in learning dynamics simulators for model-based control. Compared with off-the-shelf physics engines, a learnable simulator can quickly adapt to unseen objects, scenes, and tasks. However, existing models like interaction networks only work for fully observable systems; they also only consider pairwise interactions within a single time step, both restricting their use in practical systems. We introduce Propagation Networks (PropNet), a differentiable, learnable dynamics model that handles partially observable scenarios and enables instantaneous propagation of signals beyond pairwise interactions. With these innovations, our propagation networks not only outperform current learnable physics engines in forward simulation, but also achieves superior performance on various control tasks. Compared with existing deep reinforcement learning algorithms, model-based control with propagation networks is more accurate, efficient, and generalizable to novel, partially observable scenes and tasks. Yunzhu Li, Jiajun Wu 0001, Jun-Yan Zhu, Josh Tenenbaum, Antonio Torralba 0001, Russ Tedrake |
ICRA | 4 |
| 2019 | Differentiable Physics and Stable Modes for Tool-Use and Manipulation Planning - Extended AbtractabstractWe propose to formulate physical reasoning and manipulation planning as an optimization problem that integrates first order logic, which we call Logic-Geometric Programming. Marc Toussaint, Kelsey R. Allen, Kevin A. Smith 0001, Josh Tenenbaum |
IJCAI | 4 |
| 2019 | ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition modelsabstractWe collect a large real-world test set, ObjectNet, for object recognition with controls where object backgrounds, rotations, and imaging viewpoints are random. Most scientific experiments have controls, confounds which are removed from the data, to ensure that subjects cannot perform a task by exploiting trivial correlations in the data. Historically, large machine learning and computer vision datasets have lacked such controls. This has resulted in models that must be fine-tuned for new datasets and perform better on datasets than in real-world applications. When tested on ObjectNet, object detectors show a 40-45% drop in performance, with respect to their performance on other benchmarks, due to the controls for biases. Controls make ObjectNet robust to fine-tuning showing only small performance increases. We develop a highly automated platform that enables gathering datasets with controls by crowdsourcing image capturing and annotation. ObjectNet is the same size as the ImageNet test set (50,000 images), and by design does not come paired with a training set in order to encourage generalization. The dataset is both easier than ImageNet (objects are largely centered and unoccluded) and harder (due to the controls). Although we focus on object recognition here, data with controls can be gathered at scale using automated tools throughout machine learning to generate datasets that exercise models in new ways thus providing valuable feedback to researchers. This work opens up new avenues for research in generalizable, robust, and more human-like computer vision and in creating datasets where results are predictive of real-world performance. Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, Boris Katz |
NeurIPS | 7 |
| 2019 | Write, Execute, Assess: Program Synthesis with a REPLabstractWe present a neural program synthesis approach integrating components which write, execute, and assess code to navigate the search space of possible programs. We equip the search process with an interpreter or a read-eval-print-loop (REPL), which immediately executes partially written programs, exposing their semantics. The REPL addresses a basic challenge of program synthesis: tiny changes in syntax can lead to huge changes in semantics. We train a pair of models, a policy that proposes the new piece of code to write, and a value function that assesses the prospects of the code written so-far. At test time we can combine these models with a Sequential Monte Carlo algorithm. We apply our approach to two domains: synthesizing text editing programs and inferring 2D and 3D graphics programs. Kevin Ellis, Maxwell I. Nye, Yewen Pu, Felix Sosa, Josh Tenenbaum, Armando Solar-Lezama |
NeurIPS | 5 |
| 2019 | Visual Concept-Metaconcept LearningabstractHumans reason with concepts and metaconcepts: we recognize red and blue from visual input; we also understand that they are colors, i.e., red is an instance of color. In this paper, we propose the visual concept-metaconcept learner (VCML) for joint learning of concepts and metaconcepts from images and associated question-answer pairs. The key is to exploit the bidirectional connection between visual concepts and metaconcepts. Visual representations provide grounding cues for predicting relations between unseen pairs of concepts. Knowing that red and blue are instances of color, we generalize to the fact that green is also an instance of color since they all categorize the hue of objects. Meanwhile, knowledge about metaconcepts empowers visual concept learning from limited, noisy, and even biased data. From just a few examples of purple cubes we can understand a new color purple, which resembles the hue of the cubes instead of the shape of them. Evaluation on both synthetic and real-world datasets validates our claims. Chi Han, Jiayuan Mao, Chuang Gan 0001, Josh Tenenbaum, Jiajun Wu 0001 |
NeurIPS | 4 |
| 2019 | Finding Friend and Foe in Multi-Agent GamesabstractRecent breakthroughs in AI for multi-agent games like Go, Poker, and Dota, have seen great strides in recent years. Yet none of these games address the real-life challenge of cooperation in the presence of unknown and uncertain teammates. This challenge is a key game mechanism in hidden role games. Here we develop the DeepRole algorithm, a multi-agent reinforcement learning agent that we test on "The Resistance: Avalon", the most popular hidden role game. DeepRole combines counterfactual regret minimization (CFR) with deep value networks trained through self-play. Our algorithm integrates deductive reasoning into vector-form CFR to reason about joint beliefs and deduce partially observable actions. We augment deep value networks with constraints that yield interpretable representations of win probabilities. These innovations enable DeepRole to scale to the full Avalon game. Empirical game-theoretic methods show that DeepRole outperforms other hand-crafted and learned agents in five-player Avalon. DeepRole played with and against human players on the web in hybrid human-agent teams. We find that DeepRole outperforms human players as both a cooperator and a competitor. Jack Serrino, Max Kleiman-Weiner, David C. Parkes, Josh Tenenbaum |
NeurIPS | 4 |
| 2019 | Modeling Expectation Violation in Intuitive Physics with Coarse Probabilistic Object RepresentationsabstractFrom infancy, humans have expectations about how objects will move and interact. Even young children expect objects not to move through one another, teleport, or disappear. They are surprised by mismatches between physical expectations and perceptual observations, even in unfamiliar scenes with completely novel objects. A model that exhibits human-like understanding of physics should be similarly surprised, and adjust its beliefs accordingly. We propose ADEPT, a model that uses a coarse (approximate geometry) object-centric representation for dynamic 3D scene understanding. Inference integrates deep recognition networks, extended probabilistic physical simulation, and particle filtering for forming predictions and expectations across occlusion. We also present a new test set for measuring violations of physical expectations, using a range of scenarios derived from developmental psychology. We systematically compare ADEPT, baseline models, and human expectations on this test set. ADEPT outperforms standard network architectures in discriminating physically implausible scenes, and often performs this discrimination at the same level as people. Kevin A. Smith 0001, Lingjie Mei, Shunyu Yao 0006, Jiajun Wu 0001, Elizabeth S. Spelke, Josh Tenenbaum, Tomer D. Ullman |
NeurIPS | 6 |
| 2019 | Modeling human intuitions about liquid flow with particle-based simulationabstractHumans can easily describe, imagine, and, crucially, predict a wide variety of behaviors of liquids-splashing, squirting, gushing, sloshing, soaking, dripping, draining, trickling, pooling, and pouring-despite tremendous variability in their material and dynamical properties. Here we propose and test a computational model of how people perceive and predict these liquid dynamics, based on coarse approximate simulations of fluids as collections of interacting particles. Our model is analogous to a "game engine in the head", drawing on techniques for interactive simulations (as in video games) that optimize for efficiency and natural appearance rather than physical accuracy. In two behavioral experiments, we found that the model accurately captured people's predictions about how liquids flow among complex solid obstacles, and was significantly better than several alternatives based on simple heuristics and deep neural networks. Our model was also able to explain how people's predictions varied as a function of the liquids' properties (e.g., viscosity and stickiness). Together, the model and empirical results extend the recent proposal that human physical scene understanding for the dynamics of rigid, solid objects can be supported by approximate probabilistic simulation, to the more complex and unexplored domain of fluid dynamics. Christopher Bates, Ilker Yildirim, Josh Tenenbaum, Peter W. Battaglia |
PLoS Comput. Biol. | 3 |
| 2018 | A Computational Model of Commonsense Moral Decision MakingabstractWe introduce a computational model for building moral autonomous vehicles by learning and generalizing from human moral judgments. We draw on a cognitively inspired model of how people and young children learn moral theories from sparse and noisy data and integrate observations made from different people in different groups. The problem of moral learning for autonomous vehicles is cast as learning how to weigh the different features of the dilemma using utility calculus, with the goal of making these trade-offs reflect how people make them in a wide variety of moral dilemma. By modeling the structures of individuals and groups in a hierarchical Bayesian model, we show that an individual's moral values -- as well as a group's shared values -- can be inferred from sparse and noisy data. We evaluate our approach with data from the Moral Machine, a web application that collects human judgments on moral dilemmas involving autonomous vehicles, and show that the model rapidly and accurately infers people's preferences and can predict the difficulty of moral dilemmas from limited data. Richard Kim, Max Kleiman-Weiner, Andrés Abeliuk, Edmond Awad, Sohan Dsouza, Josh Tenenbaum, Iyad Rahwan |
AIES | 6 |
| 2018 | Grounding Compositional Hypothesis Generation in Specific Instances
Neil Bramley, Anselm Rothe, Josh Tenenbaum, Todd M. Gureckis |
CogSci | 3 |
| 2018 | Learning as program induction
Neil Bramley, Eric Schulz, Josh Tenenbaum |
CogSci | 4 |
| 2018 | Auditory scene analysis as Bayesian inference in sound source models
Maddie Cusimano, Luke B. Hewitt, Josh Tenenbaum, Josh H. McDermott |
CogSci | 3 |
| 2018 | Learning to act by integrating mental simulations and physical experiments
Ishita Dasgupta 0001, Kevin A. Smith 0001, Eric Schulz, Josh Tenenbaum, Samuel Gershman |
CogSci | 4 |
| 2018 | Tiptoeing around it: Inference from absence in potentially offensive speech
Monica A. Gates, Tess L. Veuthey, Michael Henry Tessler, Kevin A. Smith 0001, Tobias Gerstenberg, Laurie Bayet, Josh Tenenbaum |
CogSci | 7 |
| 2018 | Word learning and the acquisition of syntactic-semantic overhypotheses
Jon Gauthier, Roger Levy, Josh Tenenbaum |
CogSci | 3 |
| 2018 | What happened? Reconstructing the past through vision and sound
Tobias Gerstenberg, Max H. Siegel, Josh Tenenbaum |
CogSci | 3 |
| 2018 | Relational inductive bias for physical construction in humans and machines
Jessica B. Hamrick, Kelsey R. Allen, Victor Bapst, Tina Zhu, Kevin R. McKee, Josh Tenenbaum, Peter W. Battaglia |
CogSci | 6 |
| 2018 | Inductive Biases in the Evolution of Combinatorial Structure in Language
Matthias Hofer 0002, Josh Tenenbaum, Roger Levy |
CogSci | 2 |
| 2018 | A generative model of people's intuitive theory of emotions: inverse planning in rich social games
Sean Dae Houlihan, Max Kleiman-Weiner, Josh Tenenbaum, Rebecca Saxe |
CogSci | 3 |
| 2018 | Hierarchical Drift-Diffusion Model for Moral Dilemma: Understanding Reaction Times and Choices
Richard Kim, Niccolo Pescetelli, Max Kleiman-Weiner, Edmond Awad, Sohan Dsouza, Josh Tenenbaum, Iyad Rahwan |
CogSci | 6 |
| 2018 | The Evolution of Cooperation in Cognitively Flexible Agents
Max Kleiman-Weiner, Alejandro Vientós, David Rand, Josh Tenenbaum |
CogSci | 4 |
| 2018 | The Cognitive Mechanisms of Contractualist Moral Decision-Making
Sydney Levine, Max Kleiman-Weiner, Nick Chater, Fiery Cushman, Josh Tenenbaum |
CogSci | 5 |
| 2018 | Learning list concepts through program induction
Joshua S. Rule, Eric Schulz, Steve Piantadosi, Josh Tenenbaum |
CogSci | 4 |
| 2018 | Evidence for an Intuitive Physics Engine in the Human Brain
Sarah Schwettmann, Jason Fischer, Josh Tenenbaum, Nancy Kanwisher |
CogSci | 3 |
| 2018 | Physical Inference for Object Perception in Complex Auditory Scenes
Max H. Siegel, Josh Tenenbaum, Josh H. McDermott |
CogSci | 2 |
| 2018 | Strategies and representations in physical inference
Kevin A. Smith 0001, Josh Tenenbaum, Erin M. Anderson, Susan J. Hespos, Lance J. Rips, Chaz Firestone, Jessica B. Hamrick |
CogSci | 2 |
| 2018 | Moral Dynamics: A Computational Model of Moral Judgment
Felix Sosa, Tomer D. Ullman, Samuel Gershman, Josh Tenenbaum, Tobias Gerstenberg |
CogSci | 4 |
| 2018 | Pix3D: Dataset and Methods for Single-Image 3D Shape ModelingabstractWe study 3D shape modeling from a single image and make contributions to it in three aspects. First, we present Pix3D, a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment. Pix3D has wide applications in shape-related tasks including reconstruction, retrieval, viewpoint estimation, etc. Building such a large-scale dataset, however, is highly challenging; existing datasets either contain only synthetic data, or lack precise alignment between 2D images and 3D shapes, or only have a small number of images. Second, we calibrate the evaluation criteria for 3D shape reconstruction through behavioral studies, and use them to objectively and systematically benchmark cutting-edge reconstruction algorithms on Pix3D. Third, we design a novel model that simultaneously performs 3D reconstruction and pose estimation; our multi-task learning approach achieves state-of-the-art performance on both tasks. Xingyuan Sun, Jiajun Wu 0001, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Josh Tenenbaum, William T. Freeman |
CVPR | 7 |
| 2018 | Physical Primitive Decomposition
William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
ECCV (12) | 3 |
| 2018 | Learning Shape Priors for Single-View 3D Completion And Reconstruction
Jiajun Wu 0001, Chengkai Zhang, Xiuming Zhang, Zhoutong Zhang, William T. Freeman, Josh Tenenbaum |
ECCV (11) | 6 |
| 2018 | Seeing Tree Structure from Vibration
Tianfan Xue, Jiajun Wu 0001, Zhoutong Zhang, Chengkai Zhang, Josh Tenenbaum, William T. Freeman |
ECCV (9) | 5 |
| 2018 | Meta-Learning for Semi-Supervised Few-Shot Classification
Mengye Ren, Eleni Triantafillou, Sachin Ravi, Jake Snell, Kevin Swersky, Josh Tenenbaum, Hugo Larochelle, Richard S. Zemel |
ICLR (Poster) | 6 |
| 2018 | Augmenting Physical Simulators with Stochastic Neural Networks: Case Study of Planar Pushing and BouncingabstractAn efficient, generalizable physical simulator with universal uncertainty estimates has wide applications in robot state estimation, planning, and control. In this paper, we build such a simulator for two scenarios, planar pushing and ball bouncing, by augmenting an analytical rigid-body simulator with a neural network that learns to model uncertainty as residuals. Combining symbolic, deterministic simulators with learnable, stochastic neural nets provides us with expressiveness, efficiency, and generalizability simultaneously. Our model outperforms both purely analytical and purely learned simulators consistently on real, standard benchmarks. Compared with methods that model uncertainty using Gaussian processes, our model runs much faster, generalizes better to new object shapes, and is able to characterize the complex distribution of object trajectories. Anurag Ajay, Jiajun Wu 0001, Nima Fazeli, Maria Bauzá 0001, Leslie Pack Kaelbling, Josh Tenenbaum, Alberto Rodriguez 0003 |
IROS | 6 |
| 2018 | 3D Shape Perception from Monocular Vision, Touch, and Shape PriorsabstractPerceiving accurate 3D object shape is important for robots to interact with the physical world. Current research along this direction has been primarily relying on visual observations. Vision, however useful, has inherent limitations due to occlusions and the 2D-3D ambiguities, especially for perception with a monocular camera. In contrast, touch gets precise local shape information, though its efficiency for reconstructing the entire shape could be low. In this paper, we propose a novel paradigm that efficiently perceives accurate 3D object shape by incorporating visual and tactile observations, as well as prior knowledge of common object shapes learned from large-scale shape repositories. We use vision first, applying neural networks with learned shape priors to predict an object's 3D shape from a single-view color image. We then use tactile sensing to refine the shape; the robot actively touches the object regions where the visual prediction has high uncertainty. Our method efficiently builds the 3D shape of common objects from a color image and a small number of tactile explorations (around 10). Our setup is easy to apply and has potentials to help robots better perform grasping or manipulation tasks on real-world objects. Shaoxiong Wang, Jiajun Wu 0001, Xingyuan Sun, Wenzhen Yuan 0001, William T. Freeman, Josh Tenenbaum, Edward H. Adelson |
IROS | 6 |
| 2018 | End-to-End Differentiable Physics for Learning and ControlabstractWe present a differentiable physics engine that can be integrated as a module in deep neural networks for end-to-end learning. As a result, structured physics knowledge can be embedded into larger systems, allowing them, for example, to match observations by performing precise simulations, while achieves high sample efficiency. Specifically, in this paper we demonstrate how to perform backpropagation analytically through a physical simulator defined via a linear complementarity problem. Unlike traditional finite difference methods, such gradients can be computed analytically, which allows for greater flexibility of the engine. Through experiments in diverse domains, we highlight the system's ability to learn physical parameters from data, efficiently match and simulate observed visual behavior, and readily enable control via gradient-based planning methods. Code for the engine and experiments is included with the paper. Filipe de Avila Belbute-Peres, Kevin A. Smith 0001, Kelsey R. Allen, Josh Tenenbaum, J. Zico Kolter |
NeurIPS | 4 |
| 2018 | Learning to Exploit Stability for 3D Scene ParsingabstractHuman scene understanding uses a variety of visual and non-visual cues to perform inference on object types, poses, and relations. Physics is a rich and universal cue which we exploit to enhance scene understanding. We integrate the physical cue of stability into the learning process using a REINFORCE approach coupled to a physics engine, and apply this to the problem of producing the 3D bounding boxes and poses of objects in a scene. We first show that applying physics supervision to an existing scene understanding model increases performance, produces more stable predictions, and allows training to an equivalent performance level with fewer annotated training examples. We then present a novel architecture for 3D scene parsing named Prim R-CNN, learning to predict bounding boxes as well as their 3D size, translation, and rotation. With physics supervision, Prim R-CNN outperforms existing scene understanding approaches on this problem. Finally, we show that applying physics supervision on unlabeled real images improves real domain transfer of models training on synthetic data. Yilun Du, Hector Basevi, Ales Leonardis, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
NeurIPS | 6 |
| 2018 | Learning Libraries of Subroutines for Neurally-Guided Bayesian Program InductionabstractSuccessful approaches to program induction require a hand-engineered domain-specific language (DSL), constraining the space of allowed programs and imparting prior knowledge of the domain. We contribute a program induction algorithm that learns a DSL while jointly training a neural network to efficiently search for programs in the learned DSL. We use our model to synthesize functions on lists, edit text, and solve symbolic regression problems, showing how the model learns a domain-specific library of program components for expressing solutions to problems in the domain. Kevin Ellis, Lucas Morales, Mathias Sablé-Meyer, Armando Solar-Lezama, Josh Tenenbaum |
NeurIPS | 5 |
| 2018 | Learning to Infer Graphics Programs from Hand-Drawn ImagesabstractWe introduce a model that learns to convert simple hand drawings into graphics programs written in a subset of \LaTeX.~The model combines techniques from deep learning and program synthesis. We learn a convolutional neural network that proposes plausible drawing primitives that explain an image. These drawing primitives are a specification (spec) of what the graphics program needs to draw. We learn a model that uses program synthesis techniques to recover a graphics program from that spec. These programs have constructs like variable bindings, iterative loops, or simple kinds of conditionals. With a graphics program in hand, we can correct errors made by the deep network and extrapolate drawings. Kevin Ellis, Daniel Ritchie 0001, Armando Solar-Lezama, Josh Tenenbaum |
NeurIPS | 4 |
| 2018 | Flexible neural representation for physics predictionabstractHumans have a remarkable capacity to understand the physical dynamics of objects in their environment, flexibly capturing complex structures and interactions at multiple levels of detail. Inspired by this ability, we propose a hierarchical particle-based object representation that covers a wide variety of types of three-dimensional objects, including both arbitrary rigid geometrical shapes and deformable materials. We then describe the Hierarchical Relation Network (HRN), an end-to-end differentiable neural network based on hierarchical graph convolution, that learns to predict physical dynamics in this representation. Compared to other neural network baselines, the HRN accurately handles complex collisions and nonrigid deformations, generating plausible dynamics predictions at long time scales in novel settings, and scaling to large scene configurations. These results demonstrate an architecture with the potential to form the basis of next-generation physics predictors for use in computer vision, robotics, and quantitative cognitive science. Damian Mrowca, Chengxu Zhuang, Elias Wang, Nick Haber, Li Fei-Fei 0001, Josh Tenenbaum, Dan Yamins |
NeurIPS | 6 |
| 2018 | Learning to Share and Hide Intentions using Information RegularizationabstractLearning to cooperate with friends and compete with foes is a key component of multi-agent reinforcement learning. Typically to do so, one requires access to either a model of or interaction with the other agent(s). Here we show how to learn effective strategies for cooperation and competition in an asymmetric information game with no such model or interaction. Our approach is to encourage an agent to reveal or hide their intentions using an information-theoretic regularizer. We consider both the mutual information between goal and action given state, as well as the mutual information between goal and state. We show how to stochastically optimize these regularizers in a way that is easy to integrate with policy gradient reinforcement learning. Finally, we demonstrate that cooperative (competitive) policies learned with our approach lead to more (less) reward for a second agent in two simple asymmetric information games. DJ Strouse, Max Kleiman-Weiner, Josh Tenenbaum, Matt M. Botvinick, David J. Schwab |
NeurIPS | 3 |
| 2018 | 3D-Aware Scene Manipulation via Inverse GraphicsabstractWe aim to obtain an interpretable, expressive, and disentangled scene representation that contains comprehensive structural and textural information for each object. Previous scene representations learned by neural networks are often uninterpretable, limited to a single object, or lacking 3D knowledge. In this work, we propose 3D scene de-rendering networks (3D-SDN) to address the above issues by integrating disentangled representations for semantics, geometry, and appearance into a deep generative model. Our scene encoder performs inverse graphics, translating a scene into a structured object-wise representation. Our decoder has two components: a differentiable shape renderer and a neural texture generator. The disentanglement of semantics, geometry, and appearance supports 3D-aware scene manipulation, e.g., rotating and moving objects freely while keeping the consistent shape and texture, and changing the object appearance without affecting its shape. Experiments demonstrate that our editing scheme based on 3D-SDN is superior to its 2D counterpart. Shunyu Yao 0006, Tzu-Ming Harry Hsu, Jun-Yan Zhu, Jiajun Wu 0001, Antonio Torralba 0001, William T. Freeman, Josh Tenenbaum |
NeurIPS | 7 |
| 2018 | Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language UnderstandingabstractWe marry two powerful ideas: deep representation learning for visual recognition and language understanding, and symbolic program execution for reasoning. Our neural-symbolic visual question answering (NS-VQA) system first recovers a structural scene representation from the image and a program trace from the question. It then executes the program on the scene representation to obtain an answer. Incorporating symbolic structure as prior knowledge offers three unique advantages. First, executing programs on a symbolic space is more robust to long program traces; our model can solve complex reasoning tasks better, achieving an accuracy of 99.8% on the CLEVR dataset. Second, the model is more data- and memory-efficient: it performs well after learning on a small number of training data; it can also encode an image into a compact representation, requiring less storage than existing methods for offline question answering. Third, symbolic program execution offers full transparency to the reasoning process; we are thus able to interpret and diagnose each execution step. Kexin Yi, Jiajun Wu 0001, Chuang Gan 0001, Antonio Torralba 0001, Pushmeet Kohli, Josh Tenenbaum |
NeurIPS | 6 |
| 2018 | Learning to Reconstruct Shapes from Unseen ClassesabstractFrom a single image, humans are able to perceive the full 3D shape of an object by exploiting learned shape priors from everyday life. Contemporary single-image 3D reconstruction algorithms aim to solve this task in a similar fashion, but often end up with priors that are highly biased by training classes. Here we present an algorithm, Generalizable Reconstruction (GenRe), designed to capture more generic, class-agnostic shape priors. We achieve this with an inference network and training procedure that combine 2.5D representations of visible surfaces (depth and silhouette), spherical shape representations of both visible and non-visible surfaces, and 3D voxel-based representations, in a principled manner that exploits the causal structure of how 3D shapes give rise to 2D images. Experiments demonstrate that GenRe performs well on single-view shape reconstruction, and generalizes to diverse novel objects from categories not seen during training. Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Josh Tenenbaum, William T. Freeman, Jiajun Wu 0001 |
NeurIPS | 4 |
| 2018 | Visual Object Networks: Image Generation with Disentangled 3D RepresentationsabstractRecent progress in deep generative models has led to tremendous breakthroughs in image generation. While being able to synthesize photorealistic images, existing models lack an understanding of our underlying 3D world. Different from previous works built on 2D datasets and models, we present a new generative model, Visual Object Networks (VONs), synthesizing natural images of objects with a disentangled 3D representation. Inspired by classic graphics rendering pipelines, we unravel the image formation process into three conditionally independent factors---shape, viewpoint, and texture---and present an end-to-end adversarial learning framework that jointly models 3D shape and 2D texture. Our model first learns to synthesize 3D shapes that are indistinguishable from real shapes. It then renders the object's 2.5D sketches (i.e., silhouette and depth map) from its shape under a sampled viewpoint. Finally, it learns to add realistic textures to these 2.5D sketches to generate realistic images. The VON not only generates images that are more realistic than the state-of-the-art 2D image synthesis methods but also enables many 3D operations such as changing the viewpoint of a generated image, shape and texture editing, linear interpolation in texture and shape space, and transferring appearance across different objects and viewpoints. Jun-Yan Zhu, Zhoutong Zhang, Chengkai Zhang, Jiajun Wu 0001, Antonio Torralba 0001, Josh Tenenbaum, William T. Freeman |
NeurIPS | 6 |
| 2018 | The Variational Homoencoder: Learning to learn high capacity generative models from few examples
Luke B. Hewitt, Maxwell I. Nye, Andreea Gane, Tommi S. Jaakkola, Josh Tenenbaum |
UAI | 5 |
| 2018 | Unsupervised Learning of Latent Physical Properties Using Perception-Prediction Networks
David Zheng, Vinson Luo, Jiajun Wu 0001, Josh Tenenbaum |
UAI | 4 |
| 2018 | 3D Interpreter Networks for Viewer-Centered Wireframe Modeling
Jiajun Wu 0001, Tianfan Xue, Joseph J. Lim, Yuandong Tian, Josh Tenenbaum, Antonio Torralba 0001, William T. Freeman |
Int. J. Comput. Vis. | 5 |
| 2017 | Simulation and heuristics in flexible tool use
Kelsey R. Allen, Kevin A. Smith 0001, Josh Tenenbaum |
CogSci | 3 |
| 2017 | Learning to Learn Visual Object Categories by Integrating Deep Learning with Hierarchical Bayes
Andres Campero, Andrew Francl, Josh Tenenbaum |
CogSci | 3 |
| 2017 | Faulty Towers: A hypothetical simulation model of physical support
Tobias Gerstenberg, Kevin A. Smith 0001, Josh Tenenbaum |
CogSci | 4 |
| 2017 | Constructing Social Preferences From Anticipated Judgments: When Impartial Inequity is Fair and Why?
Max Kleiman-Weiner, Josh Tenenbaum |
CogSci | 3 |
| 2017 | Cooperative Social Intelligence: Understanding and Acting with Other
Max Kleiman-Weiner, Yibiao Zhao, Josh Tenenbaum |
CogSci | 3 |
| 2017 | One-shot Learning and Classification in Children
Eliza Kosoy, Brenden M. Lake, Josh Tenenbaum |
CogSci | 3 |
| 2017 | Thinking and Guessing: Bayesian and Empirical Models of How Humans Search
Marta Kryven, Tomer D. Ullman, William Cowan, Josh Tenenbaum |
CogSci | 4 |
| 2017 | Preschoolers and Infants Calibrate Persistence from Adult Models
Julia A. Leonard, Max Kleiman-Weiner, Josh Tenenbaum, Laura Schulz |
CogSci | 4 |
| 2017 | What's worth the effort: Ten-month-old infants infer the value of goals from the costs of actions
Shari Liu, Tomer D. Ullman, Josh Tenenbaum, Elizabeth S. Spelke |
CogSci | 3 |
| 2017 | Intuitive psychophysics: Children's exploratory play quantitatively tracks the discriminability of alternative hypotheses
Rachel Magid, Max H. Siegel, Josh Tenenbaum, Laura Schulz |
CogSci | 3 |
| 2017 | Big Data and Little Learners
John C. Trueswell, Linda B. Smith, Josh Tenenbaum, Charles Yang 0001 |
CogSci | 3 |
| 2017 | Interpreting actions by attributing compositional desires
Joey Velez-Ginorio, Max H. Siegel, Josh Tenenbaum, Julian Jara-Ettinger |
CogSci | 3 |
| 2017 | Physical problem solving: Joint planning with symbolic, geometric, and dynamic constraints
Ilker Yildirim, Tobias Gerstenberg, Basil Saeed, Marc Toussaint, Josh Tenenbaum |
CogSci | 5 |
| 2017 | Causal and compositional generative models in online perception
Ilker Yildirim, Michael Janner, Mario Belledonne, Christian Wallraven, Winrich Freiwald, Josh Tenenbaum |
CogSci | 6 |
| 2017 | Workshop proposal: Deep Learning in Computational Cognitive Science
Ilker Yildirim, Josh Tenenbaum |
CogSci | 2 |
| 2017 | Neural Scene De-renderingabstractWe study the problem of holistic scene understanding. We would like to obtain a compact, expressive, and interpretable representation of scenes that encodes information such as the number of objects and their categories, poses, positions, etc. Such a representation would allow us to reason about and even reconstruct or manipulate elements of the scene. Previous works have used encoder-decoder based neural architectures to learn image representations, however, representations obtained in this way are typically uninterpretable, or only explain a single object in the scene. In this work, we propose a new approach to learn an interpretable distributed representation of scenes. Our approach employs a deterministic rendering function as the decoder, mapping a naturally structured and disentangled scene description, which we named scene XML, to an image. By doing so, the encoder is forced to perform the inverse of the rendering operation (a.k.a. de-rendering) to transform an input image to the structured scene XML that the decoder used to produce the image. We use a object proposal based encoder that is trained by minimizing both the supervised prediction and the unsupervised reconstruction errors. Experiments demonstrate that our approach works well on scene de-rendering with two different graphics engines, and our learned representation can be easily adapted for a wide range of applications like image editing, inpainting, visual analogy-making, and image captioning. Jiajun Wu 0001, Josh Tenenbaum, Pushmeet Kohli |
CVPR | 2 |
| 2017 | Synthesizing 3D Shapes via Modeling Multi-view Depth Maps and Silhouettes with Deep Generative NetworksabstractWe study the problem of learning generative models of 3D shapes. Voxels or 3D parts have been widely used as the underlying representations to build complex 3D shapes, however, voxel-based representations suffer from high memory requirements, and parts-based models require a large collection of cached or richly parametrized parts. We take an alternative approach: learning a generative model over multi-view depth maps or their corresponding silhouettes, and using a deterministic rendering function to produce 3D shapes from these images. A multi-view representation of shapes enables generation of 3D models with fine details, as 2D depth maps and silhouettes can be modeled at a much higher resolution than 3D voxels. Moreover, our approach naturally brings the ability to recover the underlying 3D representation from depth maps of one or a few viewpoints. Experiments show that our framework can generate 3D shapes with variations and details. We also demonstrate that our model has out-of-sample generalization power for real-world tasks with occluded objects. Amir Arsalan Soltani, Jiajun Wu 0001, Tejas D. Kulkarni, Josh Tenenbaum |
CVPR | 5 |
| 2017 | Generative Modeling of Audible Shapes for Object PerceptionabstractHumans infer rich knowledge of objects from both auditory and visual cues. Building a machine of such competency, however, is very challenging, due to the great difficulty in capturing large-scale, clean data of objects with both their appearance and the sound they make. In this paper, we present a novel, open-source pipeline that generates audiovisual data, purely from 3D object shapes and their physical properties. Through comparison with audio recordings and human behavioral studies, we validate the accuracy of the sounds it generates. Using this generative model, we are able to construct a synthetic audio-visual dataset, namely Sound-20K, for object perception tasks. We demonstrate that auditory and visual information play complementary roles in object perception, and further, that the representation learned on synthetic audio-visual data can transfer to real-world scenarios. Zhoutong Zhang, Jiajun Wu 0001, Qiujia Li, Zhengjia Huang, James Traer, Josh H. McDermott, Josh Tenenbaum, William T. Freeman |
ICCV | 7 |
| 2017 | A Compositional Object-Based Approach to Learning Physical Dynamics
Michael Chang 0003, Tomer D. Ullman, Antonio Torralba 0001, Josh Tenenbaum |
ICLR (Poster) | 4 |
| 2017 | Learning to See Physics via Visual De-animationabstractWe introduce a paradigm for understanding physical scenes without human annotations. At the core of our system is a physical world representation that is first recovered by a perception module and then utilized by physics and graphics engines. During training, the perception module and the generative models learn by visual de-animation --- interpreting and reconstructing the visual information stream. During testing, the system first recovers the physical world state, and then uses the generative models for reasoning and future prediction. Even more so than forward simulation, inverting a physics or graphics engine is a computationally hard problem; we overcome this challenge by using a convolutional inversion network. Our system quickly recognizes the physical world state from appearance and motion cues, and has the flexibility to incorporate both differentiable and non-differentiable physics and graphics engines. We evaluate our system on both synthetic and real datasets involving multiple physical scenes, and demonstrate that our system performs well on both physical state estimation and reasoning problems. We further show that the knowledge learned on the synthetic dataset generalizes to constrained real images. Jiajun Wu 0001, Erika Lu, Pushmeet Kohli, William T. Freeman, Josh Tenenbaum |
NIPS | 5 |
| 2017 | MarrNet: 3D Shape Reconstruction via 2.5D Sketchesabstract3D object reconstruction from a single image is a highly under-determined problem, requiring strong prior knowledge of plausible 3D shapes. This introduces challenge for learning-based approaches, as 3D object annotations in real images are scarce. Previous work chose to train on synthetic data with ground truth 3D information, but suffered from the domain adaptation issue when tested on real data. In this work, we propose an end-to-end trainable framework, sequentially estimating 2.5D sketches and 3D object shapes. Our disentangled, two-step formulation has three advantages. First, compared to full 3D shape, 2.5D sketches are much easier to be recovered from a 2D image, and to transfer from synthetic to real data. Second, for 3D reconstruction from the 2.5D sketches, we can easily transfer the learned model on synthetic data to real images, as rendered 2.5D sketches are invariant to object appearance variations in real images, including lighting, texture, etc. This further relieves the domain adaptation problem. Third, we derive differentiable projective functions from 3D shape to 2.5D sketches, making the framework end-to-end trainable on real images, requiring no real-image annotations. Our framework achieves state-of-the-art performance on 3D shape reconstruction. Jiajun Wu 0001, Wang Yifan 0001, Tianfan Xue, Xingyuan Sun, William T. Freeman, Josh Tenenbaum |
NIPS | 6 |
| 2017 | Self-Supervised Intrinsic Image DecompositionabstractIntrinsic decomposition from a single image is a highly challenging task, due to its inherent ambiguity and the scarcity of training data. In contrast to traditional fully supervised learning approaches, in this paper we propose learning intrinsic image decomposition by explaining the input image. Our model, the Rendered Intrinsics Network (RIN), joins together an image decomposition pipeline, which predicts reflectance, shape, and lighting conditions given a single image, with a recombination function, a learned shading model used to recompose the original input based off of intrinsic image predictions. Our network can then use unsupervised reconstruction error as an additional signal to improve its intermediate representations. This allows large-scale unlabeled data to be useful during training, and also enables transferring learned knowledge to images of unseen object categories, lighting conditions, and shapes. Extensive experiments demonstrate that our method performs well on both intrinsic image decomposition and knowledge transfer. Michael Janner, Jiajun Wu 0001, Tejas D. Kulkarni, Ilker Yildirim, Josh Tenenbaum |
NIPS | 5 |
| 2017 | Shape and Material from SoundabstractHearing an object falling onto the ground, humans can recover rich information including its rough shape, material, and falling height. In this paper, we build machines to approximate such competency. We first mimic human knowledge of the physical world by building an efficient, physics-based simulation engine. Then, we present an analysis-by-synthesis approach to infer properties of the falling object. We further accelerate the process by learning a mapping from a sound wave to object properties, and using the predicted values to initialize the inference. This mapping can be viewed as an approximation of human commonsense learned from past experience. Our model performs well on both synthetic audio clips and real recordings without requiring any annotated data. We conduct behavior studies to compare human responses with ours on estimating object shape, material, and falling height from sound. Our model achieves near-human performance. Zhoutong Zhang, Qiujia Li, Zhengjia Huang, Jiajun Wu 0001, Josh Tenenbaum, William T. Freeman |
NIPS | 5 |
| 2016 | Modeling Human Ad Hoc CoordinationabstractWhether in groups of humans or groups of computer agents, collaboration is most effective between individuals who have the ability to coordinate on a joint strategy for collective action. However, in general a rational actor will only intend to coordinate if that actor believes the other group members have the same intention. This circular dependence makes rational coordination difficult in uncertain environments if communication between actors is unreliable and no prior agreements have been made. An important normative question with regard to coordination in these ad hoc settings is therefore how one can come to believe that other actors will coordinate, and with regard to systems involving humans, an important empirical question is how humans arrive at these expectations. We introduce an exact algorithm for computing the infinitely recursive hierarchy of graded beliefs required for rational coordination in uncertain environments, and we introduce a novel mechanism for multiagent coordination that uses it. Our algorithm is valid in any environment with a finite state space, and extensions to certain countably infinite state spaces are likely possible. We test our mechanism for multiagent coordination as a model for human decisions in a simple coordination game using existing experimental data. We then explore via simulations whether modeling humans in this way may improve human-agent collaboration. P. M. Krafft, Chris L. Baker, Alex Pentland, Josh Tenenbaum |
AAAI | 4 |
| 2016 | Modeling Human Understanding of Complex Intentional Action with a Bayesian Nonparametric Subgoal ModelabstractMost human behaviors consist of multiple parts, steps, or subtasks. These structures guide our ac- tion planning and execution, but when we observe others, the latent structure of their actions is typ- ically unobservable, and must be inferred in order to learn new skills by demonstration, or to as- sist others in completing their tasks. For example, an assistant who has learned the subgoal struc- ture of a colleague’s task can more rapidly rec- ognize and support their actions as they unfold. Here we model how humans infer subgoals from observations of complex action sequences using a nonparametric Bayesian model, which assumes that observed actions are generated by approxi- mately rational planning over unknown subgoal sequences. We test this model with a behavioral experiment in which humans observed different se- ries of goal-directed actions, and inferred both the number and composition of the subgoal sequences associated with each goal. The Bayesian model predicts human subgoal inferences with high ac- curacy, and significantly better than several al- ternative models and straightforward heuristics. Motivated by this result, we simulate how learn- ing and inference of subgoals can improve perfor- mance in an artificial user assistance task. The Bayesian model learns the correct subgoals from fewer observations, and better assists users by more rapidly and accurately inferring the goal of their actions than alternative approaches. Ryo Nakahashi, Chris L. Baker, Josh Tenenbaum |
AAAI | 3 |
| 2016 | Physics 101: Learning Physical Object Properties from Unlabeled Videos
Jiajun Wu 0001, Joseph J. Lim, Josh Tenenbaum, William T. Freeman |
BMVC | 4 |
| 2016 | Integrating identification and perception: A case study of familiar and unfamiliar face processing
Kelsey R. Allen, Ilker Yildirim, Josh Tenenbaum |
CogSci | 3 |
| 2016 | Natural science: Active learning in dynamic physical microworlds
Neil Bramley, Tobias Gerstenberg, Josh Tenenbaum |
CogSci | 3 |
| 2016 | Understanding "almost": Empirical and computational studies of near misses
Tobias Gerstenberg, Josh Tenenbaum |
CogSci | 2 |
| 2016 | Feature-based Joint Planning and Norm Learning in Collaborative Games
Mark K. Ho, James MacGlashan, Amy Greenwald, Michael L. Littman, Elizabeth Hilliard, Carl Trimbach, Stephen Brawner, Josh Tenenbaum, Max Kleiman-Weiner, Joseph L. Austerweil |
CogSci | 8 |
| 2016 | The Naïve Utility Calculus unifies spatial and statistical routes to preference
Julian Jara-Ettinger, Felix Sun, Laura Schulz, Josh Tenenbaum |
CogSci | 4 |
| 2016 | Coordinate to cooperate or compete: Abstract goals and joint intentions in social interaction
Max Kleiman-Weiner, Mark K. Ho, Joseph L. Austerweil, Michael L. Littman, Josh Tenenbaum |
CogSci | 5 |
| 2016 | Outcome or Strategy? A Bayesian Model of Intelligence Attribution
Marta Kryven, Tomer D. Ullman, William Cowan, Josh Tenenbaum |
CogSci | 4 |
| 2016 | Coalescing the Vapors of Human Experience into a Viable and Meaningful Comprehension
Tomer D. Ullman, Max H. Siegel, Josh Tenenbaum, Samuel Gershman |
CogSci | 3 |
| 2016 | Integrating physical reasoning and visual object recognition for fully occluded scene interpretation
Ilker Yildirim, Max H. Siegel, Josh Tenenbaum |
CogSci | 3 |
| 2016 | A Comparative Evaluation of Approximate Probabilistic Simulation and Deep Neural Networks as Accounts of Human Physical Scene Understanding
Renqiao Zhang, Jiajun Wu 0001, Chengkai Zhang, William T. Freeman, Josh Tenenbaum |
CogSci | 5 |
| 2016 | Single Image 3D Interpreter Network
Jiajun Wu 0001, Tianfan Xue, Joseph J. Lim, Yuandong Tian, Josh Tenenbaum, Antonio Torralba 0001, William T. Freeman |
ECCV (6) | 5 |
| 2016 | Inferring human intent from video by sampling hierarchical plansabstractThis paper presents a method which allows robots to infer a human's hierarchical intent from partially observed RGBD videos by imagining how the human will behave in the future. This capability is critical for creating robots which can interact socially or collaboratively with humans. We represent intent as a novel hierarchical, compositional, and probabilistic And-Or graph structure which describes a relationship between actions and plans. We infer human intent by reverse-engineering a human's decision-making and action planning processes under a Bayesian probabilistic programming framework. We present experiments from a 3D environment which demonstrate that the inferred human intent (1) matches well with human judgment, and (2) provides useful contextual cues for object tracking and action recognition. Steven Holtzen, Yibiao Zhao, Tao Gao 0004, Josh Tenenbaum, Song-Chun Zhu |
IROS | 4 |
| 2016 | Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial ModelingabstractWe study the problem of 3D object generation. We propose a novel framework, namely 3D Generative Adversarial Network (3D-GAN), which generates 3D objects from a probabilistic space by leveraging recent advances in volumetric convolutional networks and generative adversarial nets. The benefits of our model are three-fold: first, the use of an adversarial criterion, instead of traditional heuristic criteria, enables the generator to capture object structure implicitly and to synthesize high-quality 3D objects; second, the generator establishes a mapping from a low-dimensional probabilistic space to the space of 3D objects, so that we can sample objects without a reference image or CAD models, and explore the 3D object manifold; third, the adversarial discriminator provides a powerful 3D shape descriptor which, learned without supervision, has wide applications in 3D object recognition. Experiments demonstrate that our method generates high-quality 3D objects, and our unsupervisedly learned features achieve impressive performance on 3D object recognition, comparable with those of supervised learning methods. Jiajun Wu 0001, Chengkai Zhang, Tianfan Xue, William T. Freeman, Josh Tenenbaum |
NIPS | 5 |
| 2016 | Sampling for Bayesian Program LearningabstractTowards learning programs from data, we introduce the problem of sampling programs from posterior distributions conditioned on that data. Within this setting, we propose an algorithm that uses a symbolic solver to efficiently sample programs. The proposal combines constraint-based program synthesis with sampling via random parity constraints. We give theoretical guarantees on how well the samples approximate the true posterior, and have empirical results showing the algorithm is efficient in practice, evaluating our approach on 22 program learning problems in the domains of text editing and computer-aided programming. Kevin Ellis, Armando Solar-Lezama, Josh Tenenbaum |
NIPS | 3 |
| 2016 | Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic MotivationabstractLearning goal-directed behavior in environments with sparse feedback is a major challenge for reinforcement learning algorithms. One of the key difficulties is insufficient exploration, resulting in an agent being unable to learn robust policies. Intrinsically motivated agents can explore new behavior for their own sake rather than to directly solve external goals. Such intrinsic behaviors could eventually help the agent solve tasks posed by the environment. We present hierarchical-DQN (h-DQN), a framework to integrate hierarchical action-value functions, operating at different temporal scales, with goal-driven intrinsically motivated deep reinforcement learning. A top-level q-value function learns a policy over intrinsic goals, while a lower-level function learns a policy over atomic actions to satisfy the given goals. h-DQN allows for flexible goal specifications, such as functions over entities and relations. This provides an efficient space for exploration in complicated environments. We demonstrate the strength of our approach on two problems with very sparse and delayed feedback: (1) a complex discrete stochastic decision process with stochastic transitions, and (2) the classic ATARI game -- `Montezuma's Revenge'. Tejas D. Kulkarni, Karthik Narasimhan, Ardavan Saeedi, Josh Tenenbaum |
NIPS | 4 |
| 2016 | Probing the Compositionality of Intuitive FunctionsabstractHow do people learn about complex functional structure? Taking inspiration from other areas of cognitive science, we propose that this is accomplished by harnessing compositionality: complex structure is decomposed into simpler building blocks. We formalize this idea within the framework of Bayesian regression using a grammar over Gaussian process kernels. We show that participants prefer compositional over non-compositional function extrapolations, that samples from the human prior over functions are best described by a compositional model, and that people perceive compositional functions as more predictable than their non-compositional but otherwise similar counterparts. We argue that the compositional nature of intuitive functions is consistent with broad principles of human cognition. Eric Schulz, Josh Tenenbaum, David Duvenaud, Maarten Speekenbrink, Samuel Gershman |
NIPS | 2 |
| 2016 | CrossCat: A Fully Bayesian Nonparametric Method for Analyzing Heterogeneous, High Dimensional DataabstractThere is a widespread need for statistical methods that can analyze high-dimensional datasets without imposing restrictive or opaque modeling assumptions. This paper describes a domain- general data analysis method called CrossCat. CrossCat infers multiple non-overlapping views of the data, each consisting of a subset of the variables, and uses a separate nonparametric mixture to model each view. CrossCat is based on approximately Bayesian inference in a hierarchical, nonparametric model for data tables. This model consists of a Dirichlet process mixture over the columns of a data table in which each mixture component is itself an independent Dirichlet process mixture over the rows; the inner mixture components are simple parametric models whose form depends on the types of data in the table. CrossCat combines strengths of mixture modeling and Bayesian network structure learning. Like mixture modeling, CrossCat can model a broad class of distributions by positing latent variables, and produces representations that can be efficiently conditioned and sampled from for prediction. Like Bayesian networks, CrossCat represents the dependencies and independencies between variables, and thus remains accurate when there are multiple statistical signals. Inference is done via a scalable Gibbs sampling scheme; this paper shows that it works well in practice. This paper also includes empirical results on heterogeneous tabular data of up to 10 million cells, such as hospital cost and quality measures, voting records, unemployment rates, gene expression measurements, and images of handwritten digits. CrossCat infers structure that is consistent with accepted findings and common-sense knowledge in multiple domains and yields predictive accuracy competitive with generative, discriminative, and model-free alternatives. Vikash Mansinghka 0001, Patrick Shafto, Eric Jonas, Cap Petschulat, Max Gasner, Josh Tenenbaum |
J. Mach. Learn. Res. | 6 |
| 2015 | Go fishing! Responsibility judgments when cooperation breaks down
Kelsey R. Allen, Julian Jara-Ettinger, Tobias Gerstenberg, Max Kleiman-Weiner, Josh Tenenbaum |
CogSci | 5 |
| 2015 | Humans predict liquid dynamics using probabilistic simulation
Christopher Bates, Peter W. Battaglia, Ilker Yildirim, Josh Tenenbaum |
CogSci | 4 |
| 2015 | Phrase similarity in humans and machines
Samuel Gershman, Josh Tenenbaum |
CogSci | 2 |
| 2015 | How, whether, why: Causal judgments as counterfactual contrasts
Tobias Gerstenberg, Noah D. Goodman, David A. Lagnado, Josh Tenenbaum |
CogSci | 4 |
| 2015 | Responsibility judgments in voting scenarios
Tobias Gerstenberg, Joseph Y. Halpern, Josh Tenenbaum |
CogSci | 3 |
| 2015 | Language & common sense: Integrating across psychology, linguistics, and computer science
Joshua K. Hartshorne, Josh Tenenbaum |
CogSci | 2 |
| 2015 | Beliefs about desires: Children's understanding of how knowledge and preference influence choice
Julian Jara-Ettinger, Emily Lydic, Josh Tenenbaum, Laura Schulz |
CogSci | 3 |
| 2015 | The naïve utility calculus: Joint inferences about the costs and rewards of actions
Julian Jara-Ettinger, Laura Schulz, Josh Tenenbaum |
CogSci | 3 |
| 2015 | Inference of Intention and Permissibility in Moral Decision Making
Max Kleiman-Weiner, Tobias Gerstenberg, Sydney Levine, Josh Tenenbaum |
CogSci | 4 |
| 2015 | Emergent Collective Sensing in Human Groups
P. M. Krafft, Robert D. Hawkins, Alex Pentland, Noah D. Goodman, Josh Tenenbaum |
CogSci | 5 |
| 2015 | Representing and Learning a Large System of Number Concepts with Latent Predicate Networks
Joshua S. Rule, Eyal Dechter, Josh Tenenbaum |
CogSci | 3 |
| 2015 | Assessing the Perceived Predictability of Functions
Eric Schulz, Josh Tenenbaum, David N. Reshef, Maarten Speekenbrink, Samuel Gershman |
CogSci | 2 |
| 2015 | Hypothesis-Space Constraints in Causal Learning
Pedro Tsividis, Josh Tenenbaum, Laura Schulz |
CogSci | 2 |