VLDB 2026 Research / reviewers in the wild / expert
Leslie Pack Kaelbling
dblp:k/LesliePackKaelbling
· DBLP profile ↗
164ranked-venue papers
19as first author
35since 2021 · last 2025
0000-0001-6054-7145ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 163 · 19 first-author · 35 since 2021Systems, architecture and hardware · 53 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 8 since 2021Theory of computation · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Flow-based Domain Randomization for Learning and Sequencing Robotic SkillsabstractDomain randomization in reinforcement learning is an established technique for increasing the robustness of control policies learned in simulation. By randomizing properties of the environment during training, the learned policy can be robust to uncertainty along the randomized dimensions. While the environment distribution is typically specified by hand, in this paper we investigate the problem of automatically discovering this sampling distribution via entropy-regularized reward maximization of a neural sampling distribution in the form of a normalizing flow. We show that this architecture is more flexible and results in better robustness than existing approaches to learning simple parameterized sampling distributions. We demonstrate that these policies can be used to learn robust policies for contact-rich assembly tasks. Additionally, we explore how these sampling distributions, in combination with a privileged value function, can be used for out-of-distribution detection in the context of an uncertainty-aware multi-step manipulation planner. Aidan Curtis, Michael Noseworthy, Nishad Gothoskar, Sachin Chitta, Leslie Pack Kaelbling, Nicole Carey |
ICML | 7 |
| 2025 | KALM: Keypoint Abstraction Using Large Models for Object-Relative Imitation LearningabstractGeneralization to novel object configurations and instances across diverse tasks and environments is a critical challenge in robotics. Keypoint-based representations have been proven effective as a succinct representation for capturing essential object features, and for establishing a reference frame in action prediction, enabling data-efficient learning of robot skills. However, their manual design nature and reliance on additional human labels limit their scalability. In this paper, we propose KALM, a framework that leverages large pre-trained vision-language models (LMs) to automatically generate taskrelevant and cross-instance consistent keypoints. KALM distills robust and consistent keypoints across views and objects by generating proposals using LMs and verifies them against a small set of robot demonstration data. Based on the generated keypoints, we can train keypoint-conditioned policy models that predict actions in keypoint-centric frames, enabling robots to generalize effectively across varying object poses, camera views, and object instances with similar functional shapes. Our method demonstrates strong performance in the real world, adapting to different tasks and environments from only a handful of demonstrations while requiring no additional labels. Videos can be found at https://kalm-il.github.io/. Xiaolin Fang 0002, Bo-Ruei Huang, Jiayuan Mao, Jasmine Shone, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICRA | 7 |
| 2025 | One-Shot Manipulation Strategy Learning by Making Contact AnalogiesabstractWe present a novel approach, MAGIC (manipulation analogies for generalizable intelligent contacts), for one-shot learning of manipulation strategies with fast and extensive generalization to novel objects. By leveraging a reference action trajectory, MAGIC effectively identifies similar contact points and sequences of actions on novel objects to replicate a demonstrated strategy, such as using different hooks to retrieve distant objects of different shapes and sizes. Our method is based on a twostage contact-point matching process that combines global shape matching using pretrained neural features with local curvature analysis to ensure precise and physically plausible contact points. We experiment with three tasks including scooping, hanging, and hooking objects. MAGIC demonstrates superior performance over existing methods, achieving significant improvements in runtime speed and generalization to different object categories. Website: https://magic-2024.github.io/. Yuyao Liu, Jiayuan Mao, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICRA | 5 |
| 2025 | Guiding Long-Horizon Task and Motion Planning with Vision Language ModelsabstractVision-Language Models (VLM) can generate plausible high-level plans when prompted with a goal, the context, an image of the scene, and any planning constraints. However, there is no guarantee that the predicted actions are geometrically and kinematically feasible for a particular robot embodiment. As a result, many prerequisite steps such as opening drawers to access objects are often omitted in their plans. Robot task and motion planners can generate motion trajectories that respect the geometric feasibility of actions and insert physically necessary actions, but do not scale to everyday problems that require common-sense knowledge and involve large state spaces comprised of many variables. We propose VLM-TAMP, a hierarchical planning algorithm that leverages a VLM to generate both semantically-meaningful and horizon-reducing intermediate subgoals that guide a task and motion planner. When a subgoal or action cannot be refined, the VLM is queried again for replanning. We evaluate VLMTAMP on kitchen tasks where a robot must accomplish cooking goals that require performing 30-50 actions in sequence and interacting with up to 21 objects. VLM-TAMP substantially outperforms baselines that rigidly and independently execute VLM-generated action sequences, both in terms of success rates (50 to 100 % versus 0 %) and average task completion percentage (72 to 100 % versus 15 to 45 %). See project site https://zt-yang.github.io/vlm-tamp-robot/ for more information. Zhutian Yang, Caelan Reed Garrett, Dieter Fox, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICRA | 5 |
| 2024 | Generalized Planning in PDDL Domains with Pretrained Large Language ModelsabstractRecent work has considered whether large language models (LLMs) can function as planners: given a task, generate a plan. We investigate whether LLMs can serve as generalized planners: given a domain and training tasks, generate a program that efficiently produces plans for other tasks in the domain. In particular, we consider PDDL domains and use GPT-4 to synthesize Python programs. We also consider (1) Chain-of-Thought (CoT) summarization, where the LLM is prompted to summarize the domain and propose a strategy in words before synthesizing the program; and (2) automated debugging, where the program is validated with respect to the training tasks, and in case of errors, the LLM is re-prompted with four types of feedback. We evaluate this approach in seven PDDL domains and compare it to four ablations and four baselines. Overall, we find that GPT-4 is a surprisingly powerful generalized planner. We also conclude that automated debugging is very important, that CoT summarization has non-uniform impact, that GPT-4 is far superior to GPT-3.5, and that just two training tasks are often sufficient for strong generalization. Tom Silver, Soham Dan, Kavitha Srinivas, Josh Tenenbaum, Leslie Pack Kaelbling, Michael Katz 0001 |
AAAI | 5 |
| 2024 | Video Language PlanningabstractWe are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present video language planning (VLP), an algorithm that consists of a tree search procedure, where we train (i) vision-language models to serve as both policies and value functions, and (ii) text-to-video models as dynamics models. VLP takes as input a long-horizon task instruction and current image observation, and outputs a long video plan that provides detailed multimodal (video and language) specifications that describe how to complete the final task. VLP scales with increasing computation budget where more computation time results in improved video plans, and is able to synthesize long-horizon video plans across different robotics domains -- from multi-object rearrangement, to multi-camera bi-arm dexterous manipulation. Generated video plans can be translated into real robot actions via goal-conditioned policies, conditioned on each intermediate frame of the generated video. Experiments show that VLP substantially improves long-horizon task success rates compared to prior methods on both simulated and real robots (across 3 hardware platforms). Yilun Du, Sherry Yang 0001, Peter R. Florence, Fei Xia 0002, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Josh Tenenbaum, Leslie Pack Kaelbling, Andy Zeng 0001, Jonathan Tompson |
ICLR | 11 |
| 2024 | Learning Interactive Real-World SimulatorsabstractGenerative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans, robots, and other interactive agents. Applications of a real-world simulator range from controllable content creation in games and movies, to training embodied agents purely in simulation that can be directly deployed in the real world. We explore the possibility of learning a universal simulator (UniSim) of real-world interaction through generative modeling. We first make the important observation that natural datasets available for learning a real-world simulator are often rich along different axes (e.g., abundant objects in image data, densely sampled actions in robotics data, and diverse movements in navigation data). With careful orchestration of diverse datasets, each providing a different aspect of the overall experience, UniSim can emulate how humans and agents interact with the world by simulating the visual outcome of both high-level instructions such as “open the drawer” and low-level controls such as “move by x,y” from otherwise static scenes and objects. There are numerous use cases for such a real-world simulator. As an example, we use UniSim to train both high-level vision-language planners and low-level reinforcement learning policies, each of which exhibit zero-shot real-world transfer after training purely in a learned real-world simulator. We also show that other types of intelligence such as video captioning models can benefit from training with simulated experience in UniSim, opening up even wider applications. Sherry Yang 0001, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, Pieter Abbeel |
ICLR | 5 |
| 2024 | Position: Compositional Generative Modeling: A Single Model is Not All You NeedabstractLarge monolithic generative models trained on massive amounts of data have become an increasingly dominant approach in AI research. In this paper, we argue that we should instead construct large generative systems by composing smaller generative models together. We show how such a compositional generative approach enables us to learn distributions in a more data-efficient manner, enabling generalization to parts of the data distribution unseen at training time. We further show how this enables us to program and construct new generative models for tasks completely unseen at training. Finally, we show that in many cases, we can discover separate compositional components from data. Yilun Du, Leslie Pack Kaelbling |
ICML | 2 |
| 2024 | Scaling Exponents Across Parameterizations and OptimizersabstractRobust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization and optimizer choices. In this work, we propose a new perspective on parameterization by investigating a key assumption in prior work about the alignment between parameters and data and derive new theoretical results under weaker assumptions and a broader set of optimizers. Our extensive empirical investigation includes *tens of thousands* of models trained with *all combinations of* three optimizers, four parameterizations, several alignment assumptions, more than a dozen learning rates, and fourteen model sizes up to 27B parameters. We find that the best learning rate scaling prescription would often have been excluded by the assumptions in prior work. Our results show that all parameterizations, not just maximal update parameterization (muP), can achieve hyperparameter transfer; moreover, our novel per-layer learning rate prescription for standard parameterization outperforms muP. Finally, we demonstrate that an overlooked aspect of parameterization, the epsilon parameter in Adam, must be scaled correctly to avoid gradient underflow and propose *Adam-atan2*, a new numerically stable, scale-invariant version of Adam that eliminates the epsilon hyperparameter entirely. Katie Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi, Roman Novak, Peter J. Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee 0001, Jeffrey Pennington |
ICML | 9 |
| 2024 | DiMSam: Diffusion Models as Samplers for Task and Motion Planning under Partial ObservabilityabstractGenerative models such as diffusion models, excel at capturing high-dimensional distributions with diverse input modalities, e.g. robot trajectories, but are less effective at multistep constraint reasoning. Task and Motion Planning (TAMP) approaches are suited for planning multi-step autonomous robot manipulation. However, it can be difficult to apply them to domains where the environment and its dynamics are not fully known. We propose to overcome these limitations by composing diffusion models using a TAMP system. We use the learned components for constraints and samplers that are difficult to engineer in the planning model, and use a TAMP solver to search for the task plan with constraint-satisfying action parameter values. To tractably make predictions for unseen objects in the environment, we define the learned samplers and TAMP operators on learned latent embedding of changing object states. We evaluate our approach in a simulated articulated object manipulation domain and show how the combination of classical TAMP, generative modeling, and latent embedding enables multi-step constraint-based reasoning. We also apply the learned sampler in the real world. Website: https://sites.google.com/view/dimsam-tamp. Xiaolin Fang 0002, Caelan Reed Garrett, Clemens Eppner, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Dieter Fox |
IROS | 5 |
| 2024 | Embodied Uncertainty-Aware Object SegmentationabstractWe introduce uncertainty-aware object instance segmentation (UncOS) and demonstrate its usefulness for embodied interactive segmentation. To deal with uncertainty in robot perception, we propose a method for generating a hypothesis distribution of object segmentation. We obtain a set of region-factored segmentation hypotheses together with confidence estimates by making multiple queries of large pre-trained models. This process can produce segmentation results that achieve state-of-the-art performance on unseen object segmentation problems. The output can also serve as input to a belief-driven process for selecting robot actions to perturb the scene to reduce ambiguity. We demonstrate the effectiveness of this method in real-robot experiments. Website: https://sites.google.com/view/embodied-uncertain-seg. Xiaolin Fang 0002, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 2 |
| 2023 | Learning Rational Subgoals from Demonstrations and InstructionsabstractWe present a framework for learning useful subgoals that support efficient long-term planning to achieve novel goals. At the core of our framework is a collection of rational subgoals (RSGs), which are essentially binary classifiers over the environmental states. RSGs can be learned from weakly-annotated data, in the form of unsegmented demonstration trajectories, paired with abstract task descriptions, which are composed of terms initially unknown to the agent (e.g., collect-wood then craft-boat then go-across-river). Our framework also discovers dependencies between RSGs, e.g., the task collect-wood is a helpful subgoal for the task craft-boat. Given a goal description, the learned subgoals and the derived dependencies facilitate off-the-shelf planning algorithms, such as A* and RRT, by setting helpful subgoals as waypoints to the planner, which significantly improves performance-time efficiency. Project page: https://rsg.csail.mit.edu Zhezheng Luo, Jiayuan Mao, Jiajun Wu 0001, Tomás Lozano-Pérez, Josh Tenenbaum, Leslie Pack Kaelbling |
AAAI | 6 |
| 2023 | Predicate Invention for Bilevel PlanningabstractEfficient planning in continuous state and action spaces is fundamentally hard, even when the transition model is deterministic and known. One way to alleviate this challenge is to perform bilevel planning with abstractions, where a high-level search for abstract plans is used to guide planning in the original transition space. Previous work has shown that when state abstractions in the form of symbolic predicates are hand-designed, operators and samplers for bilevel planning can be learned from demonstrations. In this work, we propose an algorithm for learning predicates from demonstrations, eliminating the need for manually specified state abstractions. Our key idea is to learn predicates by optimizing a surrogate objective that is tractable but faithful to our real efficient-planning objective. We use this surrogate objective in a hill-climbing search over predicate sets drawn from a grammar. Experimentally, we show across four robotic planning environments that our learned abstractions are able to quickly solve held-out tasks, outperforming six baselines. Tom Silver, Rohan Chitnis, Nishanth Kumar, Willie McClinton, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Josh Tenenbaum |
AAAI | 6 |
| 2023 | Local Neural Descriptor Fields: Locally Conditioned Object Representations for ManipulationabstractA robot operating in a household environment will see a wide range of unique and unfamiliar objects. While a system could train on many of these, it is infeasible to predict all the objects a robot will see. In this paper, we present a method to generalize object manipulation skills acquired from a limited number of demonstrations, to novel objects from unseen shape categories. Our approach, Local Neural Descriptor Fields (L-NDF), utilizes neural descriptors defined on the local geometry of the object to effectively transfer manipulation demonstrations to novel objects at test time. In doing so, we leverage the local geometry shared between objects to produce a more general manipulation framework. We illustrate the efficacy of our approach in manipulating novel objects in novel poses - both in simulation and in the real world. Project website, videos, and code: https://elchun.github.io/lndf/. Ethan Chun, Yilun Du, Anthony Simeonov, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICRA | 5 |
| 2023 | Task-Directed Exploration in Continuous POMDPs for Robotic Manipulation of Articulated ObjectsabstractRepresenting and reasoning about uncertainty is crucial for autonomous agents acting in partially observable environments with noisy sensors. Partially observable Markov decision processes (POMDPs) serve as a general framework for representing problems in which uncertainty is an important factor. Online sample-based POMDP methods have emerged as efficient approaches to solving large POMDPs and have been shown to extend to continuous domains. However, these solutions struggle to find long-horizon plans in problems with significant uncertainty. Exploration heuristics can help guide planning, but many real-world settings contain significant task-irrelevant uncertainty that might distract from the task objective. In this paper, we propose STRUG, an online POMDP solver capable of handling domains that require long-horizon planning with significant task-relevant and task-irrelevant uncertainty. We demonstrate our solution on several temporally extended versions of toy POMDP problems as well as robotic manipulation of articulated objects using a neural perception frontend to construct a distribution of possible models. Our results show that STRUG outperforms the current sample-based online POMDP solvers on several tasks. Aidan Curtis, Leslie Pack Kaelbling, Siddarth Jain |
ICRA | 2 |
| 2023 | Visibility-Aware Navigation Among Movable ObstaclesabstractIn this paper, we examine the problem of visibility-aware robot navigation among movable obstacles (VANAMO). A variant of the well-known NAMO robotic planning problem, VANAMO puts additional visibility constraints on robot motion and object movability. This new problem formulation lifts the restrictive assumption that the map is fully visible and the object positions are fully known. We provide a formal definition of the VANAMO problem and propose the Look and Manipulate Backchaining (LAMB) algorithm for solving such problems. Lamb has a simple vision-based interface that makes it more easily transferable to real-world robot applications and scales to the large 3D environments. To evaluate Lamb, we construct a set of tasks that illustrate the complex interplay between visibility and object movability that can arise in mobile base manipulation problems in unknown environments. We show that Lamb outperforms NAMO and visibility-aware motion planning approaches as well as simple combinations of them on complex manipulation problems with partial observability. Jose Muguira-Iturralde, Aidan Curtis, Yilun Du, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 4 |
| 2023 | Compositional Foundation Models for Hierarchical PlanningabstractTo make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with the devised plan through visual-motor control. We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model which leverages multiple expert foundation model trained on language, vision and action data individually jointly together to solve long-horizon tasks. We use a large language model to construct symbolic plans that are grounded in the environment through a large video diffusion model. Generated video plans are then grounded to visual-motor control, through an inverse dynamics model that infers actions from generated videos. To enable effective reasoning within this hierarchy, we enforce consistency between the models via iterative refinement. We illustrate the efficacy and adaptability of our approach in three different long-horizon table-top manipulation tasks. Anurag Ajay, Seungwook Han, Yilun Du, Shuang Li 0013, Abhi Gupta, Tommi S. Jaakkola, Josh Tenenbaum, Leslie Pack Kaelbling, Akash Srivastava, Pulkit Agrawal 0001 |
NeurIPS | 8 |
| 2023 | What Planning Problems Can A Relational Neural Network Solve?abstractGoal-conditioned policies are generally understood to be "feed-forward" circuits, in the form of neural networks that map from the current state and the goal specification to the next action to take. However, under what circumstances such a policy can be learned and how efficient the policy will be are not well understood. In this paper, we present a circuit complexity analysis for relational neural networks (such as graph neural networks and transformers) representing policies for planning problems, by drawing connections with serialized goal regression search (S-GRS). We show that there are three general classes of planning problems, in terms of the growth of circuit width and depth as a function of the number of objects and planning horizon, providing constructive proofs. We also illustrate the utility of this analysis for designing neural networks for policy learning. Jiayuan Mao, Tomás Lozano-Pérez, Josh Tenenbaum, Leslie Pack Kaelbling |
NeurIPS | 4 |
| 2022 | Discovering State and Action Abstractions for Generalized Task and Motion PlanningabstractGeneralized planning accelerates classical planning by finding an algorithm-like policy that solves multiple instances of a task. A generalized plan can be learned from a few training examples and applied to an entire domain of problems. Generalized planning approaches perform well in discrete AI planning problems that involve large numbers of objects and extended action sequences to achieve the goal. In this paper, we propose an algorithm for learning features, abstractions, and generalized plans for continuous robotic task and motion planning (TAMP) and examine the unique difficulties that arise when forced to consider geometric and physical constraints as a part of the generalized plan. Additionally, we show that these simple generalized plans learned from only a handful of examples can be used to improve the search efficiency of TAMP solvers. Aidan Curtis, Tom Silver, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
AAAI | 5 |
| 2022 | Long-Horizon Manipulation of Unknown Objects via Task and Motion Planning with Estimated AffordancesabstractWe present a strategy for designing and building very general robot manipulation systems using a general-purpose task-and-motion planner with both engineered and learned modules that estimate properties and affordances of unknown objects. Such systems are closed-loop policies that map from RGB images, depth images, and robot joint encoder measurements to robot joint position commands. We show that this strategy leads to intelligent behaviors even without a priori knowledge regarding the set of objects, their geometries, and their affordances. We show how these modules can be flexibly composed with robot-centric primitives using the PDDLStream task and motion planning framework. Finally, we demonstrate that this strategy can enable a single policy to perform a wide variety of real-world multi-step manipulation tasks, generalizing over a broad class of objects, arrangements, and goals, without prior knowledge of the environment or re-training. Aidan Curtis, Xiaolin Fang 0002, Leslie Pack Kaelbling, Tomás Lozano-Pérez, Caelan Reed Garrett |
ICRA | 3 |
| 2022 | Fully Persistent Spatial Data Structures for Efficient Queries in Path-Dependent Motion Planning ApplicationsabstractMotion planning is a ubiquitous problem that is often a bottleneck in robotic applications. We demonstrate that motion planning problems such as minimum constraint removal, belief-space planning, and visibility-aware motion planning (VAMP) benefit from a path-dependent formulation, in which the state at a search node is represented implicitly by the path to that node. A naïve approach to computing the feasibility of a successor node in such a path-dependent formulation takes time linear in the path length to the node, in contrast to a (possibly very large) constant time for a more typical search formulation. For long-horizon plans, performing this linear-time computation, which we call the lookback, for each node becomes prohibitive. To improve upon this, we introduce the use of a fully persistent spatial data structure (FPSDS), which bounds the size of the lookback. We then focus on the application of the FPSDS in VAMP, which involves incremental geometric computations that can be accelerated by filtering configurations with bounding volumes using nearest-neighbor data structures. We demonstrate an asymptotic and practical improvement in the runtime of finding VAMP solutions in several illustrative domains. To the best of our knowledge, this is the first use of a fully persistent data structure for accelerating motion planning. Sathwik Karnik, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Gustavo Goretkin |
ICRA | 3 |
| 2022 | PG3: Policy-Guided Planning for Generalized Policy GenerationabstractA longstanding objective in classical planning is to synthesize policies that generalize across multiple problems from the same domain. In this work, we study generalized policy search-based methods with a focus on the score function used to guide the search over policies. We demonstrate limitations of two score functions --- policy evaluation and plan comparison --- and propose a new approach that overcomes these limitations. The main idea behind our approach, Policy-Guided Planning for Generalized Policy Generalization (PG3), is that a candidate policy should be used to guide planning on training problems as a mechanism for evaluating that candidate. Theoretical results in a simplified setting give conditions under which PG3 is optimal or admissible. We then study a specific instantiation of policy search where planning problems are PDDL-based and policies are lifted decision lists. Empirical results in six domains confirm that PG3 learns generalized policies more efficiently and effectively than several baselines. Ryan Yang, Tom Silver, Aidan Curtis, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IJCAI | 5 |
| 2022 | Learning Neuro-Symbolic Relational Transition Models for Bilevel PlanningabstractIn robotic domains, learning and planning are complicated by continuous state spaces, continuous action spaces, and long task horizons. In this work, we address these challenges with Neuro-Symbolic Relational Transition Models (NSRTs), a novel class of models that are data-efficient to learn, compatible with powerful robotic planning methods, and generalizable over objects. NSRTs have both symbolic and neural components, enabling a bilevel planning scheme where symbolic AI planning in an outer loop guides continuous planning with neural models in an inner loop. Experiments in four robotic planning domains show that NSRTs can be learned very data-efficiently, and then used for fast planning in new tasks that require up to 60 actions and involve many more objects than were seen during training. Rohan Chitnis, Tom Silver, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IROS | 5 |
| 2022 | Learning Object-Based State Estimators for Household RobotsabstractA robot operating in a household makes observations of multiple objects as it moves around over the course of days or weeks. The objects may be moved by inhabitants, but not completely at random. The robot may be called upon later to retrieve objects and will need a long-term object-based memory in order to know how to find them. Existing work in semantic SLAM does not attempt to capture the dynamics of object movement. In this paper, we combine some aspects of classic techniques for data-association filtering with modern attention-based neural networks to construct object-based memory systems that operate on high-dimensional observations and hypotheses. We perform end-to-end learning on labeled observation trajectories to learn both the transition and observation models. We demonstrate the system's effectiveness in maintaining memory of dynamically changing objects in both simulated environment and real images, and demonstrate improvements over classical structured approaches as well as unstructured neural approaches. Additional information available at project website: https://yilundu.github.io/obm/. Yilun Du, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IROS | 3 |
| 2022 | PDSketch: Integrated Domain Programming, Learning, and PlanningabstractThis paper studies a model learning and online planning approach towards building flexible and general robots. Specifically, we investigate how to exploit the locality and sparsity structures in the underlying environmental transition model to improve model generalization, data-efficiency, and runtime-efficiency. We present a new domain definition language, named PDSketch. It allows users to flexibly define high-level structures in the transition models, such as object and feature dependencies, in a way similar to how programmers use TensorFlow or PyTorch to specify kernel sizes and hidden dimensions of a convolutional neural network. The details of the transition model will be filled in by trainable neural networks. Based on the defined structures and learned parameters, PDSketch automatically generates domain-independent planning heuristics without additional training. The derived heuristics accelerate the performance-time planning for novel goals. Jiayuan Mao, Tomás Lozano-Pérez, Josh Tenenbaum, Leslie Pack Kaelbling |
NeurIPS | 4 |
| 2021 | GLIB: Efficient Exploration for Relational Model-Based Reinforcement Learning via Goal-Literal BabblingabstractWe address the problem of efficient exploration for transition model learning in the relational model-based reinforcement learning setting without extrinsic goals or rewards. Inspired by human curiosity, we propose goal-literal babbling (GLIB), a simple and general method for exploration in such problems. GLIB samples relational conjunctive goals that can be understood as specific, targeted effects that the agent would like to achieve in the world, and plans to achieve these goals using the transition model being learned. We provide theoretical guarantees showing that exploration with GLIB will converge almost surely to the ground truth model. Experimentally, we find GLIB to strongly outperform existing methods in both prediction and planning on a range of tasks, encompassing standard PDDL and PPDDL planning benchmarks and a robotic manipulation task implemented in the PyBullet physics simulator. Video: https://youtu.be/F6lmrPT6TOY Code: https://git.io/JIsTB Rohan Chitnis, Tom Silver, Josh Tenenbaum, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 4 |
| 2021 | Planning with Learned Object Importance in Large Problem Instances using Graph Neural NetworksabstractReal-world planning problems often involve hundreds or even thousands of objects, straining the limits of modern planners. In this work, we address this challenge by learning to predict a small set of objects that, taken together, would be sufficient for finding a plan. We propose a graph neural network architecture for predicting object importance in a single inference pass, thus incurring little overhead while greatly reducing the number of objects that must be considered by the planner. Our approach treats the planner and transition model as black boxes, and can be used with any off-the-shelf planner. Empirically, across classical planning, probabilistic planning, and robotic task and motion planning, we find that our method results in planning that is significantly faster than several baselines, including other partial grounding strategies and lifted planners. We conclude that learning to predict a sufficient set of objects for a planning problem is a simple, powerful, and general mechanism for planning in large instances. Video: https://youtu.be/FWsVJc2fvCE Code: https://git.io/JIsqX Tom Silver, Rohan Chitnis, Aidan Curtis, Josh Tenenbaum, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
AAAI | 6 |
| 2021 | A large-scale benchmark for few-shot program induction and synthesisabstractA landmark challenge for AI is to learn flexible, powerful representations from small numbers of examples. On an important class of tasks, hypotheses in the form of programs provide extreme generalization capabilities from surprisingly few examples. However, whereas large natural few-shot learning image benchmarks have spurred progress in meta-learning for deep networks, there is no comparably big, natural program-synthesis dataset that can play a similar role. This is because, whereas images are relatively easy to label from internet meta-data or annotated by non-experts, generating meaningful input-output examples for program induction has proven hard to scale. In this work, we propose a new way of leveraging unit tests and natural inputs for small programs as meaningful input-output examples for each sub-program of the overall program. This allows us to create a large-scale naturalistic few-shot program-induction benchmark and propose new challenges in this domain. The evaluation of multiple program induction and synthesis algorithms points to shortcomings of current methods and suggests multiple avenues for future work. Ferran Alet, Javier Lopez-Contreras, James Koppel, Maxwell I. Nye, Armando Solar-Lezama, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Josh Tenenbaum |
ICML | 7 |
| 2021 | Shape-Based Transfer of Generic SkillsabstractWe propose a new, data-efficient approach for skill transfer to novel objects, accounting for known categorical shape variation. A low-dimensional shape representation embedding is learned from a set of deformations, sampled between known objects within a category. This latent representation is mapped to a set of control parameters that result in successful execution of a category-level skill on that object. This method generalizes a learned manipulation policy to unseen objects with few training examples. We demonstrate this approach on pouring from cups and scooping with spatulas, where there is complex, nonlinear variation of successful control parameters across objects. Skye Thompson, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2021 | Temporal and Object Quantification NetworksabstractWe present Temporal and Object Quantification Networks (TOQ-Nets), a new class of neuro-symbolic networks with a structural bias that enables them to learn to recognize complex relational-temporal events. This is done by including reasoning layers that implement finite-domain quantification over objects and time. The structure allows them to generalize directly to input instances with varying numbers of objects in temporal sequences of varying lengths. We evaluate TOQ-Nets on input domains that require recognizing event-types in terms of complex temporal relational patterns. We demonstrate that TOQ-Nets can generalize from small amounts of data to scenarios containing more objects than were present during training and to temporal warpings of input sequences. Jiayuan Mao, Zhezheng Luo, Chuang Gan 0001, Josh Tenenbaum, Jiajun Wu 0001, Leslie Pack Kaelbling, Tomer D. Ullman |
IJCAI | 6 |
| 2021 | Learning Symbolic Operators for Task and Motion PlanningabstractRobotic planning problems in hybrid state and action spaces can be solved by integrated task and motion planners (TAMP) that handle the complex interaction between motion-level decisions and task-level plan feasibility. TAMP approaches rely on domain-specific symbolic operators to guide the task-level search, making planning efficient. In this work, we formalize and study the problem of operator learning for TAMP. Central to this study is the view that operators define a lossy abstraction of the transition model of a domain. We then propose a bottom-up relational learning method for operator learning and show how the learned operators can be used for planning in a TAMP system. Experimentally, we provide results in three domains, including long-horizon robotic planning tasks. We find our approach to substantially outperform several baselines, including three graph neural network-based model-free approaches from the recent literature. Video: https://youtu.be/iVfpX9BpBRo. Code: https://git.io/JCT0g Tom Silver, Rohan Chitnis, Josh Tenenbaum, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 4 |
| 2021 | Learning When to Quit: Meta-Reasoning for Motion PlanningabstractAnytime motion planners are widely used in robotics. However, the relationship between their solution quality and computation time is not well understood, and thus, determining when to quit planning and start execution is unclear. In this paper, we address the problem of deciding when to stop deliberation under bounded computational capacity, so called meta-reasoning, for anytime motion planning. We propose data-driven learning methods, model-based and model-free meta-reasoning, that are applicable to different environment distributions and agnostic to the choice of anytime motion planners. As a part of the framework, we design a convolutional neural network-based optimal solution predictor that predicts the optimal path length from a given 2D workspace image. We empirically evaluate the performance of the proposed methods in simulation in comparison with baselines. Yoonchang Sung, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 2 |
| 2021 | Tailoring: encoding inductive biases by optimizing unsupervised objectives at prediction timeabstractFrom CNNs to attention mechanisms, encoding inductive biases into neural networks has been a fruitful source of improvement in machine learning. Adding auxiliary losses to the main objective function is a general way of encoding biases that can help networks learn better representations. However, since auxiliary losses are minimized only on training data, they suffer from the same generalization gap as regular task losses. Moreover, by adding a term to the loss function, the model optimizes a different objective than the one we care about. In this work we address both problems: first, we take inspiration from transductive learning and note that after receiving an input but before making a prediction, we can fine-tune our networks on any unsupervised loss. We call this process tailoring, because we customize the model to each input to ensure our prediction satisfies the inductive bias. Second, we formulate meta-tailoring, a nested optimization similar to that in meta-learning, and train our models to perform well on the task objective after adapting them using an unsupervised loss. The advantages of tailoring and meta-tailoring are discussed theoretically and demonstrated empirically on a diverse set of examples. Ferran Alet, Maria Bauzá 0001, Kenji Kawaguchi, Nurullah Giray Kuru, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
NeurIPS | 6 |
| 2021 | Understanding End-to-End Model-Based Reinforcement Learning Methods as Implicit ParameterizationabstractEstimating the per-state expected cumulative rewards is a critical aspect of reinforcement learning approaches, however the experience is obtained, but standard deep neural-network function-approximation methods are often inefficient in this setting. An alternative approach, exemplified by value iteration networks, is to learn transition and reward models of a latent Markov decision process whose value predictions fit the data. This approach has been shown empirically to converge faster to a more robust solution in many cases, but there has been little theoretical study of this phenomenon. In this paper, we explore such implicit representations of value functions via theory and focused experimentation. We prove that, for a linear parametrization, gradient descent converges to global optima despite non-linearity and non-convexity introduced by the implicit representation. Furthermore, we derive convergence rates for both cases which allow us to identify conditions under which stochastic gradient descent (SGD) with this implicit representation converges substantially faster than its explicit counterpart. Finally, we provide empirical results in some simple domains that illustrate the theoretical findings. Clement Gehring, Kenji Kawaguchi, Jiaoyang Huang, Leslie Pack Kaelbling |
NeurIPS | 4 |
| 2021 | A Sufficient Statistic for Influence in Structured Multiagent EnvironmentsabstractMaking decisions in complex environments is a key challenge in artificial intelligence (AI). Situations involving multiple decision makers are particularly complex, leading to computational intractability of principled solution methods. A body of work in AI has tried to mitigate this problem by trying to distill interaction to its essence: how does the policy of one agent influence another agent? If we can find more compact representations of such influence, this can help us deal with the complexity, for instance by searching the space of influences rather than the space of policies. However, so far these notions of influence have been restricted in their applicability to special cases of interaction. In this paper we formalize influence-based abstraction (IBA), which facilitates the elimination of latent state factors without any loss in value, for a very general class of problems described as factored partially observable stochastic games (fPOSGs). On the one hand, this generalizes existing descriptions of influence, and thus can serve as the foundation for improvements in scalability and other insights in decision making in complex multiagent settings. On the other hand, since the presence of other agents can be seen as a generalization of single agent settings, our formulation of IBA also provides a sufficient statistic for decision making under abstraction for a single agent. We also give a detailed discussion of the relations to such previous works, identifying new insights and interpretations of these approaches. In these ways, this paper deepens our understanding of abstraction in a wide range of sequential decision making settings, providing the basis for new approaches and algorithms for a large class of problems. Frans A. Oliehoek, Stefan J. Witwicki, Leslie Pack Kaelbling |
J. Artif. Intell. Res. | 3 |
| 2020 | Monte Carlo Tree Search in Continuous Spaces Using Voronoi Optimistic Optimization with Regret BoundsabstractMany important applications, including robotics, data-center management, and process control, require planning action sequences in domains with continuous state and action spaces and discontinuous objective functions. Monte Carlo tree search (MCTS) is an effective strategy for planning in discrete action spaces. We provide a novel MCTS algorithm (voot) for deterministic environments with continuous action spaces, which, in turn, is based on a novel black-box function-optimization algorithm (voo) to efficiently sample actions. The voo algorithm uses Voronoi partitioning to guide sampling, and is particularly efficient in high-dimensional spaces. The voot algorithm has an instance of voo at each node in the tree. We provide regret bounds for both algorithms and demonstrate their empirical effectiveness in several high-dimensional problems including two difficult robotics planning problems. Kyungjae Lee 0001, Sungbin Lim, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 4 |
| 2020 | Few-Shot Bayesian Imitation Learning with Logical Program PoliciesabstractHumans can learn many novel tasks from a very small number (1–5) of demonstrations, in stark contrast to the data requirements of nearly tabula rasa deep learning methods. We propose an expressive class of policies, a strong but general prior, and a learning algorithm that, together, can learn interesting policies from very few examples. We represent policies as logical combinations of programs drawn from a domain-specific language (DSL), define a prior over policies with a probabilistic grammar, and derive an approximate Bayesian inference algorithm to learn policies from demonstrations. In experiments, we study six strategy games played on a 2D grid with one shared DSL. After a few demonstrations of each game, the inferred policies generalize to new game instances that differ substantially from the demonstrations. Our policy learning is 20–1,000x more data efficient than convolutional and fully convolutional policy learning and many orders of magnitude more computationally efficient than vanilla program induction. We argue that the proposed method is an apt choice for tasks that have scarce training data and feature significant, structured variation between task instances. Tom Silver, Kelsey R. Allen, Alex K. Lew, Leslie Pack Kaelbling, Josh Tenenbaum |
AAAI | 4 |
| 2020 | Elimination of All Bad Local Minima in Deep LearningabstractIn this paper, we theoretically prove that adding one special neuron per output unit eliminates all suboptimal local minima of any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function, under practical assumptions. At every local minimum of any deep neural network with these added neurons, the set of parameters of the original neural network (without added neurons) is guaranteed to be a global minimum of the original neural network. The effects of the added neurons are proven to automatically vanish at every local minimum. Moreover, we provide a novel theoretical characterization of a failure mode of eliminating suboptimal local minima via an additional theorem and several examples. This paper also introduces a novel proof technique based on the perturbable gradient basis (PGB) necessary condition of local minima, which provides new insight into the elimination of local minima and is applicable to analyze various models and transformations of objective functions beyond the elimination of local minima. Kenji Kawaguchi, Leslie Pack Kaelbling |
AISTATS | 2 |
| 2020 | Meta-learning curiosity algorithms
Ferran Alet, Martin F. Schneider, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICLR | 4 |
| 2020 | Online Replanning in Belief Space for Partially Observable Task and Motion ProblemsabstractTo solve multi-step manipulation tasks in the real world, an autonomous robot must take actions to observe its environment and react to unexpected observations. This may require opening a drawer to observe its contents or moving an object out of the way to examine the space behind it. Upon receiving a new observation, the robot must update its belief about the world and compute a new plan of action. In this work, we present an online planning and execution system for robots faced with these challenges. We perform deterministic cost-sensitive planning in the space of hybrid belief states to select likely-to-succeed observation actions and continuous control actions. After execution and observation, we replan using our new state estimate. We initially enforce that planner reuses the structure of the unexecuted tail of the last plan. This both improves planning efficiency and ensures that the overall policy does not undo its progress towards achieving the goal. Our approach is able to efficiently solve partially observable problems both in simulation and in a real-world kitchen. Caelan Reed Garrett, Chris Paxton 0001, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Dieter Fox |
ICRA | 4 |
| 2020 | Visual Prediction of Priors for Articulated Object InteractionabstractExploration in novel settings can be challenging without prior experience in similar domains. However, humans are able to build on prior experience quickly and efficiently. Children exhibit this behavior when playing with toys. For example, given a toy with a yellow and blue door, a child will explore with no clear objective, but once they have discovered how to open the yellow door, they will most likely be able to open the blue door much faster. Adults also exhibit this behaviour when entering new spaces such as kitchens. We develop a method, Contextual Prior Prediction, which provides a means of transferring knowledge between interactions in similar domains through vision. We develop agents that exhibit exploratory behavior with increasing efficiency, by learning visual features that are shared across environments, and how they correlate to actions. Our problem is formulated as a Contextual Multi-Armed Bandit where the contexts are images, and the robot has access to a parameterized action space. Given a novel object, the objective is to maximize reward with few interactions. A domain which strongly exhibits correlations between visual features and motion is kinemetically constrained mechanisms. We evaluate our method on simulated prismatic and revolute joints1. Caris Moses, Michael Noseworthy, Leslie Pack Kaelbling, Tomás Lozano-Pérez, Nicholas Roy |
ICRA | 3 |
| 2020 | Adversarially-learned Inference via an Ensemble of Discrete Undirected Graphical ModelsabstractUndirected graphical models are compact representations of joint probability distributions over random variables. To solve inference tasks of interest, graphical models of arbitrary topology can be trained using empirical risk minimization. However, to solve inference tasks that were not seen during training, these models (EGMs) often need to be re-trained. Instead, we propose an inference-agnostic adversarial training framework which produces an infinitely-large ensemble of graphical models (AGMs). The ensemble is optimized to generate data within the GAN framework, and inference is performed using a finite subset of these models. AGMs perform comparably with EGMs on inference tasks that the latter were specifically optimized for. Most importantly, AGMs show significantly better generalization to unseen inference tasks compared to EGMs, as well as deep neural architectures like GibbsNet and VAEAC which allow arbitrary conditioning. Finally, AGMs allow fast data sampling, competitive with Gibbs sampling from EGMs. Adarsh K. Jeewajee, Leslie Pack Kaelbling |
NeurIPS | 2 |
| 2019 | Adversarial Actor-Critic Method for Task and Motion Planning Problems Using Planning ExperienceabstractWe propose an actor-critic algorithm that uses past planning experience to improve the efficiency of solving robot task-and-motion planning (TAMP) problems. TAMP planners search for goal-achieving sequences of high-level operator instances specified by both discrete and continuous parameters. Our algorithm learns a policy for selecting the continuous parameters during search, using a small training set generated from the search trees of previously solved instances. We also introduce a novel fixed-length vector representation for world states with varying numbers of objects with different shapes, based on a set of key robot configurations. We demonstrate experimentally that our method learns more efficiently from less data than standard reinforcementlearning approaches and that using a learned policy to guide a planner results in the improvement of planning efficiency. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 2 |
| 2019 | Learning sparse relational transition models
Victoria Xia, Zi Wang 0004, Kelsey R. Allen, Tom Silver, Leslie Pack Kaelbling |
ICLR (Poster) | 5 |
| 2019 | Graph Element Networks: adaptive, structured computation and memoryabstractWe explore the use of graph neural networks (GNNs) to model spatial processes in which there is no a priori graphical structure. Similar to finite element analysis, we assign nodes of a GNN to spatial locations and use a computational process defined on the graph to model the relationship between an initial function defined over a space and a resulting function in the same space. We use GNNs as a computational substrate, and show that the locations of the nodes in space as well as their connectivity can be optimized to focus on the most complex parts of the space. Moreover, this representational strategy allows the learned input-output relationship to generalize over the size of the underlying space and run the same model at different levels of precision, trading computation for accuracy. We demonstrate this method on a traditional PDE problem, a physical prediction problem from robotics, and learning to predict scene images from novel viewpoints. Ferran Alet, Adarsh K. Jeewajee, Maria Bauzá 0001, Alberto Rodriguez 0003, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
ICML | 6 |
| 2019 | Combining Physical Simulators and Object-Based Networks for ControlabstractPhysics engines play an important role in robot planning and control; however, many real-world control problems involve complex contact dynamics that cannot be characterized analytically. Most physics engines therefore employ approximations that lead to a loss in precision. In this paper, we propose a hybrid dynamics model, simulator-augmented interaction networks (SAIN), combining a physics engine with an object-based neural network for dynamics modeling. Compared with existing models that are purely analytical or purely data-driven, our hybrid model captures the dynamics of interacting objects in a more accurate and data-efficient manner. Experiments both in simulation and on a real robot suggest that it also leads to better performance when used in complex control tasks. Finally, we show that our model generalizes to novel environments with varying object shapes and materials. Anurag Ajay, Maria Bauzá 0001, Jiajun Wu 0001, Nima Fazeli, Josh Tenenbaum, Alberto Rodriguez 0003, Leslie Pack Kaelbling |
ICRA | 7 |
| 2019 | Learning Quickly to Plan Quickly Using Modular Meta-LearningabstractMulti-object manipulation problems in continuous state and action spaces can be solved by planners that search over sampled values for the continuous parameters of operators. The efficiency of these planners depends critically on the effectiveness of the samplers used, but effective sampling in turn depends on details of the robot, environment, and task. Our strategy is to learn functions called speciatizers that generate values for continuous operator parameters, given a state description and values for the discrete parameters. Rather than trying to learn a single specializer for each operator from large amounts of data on a single task, we take a modular meta-learning approach. We train on multiple tasks and learn a variety of specializers that, on a new task, can be quickly adapted using relatively little data - thus, our system learns quickly to plan quickly using these specializers. We validate our approach experimentally in simulated 3D pick-and-place tasks with continuous state and action spaces. Visit http://tinyurl.com/chitnis-icra-19 for a supplementary video. Rohan Chitnis, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2019 | Omnipush: accurate, diverse, real-world dataset of pushing dynamics with RGB-D videoabstractPushing is a fundamental robotic skill. Existing work has shown how to exploit models of pushing to achieve a variety of tasks, including grasping under uncertainty, in-hand manipulation and clearing clutter. Such models, however, are approximate, which limits their applicability.Learning-based methods can reason directly from raw sensory data with accuracy, and have the potential to generalize to a wider diversity of scenarios. However, developing and testing such methods requires rich-enough datasets. In this paper we introduce Omnipush, a dataset with high variety of planar pushing behavior.In particular, we provide 250 pushes for each of 250 objects, all recorded with RGB-D and a high precision tracking system. The objects are constructed so as to systematically explore key factors that affect pushing-the shape of the object and its mass distribution-which have not been broadly explored in previous datasets, and allow to study generalization in model learning.Omnipush includes a benchmark for meta-learning dynamic models, which requires algorithms that make good predictions and estimate their own uncertainty. We also provide an RGB video prediction benchmark and propose other relevant tasks that can be suited with this dataset. Data and code are available at https://web.mit.edu/mcube/omnipush-dataset/. Maria Bauzá 0001, Ferran Alet, Yen-Chen Lin, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Phillip Isola, Alberto Rodriguez 0003 |
IROS | 5 |
| 2019 | Neural Relational Inference with Fast Modular Meta-learningabstractGraph neural networks (GNNs) are effective models for many dynamical systems consisting of entities and relations. Although most GNN applications assume a single type of entity and relation, many situations involve multiple types of interactions. Relational inference is the problem of inferring these interactions and learning the dynamics from observational data. We frame relational inference as a modular meta-learning problem, where neural modules are trained to be composed in different ways to solve many tasks. This meta-learning framework allows us to implicitly encode time invariance and infer relations in context of one another rather than independently, which increases inference capacity. Framing inference as the inner-loop optimization of meta-learning leads to a model-based approach that is more data-efficient and capable of estimating the state of entities that we do not observe directly, but whose existence can be inferred from their effect on observed entities. To address the large search space of graph neural network compositions, we meta-learn a proposal function that speeds up the inner-loop simulated annealing search within the modular meta-learning algorithm, providing two orders of magnitude increase in the size of problems that can be addressed. Ferran Alet, Erica Weng, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
NeurIPS | 4 |
| 2019 | Modeling and Planning with Macro-Actions in Decentralized POMDPsabstractDecentralized partially observable Markov decision processes (Dec-POMDPs) are general models for decentralized multi-agent decision making under uncertainty. However, they typically model a problem at a low level of granularity, where each agent's actions are primitive operations lasting exactly one time step. We address the case where each agent has macro-actions: temporally extended actions that may require different amounts of time to execute. We model macro-actions as options in a Dec-POMDP, focusing on actions that depend only on information directly available to the agent during execution. Therefore, we model systems where coordination decisions only occur at the level of deciding which macro-actions to execute. The core technical difficulty in this setting is that the options chosen by each agent no longer terminate at the same time. We extend three leading Dec-POMDP algorithms for policy generation to the macro-action case, and demonstrate their effectiveness in both standard benchmarks and a multi-robot coordination problem. The results show that our new algorithms retain agent coordination while allowing high-quality solutions to be generated for significantly longer horizons and larger state-spaces than previous Dec-POMDP methods. Furthermore, in the multi-robot domain, we show that, in contrast to most existing methods that are specialized to a particular problem class, our approach can synthesize control policies that exploit opportunities for coordination while balancing uncertainty, sensor information, and information about other agents. Christopher Amato, George Dimitri Konidaris, Leslie Pack Kaelbling, Jonathan P. How |
J. Artif. Intell. Res. | 3 |
| 2019 | Effect of Depth and Width on Local Minima in Deep LearningabstractIn this paper, we analyze the effects of depth and width on the quality of local minima, without strong overparameterization and simplification assumptions in the literature. Without any simplification assumption, for deep nonlinear neural networks with the squared loss, we theoretically show that the quality of local minima tends to improve toward the global minimum value as depth and width increase. Furthermore, with a locally induced structure on deep nonlinear neural networks, the values of local minima of neural networks are theoretically proven to be no worse than the globally optimal values of corresponding classical machine learning models. We empirically support our theoretical observation with a synthetic data set, as well as MNIST, CIFAR-10, and SVHN data sets. When compared to previous studies with strong overparameterization assumptions, the results in this letter do not require overparameterization and instead show the gradual effects of overparameterization as consequences of general results. Kenji Kawaguchi, Jiaoyang Huang, Leslie Pack Kaelbling |
Neural Comput. | 3 |
| 2019 | Every Local Minimum Value Is the Global Minimum Value of Induced Model in Nonconvex Machine LearningabstractFor nonconvex optimization in machine learning, this article proves that every local minimum achieves the globally optimal value of the perturbable gradient basis model at any differentiable point. As a result, nonconvex machine learning is theoretically as supported as convex machine learning with a handcrafted basis in terms of the loss at differentiable local minima, except in the case when a preference is given to the handcrafted basis over the perturbable gradient basis. The proofs of these results are derived under mild assumptions. Accordingly, the proven results are directly applicable to many machine learning models, including practical deep neural networks, without any modification of practical methods. Furthermore, as special cases of our general results, this article improves or complements several state-of-the-art theoretical results on deep neural networks, deep residual networks, and overparameterized deep neural networks with a unified proof technique and novel geometric insights. A special case of our results also contributes to the theoretical foundation of representation learning. Kenji Kawaguchi, Jiaoyang Huang, Leslie Pack Kaelbling |
Neural Comput. | 3 |
| 2018 | Guiding Search in Continuous State-Action Spaces by Learning an Action Sampler From Off-Target Search ExperienceabstractIn robotics, it is essential to be able to plan efficiently in high-dimensional continuous state-action spaces for long horizons. For such complex planning problems, unguided uniform sampling of actions until a path to a goal is found is hopelessly inefficient, and gradient-based approaches often fall short when the optimization manifold of a given problem is not smooth. In this paper, we present an approach that guides search in continuous spaces for generic planners by learning an action sampler from past search experience. We use a Generative Adversarial Network (GAN) to represent an action sampler, and address an important issue: search experience consists of a relatively large number of actions that are not on a solution path and a relatively small number of actions that actually are on a solution path. We introduce a new technique, based on an importance-ratio estimation method, for using samples from a non-target distribution to make GAN learning more data-efficient. We provide theoretical guarantees and empirical evaluation in three challenging continuous robot planning problems to illustrate the effectiveness of our algorithm. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 2 |
| 2018 | Selecting Representative Examples for Program SynthesisabstractProgram synthesis is a class of regression problems where one seeks a solution, in the form of a source-code program, mapping the inputs to their corresponding outputs exactly. Due to its precise and combinatorial nature, program synthesis is commonly formulated as a constraint satisfaction problem, where input-output examples are encoded as constraints and solved with a constraint solver. A key challenge of this formulation is scalability: while constraint solvers work well with a few well-chosen examples, a large set of examples can incur significant overhead in both time and memory. We describe a method to discover a subset of examples that is both small and representative: the subset is constructed iteratively, using a neural network to predict the probability of unchosen examples conditioned on the chosen examples in the subset, and greedily adding the least probable example. We empirically evaluate the representativeness of the subsets constructed by our method, and demonstrate such subsets can significantly improve synthesis time and stability. Yewen Pu, Zachery Miranda, Armando Solar-Lezama, Leslie Pack Kaelbling |
ICML | 4 |
| 2018 | Reliably Arranging Objects in Uncertain DomainsabstractA crucial challenge in robotics is achieving reliable results in spite of sensing and control uncertainty. In this work, we explore the conformant planning approach to robot manipulation. In particular, we tackle the problem of pushing multiple planar objects simultaneously to achieve a specified arrangement without external sensing. Conformant planning is a belief-state planning problem. A belief state is the set of all possible states of the world, and the goal is to find a sequence of actions that will bring an initial belief state to a goal belief state. To do forward belief-state planning, we created a deterministic belief-state transition model from supervised learning based on off-line physics simulations. We compare our method with an on-line physics-based manipulation approach and show significantly reduced planning times and increased robustness in simulated experiments. Finally, we demonstrate the success of this approach in simulations and physical robot experiments. Ariel Anders, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2018 | Finding Frequent Entities in Continuous DataabstractIn many applications that involve processing high-dimensional data, it is important to identify a small set of entities that account for a significant fraction of detections. Rather than formalize this as a clustering problem, in which all detections must be grouped into hard or soft categories, we formalize it as an instance of the frequent items or heavy hitters problem, which finds groups of tightly clustered objects that have a high density in the feature space. We show that the heavy hitters formulation generates solutions that are more accurate and effective than the clustering formulation. In addition, we present a novel online algorithm for heavy hitters, called HAC, which addresses problems in continuous space, and demonstrate its effectiveness on real video and household domains. Ferran Alet, Rohan Chitnis, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IJCAI | 3 |
| 2018 | Augmenting Physical Simulators with Stochastic Neural Networks: Case Study of Planar Pushing and BouncingabstractAn efficient, generalizable physical simulator with universal uncertainty estimates has wide applications in robot state estimation, planning, and control. In this paper, we build such a simulator for two scenarios, planar pushing and ball bouncing, by augmenting an analytical rigid-body simulator with a neural network that learns to model uncertainty as residuals. Combining symbolic, deterministic simulators with learnable, stochastic neural nets provides us with expressiveness, efficiency, and generalizability simultaneously. Our model outperforms both purely analytical and purely learned simulators consistently on real, standard benchmarks. Compared with methods that model uncertainty using Gaussian processes, our model runs much faster, generalizes better to new object shapes, and is able to characterize the complex distribution of object trajectories. Anurag Ajay, Jiajun Wu 0001, Nima Fazeli, Maria Bauzá 0001, Leslie Pack Kaelbling, Josh Tenenbaum, Alberto Rodriguez 0003 |
IROS | 5 |
| 2018 | Integrating Human-Provided Information into Belief State Representation Using Dynamic FactorizationabstractIn partially observed environments, it can be useful for a human to provide the robot with declarative information that represents probabilistic relational constraints on properties of objects in the world, augmenting the robot's sensory observations. For instance, a robot tasked with a search-and-rescue mission may be informed by the human that two victims are probably in the same room. An important question arises: how should we represent the robot's internal knowledge so that this information is correctly processed and combined with raw sensory information? In this paper, we provide an efficient belief state representation that dynamically selects an appropriate factoring, combining aspects of the belief when they are correlated through information and separating them when they are not. This strategy works in open domains, in which the set of possible objects is not known in advance, and provides significant improvements in inference time over a static factoring, leading to more efficient planning for complex partially observed tasks. We validate our approach experimentally in two open-domain planning problems: a 2D discrete gridworld task and a 3D continuous cooking task. A supplementary video can be found at http://tinyurl.com/chitnis-iros-18. Rohan Chitnis, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 2 |
| 2018 | Active Model Learning and Diverse Action Sampling for Task and Motion PlanningabstractThe objective of this work is to augment the basic abilities of a robot by learning to use new sensorimotor primitives to enable the solution of complex long-horizon problems. Solving long-horizon problems in complex domains requires flexible generative planning that can combine primitive abilities in novel combinations to solve problems as they arise in the world. In order to plan to combine primitive actions, we must have models of the preconditions and effects of those actions: under what circumstances will executing this primitive achieve some particular effect in the world? We use, and develop novel improvements on, state-of-the-art methods for active learning and sampling. We use Gaussian process methods for learning the conditions of operator effectiveness from small numbers of expensive training examples collected by experimentation on a robot. We develop adaptive sampling methods for generating diverse elements of continuous sets (such as robot configurations and object poses) during planning for solving a new task, so that planning is as efficient as possible. We demonstrate these methods in an integrated system, combining newly learned models with an efficient continuous-space robot task and motion planner to learn to solve long horizon problems more efficiently than was previously possible. Zi Wang 0004, Caelan Reed Garrett, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 3 |
| 2018 | Regret bounds for meta Bayesian optimization with an unknown Gaussian process priorabstractBayesian optimization usually assumes that a Bayesian prior is given. However, the strong theoretical guarantees in Bayesian optimization are often regrettably compromised in practice because of unknown parameters in the prior. In this paper, we adopt a variant of empirical Bayes and show that, by estimating the Gaussian process prior from offline data sampled from the same prior and constructing unbiased estimators of the posterior, variants of both GP-UCB and \emph{probability of improvement} achieve a near-zero regret bound, which decreases to a constant proportional to the observational noise as the number of offline data and the number of online evaluations increase. Empirically, we have verified our approach on challenging simulated robotic problems featuring task and motion planning. Zi Wang 0004, Leslie Pack Kaelbling |
NeurIPS | 3 |
| 2018 | Look Before You Sweep: Visibility-Aware Motion Planning
Gustavo Goretkin, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
WAFR | 2 |
| 2018 | From Skills to Symbols: Learning Symbolic Representations for Abstract High-Level PlanningabstractWe consider the problem of constructing abstract representations for planning in high-dimensional, continuous environments. We assume an agent equipped with a collection of high-level actions, and construct representations provably capable of evaluating plans composed of sequences of those actions. We first consider the deterministic planning case, and show that the relevant computation involves set operations performed over sets of states. We define the specific collection of sets that is necessary and sufficient for planning, and use them to construct a grounded abstract symbolic representation that is provably suitable for deterministic planning. The resulting representation can be expressed in PDDL, a canonical high-level planning domain language; we construct such a representation for the Playroom domain and solve it in milliseconds using an off-the-shelf planner. We then consider probabilistic planning, which we show requires generalizing from sets of states to distributions over states. We identify the specific distributions required for planning, and use them to construct a grounded abstract symbolic representation that correctly estimates the expected reward and probability of success of any plan. In addition, we show that learning the relevant probability distributions corresponds to specific instances of probabilistic density estimation and probabilistic classification. We construct an agent that autonomously learns the correct abstract representation of a computer game domain, and rapidly solves it. Finally, we apply these techniques to create a physical robot system that autonomously learns its own symbolic representation of a mobile manipulation task directly from sensorimotor data---point clouds, map locations, and joint angles---and then plans using that representation. Together, these results establish a principled link between high-level actions and abstract representations, a concrete theoretical foundation for constructing abstract representations with provable properties, and a practical mechanism for autonomously learning abstract high-level representations. George Dimitri Konidaris, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
J. Artif. Intell. Res. | 2 |
| 2017 | Learning composable models of parameterized skillsabstractThere has been a great deal of work on learning new robot skills, but very little consideration of how these newly acquired skills can be integrated into an overall intelligent system. A key aspect of such a system is compositionality: newly learned abilities have to be characterized in a form that will allow them to be flexibly combined with existing abilities, affording a (good!) combinatorial explosion in the robot's abilities. In this paper, we focus on learning models of the preconditions and effects of new parameterized skills, in a form that allows those actions to be combined with existing abilities by a generative planning and execution system. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 1 |
| 2017 | Learning to guide task and motion planning using score-space representationabstractIn this paper, we propose a learning algorithm that speeds up the search in task and motion planning problems. Our algorithm proposes solutions to three different challenges that arise in learning to improve planning efficiency: what to predict, how to represent a planning problem instance, and how to transfer knowledge from one problem instance to another. We propose a method that predicts constraints on the search space based on a generic representation of a planning problem instance, called score space, where we represent a problem instance in terms of performance of a set of solutions attempted so far. Using this representation, we transfer knowledge, in the form of constraints, from previous problems based on the similarity in score space. We design a sequential algorithm that efficiently predicts these constraints, and evaluate it in three different challenging task and motion planning problems. Results indicate that our approach perform orders of magnitudes faster than an unguided planner. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2017 | Focused model-learning and planning for non-Gaussian continuous state-action systemsabstractWe introduce a framework for model learning and planning in stochastic domains with continuous state and action spaces and non-Gaussian transition models. It is efficient because (1) local models are estimated only when the planner requires them; (2) the planner focuses on the most relevant states to the current planning problem; and (3) the planner focuses on the most informative and/or high-value actions. Our theoretical analysis shows the validity and asymptotic optimality of the proposed approach. Empirically, we demonstrate the effectiveness of our algorithm on a simulated multi-modal pushing problem. Zi Wang 0004, Stefanie Jegelka, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 3 |
| 2017 | Intelligent Robots in an Uncertain World
Leslie Pack Kaelbling |
UAI | 1 |
| 2017 | Learning to Acquire Information
Yewen Pu, Leslie Pack Kaelbling, Armando Solar-Lezama |
UAI | 2 |
| 2016 | Implicit belief-space pre-images for hierarchical planning and executionabstractWe present a method for planning and execution in very high-dimensional mixed discrete and continuous spaces in the presence of uncertainty using an implicit, factored approximation representation of pre-images and extend it to planning in belief space. We demonstrate the approach in a mobile-manipulation domain combining pushing with pick-and-place manipulation with error in sensing and manipulation. We show empirically that execution monitoring using pre-images improves computational efficiency over continual replanning, and that the hierarchical planning method it enables provides further efficiency improvements. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 1 |
| 2016 | Searching for physical objects in partially known environmentsabstractWe address the problem of a mobile manipulation robot searching for an object in a cluttered domain that is populated with an unknown number of objects in an unknown arrangement. The robot must move around its environment, looking in containers, moving occluding objects to improve its view, and reasoning about collocation of objects of different types, all in service of finding a desired object. The key contribution in reasoning is a Markov-chain Monte Carlo (MCMC) method for drawing samples of the arrangements of objects in an occluded container, conditioned on previous observations of other objects as well as spatial constraints. The key contribution in planning is a receding-horizon forward search in the space of distributions over arrangements (including number and type) of objects in the domain; to maintain tractability the search is formulated in a model that abstracts both the observations and actions available to the robot. The strategy is shown empirically to improve upon a baseline systematic search strategy, and sometimes outperforms a method from previous work. Xinkun Nie, Lawson L. S. Wong, Leslie Pack Kaelbling |
ICRA | 3 |
| 2016 | Learning to Rank for Synthesizing Planning Heuristics
Caelan Reed Garrett, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IJCAI | 2 |
| 2016 | Object-Based World Modeling in Semi-Static Environments with Dependent Dirichlet Process Mixtures
Lawson L. S. Wong, Thanard Kurutach, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IJCAI | 4 |
| 2016 | Decidability of Semi-Holonomic Prehensile Task and Motion Planning
Ashwin Deshpande, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
WAFR | 2 |
| 2015 | Planning for decentralized control of multiple robots under uncertaintyabstractThis paper presents a probabilistic framework for synthesizing control policies for general multi-robot systems that is based on decentralized partially observable Markov decision processes (Dec-POMDPs). Dec-POMDPs are a general model of decision-making where a team of agents must cooperate to optimize a shared objective in the presence of uncertainty. Dec-POMDPs also consider communication limitations, so execution is decentralized. While Dec-POMDPs are typically intractable to solve for real-world problems, recent research on the use of macro-actions in Dec-POMDPs has significantly increased the size of problem that can be practically solved. We show that, in contrast to most existing methods that are specialized to a particular problem class, our approach can synthesize control policies that exploit any opportunities for coordination that are present in the problem, while balancing uncertainty, sensor information, and information about other agents. We use three variants of a warehouse task to show that a single planner of this type can generate cooperative behavior using task allocation, direct communication, and signaling, as appropriate. This demonstrates that our algorithmic framework can automatically optimize control and communication policies for complex multi-robot systems. Christopher Amato, George Dimitri Konidaris, Gabriel Cruz, Christopher A. Maynor, Jonathan P. How, Leslie Pack Kaelbling |
ICRA | 6 |
| 2015 | Symbol Acquisition for Probabilistic High-Level Planning
George Dimitri Konidaris, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IJCAI | 2 |
| 2015 | Backward-forward search for manipulation planningabstractIn this paper we address planning problems in high-dimensional hybrid configuration spaces, with a particular focus on manipulation planning problems involving many objects. We present the hybrid backward-forward (HBF) planning algorithm that uses a backward identification of constraints to direct the sampling of the infinite action space in a forward search from the initial state towards a goal configuration. The resulting planner is probabilistically complete and can effectively construct long manipulation plans requiring both prehensile and nonprehensile actions in cluttered environments. Caelan Reed Garrett, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IROS | 3 |
| 2015 | Hierarchical planning for multi-contact non-prehensile manipulationabstractManipulation planning involves planning the combined motion of objects in the environment as well as the robot motions to achieve them. In this paper, we explore a hierarchical approach to planning sequences of non-prehensile and prehensile actions. We subdivide the planning problem into three stages (object contacts, object poses and robot contacts) and thereby reduce the size of search space that is explored. We show that this approach is more efficient than earlier strategies that search in the combined robot-object configuration space directly. Gilwoo Lee, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IROS | 3 |
| 2015 | Generalizing Over Uncertain Dynamics for Online Trajectory Generation
Albert Kim, Hongkai Dai, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ISRR (2) | 4 |
| 2015 | Bayesian Optimization with Exponential ConvergenceabstractThis paper presents a Bayesian optimization method with exponential convergence without the need of auxiliary optimization and without the delta-cover sampling. Most Bayesian optimization methods require auxiliary optimization: an additional non-convex global optimization problem, which can be time-consuming and hard to implement in practice. Also, the existing Bayesian optimization method with exponential convergence requires access to the delta-cover sampling, which was considered to be impractical. Our approach eliminates both requirements and achieves an exponential convergence rate. Kenji Kawaguchi, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
NIPS | 2 |
| 2014 | Optimizing a Start-Stop Controller Using Policy SearchabstractWe applied a policy search algorithm to the problem of optimizing a start-stop controller—a controller used in a car to turn off the vehicle’s engine, and thus save energy, when the vehicle comes to a temporary halt. We were able to improve the existing policy by approximately 12% using real driver trace data. We also experimented with using multiple policies, and found that doing so could lead to a further 8% improvement if we could determine which policy to apply at each stop. The driver’s behaviors before stopping were found to be uncorrelated with the policy that performed best; however, further experimentation showed that the driver’s behavior during the stop may be more useful, suggesting a useful direction for adding complexity to the underlying start-stop policy. Noel Hollingsworth, Jason Meyer, Ryan McGee, Jeffrey Doering, George Dimitri Konidaris, Leslie Pack Kaelbling |
AAAI | 6 |
| 2014 | Constructing Symbolic Representations for High-Level PlanningabstractWe consider the problem of constructing a symbolic description of a continuous, low-level environment for use in planning. We show that symbols that can represent the preconditions and effects of an agent's actions are both necessary and sufficient for high-level planning. This eliminates the symbol design problem when a representation must be constructed in advance, and in principle enables an agent to autonomously learn its own symbolic representations. The resulting representation can be converted into PDDL, a canonical high-level planning representation that enables very fast planning. George Dimitri Konidaris, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
AAAI | 2 |
| 2014 | Interactive Bayesian identification of kinematic mechanismsabstractThis paper addresses the problem of identifying mechanisms based on data gathered while interacting with them. We present a decision-theoretic formulation of this problem, using Bayesian filtering techniques to maintain a distributional estimate of the mechanism type and parameters. In order to reduce the amount of interaction required to arrive at a confident identification, we select actions explicitly to reduce entropy in the current estimate. We demonstrate the approach on a domain with four primitive and two composite mechanisms. The results show that this approach can correctly identify complex mechanisms including mechanisms which are difficult to model analytically. The results also show that entropy-based action selection can significantly decrease the number of actions required to gather the same information. Patrick R. Barragan, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2014 | Tracking the spin on a ping pong ball with the quaternion Bingham filterabstractA deterministic method for sequential estimation of 3-D rotations is presented. The Bingham distribution is used to represent uncertainty directly on the unit quaternion hypersphere. Quaternions avoid the degeneracies of other 3-D orientation representations, while the Bingham distribution allows tracking of large-error (high-entropy) rotational distributions. Experimental comparison to a leading EKF-based filtering approach on both synthetic signals and a ball-tracking dataset shows that the Quaternion Bingham Filter (QBF) has lower tracking error than the EKF, particularly when the state is highly dynamic. We present two versions of the QBF- suitable for tracking the state of first- and second-order rotating dynamical systems. Jared Glover, Leslie Pack Kaelbling |
ICRA | 2 |
| 2014 | Not seeing is also believing: Combining object and metric spatial informationabstractSpatial representations are fundamental to mobile robots operating in uncertain environments. Two frequently-used representations are occupancy grid maps, which only model metric information, and object-based world models, which only model object attributes. Many tasks represent space in just one of these two ways; however, because objects must be physically grounded in metric space, these two distinct layers of representation are fundamentally linked. We develop an approach that maintains these two sources of spatial information separately, and combines them on demand. We illustrate the utility and necessity of combining such information through applying our approach to a collection of motivating examples. Lawson L. S. Wong, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2014 | A constraint-based method for solving sequential manipulation planning problemsabstractIn this paper, we describe a strategy for integrated task and motion planning based on performing a symbolic search for a sequence of high-level operations, such as pick, move and place, while postponing geometric decisions. Partial plans (skeletons) in this search thus pose a geometric constraint-satisfaction problem (CSP), involving sequences of placements and paths for the robot, and grasps and locations of objects. We propose a formulation for these problems in a discretized configuration space for the robot. The resulting problems can be solved using existing methods for discrete CSP. Tomás Lozano-Pérez, Leslie Pack Kaelbling |
IROS | 2 |
| 2014 | FFRob: An Efficient Heuristic for Task and Motion Planning
Caelan Reed Garrett, Tomás Lozano-Pérez, Leslie Pack Kaelbling |
WAFR | 3 |
| 2013 | A hierarchical approach to manipulation with diverse actionsabstractWe define the Diverse Action Manipulation (DAMA) problem in which we are given a mobile robot, a set of movable objects, and a set of diverse, possibly non-prehensile manipulation actions, and the goal is to find a sequence of actions that moves each of the objects to a goal configuration. We show that the DAMA problem can be framed as a multi-modal planning problem and describe a hierarchical algorithm that takes advantage of this multi-modal nature. We also extend our earlier forward search sampling algorithm to a bi-directional version. We give results on a complicated manipulation domain and demonstrate that both new algorithms are significantly more efficient than the original, and that the hierarchical algorithm is usually much more efficient than the forward or bi-directional searches. Jennifer L. Barry, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2013 | Optimization in the now: Dynamic peephole optimization for hierarchical planningabstractFor robots to effectively interact with the real world, they will need to perform complex tasks over long time horizons. This is a daunting challenge, but recent advances using hierarchical planning [1] have been able to provide leverage on this problem. Unfortunately, this approach makes no effort to account for the execution cost of an abstract plan and often arrives at poor quality plans. This paper outlines a method for dynamically improving a hierarchical plan during execution. We frame the underlying question as one of evaluating the resource needs of an abstract operator and propose a general way to approach estimating them. We ran experiments in challenging domains and observed up to 30% reduction in execution cost when compared with a standard hierarchical planner. Dylan Hadfield-Menell, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2013 | Object placement as inverse motion planningabstractWe present an approach to robust placing that uses movable surfaces in the environment to guide a poorly grasped object into a goal pose. This problem is an instance of the inverse motion planning problem, in which we solve for a configuration of the environment that makes desired trajectories likely. To calculate the probability that an object will take a particular trajectory, we model the physics of placing as a mixture model of simple object motions. Our algorithm searches over the possible configurations of the object and environment and uses this model to choose the configuration most likely to lead to a successful place. We show that this algorithm allows the PR2 robot to execute placements that fail with traditional placing implementations. Anne Holladay, Jennifer L. Barry, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 3 |
| 2013 | Manipulation-based active search for occluded objectsabstractObject search is an integral part of daily life, and in the quest for competent mobile manipulation robots it is an unavoidable problem. Previous approaches focus on cases where objects are in unknown rooms but lying out in the open, which transforms object search into active visual search. However, in real life, objects may be in the back of cupboards occluded by other objects, instead of conveniently on a table by themselves. Extending search to occluded objects requires a more precise model and tighter integration with manipulation. We present a novel generative model for representing container contents by using object co-occurrence information and spatial constraints. Given a target object, a planner uses the model to guide an agent to explore containers where the target is likely, potentially needing to move occluding objects to enable further perception. We demonstrate the model on simulated domains and a detailed simulation involving a PR2 robot. Lawson L. S. Wong, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2013 | Foresight and reconsideration in hierarchical planning and executionabstractWe present a hierarchical planning and execution architecture that maintains the computational efficiency of hierarchical decomposition while improving optimality. It provides mechanisms for monitoring the belief state during execution and performing selective replanning to repair poor choices and take advantage of new opportunities. It also provides mechanisms for looking ahead into future plans to avoid making short-sighted choices. The effectiveness of this architecture is shown through comparative experiments in simulation and demonstrated on a real PR2 robot. Martin Levihn, Leslie Pack Kaelbling, Tomás Lozano-Pérez, Mike Stilman |
IROS | 2 |
| 2013 | Data Association for Semantic World Modeling from Partial Views
Lawson L. S. Wong, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ISRR | 2 |
| 2012 | Influence-Based Abstraction for Multiagent SystemsabstractThis paper presents a theoretical advance by which factored POSGs can be decomposed into local models. We formalize the interface between such local models as the influence agents can exert on one another; and we prove that this interface is sufficient for decoupling them. The resulting influence-based abstraction substantially generalizes previous work on exploiting weakly-coupled agent interaction structures. Therein lie several important contributions. First, our general formulation sheds new light on the theoretical relationships among previous approaches, and promotes future empirical comparisons that could come by extending them beyond the more specific problem contexts for which they were developed. More importantly, the influence-based approaches that we generalize have shown promising improvements in the scalability of planning for more restrictive models. Thus, our theoretical result here serves as the foundation for practical algorithms that we anticipate will bring similar improvements to more general planning contexts, and also into other domains such as approximate planning, decision-making in adversarial domains, and online learning. Frans A. Oliehoek, Stefan J. Witwicki, Leslie Pack Kaelbling |
AAAI | 3 |
| 2012 | Unifying perception, estimation and action for mobile manipulation via belief space planningabstractIn this paper, we describe an integrated strategy for planning, perception, state-estimation and action in complex mobile manipulation domains. The strategy is based on planning in the belief space of probability distribution over states. Our planning approach is based on hierarchical symbolic regression (pre-image back-chaining). We develop a vocabulary of fluents that describe sets of belief states, which are goals and subgoals in the planning process. We show that a relatively small set of symbolic operators lead to task-oriented perception in support of the manipulation goals. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 1 |
| 2012 | LQR-RRT*: Optimal sampling-based motion planning with automatically derived extension heuristicsabstractThe RRT* algorithm has recently been proposed as an optimal extension to the standard RRT algorithm [1]. However, like RRT, RRT* is difficult to apply in problems with complicated or underactuated dynamics because it requires the design of a two domain-specific extension heuristics: a distance metric and node extension method. We propose automatically deriving these two heuristics for RRT* by locally linearizing the domain dynamics and applying linear quadratic regulation (LQR). The resulting algorithm, LQR-RRT*, finds optimal plans in domains with complex or underactuated dynamics without requiring domain-specific design choices. We demonstrate its application in domains that are successively torque-limited, underactuated, and in belief space. Alejandro Perez, Robert Platt 0001, George Dimitri Konidaris, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 4 |
| 2012 | Non-Gaussian belief space planning: Correctness and complexityabstractWe consider the partially observable control problem where it is potentially necessary to perform complex information-gathering operations in order to localize state. One approach to solving these problems is to create plans in belief-space, the space of probability distributions over the underlying state of the system. The belief-space plan encodes a strategy for performing a task while gaining information as necessary. Unlike most approaches in the literature which rely upon representing belief state as a Gaussian distribution, we have recently proposed an approach to non-Gaussian belief space planning based on solving a non-linear optimization problem defined in terms of a set of state samples [1]. In this paper, we show that even though our approach makes optimistic assumptions about the content of future observations for planning purposes, all low-cost plans are guaranteed to gain information in a specific way under certain conditions. We show that eventually, the algorithm is guaranteed to localize the true state of the system and to reach a goal region with high probability. Although the computational complexity of the algorithm is dominated by the number of samples used to define the optimization problem, our convergence guarantee holds with as few as two samples. Moreover, we show empirically that it is unnecessary to use large numbers of samples in order to obtain good performance. Robert Platt 0001, Leslie Pack Kaelbling, Tomás Lozano-Pérez, Russ Tedrake |
ICRA | 2 |
| 2012 | Collision-free state estimationabstractIn state estimation, we often want the maximum likelihood estimate of the current state. For the commonly used joint multivariate Gaussian distribution over the state space, this can be efficiently found using a Kalman filter. However, in complex environments the state space is often highly constrained. For example, for objects within a refrigerator, they cannot interpenetrate each other or the refrigerator walls. The multivariate Gaussian is unconstrained over the state space and cannot incorporate these constraints. In particular, the state estimate returned by the unconstrained distribution may itself be infeasible. Instead, we solve a related constrained optimization problem to find a good feasible state estimate. We illustrate this for estimating collision-free configurations for objects resting stably on a 2-D surface, and demonstrate its utility in a real robot perception domain. Lawson L. S. Wong, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2011 | Hierarchical task and motion planning in the nowabstractIn this paper we outline an approach to the integration of task planning and motion planning that has the following key properties: It is aggressively hierarchical; it makes choices and commits to them in a top-down fashion in an attempt to limit the length of plans that need to be constructed, and thereby exponentially decrease the amount of search required. It operates on detailed, continuous geometric representations and does not require a-priori discretization of the state or action spaces. Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 1 |
| 2011 | DetH*: Approximate Hierarchical Solution of Large Markov Decision Processes
Jennifer L. Barry, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IJCAI | 2 |
| 2011 | Bayesian Policy Search with Policy Priors
David Wingate, Noah D. Goodman, Daniel M. Roy 0001, Leslie Pack Kaelbling, Josh Tenenbaum |
IJCAI | 4 |
| 2011 | Pre-image Backchaining in Belief Space for Mobile Manipulation
Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ISRR | 1 |
| 2011 | Efficient Planning in Non-Gaussian Belief Spaces and Its Application to Robot Grasping
Robert Platt 0001, Leslie Pack Kaelbling, Tomás Lozano-Pérez, Russ Tedrake |
ISRR | 2 |
| 2010 | Class-specific grasping of 3D objects from a single 2D imageabstractOur goal is to grasp 3D objects given a single image, by using prior 3D shape models of object classes. The shape models, defined as a collection of oriented primitive shapes centered at fixed 3D positions, can be learned from a few labeled images for each class. The 3D class model can then be used to estimate the 3D shape of a detected object, including occluded parts, from a single image. The estimated 3D shape is used as to select one of the target grasps for the object. We show that our 3D shape estimation is sufficiently accurate for a robot to successfully grasp the object, even in situations where the part to be grasped is not visible in the input image. Han-Pang Chiu, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
IROS | 3 |
| 2010 | Intelligent Interaction with the Real World
Leslie Pack Kaelbling |
ECML/PKDD (1) | 1 |
| 2009 | Learning to generate novel views of objects for class recognition
Han-Pang Chiu, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
Comput. Vis. Image Underst. | 2 |
| 2009 | Segmentation According to Natural Examples: Learning Static Segmentation from Motion SegmentationabstractThe Segmentation According to Natural Examples (SANE) algorithm learns to segment objects in static images from video training data. SANE uses background subtraction to find the segmentation of moving objects in videos. This provides object segmentation information for each video frame. The collection of frames and segmentations forms a training set that SANE uses to learn the image and shape properties of the observed motion boundaries. When presented with new static images, the trained model infers segmentations similar to the observed motion segmentations. SANE is a general method for learning environment-specific segmentation models. Because it can automatically generate training data from video, it can adapt to a new environment and new objects with relative ease, an advantage over untrained segmentation methods or those that require human-labeled training data. By using the local shape information in the training data, it outperforms a trained local boundary detector. Its performance is competitive with a trained top-down segmentation algorithm that uses global shape. The shape information it learns from one class of objects can assist the segmentation of other classes. Michael G. Ross, Leslie Pack Kaelbling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Lifted Probabilistic Inference with Counting Formulas
Brian Milch, Luke Zettlemoyer, Kristian Kersting, Michael Haimes, Leslie Pack Kaelbling |
AAAI | 5 |
| 2008 | Multi-Agent Filtering with Infinitely Nested BeliefsabstractIn partially observable worlds with many agents, nested beliefs are formed when agents simultaneously reason about the unknown state of the world and the beliefs of the other agents. The multi-agent filtering problem is to efficiently represent and update these beliefs through time as the agents act in the world. In this paper, we formally define an infinite sequence of nested beliefs about the state of the world at the current time $t$ and present a filtering algorithm that maintains a finite representation which can be used to generate these beliefs. In some cases, this representation can be updated exactly in constant time; we also present a simple approximation scheme to compact beliefs if they become too complex. In experiments, we demonstrate efficient filtering in a range of multi-agent domains. Luke Zettlemoyer, Brian Milch, Leslie Pack Kaelbling |
NIPS | 3 |
| 2007 | Action-Space Partitioning for Planning
Natalia Hernandez-Gardiol, Leslie Pack Kaelbling |
AAAI | 2 |
| 2007 | Virtual Training for Multi-View Object Class RecognitionabstractOur goal is to circumvent one of the roadblocks to using existing approaches for single-view recognition for achieving multi-view recognition, namely, the need for sufficient training data for many viewpoints. We show how to construct virtual training examples for multi-view recognition using a simple model of objects (nearly planar facades centered at fixed 3D positions). We also show how the models can be learned from a few labeled images for each class. Han-Pang Chiu, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
CVPR | 2 |
| 2007 | Grasping POMDPsabstractWe provide a method for planning under uncertainty for robotic manipulation by partitioning the configuration space into a set of regions that are closed under compliant motions. These regions can be treated as states in a partially observable Markov decision process (POMDP), which can be solved to yield optimal control policies under uncertainty. We demonstrate the approach on simple grasping problems, showing that it can construct highly robust, efficiently executable solutions Kaijen Hsiao, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ICRA | 2 |
| 2007 | Efficient Bayesian Task-Level Transfer Learning
Daniel M. Roy 0001, Leslie Pack Kaelbling |
IJCAI | 2 |
| 2007 | Learning Probabilistic Relational Dynamics for Multiple Tasks
Ashwin Deshpande, Brian Milch, Luke Zettlemoyer, Leslie Pack Kaelbling |
UAI | 4 |
| 2007 | Learning Symbolic Models of Stochastic DomainsabstractIn this article, we work towards the goal of developing agents that can learn to act in complex worlds. We develop a probabilistic, relational planning rule representation that compactly models noisy, nondeterministic action effects, and show how such rules can be effectively learned. Through experiments in simple planning domains and a 3D simulated blocks world with realistic physics, we demonstrate that this learning algorithm allows agents to effectively model world dynamics. Hanna M. Pasula, Luke Zettlemoyer, Leslie Pack Kaelbling |
J. Artif. Intell. Res. | 3 |
| 2005 | Learning Static Object Segmentation from Motion Segmentation
Michael G. Ross, Leslie Pack Kaelbling |
AAAI | 2 |
| 2005 | Learning Planning Rules in Noisy Stochastic Worlds
Luke Zettlemoyer, Hanna M. Pasula, Leslie Pack Kaelbling |
AAAI | 3 |
| 2005 | Hedged learning: regret-minimization with learning expertsabstractIn non-cooperative multi-agent situations, there cannot exist a globally optimal, yet opponent-independent learning algorithm. Regret-minimization over a set of strategies optimized for potential opponent models is proposed as a good framework for deciding how to behave in such situations. Using longer playing horizons and experts that learn as they play, the regret-minimization framework can be extended to overcome several shortcomings of earlier approaches to the problem of multi-agent learning. Yu-Han Chang, Leslie Pack Kaelbling |
ICML | 2 |
| 2004 | Representing Hierarchical POMDPs as DBNs for Multi-scale Robot LocalizationabstractWe explore the advantages of representing hierarchical partially observable Markov decision processes (H-POMDPs) as dynamic Bayesian networks (DBNs). In particular, we focus on the special case of using H-POMDPs to represent multi-resolution spatial maps for indoor robot navigation. Our results show that a DBN representation of H-POMDPs can train significantly faster than the original learning algorithm for H-POMDPs or the equivalent flat POMDP, and requires much less data. In addition, the DBN formulation can easily be extended to parameter tying and factoring of variables, which further reduces the time and sample complexity. This enables us to apply H-POMDP methods to much larger problems than previously possible. Georgios Theocharous, Kevin Murphy 0002, Leslie Pack Kaelbling |
ICRA | 3 |
| 2004 | Learning distributed control for modular robotsabstractWe propose to automate controller design for distributed modular robots. In this paper, we present some initial experiments with learning distributed controllers for synthesizing compliant locomotion gaits for modular, self-reconfigurable robots. We use both centralized and distributed policy search and find that the learning approach is promising, as locomotion tasks are learnt well. We also find that the additive nature of the robotic platforms can help speed up learning if we increase the robot size incrementally. Paulina Varshavskaya, Leslie Pack Kaelbling, Daniela Rus |
IROS | 2 |
| 2004 | Learning Probabilistic Relational Planning Rules
Hanna M. Pasula, Luke Zettlemoyer, Leslie Pack Kaelbling |
KR | 3 |
| 2003 | All learning is Local: Multi-agent Learning in Global Reward GamesabstractIn large multiagent games, partial observability, coordination, and credit assignment persistently plague attempts to design good learning algo- rithms. We provide a simple and efficient algorithm that in part uses a linear system to model the world from a single agent’s limited per- spective, and takes advantage of Kalman filtering to allow an agent to construct a good training signal and learn an effective policy. Yu-Han Chang, Tracey Ho, Leslie Pack Kaelbling |
NIPS | 3 |
| 2003 | Envelope-based Planning in Relational MDPsabstractA mobile robot acting in the world is faced with a large amount of sen- sory data and uncertainty in its action outcomes. Indeed, almost all in- teresting sequential decision-making domains involve large state spaces and large, stochastic action sets. We investigate a way to act intelli- gently as quickly as possible in domains where finding a complete policy would take a hopelessly long time. This approach, Relational Envelope- based Planning (REBP) tackles large, noisy problems along two axes. First, describing a domain as a relational MDP (instead of as an atomic or propositionally-factored MDP) allows problem structure and dynam- ics to be captured compactly with a small set of probabilistic, relational rules. Second, an envelope-based approach to planning lets an agent be- gin acting quickly within a restricted part of the full state space and to judiciously expand its envelope as resources permit. Natalia Hernandez-Gardiol, Leslie Pack Kaelbling |
NIPS | 2 |
| 2003 | Approximate Planning in POMDPs with Macro-Actions
Georgios Theocharous, Leslie Pack Kaelbling |
NIPS | 2 |
| 2003 | A Dynamical Model of Visually-Guided Steering, Obstacle Avoidance, and Route Selection
Brett R. Fajen, William H. Warren, Selim Temizer, Leslie Pack Kaelbling |
Int. J. Comput. Vis. | 4 |
| 2002 | Effective Reinforcement Learning for Mobile RobotsabstractProgramming mobile robots can be a long, time-consuming process. Specifying the low-level mapping from sensors to actuators is prone to programmer misconceptions, and debugging such a mapping can be tedious. The idea of having a robot learn how to accomplish a task, rather than being told explicitly, is an appealing one. It seems easier and much more intuitive for the programmer to specify what the robot should be doing, and to let it learn the fine details of how to do it. In this paper, we introduce a framework for reinforcement learning on mobile robots and describe our experiments using it to learn simple tasks. William D. Smart, Leslie Pack Kaelbling |
ICRA | 2 |
| 2002 | The Thing that we Tried Didn't Work very Well: Deictic Representation in Reinforcement Learning
Sarah Finney, Natalia Hernandez-Gardiol, Leslie Pack Kaelbling, Tim Oates 0001 |
UAI | 3 |
| 2002 | Learning Geometrically-Constrained Hidden Markov Models for Robot Navigation: Bridging the Topological-Geometrical GapabstractHidden Markov models (HMMs) and partially observable Markov decision processes (POMDPs) provide useful tools for modeling dynamical systems. They are particularly useful for representing the topology of environments such as road networks and office buildings, which are typical for robot navigation and planning. The work presented here describes a formal framework for incorporating readily available odometric information and geometrical constraints into both the models and the algorithm that learns them. By taking advantage of such information, learning HMMs/POMDPs can be made to generate better solutions and require fewer iterations, while being robust in the face of data reduction. Experimental results, obtained from both simulated and real robot data, demonstrate the effectiveness of the approach. Hagit Shatkay, Leslie Pack Kaelbling |
J. Artif. Intell. Res. | 2 |
| 2001 | Playing is believing: The role of beliefs in multi-agent learningabstractWe propose a new classification for multi-agent learning algorithms, with each league of players characterized by both their possible strategies and possible beliefs. Using this classification, we review the optimality of ex- isting algorithms, including the case of interleague play. We propose an incremental improvement to the existing algorithms that seems to achieve average payoffs that are at least the Nash equilibrium payoffs in the long- run against fair opponents. Yu-Han Chang, Leslie Pack Kaelbling |
NIPS | 2 |
| 2000 | State-based Classification of Finger Gestures from Electromyographic Signals
Peter Ju, Leslie Pack Kaelbling, Yoram Singer |
ICML | 2 |
| 2000 | Practical Reinforcement Learning in Continuous Spaces
William D. Smart, Leslie Pack Kaelbling |
ICML | 2 |
| 2000 | Adaptive Importance Sampling for Estimation in Structured Domains
Luis E. Ortiz, Leslie Pack Kaelbling |
UAI | 2 |
| 2000 | Learning to Cooperate via Policy Search
Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau, Leslie Pack Kaelbling |
UAI | 4 |
| 1999 | Learning Policies with External Memory
Leonid Peshkin, Nicolas Meuleau, Leslie Pack Kaelbling |
ICML | 3 |
| 1999 | Multi-Value-Functions: Efficient Automatic Action Hierarchies for Multiple Goal MDPs
Andrew W. Moore 0001, Leemon Baird, Leslie Pack Kaelbling |
IJCAI | 3 |
| 1999 | Solving POMDPs by Searching the Space of Finite Policies
Nicolas Meuleau, Kee-Eung Kim, Leslie Pack Kaelbling, Anthony R. Cassandra |
UAI | 3 |
| 1999 | Learning Finite-State Controllers for Partially Observable Environments
Nicolas Meuleau, Leonid Peshkin, Kee-Eung Kim, Leslie Pack Kaelbling |
UAI | 4 |
| 1999 | Accelerating EM: An Empirical Study
Luis E. Ortiz, Leslie Pack Kaelbling |
UAI | 2 |
| 1998 | Heading in the Right Direction
Hagit Shatkay, Leslie Pack Kaelbling |
ICML | 2 |
| 1998 | Hierarchical Solution of Markov Decision Processes using Macro-actions
Milos Hauskrecht, Nicolas Meuleau, Leslie Pack Kaelbling, Thomas L. Dean, Craig Boutilier |
UAI | 3 |
| 1998 | Planning and Acting in Partially Observable Stochastic Domains
Leslie Pack Kaelbling, Michael L. Littman, Anthony R. Cassandra |
Artif. Intell. | 1 |
| 1997 | Learning Topological Maps with Weak Local Odometric Information
Hagit Shatkay, Leslie Pack Kaelbling |
IJCAI (2) | 2 |
| 1996 | Acting under uncertainty: discrete Bayesian models for mobile-robot navigationabstractDiscrete Bayesian models have been used to model uncertainty for mobile-robot navigation, but the question of how actions should be chosen remains largely unexplored. This paper presents the optimal solution to the problem, formulated as a partially observable Markov decision process. Since solving for the optimal control policy is intractable, in general, it goes on to explore a variety of heuristic control strategies. The control strategies are compared experimentally, both in simulation and in runs on a robot. Anthony R. Cassandra, Leslie Pack Kaelbling, James Kurien |
IROS | 2 |
| 1996 | On reinforcement learning for robotsabstractIn the paper the author considers: what reinforcement learning is; what makes reinforcement learning difficult; biasing reinforcement learning; and integrating reinforcement learning into agent architectures. Leslie Pack Kaelbling |
IROS | 1 |
| 1996 | Reinforcement Learning: A SurveyabstractThis paper surveys the field of reinforcement learning from a computer-science perspective. It is written to be accessible to researchers familiar with machine learning. Both the historical basis of the field and a broad selection of current work are summarized. Reinforcement learning is the problem faced by an agent that learns behavior through trial-and-error interactions with a dynamic environment. The work described here has a resemblance to work in psychology, but differs considerably in the details and in the use of the word ``reinforcement.'' The paper discusses central issues of reinforcement learning, including trading off exploration and exploitation, establishing the foundations of the field via Markov decision theory, learning from delayed reinforcement, constructing empirical models to accelerate learning, making use of generalization and hierarchy, and coping with hidden state. It concludes with a survey of some implemented systems and an assessment of the practical utility of current methods for reinforcement learning. Leslie Pack Kaelbling, Michael L. Littman, Andrew W. Moore 0001 |
J. Artif. Intell. Res. | 1 |
| 1996 | Introduction
Leslie Pack Kaelbling |
Mach. Learn. | 1 |
| 1995 | Learning Policies for Partially Observable Environments: Scaling Up
Michael L. Littman, Anthony R. Cassandra, Leslie Pack Kaelbling |
ICML | 3 |
| 1995 | On the Complexity of Solving Markov Decision Problems
Michael L. Littman, Thomas L. Dean, Leslie Pack Kaelbling |
UAI | 3 |
| 1995 | Learning Dynamics: System Identification for Perceptually Challenged Agents
Kenneth Basye, Thomas L. Dean, Leslie Pack Kaelbling |
Artif. Intell. | 3 |
| 1995 | Planning under Time Constraints in Stochastic Domains
Thomas L. Dean, Leslie Pack Kaelbling, Jak Kirman, Ann E. Nicholson |
Artif. Intell. | 2 |
| 1995 | A Situated View of Representation and Control
Stanley J. Rosenschein, Leslie Pack Kaelbling |
Artif. Intell. | 2 |
| 1995 | Inferring Finite Automata with Stochastic Output Functions and an Application to Map Learning
Thomas L. Dean, Dana Angluin, Kenneth Basye, Shlomo Argamon, Leslie Pack Kaelbling, Evangelos Kokkevis, Oded Maron |
Mach. Learn. | 5 |
| 1994 | Acting Optimally in Partially Observable Stochastic Domains
Anthony R. Cassandra, Leslie Pack Kaelbling, Michael L. Littman |
AAAI | 2 |
| 1994 | Learning and intelligent Agents
Leslie Pack Kaelbling |
ECAI | 1 |
| 1994 | Associative Reinforcement Learning: Functions in k-DNF
Leslie Pack Kaelbling |
Mach. Learn. | 1 |
| 1994 | Associative Reinforcement Learning: A Generate and Test Algorithm
Leslie Pack Kaelbling |
Mach. Learn. | 1 |
| 1993 | Planning With Deadlines in Stochastic Domains
Thomas L. Dean, Leslie Pack Kaelbling, Jak Kirman, Ann E. Nicholson |
AAAI | 2 |
| 1993 | Hierarchical Learning in Stochastic Domains: Preliminary Results
Leslie Pack Kaelbling |
ICML | 1 |
| 1993 | Learning to Achieve Goals
Leslie Pack Kaelbling |
IJCAI | 1 |
| 1993 | Deliberation Scheduling for Time-Critical Sequential Decision Making
Thomas L. Dean, Leslie Pack Kaelbling, Jak Kirman, Ann E. Nicholson |
UAI | 2 |
| 1992 | Inferring Finite Automata with Stochastic Output Functions and an Application to Map Learning
Thomas L. Dean, Dana Angluin, Kenneth Basye, Shlomo Argamon, Leslie Pack Kaelbling, Evangelos Kokkevis, Oded Maron |
AAAI | 5 |
| 1991 | Input Generalization in Delayed Reinforcement Learning: An Algorithm and Performance Comparisons
Leslie Pack Kaelbling |
IJCAI | 2 |
| 1990 | Learning Functions in k-DNF from Reinforcement
Leslie Pack Kaelbling |
ML | 1 |
| 1989 | A Formal Framework for Learning in Embedded Systems
Leslie Pack Kaelbling |
ML | 1 |
| 1988 | Goals as Parallel Program Specifications
Leslie Pack Kaelbling |
AAAI | 1 |
| 1986 | The Synthesis of Digital Machines With Provable Epistemic Properties
Stanley J. Rosenschein, Leslie Pack Kaelbling |
TARK | 2 |