EDBT 2026 Demo / reviewers in the wild / expert
Mrinal Kalakrishnan
dblp:46/4195
· DBLP profile ↗
30ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-4292-9857ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 5 since 2021Systems, architecture and hardware · 21 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
20 papers |
Robot manipulation · 28% Transfer learning and domain adaptation · 20% Motion planning and robot control · 14% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% |
Topics — the 30 heaviest of 49, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation
grasping |
2.0 | 6 | 2020 | Action Image Representation: Learning Scalable Deep Grasping Policies with Zero Real World Data · ICRA 2020 Learning Probabilistic Multi-Modal Actor Models for Vision-Based Robotic Grasping · ICRA 2019 Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019 |
Robotics › Robot manipulation
learning from demonstration |
1.2 | 2 | 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024 Watch, Try, Learn: Meta-Learning from Demonstrations and Rewards · ICLR 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
1.0 | 3 | 2019 | Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019 Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018 |
Computer vision › 3D vision › 3d object detection
3d object localization |
0.9 | 1 | 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025 |
Computer vision › Vision and language › visual grounding
referential grounding |
0.9 | 1 | 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025 |
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Computer vision › 3D vision
environmental understanding |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Robotics › Robot navigation and mapping
social navigation |
0.8 | 1 | 2024 | Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.6 | 3 | 2020 | Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019 Action Image Representation: Learning Scalable Deep Grasping Policies with Zero Real World Data · ICRA 2020 Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018 |
Machine learning › Transfer learning and domain adaptation › sim-to-real transfer
domain randomization |
0.5 | 2 | 2020 | Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation Networks · CVPR 2019 Action Image Representation: Learning Scalable Deep Grasping Policies with Zero Real World Data · ICRA 2020 |
Robotics › Motion planning and robot control
robot learning |
0.5 | 2 | 2017 | Path integral guided policy search · ICRA 2017 Learning objective functions for manipulation · ICRA 2013 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.4 | 1 | 2020 | Watch, Try, Learn: Meta-Learning from Demonstrations and Rewards · ICLR 2020 |
Robotics › Motion planning and robot control
motion planning |
0.4 | 3 | 2013 | Learning objective functions for manipulation · ICRA 2013 STOMP: Stochastic trajectory optimization for motion planning · ICRA 2011 Combining planning techniques for manipulation using realtime perception · ICRA 2010 |
Robotics › Robot manipulation › grasping
vision-based grasping |
0.4 | 1 | 2019 | Learning Probabilistic Multi-Modal Actor Models for Vision-Based Robotic Grasping · ICRA 2019 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation |
0.3 | 1 | 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 |
Robotics › Robot manipulation › grasping
instance grasping |
0.3 | 1 | 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from Simulation · ICRA 2018 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
pixel-level domain adaptation |
0.3 | 1 | 2018 | Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping · ICRA 2018 |
Machine learning › Reinforcement learning › policy search
guided policy search |
0.3 | 1 | 2017 | Path integral guided policy search · ICRA 2017 |
Robotics › Motion planning and robot control
robot control |
0.3 | 2 | 2013 | Learning objective functions for manipulation · ICRA 2013 Fast, robust quadruped locomotion over challenging terrain · ICRA 2010 |
Robotics › Robot navigation and mapping
embodied AI simulation |
0.2 | 1 | 2024 | Habitat 3.0: A Co-Habitat for Humans, Avatars, and Robots · ICLR 2024 |
Robotics › Robot navigation and mapping › mobile robot navigation
indoor navigation |
0.2 | 1 | 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024 |
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
zero-shot sim-to-real transfer |
0.2 | 1 | 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024 |
Robotics › Motion planning and robot control
trajectory optimization |
0.2 | 2 | 2010 | Combining planning techniques for manipulation using realtime perception · ICRA 2010 Fast, robust quadruped locomotion over challenging terrain · ICRA 2010 |
Robotics › Motion planning and robot control › robot control
inverse kinematics |
0.2 | 1 | 2013 | Learning objective functions for manipulation · ICRA 2013 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.2 | 1 | 2013 | Learning objective functions for manipulation · ICRA 2013 |
Robotics › Motion planning and robot control › robot calibration
kinematic error modeling |
0.2 | 1 | 2013 | Learning task error models for manipulation · ICRA 2013 |
Robotics › Motion planning and robot control › motion planning
optimization-based motion planning |
0.2 | 1 | 2013 | Learning objective functions for manipulation · ICRA 2013 |
Robotics › Robot manipulation
compliant manipulation |
0.1 | 1 | 2012 | Learning Force Control Policies for Compliant Robotic Manipulation · ICML 2012 |
Robotics › Robot manipulation › grasping › grasp planning
grasp selection |
0.1 | 1 | 2012 | Template-based learning of grasp selection · ICRA 2012 |
Robotics › Motion planning and robot control › robot learning › sensorimotor learning
motor skill learning |
0.1 | 1 | 2011 | Skill learning and task outcome prediction for manipulation · ICRA 2011 |
Methods — techniques the papers use, named apart from their topics
humanoid simulation · 1.5VR interface · 1.5reinforcement learning · 0.9self-supervised learning · 0.9masked prediction · 0.9foundation model features · 0.9domain randomization · 0.8learned policy · 0.8learned policies · 0.8large language model evaluation · 0.8foundation model · 0.8network integration · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3DabstractWe present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state-of-the-art on standard referential grounding benchmarks and showcases robust generalization capabilities. Notably, LOCATE 3D operates directly on sensor observation streams (posed RGB-D frames), enabling real-world deployment on robots and AR devices. Key to our approach is 3D-JEPA, a novel self-supervised learning (SSL) algorithm applicable to sensor point clouds. It takes as input a 3D pointcloud featurized using 2D foundation models (CLIP, DINO). Subsequently, masked prediction in latent space is employed as a pretext task to aid the self-supervised learning of contextualized pointcloud features. Once trained, the 3D-JEPA encoder is finetuned alongside a language-conditioned decoder to jointly predict 3D masks and bounding boxes. Additionally, we introduce LOCATE 3D DATASET, a new dataset for 3D referential grounding, spanning multiple capture setups with over 130K annotations. This enables a systematic study of generalization capabilities as well as a stronger model. Code, models and dataset can be found at the project website: locate3d.atmeta.com Paul McVay, Sergio Arnaud, Ada Martin, Arjun Majumdar, Krishna Murthy Jatavallabhula, Phillip Thomas, Ruslan Partsey, Daniel Dugas, Abha Gejji, Alexander Sax, Vincent-Pierre Berges, Mikael Henaff, Ang Cao, Ishita Prasad, Mrinal Kalakrishnan, Michael G. Rabbat, Nicolas Ballas, Mido Assran, Oleksandr Maksymets, Aravind Rajeswaran |
ICML | 16 |
| 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation ModelsabstractWe present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models. Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran |
CVPR | 20 |
| 2024 | Habitat 3.0: A Co-Habitat for Humans, Avatars, and RobotsabstractWe present Habitat 3.0: a simulation platform for studying collaborative human-robot tasks in home environments. Habitat 3.0 offers contributions across three dimensions: (1) Accurate humanoid simulation: addressing challenges in modeling complex deformable bodies and diversity in appearance and motion, all while ensuring high simulation speed. (2) Human-in-the-loop infrastructure: enabling real human interaction with simulated robots via mouse/keyboard or a VR interface, facilitating evaluation of robot policies with human input. (3) Collaborative tasks: studying two collaborative tasks, Social Navigation and Social Rearrangement. Social Navigation investigates a robot's ability to locate and follow humanoid avatars in unseen environments, whereas Social Rearrangement addresses collaboration between a humanoid and robot while rearranging a scene. These contributions allow us to study end-to-end learned and heuristic baselines for human-robot collaboration in-depth, as well as evaluate them with humans in the loop. Our experiments demonstrate that learned robot policies lead to efficient task completion when collaborating with unseen humanoid agents and human partners that might exhibit behaviors that the robot has not seen before. Additionally, we observe emergent behaviors during collaborative task execution, such as the robot yielding space when obstructing a humanoid agent, thereby allowing the effective completion of the task by the humanoid agent. Furthermore, our experiments using the human-in-the-loop tool demonstrate that our automated evaluation with humanoids can provide an indication of the relative ordering of different policies when evaluated with real human collaborators. Habitat 3.0 unlocks interesting new features in simulators for Embodied AI, and we hope it paves the way for a new frontier of embodied human-AI interaction capabilities. For more details and visualizations, visit: https://aihabitat.org/habitat3. Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander Clegg, Michal Hlavac, So Yeon Min, Vladimir Vondrus, Théophile Gervet, Vincent-Pierre Berges, John M. Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakrishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, Roozbeh Mottaghi |
ICLR | 17 |
| 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?abstractWe present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluation using three different robots and two different policy learning paradigms. From this e ort, we can arrive at three insights: 1) the performance trends of PVRs in the simulation are generally indicative of their trends in the real world, 2) the use of PVRs enables a first-of-its-kind result with indoor ImageNav (zero-shot transfer to a held-out scene in the real world), and 3) the benefits from variations in PVRs, primarily data-augmentation and fine-tuning, also transfer to the real-world performance. See project website1for additional details and visuals. Sneha Silwal, Karmesh Yadav, Tingfan Wu, Jay Vakil, Arjun Majumdar, Sergio Arnaud, Vincent-Pierre Berges, Dhruv Batra, Aravind Rajeswaran, Mrinal Kalakrishnan, Franziska Meier, Oleksandr Maksymets |
ICRA | 11 |
| 2023 | USA-Net: Unified Semantic and Affordance Representations for Robot MemoryabstractIn order for robots to follow open-ended instructions like “go open the brown cabinet over the sink,” they require an understanding of both the scene geometry and the semantics of their environment. Robotic systems often handle these through separate pipelines, sometimes using very different representation spaces, which can be suboptimal when the two objectives conflict. In this work, we present USA-Net, a simple method for constructing a world representation that encodes both the semantics and spatial affordances of a scene in a differentiable map. This allows us to build a gradient-based planner which can navigate to locations in the scene specified using open-ended vocabulary. We use this planner to consistently generate trajectories which are both shorter 5-10% shorter and 10-30% closer to our goal query in CLIP embedding space than paths from comparable grid-based planners which don't leverage gradient information. To our knowledge, this is the first end-to-end differentiable planner optimizes for both semantics and affordance in a single implicit map. Code and visuals are available at our website: usa.bolte.cc Benjamin Bolte, Austin S. Wang, Jimmy Yang, Mustafa Mukadam, Mrinal Kalakrishnan, Chris Paxton 0001 |
IROS | 5 |
| 2020 | Watch, Try, Learn: Meta-Learning from Demonstrations and Rewards
Allan Zhou, Eric Jang, Daniel Kappler, Mohi Khansari, Paul Wohlhart, Mrinal Kalakrishnan, Sergey Levine, Chelsea Finn |
ICLR | 8 |
| 2020 | Action Image Representation: Learning Scalable Deep Grasping Policies with Zero Real World DataabstractThis paper introduces Action Image, a new grasp proposal representation that allows learning an end-to-end deep-grasping policy. Our model achieves 84% grasp success on 172 real world objects while being trained only in simulation on 48 objects with just naive domain randomization. Similar to computer vision problems, such as object detection, Action Image builds on the idea that object features are invariant to translation in image space. Therefore, grasp quality is invariant when evaluating the object-gripper relationship; a successful grasp for an object depends on its local context, but is independent of the surrounding environment. Action Image represents a grasp proposal as an image and uses a deep convolutional network to infer grasp quality. We show that by using an Action Image representation, trained networks are able to extract local, salient features of grasping tasks that generalize across different objects and environments. We show that this representation works on a variety of inputs, including color images (RGB), depth images (D), and combined color-depth (RGB-D). Our experimental results demonstrate that networks utilizing an Action Image representation exhibit strong domain transfer between training on simulated data and inference on real-world sensor streams. Finally, our experiments show that a network trained with Action Image improves grasp success (84% vs. 53%) over a baseline model with the same structure, but using actions encoded as vectors. Mohi Khansari, Daniel Kappler, Jianlan Luo, Jeffrey T. Bingham, Mrinal Kalakrishnan |
ICRA | 5 |
| 2019 | Sim-To-Real via Sim-To-Sim: Data-Efficient Robotic Grasping via Randomized-To-Canonical Adaptation NetworksabstractReal world data, especially in the domain of robotics, is notoriously costly to collect. One way to circumvent this can be to leverage the power of simulation to produce large amounts of labelled data. However, training models on simulated images does not readily transfer to real-world ones. Using domain adaptation methods to cross this "reality gap" requires a large amount of unlabelled real-world data, whilst domain randomization alone can waste modeling power. In this paper, we present Randomized-to-Canonical Adaptation Networks (RCANs), a novel approach to crossing the visual reality gap that uses no real-world data. Our method learns to translate randomized rendered images into their equivalent non-randomized, canonical versions. This in turn allows for real images to also be translated into canonical sim images. We demonstrate the effectiveness of this sim-to-real approach by training a vision-based closed-loop grasping reinforcement learning agent in simulation, and then transferring it to the real world to attain 70% zero-shot grasp success on unseen objects, a result that almost doubles the success of learning the same task directly on domain randomization alone. Additionally, by joint finetuning in the real-world with only 5,000 real-world grasps, our method achieves 91%, attaining comparable performance to a state-of-the-art system trained with 580,000 real-world grasps, resulting in a reduction of real-world data by more than 99%. Stephen James, Paul Wohlhart, Mrinal Kalakrishnan, Dmitry Kalashnikov, Alex Irpan, Julian Ibarz, Sergey Levine, Raia Hadsell, Konstantinos Bousmalis |
CVPR | 3 |
| 2019 | Learning Probabilistic Multi-Modal Actor Models for Vision-Based Robotic GraspingabstractMany previous works approach vision-based robotic grasping by training a value network that evaluates grasp proposals. These approaches require an optimization process at run-time to infer the best action from the value network. As a result, the inference time grows exponentially as the dimension of action space increases. We propose an alternative method, by directly training a neural density model to approximate the conditional distribution of successful grasp poses from the input images. We construct a neural network that combines Gaussian mixture and normalizing flows, which is able to represent multi-modal, complex probability distributions. We demonstrate on both simulation and real robot that the proposed actor model achieves similar performance compared to the value network using the Cross-Entropy Method (CEM) for inference, on top-down grasping with a 4 dimensional action space. Our actor model reduces the inference time by 3 times compared to the state-of-the-art CEM method. We believe that actor models will play an important role when scaling up these approaches to higher dimensional action spaces. Mengyuan Yan, Adrian Li, Mrinal Kalakrishnan, Peter Pastor |
ICRA | 3 |
| 2018 | Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic GraspingabstractInstrumenting and collecting annotated visual grasping datasets to train modern machine learning algorithms can be extremely time-consuming and expensive. An appealing alternative is to use off-the-shelf simulators to render synthetic data for which ground-truth annotations are generated automatically. Unfortunately, models trained purely on simulated data often fail to generalize to the real world. We study how randomized simulated environments and domain adaptation methods can be extended to train a grasping system to grasp novel objects from raw monocular RGB images. We extensively evaluate our approaches with a total of more than 25,000 physical test grasps, studying a range of simulation conditions and domain adaptation methods, including a novel extension of pixel-level domain adaptation that we term the GraspGAN. We show that, by using synthetic data and domain adaptation, we are able to reduce the number of real-world samples needed to achieve a given level of performance by up to 50 times, using only randomly generated simulated objects. We also show that by using only unlabeled real-world data and our GraspGAN methodology, we obtain real-world grasping performance without any real-world labels that is similar to that achieved with 939,777 labeled real-world samples. Konstantinos Bousmalis, Alex Irpan, Paul Wohlhart, Matthew Kelcey, Mrinal Kalakrishnan, Laura Downs, Julian Ibarz, Peter Pastor, Kurt Konolige, Sergey Levine, Vincent Vanhoucke |
ICRA | 6 |
| 2018 | Multi-Task Domain Adaptation for Deep Learning of Instance Grasping from SimulationabstractLearning-based approaches to robotic manipulation are limited by the scalability of data collection and accessibility of labels. In this paper, we present a multi-task domain adaptation framework for instance grasping in cluttered scenes by utilizing simulated robot experiments. Our neural network takes monocular RGB images and the instance segmentation mask of a specified target object as inputs, and predicts the probability of successfully grasping the specified object for each candidate motor command. The proposed transfer learning framework trains a model for instance grasping in simulation and uses a domain-adversarial loss to transfer the trained model to real robots using indiscriminate grasping data, which is available both in simulation and the real world. We evaluate our model in real-world robot experiments, comparing it with alternative model architectures as well as an indiscriminate grasping baseline. Kuan Fang, Stefan Hinterstoißer, Silvio Savarese, Mrinal Kalakrishnan |
ICRA | 5 |
| 2017 | Path integral guided policy searchabstractWe present a policy search method for learning complex feedback control policies that map from high-dimensional sensory inputs to motor torques, for manipulation tasks with discontinuous contact dynamics. We build on a prior technique called guided policy search (GPS), which iteratively optimizes a set of local policies for specific instances of a task, and uses these to train a complex, high-dimensional global policy that generalizes across task instances. We extend GPS in the following ways: (1) we propose the use of a model-free local optimizer based on path integral stochastic optimal control (PI2), which enables us to learn local policies for tasks with highly discontinuous contact dynamics; and (2) we enable GPS to train on a new set of task instances in every iteration by using on-policy sampling: this increases the diversity of the instances that the policy is trained on, and is crucial for achieving good generalization. We show that these contributions enable us to learn deep neural network policies that can directly perform torque control from visual input. We validate the method on a challenging door opening task and a pick-and-place task, and we demonstrate that our approach substantially outperforms the prior LQR-based local policy optimizer on these tasks. Furthermore, we show that on-policy sampling significantly increases the generalization ability of these policies. Yevgen Chebotar, Mrinal Kalakrishnan, Ali Yahya, Adrian Li, Stefan Schaal, Sergey Levine |
ICRA | 2 |
| 2017 | Collective robot reinforcement learning with distributed asynchronous guided policy searchabstractPolicy search methods and, more broadly, reinforcement learning can enable robots to learn highly complex and general skills that may allow them to function amid the complexity and diversity of the real world. However, training a policy that generalizes well across a wide range of real-world conditions requires far greater quantity and diversity of experience than is practical to collect with a single robot. Fortunately, it is possible for multiple robots to share their experience with one another, and thereby, learn a policy collectively. In this work, we explore distributed and asynchronous policy learning as a means to achieve generalization and improved training times on challenging, real-world manipulation tasks. We propose a distributed and asynchronous version of guided policy search and use it to demonstrate collective policy learning on a vision-based door opening task using four robots. We describe how both policy learning and data collection can be conducted in parallel across multiple robots, and present a detailed empirical evaluation of our system. Our results indicate that distributed learning significantly improves training time, and that parallelizing policy learning and data collection substantially improves utilization. We also demonstrate that we can achieve substantial generalization on a challenging real-world door opening task. Ali Yahya, Adrian Li, Mrinal Kalakrishnan, Yevgen Chebotar, Sergey Levine |
IROS | 3 |
| 2013 | Learning objective functions for manipulationabstractWe present an approach to learning objective functions for robotic manipulation based on inverse reinforcement learning. Our path integral inverse reinforcement learning algorithm can deal with high-dimensional continuous state-action spaces, and only requires local optimality of demonstrated trajectories. We use L1regularization in order to achieve feature selection, and propose an efficient algorithm to minimize the resulting convex objective function. We demonstrate our approach by applying it to two core problems in robotic manipulation. First, we learn a cost function for redundancy resolution in inverse kinematics. Second, we use our method to learn a cost function over trajectories, which is then used in optimization-based motion planning for grasping and manipulation tasks. Experimental results show that our method outperforms previous algorithms in high-dimensional settings. Mrinal Kalakrishnan, Peter Pastor, Ludovic Righetti, Stefan Schaal |
ICRA | 1 |
| 2013 | Learning task error models for manipulationabstractPrecise kinematic forward models are important for robots to successfully perform dexterous grasping and manipulation tasks, especially when visual servoing is rendered infeasible due to occlusions. A lot of research has been conducted to estimate geometric and non-geometric parameters of kinematic chains to minimize reconstruction errors. However, kinematic chains can include non-linearities, e.g. due to cable stretch and motor-side encoders, that result in significantly different errors for different parts of the state space. Previous work either does not consider such non-linearities or proposes to estimate non-geometric parameters of carefully engineered models that are robot specific. We propose a data-driven approach that learns task error models that account for such unmodeled non-linearities. We argue that in the context of grasping and manipulation, it is sufficient to achieve high accuracy in the task relevant state space. We identify this relevant state space using previously executed joint configurations and learn error corrections for those. Therefore, our system is developed to generate subsequent executions that are similar to previous ones. The experiments show that our method successfully captures the non-linearities in the head kinematic chain (due to a counterbalancing spring) and the arm kinematic chains (due to cable stretch) of the considered experimental platform, see Fig. 1. The feasibility of the presented error learning approach has also been evaluated in independent DARPA ARM-S testing contributing to successfully complete 67 out of 72 grasping and manipulation tasks. Peter Pastor, Mrinal Kalakrishnan, Jonathan Binney, Jonathan Kelly, Ludovic Righetti, Gaurav S. Sukhatme, Stefan Schaal |
ICRA | 2 |
| 2013 | Probabilistic object tracking using a range cameraabstractWe address the problem of tracking the 6-DoF pose of an object while it is being manipulated by a human or a robot. We use a dynamic Bayesian network to perform inference and compute a posterior distribution over the current object pose. Depending on whether a robot or a human manipulates the object, we employ a process model with or without knowledge of control inputs. Observations are obtained from a range camera. As opposed to previous object tracking methods, we explicitly model self-occlusions and occlusions from the environment, e.g, the human or robotic hand. This leads to a strongly non-linear observation model and additional dependencies in the Bayesian network. We employ a Rao-Blackwellised particle filter to compute an estimate of the object pose at every time step. In a set of experiments, we demonstrate the ability of our method to accurately and robustly track the object pose in real-time while it is being manipulated by a human or a robot. Manuel Wüthrich, Peter Pastor, Mrinal Kalakrishnan, Jeannette Bohg, Stefan Schaal |
IROS | 3 |
| 2012 | Learning Force Control Policies for Compliant Robotic Manipulation
Mrinal Kalakrishnan, Ludovic Righetti, Peter Pastor, Stefan Schaal |
ICML | 1 |
| 2012 | Template-based learning of grasp selectionabstractThe ability to grasp unknown objects is an important skill for personal robots, which has been addressed by many present and past research projects, but still remains an open problem. A crucial aspect of grasping is choosing an appropriate grasp configuration, i.e. the 6d pose of the hand relative to the object and its finger configuration. Finding feasible grasp configurations for novel objects, however, is challenging because of the huge variety in shape and size of these objects. Moreover, possible configurations also depend on the specific kinematics of the robotic arm and hand in use. In this paper, we introduce a new grasp selection algorithm able to find object grasp poses based on previously demonstrated grasps. Assuming that objects with similar shapes can be grasped in a similar way, we associate to each demonstrated grasp a grasp template. The template is a local shape descriptor for a possible grasp pose and is constructed using 3d information from depth sensors. For each new object to grasp, the algorithm then finds the best grasp candidate in the library of templates. The grasp selection is also able to improve over time using the information of previous grasp attempts to adapt the ranking of the templates. We tested the algorithm on two different platforms, the Willow Garage PR2 and the Barrett WAM arm which have very different hands. Our results show that the algorithm is able to find good grasp configurations for a large set of objects from a relatively small set of demonstrations, and does indeed improve its performance over time. Peter Pastor, Mrinal Kalakrishnan, Ludovic Righetti, Tamim Asfour, Stefan Schaal |
ICRA | 3 |
| 2011 | STOMP: Stochastic trajectory optimization for motion planningabstractWe present a new approach to motion planning using a stochastic trajectory optimization framework. The approach relies on generating noisy trajectories to explore the space around an initial (possibly infeasible) trajectory, which are then combined to produced an updated trajectory with lower cost. A cost function based on a combination of obstacle and smoothness cost is optimized in each iteration. No gradient information is required for the particular optimization algorithm that we use and so general costs for which derivatives may not be available (e.g. costs corresponding to constraints and motor torques) can be included in the cost function. We demonstrate the approach both in simulation and on a mobile manipulation system for unconstrained and constrained tasks. We experimentally show that the stochastic nature of STOMP allows it to overcome local minima that gradient-based methods like CHOMP can get stuck in. Mrinal Kalakrishnan, Sachin Chitta, Evangelos A. Theodorou, Peter Pastor, Stefan Schaal |
ICRA | 1 |
| 2011 | Skill learning and task outcome prediction for manipulationabstractLearning complex motor skills for real world tasks is a hard problem in robotic manipulation that often requires painstaking manual tuning and design by a human expert. In this work, we present a Reinforcement Learning based approach to acquiring new motor skills from demonstration. Our approach allows the robot to learn fine manipulation skills and significantly improve its success rate and skill level starting from a possibly coarse demonstration. Our approach aims to incorporate task domain knowledge, where appropriate, by working in a space consistent with the constraints of a specific task. In addition, we also present an approach to using sensor feedback to learn a predictive model of the task outcome. This allows our system to learn the proprioceptive sensor feedback needed to monitor subsequent executions of the task online and abort execution in the event of predicted failure. We illustrate our approach using two example tasks executed with the PR2 dual-arm robot: a straight and accurate pool stroke and a box flipping task using two chopsticks as tools. Peter Pastor, Mrinal Kalakrishnan, Sachin Chitta, Evangelos A. Theodorou, Stefan Schaal |
ICRA | 2 |
| 2011 | Learning force control policies for compliant manipulationabstractDeveloping robots capable of fine manipulation skills is of major importance in order to build truly assistive robots. These robots need to be compliant in their actuation and control in order to operate safely in human environments. Manipulation tasks imply complex contact interactions with the external world, and involve reasoning about the forces and torques to be applied. Planning under contact conditions is usually impractical due to computational complexity, and a lack of precise dynamics models of the environment. We present an approach to acquiring manipulation skills on compliant robots through reinforcement learning. The initial position control policy for manipulation is initialized through kinesthetic demonstration. We augment this policy with a force/torque profile to be controlled in combination with the position trajectories. We use the Policy Improvement with Path Integrals (PI2) algorithm to learn these force/torque profiles by optimizing a cost function that measures task success. We demonstrate our approach on the Barrett WAM robot arm equipped with a 6-DOF force/torque sensor on two different manipulation tasks: opening a door with a lever door handle, and picking up a pen off the table. We show that the learnt force control policies allow successful, robust execution of the tasks. Mrinal Kalakrishnan, Ludovic Righetti, Peter Pastor, Stefan Schaal |
IROS | 1 |
| 2011 | Online movement adaptation based on previous sensor experiencesabstractPersonal robots can only become widespread if they are capable of safely operating among humans. In uncertain and highly dynamic environments such as human households, robots need to be able to instantly adapt their behavior to unforseen events. In this paper, we propose a general framework to achieve very contact-reactive motions for robotic grasping and manipulation. Associating stereotypical movements to particular tasks enables our system to use previous sensor experiences as a predictive model for subsequent task executions. We use dynamical systems, named Dynamic Movement Primitives (DMPs), to learn goal-directed behaviors from demonstration. We exploit their dynamic properties by coupling them with the measured and predicted sensor traces. This feedback loop allows for online adaptation of the movement plan. Our system can create a rich set of possible motions that account for external perturbations and perception uncertainty to generate truly robust behaviors. As an example, we present an application to grasping with the WAM robot arm. Peter Pastor, Ludovic Righetti, Mrinal Kalakrishnan, Stefan Schaal |
IROS | 3 |
| 2011 | Learning motion primitive goals for robust manipulationabstractApplying model-free reinforcement learning to manipulation remains challenging for several reasons. First, manipulation involves physical contact, which causes discontinuous cost functions. Second, in manipulation, the end-point of the movement must be chosen carefully, as it represents a grasp which must be adapted to the pose and shape of the object. Finally, there is uncertainty in the object pose, and even the most carefully planned movement may fail if the object is not at the expected position. To address these challenges we 1) present a simplified, computationally more efficient version of our model-free reinforcement learning algorithm PI2; 2) extend PI2so that it simultaneously learns shape parameters and goal parameters of motion primitives; 3) use shape and goal learning to acquire motion primitives that are robust to object pose uncertainty. We evaluate these contributions on a manipulation platform consisting of a 7-DOF arm with a 4-DOF hand. Freek Stulp, Evangelos A. Theodorou, Mrinal Kalakrishnan, Peter Pastor, Ludovic Righetti, Stefan Schaal |
IROS | 3 |
| 2010 | Fast, robust quadruped locomotion over challenging terrainabstractWe present a control architecture for fast quadruped locomotion over rough terrain. We approach the problem by decomposing it into many sub-systems, in which we apply state-of-the-art learning, planning, optimization and control techniques to achieve robust, fast locomotion. Unique features of our control strategy include: (1) a system that learns optimal foothold choices from expert demonstration using terrain templates, (2) a body trajectory optimizer based on the Zero-Moment Point (ZMP) stability criterion, and (3) a floating-base inverse dynamics controller that, in conjunction with force control, allows for robust, compliant locomotion over unperceived obstacles. We evaluate the performance of our controller by testing it on the LittleDog quadruped robot, over a wide variety of rough terrain of varying difficulty levels. We demonstrate the generalization ability of this controller by presenting test results from an independent external test team on terrains that have never been shown to us. Mrinal Kalakrishnan, Jonas Buchli, Peter Pastor, Michael N. Mistry, Stefan Schaal |
ICRA | 1 |
| 2010 | Combining planning techniques for manipulation using realtime perceptionabstractWe present a novel combination of motion planning techniques to compute motion plans for robotic arms. We compute plans that move the arm as close as possible to the goal region using sampling-based planning and then switch to a trajectory optimization technique for the last few centimeters necessary to reach the goal region. This combination allows fast computation and safe execution of motion plans even when the goals are very close to objects in the environment. The system incorporates realtime sensory inputs and correctly deals with occlusions that can occur when robot body parts block the sensor view of the environment. The system is tested on a 7 degree-of-freedom robot arm with sensory input from a tilting laser scanner that provides 3D information about the environment. Ioan Alexandru Sucan, Mrinal Kalakrishnan, Sachin Chitta |
ICRA | 2 |
| 2009 | Compliant quadruped locomotion over rough terrainabstractMany critical elements for statically stable walking for legged robots have been known for a long time, including stability criteria based on support polygons, good foothold selection, recovery strategies to name a few. All these criteria have to be accounted for in the planning as well as the control phase. Most legged robots usually employ high gain position control, which means that it is crucially important that the planned reference trajectories are a good match for the actual terrain, and that tracking is accurate. Such an approach leads to conservative controllers, i.e. relatively low speed, ground speed matching, etc. Not surprisingly such controllers are not very robust - they are not suited for the real world use outside of the laboratory where the knowledge of the world is limited and error prone. Thus, to achieve robust robotic locomotion in the archetypical domain of legged systems, namely complex rough terrain, where the size of the obstacles are in the order of leg length, additional elements are required. A possible solution to improve the robustness of legged locomotion is to maximize the compliance of the controller. While compliance is trivially achieved by reduced feedback gains, for terrain requiring precise foot placement (e.g. climbing rocks, walking over pegs or cracks) compliance cannot be introduced at the cost of inferior tracking. Thus, model-based control and - in contrast to passive dynamic walkers - active balance control is required. To achieve these objectives, in this paper we add two crucial elements to legged locomotion, i.e., floating-base inverse dynamics control and predictive force control, and we show that these elements increase robustness in face of unknown and unanticipated perturbations (e.g. obstacles). Furthermore, we introduce a novel line-based COG trajectory planner, which yields a simpler algorithm than traditional polygon based methods and creates the appropriate input to our control system.We show results from both simulation and real world of a robotic dog walking over non-perceived obstacles and rocky terrain. The results prove the effectivity of the inverse dynamics/force controller. The presented results show that we have all elements needed for robust all-terrain locomotion, which should also generalize to other legged systems, e.g., humanoid robots. Jonas Buchli, Mrinal Kalakrishnan, Michael N. Mistry, Peter Pastor, Stefan Schaal |
IROS | 2 |
| 2009 | Learning locomotion over rough terrain using terrain templatesabstractWe address the problem of foothold selection in robotic legged locomotion over very rough terrain. The difficulty of the problem we address here is comparable to that of human rock-climbing, where foot/hand-hold selection is one of the most critical aspects. Previous work in this domain typically involves defining a reward function over footholds as a weighted linear combination of terrain features. However, a significant amount of effort needs to be spent in designing these features in order to model more complex decision functions, and hand-tuning their weights is not a trivial task. We propose the use of terrain templates, which are discretized height maps of the terrain under a foothold on different length scales, as an alternative to manually designed features. We describe an algorithm that can simultaneously learn a small set of templates and a foothold ranking function using these templates, from expert-demonstrated footholds. Using the LittleDog quadruped robot, we experimentally show that the use of terrain templates can produce complex ranking functions with higher performance than standard terrain features, and improved generalization to unseen terrain. Mrinal Kalakrishnan, Jonas Buchli, Peter Pastor, Stefan Schaal |
IROS | 1 |
| 2009 | Integrative disease classification based on cross-platform microarray dataabstractBACKGROUND: Disease classification has been an important application of microarray technology. However, most microarray-based classifiers can only handle data generated within the same study, since microarray data generated by different laboratories or with different platforms can not be compared directly due to systematic variations. This issue has severely limited the practical use of microarray-based disease classification. RESULTS: In this study, we tested the feasibility of disease classification by integrating the large amount of heterogeneous microarray datasets from the public microarray repositories. Cross-platform data compatibility is created by deriving expression log-rank ratios within datasets. One may then compare vectors of log-rank ratios across datasets. In addition, we systematically map textual annotations of datasets to concepts in Unified Medical Language System (UMLS), permitting quantitative analysis of the phenotype "distance" between datasets and automated construction of disease classes. We design a new classification approach named ManiSVM, which integrates Manifold data transformation with SVM learning to exploit the data properties. Using the leave one dataset out cross validation, ManiSVM achieved the overall accuracy of 70.7% (68.6% precision and 76.9% recall) with many disease classes achieving the accuracy higher than 80%. CONCLUSION: Our results not only demonstrated the feasibility of the integrated disease classification approach, but also showed that the classification accuracy increases with the number of homogenous training datasets. Thus, the power of the integrative approach will increase with the continuous accumulation of microarray data in public repositories. Our study shows that automated disease diagnosis can be an important and promising application of the enormous amount of costly to generate, yet freely available, public microarray data. Chun-Chi Liu, Jianjun Hu, Mrinal Kalakrishnan, Xianghong Jasmine Zhou |
BMC Bioinform. | 3 |
| 2008 | Bayesian Kernel Shaping for Learning ControlabstractIn kernel-based regression learning, optimizing each kernel individually is useful when the data density, curvature of regression surfaces (or decision boundaries) or magnitude of output noise (i.e., heteroscedasticity) varies spatially. Unfortunately, it presents a complex computational problem as the danger of overfitting is high and the individual optimization of every kernel in a learning system may be overly expensive due to the introduction of too many open learning parameters. Previous work has suggested gradient descent techniques or complex statistical hypothesis methods for local kernel shaping, typically requiring some amount of manual tuning of meta parameters. In this paper, we focus on nonparametric regression and introduce a Bayesian formulation that, with the help of variational approximations, results in an EM-like algorithm for simultaneous estimation of regression and kernel parameters. The algorithm is computationally efficient (suitable for large data sets), requires no sampling, automatically rejects outliers and has only one prior to be specified. It can be used for nonparametric regression with local polynomials or as a novel method to achieve nonstationary regression with Gaussian Processes. Our methods are particularly useful for learning control, where reliable estimation of local tangent planes is essential for adaptive controllers and reinforcement learning. We evaluate our methods on several synthetic data sets and on an actual robot which learns a task-level control law. Jo-Anne Ting, Mrinal Kalakrishnan, Sethu Vijayakumar, Stefan Schaal |
NIPS | 2 |
| 2008 | An Integrative Network Approach to Map the Transcriptome to the Phenome
Michael R. Mehan, Juan Nunez-Iglesias, Mrinal Kalakrishnan, Michael S. Waterman, Xianghong Jasmine Zhou |
RECOMB | 3 |