EDBT 2026 Demo / reviewers in the wild / expert
Tao Chen 0046
dblp:69/510-46
· DBLP profile ↗
13ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0003-0190-3853ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vegetable Peeling: A Case Study in Constrained Dexterous ManipulationabstractRecent studies have made significant progress in addressing dexterous manipulation problems, particularly in inhand object reorientation. However, there are few existing works that explore the potential utilization of developed dexterous manipulation controllers for downstream tasks. In this study, we focus on constrained dexterous manipulation for food peeling. Food peeling presents various constraints on the reorientation controller, such as the requirement for the hand to securely hold the object after reorientation for peeling. We propose a simple system for learning a reorientation controller that facilitates the subsequent peeling task. Videos are available at: https://taochenshh.github.io/projects/veg-peeling. Tao Chen 0046, Eric Cousineau, Naveen Kuppuswamy, Pulkit Agrawal 0001 |
ICRA | 1 |
| 2024 | Lifelong Robot Learning with Human Assisted Language PlannersabstractLarge Language Models (LLMs) have been shown to act like planners that can decompose high-level instructions into a sequence of executable instructions. However, current LLM-based planners are only able to operate with a fixed set of skills. We overcome this critical limitation and present a method for using LLM-based planners to query new skills and teach robots these skills in a data and time-efficient manner for rigid object manipulation. Our system can re-use newly acquired skills for future tasks, demonstrating the potential of open world and lifelong learning. We evaluate the proposed framework on multiple tasks in simulation and the real world. Videos are available at: https://sites.google.com/mit.edu/halp-robot-learning Meenal Parakh, Alisha Fong, Anthony Simeonov, Tao Chen 0046, Abhishek Gupta 0004, Pulkit Agrawal 0001 |
ICRA | 4 |
| 2024 | Learning Multimodal Behaviors from Scratch with Diffusion Policy GradientabstractDeep reinforcement learning (RL) algorithms typically parameterize the policy as a deep network that outputs either a deterministic action or a stochastic one modeled as a Gaussian distribution, hence restricting learning to a single behavioral mode. Meanwhile, diffusion models emerged as a powerful framework for multimodal learning. However, the use of diffusion policies in online RL is hindered by the intractability of policy likelihood approximation, as well as the greedy objective of RL methods that can easily skew the policy to a single mode. This paper presents Deep Diffusion Policy Gradient (DDiffPG), a novel actor-critic algorithm that learns from scratch multimodal policies parameterized as diffusion models while discovering and maintaining versatile behaviors. DDiffPG explores and discovers multiple modes through off-the-shelf unsupervised clustering combined with novelty-based intrinsic motivation. DDiffPG forms a multimodal training batch and utilizes mode-specific Q-learning to mitigate the inherent greediness of the RL objective, ensuring the improvement of the diffusion policy across all modes. Our approach further allows the policy to be conditioned on mode-specific embeddings to explicitly control the learned modes. Empirical studies validate DDiffPG's capability to master multimodal behaviors in complex, high-dimensional continuous control tasks with sparse rewards, also showcasing proof-of-concept dynamic online replanning when navigating mazes with unseen obstacles. Our project page is available at https://supersglzc.github.io/projects/ddiffpg/. Steven Li, Rickmer Krohn, Tao Chen 0046, Anurag Ajay, Pulkit Agrawal 0001, Georgia Chalvatzaki |
NeurIPS | 3 |
| 2023 | DexDeform: Dexterous Deformable Object Manipulation with Human Demonstrations and Differentiable Physics
Zhiao Huang, Tao Chen 0046, Tao Du 0001, Hao Su 0001, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 3 |
| 2023 | Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel SimulationabstractReinforcement learning is time-consuming for complex tasks due to the need for large amounts of training data. Recent advances in GPU-based simulation, such as Isaac Gym, have sped up data collection thousands of times on a commodity GPU. Most prior works have used on-policy methods like PPO due to their simplicity and easy-to-scale nature. Off-policy methods are more sample-efficient, but challenging to scale, resulting in a longer wall-clock training time. This paper presents a novel Parallel Q-Learning (PQL) scheme that outperforms PPO in terms of wall-clock time and maintains superior sample efficiency. The driving force lies in the parallelization of data collection, policy function learning, and value function learning. Different from prior works on distributed off-policy learning, such as Apex, our scheme is designed specifically for massively parallel GPU-based simulation and optimized to work on a single workstation. In experiments, we demonstrate the capability of scaling up Q-learning methods to tens of thousands of parallel environments and investigate important factors that can affect learning speed, including the number of parallel environments, exploration strategies, batch size, GPU models, etc. The code is available at https://github.com/Improbable-AI/pql. Zechu Li, Tao Chen 0046, Zhang-Wei Hong, Anurag Ajay, Pulkit Agrawal 0001 |
ICML | 2 |
| 2023 | TactoFind: A Tactile Only System for Object RetrievalabstractWe study the problem of object retrieval in scenarios where visual sensing is absent, object shapes are unknown beforehand and objects can move freely, like grabbing objects out of a drawer. Successful solutions require localizing free objects, identifying specific object instances, and then grasping the identified objects, only using touch feedback. Unlike vision, where cameras can observe the entire scene, touch sensors are local and only observe parts of the scene that are in contact with the manipulator. Moreover, information gathering via touch sensors necessitates applying forces on the touched surface which may disturb the scene itself. Reasoning with touch, therefore, requires careful exploration and integration of information over time - a challenge we tackle. We present a system capable of using sparse tactile feedback from fingertip touch sensors on a dexterous hand to localize, identify and grasp novel objects without any visual feedback. Videos are available at https://sites.google.com/view/tactofind. Sameer Pai, Tao Chen 0046, Megha Tippur, Edward H. Adelson, Abhishek Gupta 0004, Pulkit Agrawal 0001 |
ICRA | 2 |
| 2023 | Breadcrumbs to the Goal: Supervised Goal Selection from Human-in-the-Loop Feedback
Marcel Torne, Max Balsells, Samedh Desai, Tao Chen 0046, Pulkit Agrawal 0001, Abhishek Gupta 0004 |
NeurIPS | 5 |
| 2022 | Topological Experience Replay
Zhang-Wei Hong, Tao Chen 0046, Yen-Chen Lin, Joni Pajarinen, Pulkit Agrawal 0001 |
ICLR | 2 |
| 2022 | HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object HandoversabstractWe introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics. We analyze the performance of a set of baselines and show a correlation with a real-world evaluation.11Code is open sourced at https://handover-sim.github.io. Yu-Wei Chao, Chris Paxton 0001, Yu Xiang 0001, Wei Yang 0019, Balakumar Sundaralingam, Tao Chen 0046, Adithyavairavan Murali, Maya Cakmak, Dieter Fox |
ICRA | 6 |
| 2022 | Pre-Trained Language Models for Interactive Decision-MakingabstractLanguage model (LM) pre-training is useful in many language processing tasks. But can pre-trained LMs be further leveraged for more general machine learning problems? We propose an approach for using LMs to scaffold learning and generalization in general sequential decision-making problems. In this approach, goals and observations are represented as a sequence of embeddings, and a policy network initialized with a pre-trained LM predicts the next action. We demonstrate that this framework enables effective combinatorial generalization across different environments and supervisory modalities. We begin by assuming access to a set of expert demonstrations, and show that initializing policies with LMs and fine-tuning them via behavior cloning improves task completion rates by 43.6% in the VirtualHome environment. Next, we integrate an active data gathering procedure in which agents iteratively interact with the environment, relabel past "failed" experiences with new goals, and update their policies in a self-supervised loop. Active data gathering further improves combinatorial generalization, outperforming the best baseline by 25.1%. Finally, we explain these results by investigating three possible factors underlying the effectiveness of the LM-based policy. We find that sequential input representations (vs. fixed-dimensional feature vectors) and LM-based weight initialization are both important for generalization. Surprisingly, however, the format of the policy inputs encoding (e.g. as a natural language string vs. an arbitrary sequential encoding) has little influence. Together, these results suggest that language modeling induces representations that are useful for modeling not just language, but also goals and plans; these representations can aid learning and generalization even outside of language processing. Shuang Li 0013, Xavier Puig, Chris Paxton 0001, Yilun Du, Clinton Wang, Linxi Fan, Tao Chen 0046, De-An Huang, Ekin Akyürek, Anima Anandkumar, Jacob Andreas, Igor Mordatch, Antonio Torralba 0001, Yuke Zhu |
NeurIPS | 7 |
| 2021 | Residual Model Learning for Microrobot ControlabstractA majority of microrobots are constructed using compliant materials that are difficult to model analytically, limiting the utility of traditional model-based controllers. Challenges in data collection on microrobots and large errors between simulated models and real robots make current model-based learning and sim-to-real transfer methods difficult to apply. We propose a novel framework residual model learning (RML) that leverages approximate models to substantially reduce the sample complexity associated with learning an accurate robot model. We show that using RML, we can learn a model of the Harvard Ambulatory MicroRobot (HAMR) using just 12 seconds of passively collected interaction data. The learned model is accurate enough to be leveraged as "proxy-simulator" for learning walking and turning behaviors using model-free reinforcement learning algorithms. RML provides a general framework for learning from extremely small amounts of interaction data, and our experiments with HAMR clearly demonstrate that RML substantially outperforms existing techniques. Joshua Gruenstein, Tao Chen 0046, Neel Doshi, Pulkit Agrawal 0001 |
ICRA | 2 |
| 2019 | Learning Exploration Policies for Navigation
Tao Chen 0046, Saurabh Gupta 0001, Abhinav Gupta 0001 |
ICLR (Poster) | 1 |
| 2018 | Hardware Conditioned Policies for Multi-Robot Transfer LearningabstractDeep reinforcement learning could be used to learn dexterous robotic policies but it is challenging to transfer them to new robots with vastly different hardware properties. It is also prohibitively expensive to learn a new policy from scratch for each robot hardware due to the high sample complexity of modern state-of-the-art algorithms. We propose a novel approach called Hardware Conditioned Policies where we train a universal policy conditioned on a vector representation of robot hardware. We considered robots in simulation with varied dynamics, kinematic structure, kinematic lengths and degrees-of-freedom. First, we use the kinematic structure directly as the hardware encoding and show great zero-shot transfer to completely novel robots not seen during training. For robots with lower zero-shot success rate, we also demonstrate that fine-tuning the policy network is significantly more sample-efficient than training a model from scratch. In tasks where knowing the agent dynamics is important for success, we learn an embedding for robot hardware and show that policies conditioned on the encoding of hardware tend to generalize and transfer well. Videos of experiments are available at: https://sites.google.com/view/robot-transfer-hcp. Tao Chen 0046, Adithyavairavan Murali, Abhinav Gupta 0001 |
NeurIPS | 1 |