Tongzhou Mu

dblp:183/0943 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-4384-2526ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Reinforcement learning · 42% Robot manipulation · 30% Motion planning and robot control · 12%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
2.232025
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025
When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning? · AAAI 2025
State Alignment-based Imitation Learning · ICLR 2020
Machine learning › Reinforcement learning › reward learning
dense reward learning
1.022025
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks · ICLR 2024
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning · ICML 2025
Machine learning › Reinforcement learning
reward learning
1.022025
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks · ICLR 2024
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning · ICML 2025
Robotics › Robot manipulation
dexterous manipulation
0.912025
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025
Robotics › Motion planning and robot control › motion planning › manipulation planning
long-horizon manipulation
0.912025
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning · ICML 2025
Robotics › Robot manipulation › learning from demonstration
reinforcement learning from demonstration
0.912025
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning · ICML 2025
Machine learning › Reinforcement learning › policy optimization
residual policy learning
0.912025
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning
0.912025
When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning? · AAAI 2025
Computer vision › 3D vision › range sensing
depth sensing
0.712023
Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation · IEEE Trans. Robotics 2023
Robotics › Motion planning and robot control › robot learning
manipulation skill learning
0.712023
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills · ICLR 2023
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.712023
Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation · IEEE Trans. Robotics 2023
Robotics › Robot manipulation
visuomotor control
0.712023
On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline · ICML 2023
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent
0.522017
Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk Minimization · IJCAI 2017
Adaptive Variance Reducing for Stochastic Gradient Descent · IJCAI 2016
Mathematical optimization
stochastic optimization
0.522017
Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk Minimization · IJCAI 2017
Adaptive Variance Reducing for Stochastic Gradient Descent · IJCAI 2016
Natural language and speech › Language models and text generation
compositional generalization
0.412020
Refactoring Policy for Compositional Generalizability using Self-Supervised Object Proposals · NeurIPS 2020
Machine learning › Reinforcement learning
policy learning
0.412020
Refactoring Policy for Compositional Generalizability using Self-Supervised Object Proposals · NeurIPS 2020
Robotics › Motion planning and robot control
robot learning
0.422023
Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation · IEEE Trans. Robotics 2023
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills · ICLR 2023
Machine learning › Learning theory
empirical risk minimization
0.312017
Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk Minimization · IJCAI 2017
Machine learning › Optimization for machine learning
large-scale optimization
0.312017
Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk Minimization · IJCAI 2017
Robotics › Robot manipulation
grasping
0.312025
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model · ICLR 2025
Machine learning › Reinforcement learning › deep reinforcement learning
visual policy learning
0.312025
When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning? · AAAI 2025
Mathematical optimization › stochastic optimization
variance reduction
0.212016
Adaptive Variance Reducing for Stochastic Gradient Descent · IJCAI 2016

Methods — techniques the papers use, named apart from their topics

world model learning · 0.9state-to-visual DAgger · 0.9residual policy learning · 0.9reinforcement learning · 0.9empirical comparison · 0.9diffusion policy · 0.9demonstration augmentation · 0.9controlled exploration · 0.9behavior transformer · 0.9demonstration learning · 0.8multi-momentum acceleration · 0.3doubly stochastic gradient · 0.3adaptive variance reduction · 0.2
YearPublicationVenuePosition
2025 When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning?
abstract
Learning policies from high-dimensional visual inputs, such as pixels and point clouds, is crucial in various applications. Visual reinforcement learning is a promising approach that directly trains policies from visual observations, although it faces challenges in sample efficiency and computational costs. This study conducts an empirical comparison of State-to-Visual DAgger — a two-stage framework that initially trains a state policy before adopting online imitation to learn a visual policy — and Visual RL across a diverse set of tasks. We evaluate both methods across 16 tasks from three benchmarks, focusing on their asymptotic performance, sample efficiency, and computational costs. Surprisingly, our findings reveal that State-to-Visual DAgger does not universally outperform Visual RL but shows significant advantages in challenging tasks, offering more consistent performance. In contrast, its benefits in sample efficiency are less pronounced, although it often reduces the overall wall-clock time required for training. Based on our findings, we provide recommendations for practitioners and hope that our results contribute valuable perspectives for future research in visual policy learning.
Tongzhou Mu, Zhaoyang Li 0007, Stanislaw Wiktor Strzelecki, Xiu Yuan, Yunchao Yao, Litian Liang, Hao Su 0001
AAAI1
2025 Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
abstract
Recent advancements in robot learning have used imitation learning with large models and extensive demonstrations to develop effective policies. However, these models are often limited by the quantity quality, and diversity of demonstrations. This paper explores improving offline-trained imitation learning models through online interactions with the environment. We introduce Policy Decorator, which uses a model-agnostic residual policy to refine large imitation learning models during online interactions. By implementing controlled exploration strategies, Policy Decorator enables stable, sample-efficient online learning. Our evaluation spans eight tasks across two benchmarks—ManiSkill and Adroit—and involves two state-of-the-art imitation learning models (Behavior Transformer and Diffusion Policy). The results show Policy Decorator effectively improves the offline-trained policies and preserves the smooth motion of imitation learning models, avoiding the erratic behaviors of pure RL policies. See our [project page](https://policydecorator.github.io/) for videos.
Xiu Yuan, Tongzhou Mu, Stone Tao, Yunhao Fang, Mengke Zhang, Hao Su 0001
ICLR2
2025 Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
abstract
Long-horizon tasks in robotic manipulation present significant challenges in reinforcement learning (RL) due to the difficulty of designing dense reward functions and effectively exploring the expansive state-action space. However, despite a lack of dense rewards, these tasks often have a multi-stage structure, which can be leveraged to decompose the overall objective into manageable sub-goals. In this work, we propose DEMO$^3$, a framework that exploits this structure for efficient learning from visual inputs. Specifically, our approach incorporates multi-stage dense reward learning, a bi-phasic training scheme, and world model learning into a carefully designed demonstration-augmented RL framework that strongly mitigates the challenge of exploration in long-horizon tasks. Our evaluations demonstrate that our method improves data-efficiency by an average of 40% and by 70% on particularly difficult tasks compared to state-of-the-art approaches. We validate this across 16 sparse-reward tasks spanning four domains, including challenging humanoid visual control tasks using as few as five demonstrations.
Adrià López Escoriza, Nicklas Hansen 0001, Stone Tao, Tongzhou Mu, Hao Su 0001
ICML4
2024 DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
abstract
The success of many RL techniques heavily relies on human-engineered dense rewards, which typically demands substantial domain expertise and extensive trial and error. In our work, we propose **DrS** (**D**ense **r**eward learning from **S**tages), a novel approach for learning *reusable* dense rewards for multi-stage tasks in a data-driven manner. By leveraging the stage structures of the task, DrS learns a high-quality dense reward from sparse rewards and demonstrations if given. The learned rewards can be *reused* in unseen tasks, thus reducing the human effort for reward engineering. Extensive experiments on three physical robot manipulation task families with 1000+ task variants demonstrate that our learned rewards can be reused in unseen tasks, resulting in improved performance and sample efficiency of RL algorithms. The learned rewards even achieve comparable performance to human-engineered rewards on some tasks. See our [project page](https://sites.google.com/view/iclr24drs) for more details.
Tongzhou Mu, Minghua Liu, Hao Su 0001
ICLR1
2023 ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Jiayuan Gu, Fanbo Xiang, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen 0019, Hao Su 0001
ICLR6
2023 On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline
abstract
In this paper, we examine the effectiveness of pre-training for visuo-motor control tasks. We revisit a simple Learning-from-Scratch (LfS) baseline that incorporates data augmentation and a shallow ConvNet, and find that this baseline is surprisingly competitive with recent approaches (PVR, MVP, R3M) that leverage frozen visual representations trained on large-scale vision datasets – across a variety of algorithms, task domains, and metrics in simulation and on a real robot. Our results demonstrate that these methods are hindered by a significant domain gap between the pre-training datasets and current benchmarks for visuo-motor control, which is alleviated by finetuning. Based on our findings, we provide recommendations for future research in pre-training for control and hope that our simple yet strong baseline will aid in accurately benchmarking progress in this area. Code: https://github.com/gemcollector/learning-from-scratch.
Nicklas Hansen 0001, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu, Aravind Rajeswaran, Hao Su 0001, Huazhe Xu, Xiaolong Wang 0004
ICML4
2023 Abstract-to-Executable Trajectory Translation for One-Shot Task Generalization
abstract
Training long-horizon robotic policies in complex physical environments is essential for many applications, such as robotic manipulation. However, learning a policy that can generalize to unseen tasks is challenging. In this work, we propose to achieve one-shot task generalization by decoupling plan generation and plan execution. Specifically, our method solves complex long-horizon tasks in three steps: build a paired abstract environment by simplifying geometry and physics, generate abstract trajectories, and solve the original task by an abstract-to-executable trajectory translator. In the abstract environment, complex dynamics such as physical manipulation are removed, making abstract trajectories easier to generate. However, this introduces a large domain gap between abstract trajectories and the actual executed trajectories as abstract trajectories lack low-level details and are not aligned frame-to-frame with the executed trajectory. In a manner reminiscent of language translation, our approach leverages a seq-to-seq model to overcome the large domain gap between the abstract and executable trajectories, enabling the low-level policy to follow the abstract trajectory. Experimental results on various unseen long-horizon tasks with different robot embodiments demonstrate the practicability of our methods to achieve one-shot task generalization.
Stone Tao, Tongzhou Mu, Zhiao Huang, Yuzhe Qin, Hao Su 0001
ICML3
2023 Close the Optical Sensing Domain Gap by Physics-Grounded Active Stereo Sensor Simulation
abstract
In this article, we focus on the simulation of active stereovision depth sensors, which are popular in both academic and industry communities. Inspired by the underlying mechanism of the sensors, we designed a fully physics-grounded simulation pipeline that includes material acquisition, ray-tracing-based infrared (IR) image rendering, IR noise simulation, and depth estimation. The pipeline is able to generate depth maps with material-dependent error patterns similar to a real depth sensor in real time. We conduct real experiments to show that perception algorithms and reinforcement learning policies trained in our simulation platform could transfer well to the real-world test cases without any fine-tuning. Furthermore, due to the high degree of realism of this simulation, our depth sensor simulator can be used as a convenient testbed to evaluate the algorithm performance in the real world, which will largely reduce the human effort in developing robotic algorithms. The entire pipeline has been integrated into the SAPIEN simulator and is open-sourced to promote the research of vision and robotics communities.
Xiaoshuai Zhang, Rui Chen 0019, Ang Li 0010, Fanbo Xiang, Yuzhe Qin, Jiayuan Gu, Zhan Ling, Minghua Liu, Peiyu Zeng, Songfang Han, Zhiao Huang, Tongzhou Mu, Jing Xu 0011, Hao Su 0001
IEEE Trans. Robotics12
2020 State Alignment-based Imitation Learning
Fangchen Liu, Zhan Ling, Tongzhou Mu, Hao Su 0001
ICLR3
2020 Refactoring Policy for Compositional Generalizability using Self-Supervised Object Proposals
abstract
We study how to learn a policy with compositional generalizability. We propose a two-stage framework, which refactorizes a high-reward teacher policy into a generalizable student policy with strong inductive bias. Particularly, we implement an object-centric GNN-based student policy, whose input objects are learned from images through self-supervised learning. Empirically, we evaluate our approach on four difficult tasks that require compositional generalizability, and achieve superior performance compared to baselines.
Tongzhou Mu, Jiayuan Gu, Zhiwei Jia, Hao Tang 0008, Hao Su 0001
NeurIPS1
2017 Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk Minimization
abstract
Nowadays, algorithms with fast convergence, small memory footprints, and low per-iteration complexity are particularly favorable for artificial intelligence applications. In this paper, we propose a doubly stochastic algorithm with a novel accelerating multi-momentum technique to solve large scale empirical risk minimization problem for learning tasks. While enjoying a provably superior convergence rate, in each iteration, such algorithm only accesses a mini batch of samples and meanwhile updates a small block of variable coordinates, which substantially reduces the amount of memory reference when both the massive sample size and ultra-high dimensionality are involved. Specifically, to obtain an ε-accurate solution, our algorithm requires only O(log(1/ε)/sqrt(ε)) overall computation for the general convex case and O((n+sqrt{nκ})log(1/ε)) for the strongly convex case. Empirical studies on huge scale datasets are conducted to illustrate the efficiency of our method in practice.
Zebang Shen, Hui Qian 0001, Tongzhou Mu, Chao Zhang 0029
IJCAI3
2016 Adaptive Variance Reducing for Stochastic Gradient Descent
Zebang Shen, Hui Qian 0001, Tongzhou Mu
IJCAI4