Chuning Zhu

dblp:295/9468 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 50% Motion planning and robot control · 18% Transfer learning and domain adaptation · 16%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
2.352024
Distributional Successor Features Enable Zero-Shot Policy Optimization · NeurIPS 2024
RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior Predictability · NeurIPS 2023
Model-Based Reinforcement Learning via Latent-Space Collocation · ICML 2021
Robotics › Motion planning and robot control › robot control
model-based control
0.812024
ASID: Active Exploration for System Identification in Robotic Manipulation · ICLR 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning · ICLR 2024
Machine learning › Reinforcement learning › offline reinforcement learning
return-conditioned supervised learning
0.812024
Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning · ICLR 2024
Robotics › Motion planning and robot control
robot control
0.812024
ASID: Active Exploration for System Identification in Robotic Manipulation · ICLR 2024
Robotics › Motion planning and robot control
robot learning
0.812024
ASID: Active Exploration for System Identification in Robotic Manipulation · ICLR 2024
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.812024
ASID: Active Exploration for System Identification in Robotic Manipulation · ICLR 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
successor features
0.812024
Distributional Successor Features Enable Zero-Shot Policy Optimization · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
trajectory stitching
0.812024
Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning · ICLR 2024
Machine learning › Transfer learning and domain adaptation
cross-task transfer
0.712023
Self-Supervised Reinforcement Learning that Transfers using Random Features · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning
0.712023
RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior Predictability · NeurIPS 2023
Machine learning › Reinforcement learning
model-free reinforcement learning
0.712023
Self-Supervised Reinforcement Learning that Transfers using Random Features · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.712023
Self-Supervised Reinforcement Learning that Transfers using Random Features · NeurIPS 2023
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
latent space planning
0.512021
Model-Based Reinforcement Learning via Latent-Space Collocation · ICML 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-horizon planning
0.512021
Model-Based Reinforcement Learning via Latent-Space Collocation · ICML 2021
Machine learning › Reinforcement learning
exploration
0.212024
ASID: Active Exploration for System Identification in Robotic Manipulation · ICLR 2024
Robotics › Motion planning and robot control › robot control
model predictive control
0.212023
Self-Supervised Reinforcement Learning that Transfers using Random Features · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
0.212023
RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior Predictability · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.212023
RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior Predictability · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

system identification · 0.8simulation · 0.8return-conditioned supervised learning · 0.8model-based forward sampling · 0.8dynamic programming · 0.8diffusion model · 0.8active exploration · 0.8self-supervised pretraining · 0.7random features · 0.7model predictive control · 0.7
YearPublicationVenuePosition
2024 ASID: Active Exploration for System Identification in Robotic Manipulation
abstract
Model-free control strategies such as reinforcement learning have shown the ability to learn control strategies without requiring an accurate model or simulator of the world. While this is appealing due to the lack of modeling requirements, such methods can be sample inefficient, making them impractical in many real-world domains. On the other hand, model-based control techniques leveraging accurate simulators can circumvent these challenges and use a large amount of cheap simulation data to learn controllers that can effectively transfer to the real world. The challenge with such model-based techniques is the requirement for an extremely accurate simulation, requiring both the specification of appropriate simulation assets and physical parameters. This requires considerable human effort to design for every environment being considered. In this work, we propose a learning system that can leverage a small amount of real-world data to autonomously refine a simulation model and then plan an accurate control strategy that can be deployed in the real world. Our approach critically relies on utilizing an initial (possibly inaccurate) simulator to design effective exploration policies that, when deployed in the real world, collect high-quality data. We demonstrate the efficacy of this paradigm in identifying articulation, mass, and other physical parameters in several challenging robotic manipulation tasks, and illustrate that only a small amount of real-world data can allow for effective sim-to-real transfer.
Marius Memmel, Andrew J. Wagenmaker, Chuning Zhu, Dieter Fox, Abhishek Gupta 0004
ICLR3
2024 Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning
abstract
Off-policy dynamic programming (DP) techniques such as $Q$-learning have proven to be important in sequential decision-making problems. In the presence of function approximation, however, these techniques often diverge due to the absence of Bellman completeness in the function classes considered, a crucial condition for the success of DP-based methods. In this paper, we show how off-policy learning techniques based on return-conditioned supervised learning (RCSL) are able to circumvent these challenges of Bellman completeness, converging under significantly more relaxed assumptions inherited from supervised learning. We prove there exists a natural environment in which if one uses two-layer multilayer perceptron as the function approximator, the layer width needs to grow *linearly* with the state space size to satisfy Bellman completeness while a constant layer width is enough for RCSL. These findings take a step towards explaining the superior empirical performance of RCSL methods compared to DP-based methods in environments with near-optimal datasets. Furthermore, in order to learn from sub-optimal datasets, we propose a simple framework called MBRCSL, granting RCSL methods the ability of dynamic programming to stitch together segments from distinct trajectories. MBRCSL leverages learned dynamics models and forward sampling to accomplish trajectory stitching while avoiding the need for Bellman completeness that plagues all dynamic programming algorithms. We propose both theoretical analysis and experimental evaluation to back these claims, outperforming state-of-the-art model-free and model-based offline RL algorithms across several simulated robotics problems.
Zhaoyi Zhou, Chuning Zhu, Runlong Zhou, Qiwen Cui, Abhishek Gupta 0004, Simon S. Du
ICLR2
2024 Distributional Successor Features Enable Zero-Shot Policy Optimization
abstract
Intelligent agents must be generalists, capable of quickly adapting to various tasks. In reinforcement learning (RL), model-based RL learns a dynamics model of the world, in principle enabling transfer to arbitrary reward functions through planning. However, autoregressive model rollouts suffer from compounding error, making model-based RL ineffective for long-horizon problems. Successor features offer an alternative by modeling a policy's long-term state occupancy, reducing policy evaluation under new rewards to linear regression. Yet, policy optimization with successor features can be challenging. This work proposes a novel class of models, i.e., Distributional Successor Features for Zero-Shot Policy Optimization (DiSPOs), that learn a distribution of successor features of a stationary dataset's behavior policy, along with a policy that acts to realize different successor features within the dataset. By directly modeling long-term outcomes in the dataset, DiSPOs avoid compounding error while enabling a simple scheme for zero-shot policy optimization across reward functions. We present a practical instantiation of DiSPOs using diffusion models and show their efficacy as a new class of transferable models, both theoretically and empirically across various simulated robotics problems. Videos and code are available at https://weirdlabuw.github.io/dispo/.
Chuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du, Abhishek Gupta 0004
NeurIPS1
2023 Self-Supervised Reinforcement Learning that Transfers using Random Features
abstract
Model-free reinforcement learning algorithms have exhibited great potential in solving single-task sequential decision-making problems with high-dimensional observations and long horizons, but are known to be hard to generalize across tasks. Model-based RL, on the other hand, learns task-agnostic models of the world that naturally enables transfer across different reward functions, but struggles to scale to complex environments due to the compounding error. To get the best of both worlds, we propose a self-supervised reinforcement learning method that enables the transfer of behaviors across tasks with different rewards, while circumventing the challenges of model-based RL. In particular, we show self-supervised pre-training of model-free reinforcement learning with a number of random features as rewards allows implicit modeling of long-horizon environment dynamics. Then, planning techniques like model-predictive control using these implicit models enable fast adaptation to problems with new reward functions. Our method is self-supervised in that it can be trained on offline datasets without reward labels, but can then be quickly deployed on new tasks. We validate that our proposed method enables transfer across tasks on a variety of manipulation and locomotion domains in simulation, opening the door to generalist decision-making agents.
Boyuan Chen 0003, Chuning Zhu, Pulkit Agrawal 0001, Kaiqing Zhang, Abhishek Gupta 0004
NeurIPS2
2023 RePo: Resilient Model-Based Reinforcement Learning by Regularizing Posterior Predictability
abstract
Visual model-based RL methods typically encode image observations into low-dimensional representations in a manner that does not eliminate redundant information. This leaves them susceptible to spurious variations -- changes in task-irrelevant components such as background distractors or lighting conditions. In this paper, we propose a visual model-based RL method that learns a latent representation resilient to such spurious variations. Our training objective encourages the representation to be maximally predictive of dynamics and reward, while constraining the information flow from the observation to the latent representation. We demonstrate that this objective significantly bolsters the resilience of visual model-based RL methods to visual distractors, allowing them to operate in dynamic environments. We then show that while the learned encoder is able to operate in dynamic environments, it is not invariant under significant distribution shift. To address this, we propose a simple reward-free alignment procedure that enables test time adaptation of the encoder. This allows for quick adaptation to widely differing environments without having to relearn the dynamics and policy. Our effort is a step towards making model-based RL a practical and useful tool for dynamic, diverse domains and we show its effectiveness in simulation tasks with significant spurious variations.
Chuning Zhu, Max Simchowitz, Siri Gadipudi, Abhishek Gupta 0004
NeurIPS1
2021 Model-Based Reinforcement Learning via Latent-Space Collocation
abstract
The ability to plan into the future while utilizing only raw high-dimensional observations, such as images, can provide autonomous agents with broad and general capabilities. However, realistic tasks require performing temporally extended reasoning, and cannot be solved with only myopic, short-sighted planning. Recent work in model-based reinforcement learning (RL) has shown impressive results on tasks that require only short-horizon reasoning. In this work, we study how the long-horizon planning abilities can be improved with an algorithm that optimizes over sequences of states, rather than actions, which allows better credit assignment. To achieve this, we draw on the idea of collocation and adapt it to the image-based setting by leveraging probabilistic latent variable models, resulting in an algorithm that optimizes trajectories over latent variables. Our latent collocation method (LatCo) provides a general and effective visual planning approach, and significantly outperforms prior model-based approaches on challenging visual control tasks with sparse rewards and long-term goals. See the videos on the supplementary website \url{https://sites.google.com/view/latco-mbrl/.}
Oleh Rybkin, Chuning Zhu, Anusha Nagabandi, Kostas Daniilidis, Igor Mordatch, Sergey Levine
ICML2