Tonghe Zhang

dblp:357/2661 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 39% Motion planning and robot control · 31% Legged, aerial and field robots · 21%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot learning › robot policy learning
flow matching policies
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Robotics › Legged, aerial and field robots › legged robots
humanoid locomotion
0.912025
Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands · ICRA 2025
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Robotics › Motion planning and robot control
robot control
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Robotics › Motion planning and robot control
robot learning
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
entropic risk measure
0.812024
Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.812024
Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024
Machine learning › Reinforcement learning › reinforcement learning theory
Provably efficient RL
0.812024
Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024
Machine learning › Learning theory › online learning
regret bounds
0.812024
Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning
0.812024
Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024
Machine learning › Reinforcement learning
imitation learning
0.312025
Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands · ICRA 2025

Methods — techniques the papers use, named apart from their topics

wasserstein divergence · 0.9online reinforcement learning · 0.9hybrid internal model · 0.9flow matching · 0.9denoising steps · 0.9curiosity bonus · 0.9change-of-measure · 0.8beta vectors · 0.8
YearPublicationVenuePosition
2025 Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands
abstract
While it is relatively easier to train humanoid robots to mimic specific locomotion skills, it is more challenging to learn from various motions and adhere to continuously changing commands. These robots must accurately track motion instructions, seamlessly transition between a variety of movements, and master intermediate motions not present in their reference data. In this work, we propose a novel approach that integrates human-like motion transfer with precise velocity tracking by a series of improvements to classical imitation learning. To enhance generalization, we employ the Wasserstein divergence criterion (WGAN-div). Furthermore, a Hybrid Internal Model provides structured estimates of hidden states and velocity to enhance mobile stability and environment adaptability, while a curiosity bonus fosters exploration. Our comprehensive method promises highly human-like locomotion that adapts to varying velocity requirements, direct generalization to unseen motions and multitasking, as well as zero-shot transfer to the simulator and the real world across different terrains. These advancements are validated through simulations across various robot models and extensive real-world experiments.
Huaxing Huang, Wenhao Cui, Tonghe Zhang, Shengtao Li, Jinchao Han, Bangyu Qin, Tianchu Zhang, Ziyang Tang, Chenxu Hu, Shipu Zhang, Zheyuan Jiang
ICRA3
2025 ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
abstract
We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a flow policy’s deterministic path, converting the flow into a discrete-time Markov Process for exact and straightforward likelihood computation. This conversion facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune diverse flow model variants stably, including Rectified Flow [34] and Shortcut Models [18], particularly at very few or even one denoising step. We benchmark ReinFlow in representative locomotion and manipulation tasks, including long- horizon planning with visual input and sparse reward. The episode reward of Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning in challenging legged locomotion tasks while saving denoising steps and 82.63% of wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42]. The success rate of the Shortcut Model policies in state and visual manipulation tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow at four or even one denoising step, whose performance is comparable to fine-tuned DDIM policies while saving computation time for an average of 23.20% . Code, model, and checkpoints available on the project website: https://reinflow.github.io/
Tonghe Zhang, Chao Yu 0005, Sichang Su, Yu Wang 0002
NeurIPS1
2024 Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation
abstract
This work pioneers regret analysis of risk-sensitive reinforcement learning in partially observable environments with hindsight observation, addressing a gap in theoretical exploration. We introduce a novel formulation that integrates hindsight observations into a Partially Observable Markov Decision Process (POMDP) framework, where the goal is to optimize accumulated reward under the entropic risk measure. We develop the first provably efficient RL algorithm tailored for this setting. We also prove by rigorous analysis that our algorithm achieves polynomial regret $\tilde{O}\left(\frac{e^{|{\gamma}|H}-1}{|{\gamma}|H}H^2\sqrt{KHS^2OA}\right)$, which outperforms or matches existing upper bounds when the model degenerates to risk-neutral or fully observable settings. We adopt the method of change-of-measure and develop a novel analytical tool of beta vectors to streamline mathematical derivations. These techniques are of particular interest to the theoretical study of reinforcement learning.
Tonghe Zhang, Yu Chen 0074, Longbo Huang
ICML1