VLDB 2026 Research / reviewers in the wild / expert
Tonghe Zhang
dblp:357/2661
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 39% Motion planning and robot control · 31% Legged, aerial and field robots · 21% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Motion planning and robot control › robot learning › robot policy learning
flow matching policies |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Robotics › Legged, aerial and field robots › legged robots
humanoid locomotion |
0.9 | 1 | 2025 | Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands · ICRA 2025 |
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Robotics › Motion planning and robot control
robot control |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning › safe reinforcement learning › risk-sensitive reinforcement learning
entropic risk measure |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning
partially observable reinforcement learning |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning theory
Provably efficient RL |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Learning theory › online learning
regret bounds |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning › safe reinforcement learning
risk-sensitive reinforcement learning |
0.8 | 1 | 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight Observation · ICML 2024 |
Machine learning › Reinforcement learning
imitation learning |
0.3 | 1 | 2025 | Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing Commands · ICRA 2025 |
Methods — techniques the papers use, named apart from their topics
wasserstein divergence · 0.9online reinforcement learning · 0.9hybrid internal model · 0.9flow matching · 0.9denoising steps · 0.9curiosity bonus · 0.9change-of-measure · 0.8beta vectors · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Think on Your Feet: Seamless Transition Between Human-Like Locomotion in Response to Changing CommandsabstractWhile it is relatively easier to train humanoid robots to mimic specific locomotion skills, it is more challenging to learn from various motions and adhere to continuously changing commands. These robots must accurately track motion instructions, seamlessly transition between a variety of movements, and master intermediate motions not present in their reference data. In this work, we propose a novel approach that integrates human-like motion transfer with precise velocity tracking by a series of improvements to classical imitation learning. To enhance generalization, we employ the Wasserstein divergence criterion (WGAN-div). Furthermore, a Hybrid Internal Model provides structured estimates of hidden states and velocity to enhance mobile stability and environment adaptability, while a curiosity bonus fosters exploration. Our comprehensive method promises highly human-like locomotion that adapts to varying velocity requirements, direct generalization to unseen motions and multitasking, as well as zero-shot transfer to the simulator and the real world across different terrains. These advancements are validated through simulations across various robot models and extensive real-world experiments. Huaxing Huang, Wenhao Cui, Tonghe Zhang, Shengtao Li, Jinchao Han, Bangyu Qin, Tianchu Zhang, Ziyang Tang, Chenxu Hu, Shipu Zhang, Zheyuan Jiang |
ICRA | 3 |
| 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement LearningabstractWe propose ReinFlow, a simple yet effective online reinforcement learning (RL)
framework that fine-tunes a family of flow matching policies for continuous robotic
control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a
flow policy’s deterministic path, converting the flow into a discrete-time Markov
Process for exact and straightforward likelihood computation. This conversion
facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune
diverse flow model variants stably, including Rectified Flow [34] and Shortcut
Models [18], particularly at very few or even one denoising step. We benchmark
ReinFlow in representative locomotion and manipulation tasks, including long-
horizon planning with visual input and sparse reward. The episode reward of
Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning
in challenging legged locomotion tasks while saving denoising steps and 82.63% of
wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42].
The success rate of the Shortcut Model policies in state and visual manipulation
tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow
at four or even one denoising step, whose performance is comparable to fine-tuned
DDIM policies while saving computation time for an average of 23.20% . Code,
model, and checkpoints available on the project website: https://reinflow.github.io/ Tonghe Zhang, Chao Yu 0005, Sichang Su, Yu Wang 0002 |
NeurIPS | 1 |
| 2024 | Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight ObservationabstractThis work pioneers regret analysis of risk-sensitive reinforcement learning in partially observable environments with hindsight observation, addressing a gap in theoretical exploration. We introduce a novel formulation that integrates hindsight observations into a Partially Observable Markov Decision Process (POMDP) framework, where the goal is to optimize accumulated reward under the entropic risk measure. We develop the first provably efficient RL algorithm tailored for this setting. We also prove by rigorous analysis that our algorithm achieves polynomial regret $\tilde{O}\left(\frac{e^{|{\gamma}|H}-1}{|{\gamma}|H}H^2\sqrt{KHS^2OA}\right)$, which outperforms or matches existing upper bounds when the model degenerates to risk-neutral or fully observable settings. We adopt the method of change-of-measure and develop a novel analytical tool of beta vectors to streamline mathematical derivations. These techniques are of particular interest to the theoretical study of reinforcement learning. Tonghe Zhang, Yu Chen 0074, Longbo Huang |
ICML | 1 |