VLDB 2026 Research / reviewers in the wild / expert
Sichang Su
dblp:373/5072
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Motion planning and robot control · 75% Legged, aerial and field robots · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Motion planning and robot control › robot learning › robot policy learning
flow matching policies |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Robotics › Motion planning and robot control
robot control |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
online reinforcement learning · 0.9flow matching · 0.9denoising steps · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement LearningabstractWe propose ReinFlow, a simple yet effective online reinforcement learning (RL)
framework that fine-tunes a family of flow matching policies for continuous robotic
control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a
flow policy’s deterministic path, converting the flow into a discrete-time Markov
Process for exact and straightforward likelihood computation. This conversion
facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune
diverse flow model variants stably, including Rectified Flow [34] and Shortcut
Models [18], particularly at very few or even one denoising step. We benchmark
ReinFlow in representative locomotion and manipulation tasks, including long-
horizon planning with visual input and sparse reward. The episode reward of
Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning
in challenging legged locomotion tasks while saving denoising steps and 82.63% of
wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42].
The success rate of the Shortcut Model policies in state and visual manipulation
tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow
at four or even one denoising step, whose performance is comparable to fine-tuned
DDIM policies while saving computation time for an average of 23.20% . Code,
model, and checkpoints available on the project website: https://reinflow.github.io/ Tonghe Zhang, Chao Yu 0005, Sichang Su, Yu Wang 0002 |
NeurIPS | 3 |