Sichang Su

dblp:373/5072 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Motion planning and robot control · 75% Legged, aerial and field robots · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot learning › robot policy learning
flow matching policies
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Robotics › Legged, aerial and field robots › legged robots
legged robot locomotion
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Robotics › Motion planning and robot control
robot control
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025
Robotics › Motion planning and robot control
robot learning
0.912025
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

online reinforcement learning · 0.9flow matching · 0.9denoising steps · 0.9
YearPublicationVenuePosition
2025 ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
abstract
We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a flow policy’s deterministic path, converting the flow into a discrete-time Markov Process for exact and straightforward likelihood computation. This conversion facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune diverse flow model variants stably, including Rectified Flow [34] and Shortcut Models [18], particularly at very few or even one denoising step. We benchmark ReinFlow in representative locomotion and manipulation tasks, including long- horizon planning with visual input and sparse reward. The episode reward of Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning in challenging legged locomotion tasks while saving denoising steps and 82.63% of wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42]. The success rate of the Shortcut Model policies in state and visual manipulation tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow at four or even one denoising step, whose performance is comparable to fine-tuned DDIM policies while saving computation time for an average of 23.20% . Code, model, and checkpoints available on the project website: https://reinflow.github.io/
Tonghe Zhang, Chao Yu 0005, Sichang Su, Yu Wang 0002
NeurIPS3