Junxi Jin

dblp:393/1782 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 39% Robot manipulation · 30% Motion planning and robot control · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot learning › robot policy learning
flow matching policies
1.012026
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026
Machine learning › Reinforcement learning
offline reinforcement learning
1.012026
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.012026
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026
Machine learning › Reinforcement learning
imitation learning
0.312026
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026

Methods — techniques the papers use, named apart from their topics

offline reinforcement learning · 1.0flow matching · 1.0bias-variance trade-off · 1.0
YearPublicationVenuePosition
2026 Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
abstract
Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these models on complex downstream tasks is unsatisfactory. One important reason is that these models rely solely on the post-training paradigm of imitation learning, which makes it difficult to have a deeper understanding of the distribution properties of data quality, which is exactly what Reinforcement Learning (RL) excels at. In this paper, we theoretically propose an offline RL post-training objective for VLA flow models and induce an efficient and feasible offline RL fine-tuning algorithm −− Adaptive Reinforced Flow Matching (ARFM). By introducing an adaptively adjusted scaling factor in the VLA flow model loss, we construct a principled bias-variance trade-off objective function to optimally control the impact of RL signal on flow loss. ARFM adaptively balances RL advantage preservation and flow loss gradient variance control, resulting in a more stable and efficient fine-tuning process. Extensive simulation and real-world experimental results show that ARFM exhibits excellent generalization, robustness, few-shot learning, and continuous learning performance.
Hongyin Zhang 0001, Junxi Jin, Qixin Zeng, Hongchao Lu
AAAI3