VLDB 2026 Research / reviewers in the wild / expert
Hongyin Zhang 0001
dblp:216/9018-1
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-9216-8403ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow ModelsabstractVision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these models on complex downstream tasks is unsatisfactory. One important reason is that these models rely solely on the post-training paradigm of imitation learning, which makes it difficult to have a deeper understanding of the distribution properties of data quality, which is exactly what Reinforcement Learning (RL) excels at. In this paper, we theoretically propose an offline RL post-training objective for VLA flow models and induce an efficient and feasible offline RL fine-tuning algorithm −− Adaptive Reinforced Flow Matching (ARFM). By introducing an adaptively adjusted scaling factor in the VLA flow model loss, we construct a principled bias-variance trade-off objective function to optimally control the impact of RL signal on flow loss. ARFM adaptively balances RL advantage preservation and flow loss gradient variance control, resulting in a more stable and efficient fine-tuning process. Extensive simulation and real-world experimental results show that ARFM exhibits excellent generalization, robustness, few-shot learning, and continuous learning performance. Hongyin Zhang 0001, Junxi Jin, Qixin Zeng, Hongchao Lu |
AAAI | 1 |
| 2025 | GEVRM: Goal-Expressive Video Generation Model For Robust Visual ManipulationabstractWith the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the inevitable external perturbations encountered during deployment. These perturbations introduce unforeseen state information to the VLA, resulting in inaccurate actions and consequently, a significant decline in generalization performance. The classic internal model control (IMC) principle demonstrates that a closed-loop system with an internal model that includes external input signals can accurately track the reference input and effectively offset the disturbance. We propose a novel closed-loop VLA method GEVRM that integrates the IMC principle to enhance the robustness of robot visual manipulation. The text-guided video generation model in GEVRM can generate highly expressive future visual planning goals. Simultaneously, we evaluate perturbations by simulating responses, which are called internal embeddings and optimized through prototype contrastive learning. This allows the model to implicitly infer and distinguish perturbations from the external environment. The proposed GEVRM achieves state-of-the-art performance on both standard and perturbed CALVIN benchmarks and shows significant improvements in realistic robot tasks. Hongyin Zhang 0001, Pengxiang Ding, Shangke Lyu |
ICLR | 1 |
| 2025 | ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement LearningabstractVision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning robust policy models from mixed-quality data. In this paper, we introduce Reinforced robot GPT (ReinboT), a novel end-to-end VLA model that integrates the RL principle of maximizing cumulative reward. ReinboT achieves a deeper understanding of the data quality distribution by predicting dense returns that capture the nuances of manipulation tasks. The dense return prediction capability enables the robot to generate more robust decision-making actions, oriented towards maximizing future benefits. Extensive experiments show that ReinboT achieves state-of-the-art performance on the CALVIN mixed-quality dataset and exhibits superior few-shot learning and out-of-distribution generalization capabilities in real-world tasks. Hongyin Zhang 0001, Zifeng Zhuang, Han Zhao 0008, Pengxiang Ding, Hongchao Lu |
ICML | 1 |
| 2025 | Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot LearningabstractThis paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUARTOnline, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference at 50 Hz in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65 %. Our project page is https://quart-online.github.io. Xinyang Tong, Pengxiang Ding, Yiguo Fan, Can Cui 0008, Han Zhao 0008, Hongyin Zhang 0001, Yonghao Dang, Siteng Huang, Shangke Lyu |
ICRA | 9 |
| 2025 | VSG: Rapid Adaptation in Autonomous Driving via Vehicle Skill GraphabstractThe ability to rapidly adapt to unseen scenarios for safe driving has long been a core challenge in autonomous driving. Rule-based methods are heavily reliant on labeled data and suffer from data biases. Many current reinforcement learning(RL)-based methods, on the other hand, are restricted to certain training scenarios, making it difficult for them to adapt to more complex and diverse traffic scenarios. In contrast, human drivers can quickly adapt to new driving situations based on their accumulated driving skills. Inspired by this, we propose the Vehicle Skill Graph (VSG), a novel framework for autonomous driving decision-making. By accumulating 1,000 diverse skills using RL and adopting knowledge graph embedding (KGE) techniques, we build a skill graph that offers a structured understanding of driving knowledge and discovers the potential relations between driving skills and new traffic scenarios. This enables rapid adaptation to new driving environments and addresses the issue of training scenario dependency in RL learning-based approaches. Experimental results demonstrate that VSG effectively captures the latent connections between traffic scenes and driving skills, facilitating efficient sequential decision-making in complex driving situations. Hongyin Zhang 0001, Xuetao Zhang 0001 |
IV | 2 |
| 2024 | A Real-World Quadrupedal Locomotion Benchmark for Offline Reinforcement LearningabstractOnline reinforcement learning (RL) methods are often data-inefficient or unreliable, making them difficult to train on real robotic hardware, especially quadruped robots. So learning robotic tasks from pre-collected data is a promising direction. Agile and stable legged locomotion remains an open issue in its general form. Analogous to the rapid progress of supervised learning in recent years, the combination of offline reinforcement learning (ORL) and realistic datasets has the potential to make breakthroughs in this challenging field. To facilitate the ORL research for real-world applications, we benchmark ten ORL algorithms in the realistic quadrupedal locomotion dataset. The dataset is collected by the classical model predictive control (MPC) method, rather than the online RL method commonly utilized by previous ORL benchmarks. Extensive experimental results show that the best-performing ORL algorithms can achieve competitive performance compared with the online RL, and even surpass it in some tasks. However, there is still a gap between the learning-based methods and classical MPC, especially in terms of stability and task response accuracy. Our benchmark can provide a fertile ground for future application-oriented ORL research. Hongyin Zhang 0001, Shuyu Yang |
IJCNN | 1 |
| 2023 | Design from Policies: Conservative Test-Time Adaptation for Offline Policy OptimizationabstractIn this work, we decouple the iterative bi-level offline RL (value estimation and policy extraction) from the offline training phase, forming a non-iterative bi-level paradigm and avoiding the iterative error propagation over two levels. Specifically, this non-iterative paradigm allows us to conduct inner-level optimization (value estimation) in training, while performing outer-level optimization (policy extraction) in testing. Naturally, such a paradigm raises three core questions that are not fully answered by prior non-iterative offline RL counterparts like reward-conditioned policy: (q1) What information should we transfer from the inner-level to the outer-level? (q2) What should we pay attention to when exploiting the transferred information for safe/confident outer-level optimization? (q3) What are the benefits of concurrently conducting outer-level optimization during testing? Motivated by model-based optimization (MBO), we propose DROP (design from policies), which fully answers the above questions. Specifically, in the inner-level, DROP decomposes offline data into multiple subsets, and learns an MBO score model (a1). To keep safe exploitation to the score model in the outer-level, we explicitly learn a behavior embedding and introduce a conservative regularization (a2). During testing, we show that DROP permits deployment adaptation, enabling an adaptive inference across states (a3). Empirically, we evaluate DROP on various tasks, showing that DROP gains comparable or better performance compared to prior methods. Hongyin Zhang 0001, Zifeng Zhuang, Yachen Kang |
NeurIPS | 2 |
| 2022 | DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement Learning
Hongyin Zhang 0001 |
ICLR | 2 |
| 2021 | Hierarchical Terrain-Aware Control for Quadrupedal Locomotion by Combining Deep Reinforcement Learning and Optimal ControlabstractQuadruped robots possess advantages on different terrains over other types of mobile robots by virtue of their flexible choices of foothold points. It is crucial to integrate terrain perception with motion planning to exploit the potential of quadruped robots. We propose a novel hierarchical terrain-aware control (HTC) framework, which leverages deep reinforcement learning (DRL) for the high-level planner and optimal control for the low-level controller. In general, traditional control methods yield better stability by using an optimization algorithm. In addition, DRL is able to offer more adaptive behavior. Our approach makes full use of the advantages of these two methods and possesses better adaptability and stability in challenging natural environments. Furthermore, the global height map of the terrain serves as visual information for the DRL, which determines the desired footholds for the robot’s leg swings and body postures. Optimal control calculates the torque of the joints on the standing legs to maintain body balance. Our method is tested on various terrains both simulated and real environments. The experimental results show that HTC can effectively enhance the adaptability of the quadruped robot by coordinating body posture. Qingfeng Yao, Jilong Wang 0008, Shuyu Yang, Hongyin Zhang 0001, Yinuo Wang 0001, Frank Zhengqing Wu |
IROS | 5 |
| 2021 | Terrain-Aware Risk-Assessment-Network-Aided Deep Reinforcement Learning for Quadrupedal Locomotion in Tough TerrainabstractWhen it comes to the control system of quadruped robots, deep reinforcement learning (DRL) is considered to be a promising solution. Despite years of development in this field, difficulties remain in guaranteeing the action stability of DRL-based quadruped robots’ locomotion, especially in tough terrain. In this paper, a terrain-aware teacher-student controller integrating a risk assessment network (RAN) is proposed to alleviate this problem. During the training phase, the RAN can evaluate the risk level of historical observation or current state and further guide the update of the policy, thereby assisting the policy in selecting better actions and avoid risky ones. Furthermore, the real-time elevation map is transmitted to the controller as visual information, so that it can perceive the terrain to produce higher performance locomotion. With the aforementioned configuration, we enable a robot to traverse various challenging terrain in simulation and bound or trot stably in the real environment. Hongyin Zhang 0001, Jilong Wang 0008, Frank Zhengqing Wu, Yinuo Wang 0001 |
IROS | 1 |
| 2020 | Preferential Experience Collection with Frequency based Intrinsic Reward for Deep Reinforcement LearningabstractExperience replay plays a key role in the success of deep reinforcement learning (DRL), where the experience collection is defined as the process of placing experience tuples into a replay buffer. Although this deterministic experience collection works well for the posterior replay, there is still a potential to improve the efficiency by using a preferential way. In this paper, we propose a preferential experience collection (PEC) method for experience replay in model-free DRL, where the reward is composed of extrinsic and intrinsic components. Since designing a good extrinsic reward is difficult especially in a complicated environment, this paper considers to establish a novel self-supervised intrinsic reward (IR) for model-free DRL. By looking into the frequency domain, the intrinsic reward is designed such that the agent can understand the behavior itself to some extent in addition to completing the task. Extensive experimental results demonstrate that our method can improve the efficiency and stability of the model-free DRL. Hongyin Zhang 0001, Qiangxing Tian, Kaichen Wei |
ICTAI | 1 |