EDBT 2026 Demo / reviewers in the wild / expert
Han Zhao 0008
dblp:03/3520-8
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-3948-0649ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot PerceiverabstractRecent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispersed. To guide the visual attention grounding on the correct target, we propose ReconVLA, a reconstructive VLA model with an implicit grounding paradigm. Conditioned on the model's visual outputs, a diffusion transformer aims to reconstruct the gaze region of the image, which corresponds to the target manipulated objects. This process prompts the VLA model to learn fine-grained representations and accurately allocate visual attention, thus effectively leveraging task-specific visual information and conducting precise manipulation. Moreover, we curate a large-scale pretraining dataset comprising over 100k trajectories and 2 million data samples from open-source robotic datasets, further boosting the model’s generalization in visual reconstruction. Extensive experiments in simulation and the real world demonstrate the superiority of our implicit grounding method, showcasing its capabilities of precise manipulation and generalization. Wenxuan Song, Han Zhao 0008, Pengxiang Ding, Haodong Yan, Haoang Li |
AAAI | 3 |
| 2026 | VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelabstractVision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how to effectively bridge vision-language (VL) representations to action (A). We introduce VLA-Adapter, a novel paradigm designed to reduce the reliance of VLA models on large-scale VLMs and extensive pre-training. To this end, we first systematically analyze the effectiveness of various VL conditions and present key findings on which conditions are essential for bridging perception and action spaces. Based on these insights, we propose a lightweight Policy module with Bridge Attention, which autonomously injects the optimal condition into the action space. In this way, our method achieves high performance using only a 0.5B-parameter backbone, without any robotic data pre-training. Extensive experiments on both simulated and real-world robotic benchmarks show that VLA-Adapter not only achieves state-of-the-art level performance, but also offers the fast inference speed reported to date. Furthermore, thanks to the proposed advanced bridging paradigm, VLA-Adapter enables the training of a powerful VLA model on a single consumer-grade GPU, greatly lowering the barrier to deploying VLA model. Yihao Wang 0006, Pengxiang Ding, Can Cui 0008, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao 0008, Pengxu Hou, Siteng Huang, Ru Zhang 0002 |
AAAI | 8 |
| 2025 | Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceabstractIn recent years, applying multi-modal large language models (MLLMs) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, MLLMs comprise the well-known Transformer network, which has a less efficient quadratic computation complexity. In this study, we introduce Cobra, a multi-modal large-scale language model built upon a state-space model, which has demonstrated significant potential in efficiently handling long sequences with fast inference and linear scalability concerning sequence length. Specifically, Cobra involves replacing Transformer-based backbone models (e.g., LLaMA or Phi) with pre-trained Mamba language models. We then empirically explore effective strategies for aligning visual and textual modalities and integrating various pre-trained Mamba model variants with visual encoders. Experiments across various multi-modal benchmarks demonstrate that: (i) Cobra performs 3× ∼ 4× faster than the most computationally efficient state-of-the-art methods, e.g., LLaVA-Phi and MobileVLM v2. Additionally, its performance is significantly enhanced thanks to the implementation of linear sequential modeling. (ii) Cobra fine-tunes a small parameter (∼48% of model parameters), leading to a significant improvement in overall performance compared to LLaVA. Han Zhao 0008, Min Zhang 0068, Pengxiang Ding, Siteng Huang |
AAAI | 1 |
| 2025 | VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot ManipulationabstractVision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human interaction. A typical approach of incorporating speech modality into VLA necessitates a separate speech recognition system to transcribe spoken instructions into text. Such a cascading pipeline raises two major concerns for robotic systems. First, the entire model grows in size and complexity, potentially resulting in redundant computations and increased memory consumption. Second, the transcription procedure would lose non-semantic information in the raw speech, such as voiceprint, which is crucial for a robot to successfully understand and complete customized tasks. To this end, we propose VLAS, the fisrt end-to-end policy model that seamlessly integrates speech modality for robot manipulation. We present a three-stage speech instruction tuning strategy leveraging multimodal datasets, including our manually curated SQA and CSI datasets. Furthermore, to facilitate personalized operations, we develop a voice retrieval-augmented generation (RAG) approach to enhance the robot's performance in tasks requiring individual-specific knowledge. Experimental results show that the proposed VLAS, following either textual or speech instructions, can achieve performance comparable to traditional VLAs on the CALVIN benchmark. In addition, we created a benchmark consisting of customization tasks, where our VLAS demonstrates absolute superiority by fully leveraging the auxiliary information in speech. Pengxiang Ding, Min Zhang 0068, Zhefei Gong, Shuanghao Bai, Han Zhao 0008 |
ICLR | 6 |
| 2025 | ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement LearningabstractVision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning robust policy models from mixed-quality data. In this paper, we introduce Reinforced robot GPT (ReinboT), a novel end-to-end VLA model that integrates the RL principle of maximizing cumulative reward. ReinboT achieves a deeper understanding of the data quality distribution by predicting dense returns that capture the nuances of manipulation tasks. The dense return prediction capability enables the robot to generate more robust decision-making actions, oriented towards maximizing future benefits. Extensive experiments show that ReinboT achieves state-of-the-art performance on the CALVIN mixed-quality dataset and exhibits superior few-shot learning and out-of-distribution generalization capabilities in real-world tasks. Hongyin Zhang 0001, Zifeng Zhuang, Han Zhao 0008, Pengxiang Ding, Hongchao Lu |
ICML | 3 |
| 2025 | Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot LearningabstractThis paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUARTOnline, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference at 50 Hz in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65 %. Our project page is https://quart-online.github.io. Xinyang Tong, Pengxiang Ding, Yiguo Fan, Can Cui 0008, Han Zhao 0008, Hongyin Zhang 0001, Yonghao Dang, Siteng Huang, Shangke Lyu |
ICRA | 8 |
| 2025 | MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action ModelsabstractDeveloping versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinforcement learning (RL) for fine-tuning large-scale VLA models with a large amount of mixed-quality data. MoRE integrates multiple low-rank adaptation modules as distinct experts within a dense multi-modal large language model (MLLM), forming a sparse-activated mixture-of-experts model. This design enables the model to effectively adapt to a wide array of downstream tasks. Moreover, we employ a reinforcement learning-based training objective to train our model as a Q-function after deeply exploring the structural properties of our tasks. Effective learning from automatically collected mixed-quality data enhances data efficiency and model performance. Extensive experiments demonstrate that MoRE outperforms all baselines across six different skills and exhibits superior generalization capabilities in out-of-distribution scenarios. We further validate our method in real-world scenarios, confirming the practicality of our approach and laying a solid foundation for future research on multi-task learning in quadruped robots. Han Zhao 0008, Wenxuan Song, Xinyang Tong, Pengxiang Ding, Xuelian Cheng, ZongYuan Ge |
ICRA | 1 |
| 2025 | PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel DecodingabstractVision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in VLA models with increased chunking sizes. This reduces the inference efficiency. Therefore, accelerating VLA integrated with action chunking is an urgent need. To tackle this problem, we propose PD-VLA, the first parallel decoding framework for VLA models integrated with action chunking. Our framework reformulates autoregressive decoding as a nonlinear system solved by parallel fixed-point iterations. This approach preserves model performance with mathematical guarantees while significantly improving decoding speed. In addition, it enables training-free acceleration without architectural changes, as well as seamless synergy with existing acceleration techniques. Extensive simulations validate that our PD-VLA maintains competitive success rates while achieving 2.52× execution frequency on manipulators (with 7 degrees of freedom) compared with the fundamental VLA model. Furthermore, we experimentally identify the most effective settings for acceleration. Finally, real-world experiments validate its high applicability across different tasks. Wenxuan Song, Pengxiang Ding, Han Zhao 0008, Zhide Zhong, ZongYuan Ge, Jun Ma 0008, Haoang Li |
IROS | 4 |
| 2025 | SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial ReasoningabstractDespite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either require specialized sensors or fail to effectively exploit depth information for higher-order reasoning. To this end, we propose a novel Spatial Sense and Reasoning method, dubbed SSR, a novel framework that transforms raw depth data into structured, interpretable textual rationales. These textual rationales serve as meaningful intermediate representations to significantly enhance spatial reasoning capabilities. Additionally, we leverage knowledge distillation to compress the generated rationales into compact latent embeddings, which facilitate resource-efficient and plug-and-play integration into existing VLMs without retraining. To enable comprehensive evaluation, we introduce a new dataset named SSR-CoT, a million-scale visual-language reasoning dataset enriched with intermediate spatial reasoning annotations, and present SSRBench, a comprehensive multi-task benchmark. Extensive experiments on multiple benchmarks demonstrate SSR substantially improves depth utilization and enhances spatial reasoning, thereby advancing VLMs toward more human-like multi-modal understanding. Project page: [https://yliu-cs.github.io/SSR](https://yliu-cs.github.io/SSR). Yang Liu 0358, Xiaomin Yu, Pengxiang Ding, Han Zhao 0008, Siteng Huang |
NeurIPS | 5 |
| 2025 | A nearly optimal adaptive saturation function tuning method for quasi-sliding mode control based on integral reinforcement learning
Lei Guo 0011, Wenbo Xiong, Han Zhao 0008, Dongming Gan |
Neurocomputing | 3 |
| 2024 | QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Pengxiang Ding, Han Zhao 0008, Wenxuan Song, Min Zhang 0068, Siteng Huang, Ningxi Yang |
ECCV (5) | 2 |
| 2024 | PiTe: Pixel-Temporal Alignment for Large Video-Language Model
Yang Liu 0358, Pengxiang Ding, Siteng Huang, Min Zhang 0068, Han Zhao 0008 |
ECCV (5) | 5 |
| 2024 | GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped RobotabstractMulti-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in collecting training datasets. In this paper, we propose GeRM (Generalist Robotic Model). We utilize offline reinforcement learning to optimize data utilization strategies to learn from both demonstrations and sub-optimal data, thus surpassing the limitations of human demonstrations. Thereafter, we employ a transformer-based VLA network to process multi-modal inputs and output actions. By introducing the Mixture-of-Experts structure, GeRM allows faster inference speed with higher whole model capacity, and thus resolves the issue of limited RL parameters, enhancing model performance in multi-task learning while controlling computational costs. Through a series of experiments, we demonstrate that GeRM outperforms other methods across all tasks, while also validating its efficiency in both training and inference processes. Additionally, we uncover its potential to acquire emergent skills. Additionally, we contribute the QUARD-Auto dataset, collected automatically to support our training approach and foster advancements in multi-task quadruped robot learning. This work presents a new paradigm for reducing the cost of collecting robot data and driving progress in the multi-task learning community.You can reach our project and video through the link: https://songwxuan.github.io/GeRM/. Wenxuan Song, Han Zhao 0008, Pengxiang Ding, Can Cui 0008, Shangke Lyu, Yaning Fan |
IROS | 2 |
| 2023 | A Composite Control Strategy for Quadruped Robot by Integrating Reinforcement Learning and Model-Based ControlabstractLocomotion in the wild requires the quadruped robot to have strong capabilities in adaptation and robustness. The deep reinforcement learning (DRL) exhibits the huge potential in environmental adaptability, while its stability issues remain open. On the other hand, the quadruped robot dynamic model contains a lot of useful information that is beneficial to the robust control. The combination of DRL with model-based control may take both strengths and hold promises in better robustness. In this paper, the DRL and the proposed model-based controller are firmly integrated in a novel manner such that the proposed model-based controller is able to rectify the gait commands generated by DRL based on the system dynamic model so as to enhance the robustness of the quadruped robot against the external disturbances. Besides, a potential energy function is introduced to achieve the compliant contact. The stability of the proposed method is ensured in terms of passivity analysis. Several physical experiments are carried out to verify the performance of the proposed method. Shangke Lyu, Han Zhao 0008 |
IROS | 2 |
| 2023 | Online adaptive optimal control algorithm based on synchronous integral reinforcement learning with explorations
Lei Guo 0011, Han Zhao 0008 |
Neurocomputing | 2 |