Hanbo Zhang

dblp:119/1807 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 15 since 2021Systems, architecture and hardware · 10 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DCAM U-Net: A 3D U-Net Model with 3D Channel-Spatial Attention and Feature Fusion for Subsurface Void Segment in Ground Penetrating Radar Data
Peiwen Yao, Hanbo Zhang, Dandan Cheng, Guoliang Xin, Liyong Zhang
KSEM (2)3
2026 MInCo: Mitigating conflicting objectives in distracted visual model-based reinforcement learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
Knowl. Based Syst.2
2026 FSG-Zero: An efficient end-to-end fusion framework for robust continual adaptation in EEG seizure prediction
Hanbo Zhang, Jincan Zhang, Wenna Chen, Ganqin Du
Knowl. Based Syst.1
2025 ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained Environments
abstract
We introduce ProcWORLD, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM).ProcWORLD features a wide range of challenging embodied navigation and object manipulation tasks, covering 16 task types, 5,000 rooms, and over 10 million evaluation trajectories with diverse data distribution.ProcWORLD supports configurable observation modes, ranging from text-only descriptions to vision-only observations.It enables text-based actions to control the agent following language instructions.ProcWORLD has presented significant challenges for LLMs and VLMs: (1) active information gathering given partial observations for disambiguation; (2) simultaneous localization and decision-making by tracking the spatio-temporal state-action distribution; (3) constrained reasoning with dynamic states subject to physical reachability.Our extensive evaluation of 15 foundation models and 5 reasoning algorithms (with over 1 million rollouts) indicates larger models perform better.However, ProcWORLD remains highly challenging for existing state-of-the-art models and in-context learning methods due to constrained reachability and the need of combinatorial spatial reasoning.
Xinghang Li, Zhengshen Zhang, Jirong Liu, Xiao Ma 0006, Hanbo Zhang, Tao Kong, Huaping Liu 0001
EMNLP6
2025 Towards Extrinsic Dexterity Grasping in Unrestricted Environments
abstract
Grasping large and flat objects (e.g., a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Prior research has exploited environmental interactions through Extrinsic Dexterity, utilizing external structures such as walls or table edges to facilitate object grasping. However, they are confined to task-specific policies while neglecting semantic perception and planning to identify optimal pre-grasp configurations. This limits their operational versatility, impeding effective adaptation to varied extrinsic dexterity constraints. In this work, we present ExDiff, a robot manipulation approach for extrinsic dexterity grasping in unrestricted environments. It utilizes Vision-Language Models (VLMs) to perceive the environmental state and generate instructions, followed by a Goal-Conditioned Action Diffusion (GCAD) model to predict the sequence of low-level actions. This diffusion model learns the low-level policy, conditioned on high-level instructions and cumulative rewards, which improves the generation of robot actions. Simulation experiments and real-world deployment results demonstrate that ExDiff effectively performs ungraspable tasks and generalizes to previously unseen target objects and scenes. Videos at - https://exdiff.github.io/index.html
Chengzhong Ma, Houxue Yang, Hanbo Zhang, Zeyang Liu 0001, Xuguang Lan, Nanning Zheng 0001
IROS3
2025 RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
abstract
Multimodal Large Language Models (MLLMs) increasingly excel at perception,understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RT V-Bench, a fine-grained benchmark for MLLM real-time video analysis. RTV-Bench includes three key principles: (1) Multi-Timestamp Question Answering (MTQA), where answers evolve with scene changes; (2) Hierarchical Question Structure, combining basic and advanced queries; and (3) Multi-dimensional Evaluation, assessing the ability of continuous perception, understanding, and reasoning. RTV-Bench contains 552 diverse videos (167.2 hours) and 4,631 high-quality QA pairs. We evaluated leading MLLMs, including proprietary (GPT-4o, Gemini 2.0), open-source offline (Qwen2.5-VL, VideoLLaMA3), and open-source real-time (VITA-1.5, InternLM-XComposer2.5-OmniLive) models. Experiment results show open-source real-time models largely outperform offline ones but still trail top proprietary models. Our analysis also reveals that larger model size or higher frame sampling rates do not significantly boost RTV-Bench performance, sometimes causing slight decreases.This underscores the need for better model architectures optimized for video stream processing and long sequences to advance real-time video analysis with MLLMs.
Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Hanqian Li, Linghao Zhang, Shikang Wang, Hanbo Zhang, Xuming Hu
NeurIPS12
2025 Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
abstract
We present Chain-of-Action (CoA), a novel visuomotor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-level Chain-of-Thought (CoT) process. This process is unified within a single autoregressive structure: (1) the first token corresponds to a stable keyframe action that encodes the task-specific goals; and (2) subsequent action tokens are generated autoregressively, conditioned on the initial keyframe and previously predicted actions. This backward action reasoning enforces a global-to-local structure, allowing each local action to be tightly constrained by the final goal. To further realize the action reasoning structure, CoA incorporates four complementary designs: continuous action token representation; dynamic stopping for variable-length trajectory generation; reverse temporal ensemble; and multi-token prediction to balance action chunk modeling with global structure. As a result, CoA gives strong spatial generalization capabilities while preserving the flexibility and simplicity of a visuomotor policy. Empirically, we observe that CoA outperforms representative imitation learning algorithms such as ACT and Diffusion Policy across 60 RLBench tasks and 8 real-world tasks.
Wenbo Zhang 0009, Tianrun Hu, Hanbo Zhang, Yanyuan Qiao, Yuchu Qin, Yang Li 0184, Jiajun Liu 0004, Tao Kong, Lingqiao Liu
NeurIPS3
2024 Vision-Language Foundation Models as Effective Robot Imitators
abstract
Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on robotics data. To this end, we derive a simple and novel vision-language manipulation framework, dubbed RoboFlamingo, built upon the open-source VLMs, OpenFlamingo. Unlike prior works, RoboFlamingo utilizes pre-trained VLMs for single-step vision-language comprehension, models sequential history information with an explicit policy head, and is slightly fine-tuned by imitation learning only on language-conditioned manipulation datasets. Such a decomposition provides RoboFlamingo the flexibility for open-loop control and deployment on low-performance platforms. By exceeding the state-of-the-art performance with a large margin on the tested benchmark, we show RoboFlamingo can be an effective and competitive alternative to adapt VLMs to robot control. Our extensive experimental results also reveal several interesting conclusions regarding the behavior of different pre-trained VLMs on manipulation tasks. We believe RoboFlamingo has the potential to be a cost-effective and easy-to-use solution for robotics manipulation, empowering everyone with the ability to fine-tune their own robotics policy. Our code will be made public upon acceptance.
Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Chilam Cheang, Ya Jing, Weinan Zhang 0001, Huaping Liu 0001, Tao Kong
ICLR3
2024 Towards Unified Interactive Visual Grounding in The Wild
abstract
Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user’s input by active information gathering. Previous approaches often rely on predefined templates to ask disambiguation questions, resulting in performance reduction in realistic interactive scenarios. In this paper, we propose TiO, an end-to-end system for interactive visual grounding in human-robot interaction. Benefiting from a unified formulation of visual dialog and grounding, our method can be trained on a joint of extensive public data, and show superior generality to diversified and challenging open-world scenarios. In the experiments, we validate TiO on GuessWhat?! and InViG benchmarks, setting new state-of-the-art performance by a clear margin. Moreover, we conduct HRI experiments on the carefully selected 150 challenging scenes as well as real-robot platforms. Results show that our method demonstrates superior generality to diversified visual and language inputs with a high success rate. Codes and demos are available on https://jxu124.github.io/TiO/.
Hanbo Zhang, Qingyi Si, Xuguang Lan, Tao Kong
ICRA2
2024 A Large-Area LTPS-TFT-Based Bi-directional Biomedical Interface with Process-Invariant In-pixel Biopotential-to-Digital Converters
abstract
In this paper, we demonstrate a bi-directional biomedical process-invariant pixel interface based on the LTPS-TFT that integrates a front-end amplifier, an analog-to-digital converter and a stimulator. The pixel interface can convert biopotential into digital signals. This near-sensor signal processing can avoid the interference of motion artifacts, and the all-digital transfer has a high noise tolerance. Under the process variation of 1×-1.4× threshold voltage and ±10% mobility fluctuations, the DC gain of the operational amplifier changes from 58.83 dB to 57.21 dB, and 39.09 dB to 37.88 dB for the front-end amplifier. Compared with the single-stage amplifier, the stability has been improved by 13.2×. The proposed pseudo differential VCO-based ADC can effectively eliminate second-order non-linearity which achieves the best performance in state-of-the-art TFT-ADCs. When the threshold voltage and mobility fluctuate, ENOB changes from 7.36 bit to 7.30 bit (OSR=64) and 11.52 bit to 11.14 bit (OSR=256). The change rate does not exceed 4%.
Hanbo Zhang, Yuqing Lou, Zhihang Zhang, Yongfu Li 0002, Fakhrul Z. Rokhani, Guoxing Wang, Jian Zhao 0004
ISCAS1
2024 What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
abstract
Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong, Ruihua Song. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Tao Kong, Ruihua Song
NAACL-HLT2
2024 Optimal bipartite graph matching-based goal selection for policy-based hindsight learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan
Neurocomputing2
2023 Towards Open-World Interactive Disambiguation for Robotic Grasping
abstract
Language-based communications are essential in human-robot interaction, especially for the majority of non-expert users. In this paper, we present SeeAsk, an open-world interactive visual grounding system to grasp specified targets with ambiguous natural language instructions. The main contribution of SeeAsk is that it can robustly handle open-world scenes in terms of both open-set objects and open-vocabulary interactions. Specifically, our SeeAsk is built upon modern large-scale vision-language pre-trained models and traditional decision-making process, and shows promising results to be deployed in real-world scenarios. SeeAsk outperforms previous state-of-the-art algorithms with a clear margin in terms of not only success rate but also asking smarter and more informative questions. User studies also demonstrate its advantages over previous works.
Yuchen Mo, Hanbo Zhang, Tao Kong
ICRA2
2023 Towards a Generic Framework for Mechanism-guided Deep Learning for Manufacturing Applications
abstract
Manufacturing data analytics tasks are traditionally undertaken with Mechanism Models (MMs), which are domain-specific mathematical equations modeling the underlying physical or chemical processes of the tasks. Recently, Deep Learning (DL) has been increasingly applied to manufacturing. MMs and DL have their individual pros and cons, motivating the development of Mechanism-guided Deep Learning Models (MDLMs) that combine the two. Existing MDLMs are often tailored to specific tasks or types of MMs, and can fail to effectively 1) utilize interconnections of multiple input examples, 2) adaptively self-correct prediction errors with error bounding, and 3) ensemble multiple MMs. In this work, we propose a generic, task-agnostic MDLM framework that can embed one or more MMs in deep networks, and address the 3 aforementioned issues. We present 2 diverse use cases where we experimentally demonstrate the effectiveness and efficiency of our models.
Hanbo Zhang, Jiangxin Li, Peng Wang 0027, Themis Palpanas, Chen Wang 0018, Wei Wang 0009, Haoxuan Zhou, Jianwei Song, Wen Lu 0002
KDD1
2023 Shapelet Based Two-Step Time Series Positive and Unlabeled Learning
Hanbo Zhang, Peng Wang 0027, Wei Wang 0009
J. Comput. Sci. Technol.1
2021 REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point Clouds
abstract
Reliable robotic grasping in unstructured environments is a crucial but challenging task. The main problem is to generate the optimal grasp of novel objects from partial noisy observations. This paper presents an end-to-end grasp detection network taking one single-view point cloud as input to tackle the problem. Our network includes three stages: Score Network (SN), Grasp Region Network (GRN), and Refine Network (RN). Specifically, SN regresses point grasp confidence and selects positive points with high confidence. Then GRN conducts grasp proposal prediction on the selected positive points. RN generates more accurate grasps by refining proposals predicted by GRN. To further improve the performance, we propose a grasp anchor mechanism, in which grasp anchors with assigned gripper orientations are introduced to generate grasp proposals. Experiments demonstrate that REGNet achieves a success rate of 79.34% and a completion rate of 96% in real-world clutter, which significantly outperforms several state-of-the-art point-cloud based methods, including GPD, PointNetGPD, and S4G. The code is available at https://github.com/zhaobinglei/REGNet for 3D Grasping.
Binglei Zhao 0002, Hanbo Zhang, Xuguang Lan, Nanning Zheng 0001
ICRA2
2021 Hindsight Trust Region Policy Optimization
abstract
Reinforcement Learning (RL) with sparse rewards is a major challenge. We pro- pose Hindsight Trust Region Policy Optimization (HTRPO), a new RL algorithm that extends the highly successful TRPO algorithm with hindsight to tackle the challenge of sparse rewards. Hindsight refers to the algorithm’s ability to learn from information across goals, including past goals not intended for the current task. We derive the hindsight form of TRPO, together with QKL, a quadratic approximation to the KL divergence constraint on the trust region. QKL reduces variance in KL divergence estimation and improves stability in policy updates. We show that HTRPO has similar convergence property as TRPO. We also present Hindsight Goal Filtering (HGF), which further improves the learning performance for suitable tasks. HTRPO has been evaluated on various sparse-reward tasks, including Atari games and simulated robot control. Experimental results show that HTRPO consistently outperforms TRPO, as well as HPG, a state-of-the-art policy 14 gradient algorithm for RL with sparse rewards.
Hanbo Zhang, Cedar Site Bai, Xuguang Lan, David Hsu, Nanning Zheng 0001
IJCAI1
2021 A Real-Time Robotic Grasping Approach With Oriented Anchor Box
abstract
Grasping is an essential skill for robots to interact with humans and the environment. In this paper, we build a vision-based, robust, and real-time robotic grasping approach with fully convolutional neural network. The main component of our approach is a grasp detection network with oriented anchor boxes as detection priors. Because the orientation of detected grasps is significant, which determines the rotation angle configuration of the gripper, we propose the orientation anchor box mechanism to regress grasp angle based on predefined assumption instead of classification or regression without any priors. With oriented anchor boxes, the grasps can be predicted more accurately and efficiently. Besides, to accelerate the network training and further improve the performance of angle regression, angle matching is proposed during training instead of Jaccard index matching. Fivefold cross-validation results demonstrate that our proposed algorithm achieves an accuracy of 98.8% and 97.8% in image-wise split and object-wise split, respectively, and the speed of our detection algorithm is 67 frames per second (FPS) with GTX 1080Ti, outperforming all the current state-of-the-art grasp detection algorithms on Cornell Dataset both in speed and accuracy. Robotic experiments demonstrate the robustness and generalization ability in unseen objects and real-world environment, with the average success rate of 90.0% and 84.2% of familiar things and unseen things, respectively, on Baxter robot platform.
Hanbo Zhang, Xinwen Zhou, Xuguang Lan, Jin Li 0011, Nanning Zheng 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2021 ELIS++: a shapelet learning approach for accurate and efficient time series classification
abstract
Abstract In recent years, time series classification with shapelets, due to the high accuracy and good interpretability, has attracted considerable interests. These approaches extract or learn shapelets from the training time series. Although they can achieve higher accuracy than other approaches, there still confront some challenges. First, they may suffer from low accuracy in the case of small training dataset. Second, they must manually set some parameters, like the number of shapelets and the length of each shapelet beforehand, and some hyper-parameters, like learning rate and regulation weight, which are difficult to set without prior knowledge. Third, extracting or learning shapelets incurs a huge computation cost, due to the huge search space. In this paper, we extend our previous shapelet learning approach ELIS to ELIS++. To improve the accuracy on the small training dataset, we propose a data augmentation approach. To learn the higher quality shapelets, based on the PAA shapelet candidates search technique proposed in ELIS, ELIS++ first propose a novel entropy-based approach shapelet candidate selection mechanism to discover shapelet candidates, and then applies the logistic regression model to adjust shapelets.To avoid setting other parameters manually, we propose a Bayesian Optimization based approach. Moreover, two techniques are proposed to improve the efficiency, coarse-grained shapelet adjustment and SIMD-based parallel computation. We conduct extensive experiments on 35 UCR datasets, and results verify the effectiveness and efficiency of ELIS++.
Hanbo Zhang, Peng Wang 0027, Zicheng Fang, Zeyu Wang 0007, Wei Wang 0009
World Wide Web1
2020 Autonomous Tool Construction with Gated Graph Neural Network
abstract
Autonomous tool construction is a significant but challenging task in robotics. This task can be interpreted as when given a reference tool, selecting some available candidate parts to reconstruct it. Most of the existing works perform tool construction in the form of action part and grasp part, which is only a specific construction pattern and limits its application to some extent. In general scenarios, a tool can be constructed in various patterns with different part pairs. Therefore, whether a part pair is most suitable for constructing the tool depends not only on itself, but on other parts in the same scene. To solve this problem, we construct a Gated Graph Neural Network (GGNN) to model the relations between all part pairs, so that we can select the candidate parts in consideration of the global information. Afterwards, we embed the constructed GGNN into a RCNN-like structure to finally accomplish tool construction. The whole model will be named Tool Construction Graph RCNN (TC-GRCNN). In addition, we develop a mechanism that can generate large-scale training and testing data in simulation environments, by which we can save the time of data collection and annotation. Finally, the proposed model is deployed on the physical robot. The experiment results show that TC-GRCNN can perform well in the general scenarios of tool construction.
Chenjie Yang, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001
ICRA3
2020 Visual manipulation relationship recognition in object-stacking scenes
Hanbo Zhang, Xuguang Lan, Xinwen Zhou, Nanning Zheng 0001
Pattern Recognit. Lett.1
2019 Task-oriented Grasping in Object Stacking Scenes with CRF-based Semantic Model
abstract
In task-oriented grasping, the robot is supposed to manipulate the objects in a task-compatible manner, which is more important but more challenging than just stably grasping. However, most of existing works perform task-oriented grasping only in single object scenes. This greatly limits their practical application in real world scenes, in which there are usually multiple stacked objects with serious overlaps and occlusions. To perform task-oriented grasping in object stacking scenes, in this paper, we firstly build a synthetic dataset named Object Stacking Grasping Dataset (OSGD) for task-oriented grasping in object stacking scenes. Secondly, a Conditional Random Field (CRF) is constructed to model the semantic contents in object regions. The modelled semantic contents can be illustrated as incompatibility of task labels and continuity of task regions. This proposed approach can greatly reduce the interference of overlaps and occlusions in object stacking scenes. To embed the CRF-based semantic model into our grasp detection network, we implement the inference process of CRFs as a RNN so that the whole model, Task-oriented Grasping CRFs (TOG-CRFs) can be trained end to end. Finally, in object stacking scenes, the constructed model can help robot achieve 69.4% success rate for task-oriented grasping.
Chenjie Yang, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001
IROS3
2019 A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes
abstract
Autonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to grasp; (2) it is hard to deal with the situations in which the target is hidden or covered by other objects. In this paper, we propose a multi-task convolutional neural network for autonomous robotic grasping, which can help the robot find the target, make the plan for grasping and finally grasp the target step by step in object stacking scenes. We integrate vision-based robotic grasping detection and visual manipulation relationship reasoning in one single deep network and build the autonomous robotic grasping system. Experimental results demonstrate that with our model, Baxter robot can autonomously grasp the target with a success rate of 90.6%, 71.9% and 59.4% in object cluttered scenes, familiar stacking scenes and complex stacking scenes respectively.
Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Lipeng Wan 0003, Chenjie Yang, Nanning Zheng 0001
IROS1
2019 ROI-based Robotic Grasp Detection for Object Overlapping Scenes
abstract
Grasp detection considering the affiliations between grasps and their owner in object overlapping scenes is a necessary and challenging task for the practical use of the robotic grasping approach. In this paper, a robotic grasp detection algorithm named ROI-GD is proposed to provide a feasible solution to this problem based on Region of Interest (ROI), which is the region proposal for objects. ROI-GD uses features from ROIs to detect grasps instead of the whole scene. It has two stages: the first stage is to provide ROIs in the input image and the second-stage is the grasp detector based on ROI features. We also contribute a multi-object grasp dataset, (a) which is much larger than Cornell Grasp Dataset, by labeling Visual Manipulation Relationship Dataset. Experimental results demonstrate that ROI-GD performs much better in object overlapping scenes and at the meantime, remains comparable with state-of-the-art grasp detection algorithms on Cornell Grasp Dataset and Jacquard Dataset. Robotic experiments demonstrate that ROI-GD can help robots grasp the target in single-object and multi-object scenes with the overall success rates of 92.5% and 83.8% respectively.
Hanbo Zhang, Xuguang Lan, Cedar Site Bai, Xinwen Zhou, Nanning Zheng 0001
IROS1
2018 Fully Convolutional Grasp Detection Network with Oriented Anchor Box
abstract
In this paper, we present a real-time approach to predict multiple grasping poses for a parallel-plate robotic gripper using RGB images. A model with oriented anchor box mechanism is proposed and a new matching strategy is used during the training process. An end-to-end fully convolutional neural network is employed in our work. The network consists of two parts: the feature extractor and multi-grasp predictor. The feature extractor is a deep convolutional neural network. The multi-grasp predictor regresses grasp rectangles from predefined oriented rectangles, called oriented anchor boxes, and classifies the rectangles into graspable and ungraspable. On the standard Cornell Grasp Dataset, our model achieves an accuracy of 97.74% and 96.61% on image-wise split and object-wise split respectively, and outperforms the latest state-of-the-art approach by 1.74% on image-wise split and 0.51% on object-wise split.
Xinwen Zhou, Xuguang Lan, Hanbo Zhang, Nanning Zheng 0001
IROS3