Haiming Gao

dblp:71/2551 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot navigation and mapping · 42% Segmentation and scene understanding · 20% 3D vision · 13%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
localization
0.912025
ERPoT: Effective and Reliable Pose Tracking for Mobile Robots Using Lightweight Polygon Maps · IEEE Trans. Robotics 2025
Computer vision › 3D vision › pose estimation
pose tracking
0.912025
ERPoT: Effective and Reliable Pose Tracking for Mobile Robots Using Lightweight Polygon Maps · IEEE Trans. Robotics 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing · NeurIPS 2024
Robotics › Robot navigation and mapping › place recognition
visual place recognition
0.812024
EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing · NeurIPS 2024
Robotics › Autonomous driving › perception › vision-based perception
lane detection
0.712023
PriorLane: A Prior Knowledge Enhanced Lane Detection Approach Based on Transformer · ICRA 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
PriorLane: A Prior Knowledge Enhanced Lane Detection Approach Based on Transformer · ICRA 2023
Computer vision › Segmentation and scene understanding › semantic segmentation
transformer-based segmentation
0.712023
PriorLane: A Prior Knowledge Enhanced Lane Detection Approach Based on Transformer · ICRA 2023
Robotics › Robot navigation and mapping › localization › range-based localization
LiDAR localization
0.312025
ERPoT: Effective and Reliable Pose Tracking for Mobile Robots Using Lightweight Polygon Maps · IEEE Trans. Robotics 2025
Computer vision › Image recognition and object detection
image retrieval
0.212024
EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

point-polygon matching · 0.9obstacle selection · 0.9ground removal · 0.9second-order feature pooling · 0.8dynamic power normalization · 0.8centroid-free probing · 0.8vision transformer · 0.7knowledge embedding alignment · 0.7
YearPublicationVenuePosition
2026 Transparent fault diagnosis for magnetic control circuit Breakers: A ViT-Based unsupervised approach
Kunquan Chen, Haoqing Wang, Fengchao Wang, Haiming Gao, Yiran Xia, Shude Zhao, Yakui Liu
Expert Syst. Appl.4
2025 Optimal Scheduling of a Dual-Arm Robot for Efficient Strawberry Harvesting in Plant Factories
abstract
Plant factory cultivation is widely recognized for its ability to optimize resource use and boost crop yields. To further increase the efficiency in these environments, we propose a mixed-integer linear programming (MILP) framework that systematically schedules and coordinates dual-arm harvesting tasks, minimizing the overall harvesting makespan based on pre-mapped fruit locations. Specifically, we focus on a specialized dual-arm harvesting robot and employ pose coverage analysis of its end effector to maximize picking reachability. Additionally, we compare the performance of the dual-arm configuration with that of a single-arm vehicle, demonstrating that the dual-arm system can nearly double efficiency when fruit densities are roughly equal on both sides. Extensive simulations show a 10–20% increase in throughput and a significant reduction in the number of stops compared to non-optimized methods. These results underscore the advantages of an optimal scheduling approach in improving the scalability and efficiency of robotic harvesting in plant factories.
Yuankai Zhu, Wenwu Lu, Guoqiang Ren, Haiming Gao, Stavros G. Vougioukas, Yibin Ying
IROS4
2025 M3CS: Multi-Target Masked Point Modeling With Learnable Codebook and Siamese Decoders
abstract
Masked point modeling has become a promising scheme of self-supervised pre-training for point clouds. Existing methods reconstruct either the masked points or related features as the objective of pre-training. However, considering the diversity of downstream tasks, it is necessary for the model to have both low- and high-level representation modeling capabilities during pre-training. It enables the capture of both geometric details and semantic contexts. To this end, M3CS is proposed to endow the model with the above abilities. Specifically, with the masked point cloud as input, M3CS introduces two decoders to reconstruct masked representations and the masked points simultaneously. While an extra decoder doubles parameters for the decoding process and may lead to overfitting, we propose siamese decoders to keep the number of learnable parameters unchanged. Further, we propose an online codebook projecting continuous tokens into discrete ones before reconstructing masked points. In such a way, we can compel the decoder to take effect through the combinations of tokens rather than remembering each token. Comprehensive experiments show that M3CS achieves superior performance across both classification and segmentation tasks, outperforming existing methods that are also single-modality and single-scale.
Qibo Qiu, Honghui Yang, Haochao Ying, Haiming Gao, Wenxiao Wang 0001, Xiaofei He 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 ERPoT: Effective and Reliable Pose Tracking for Mobile Robots Using Lightweight Polygon Maps
abstract
This paper presents an effective and reliable pose tracking solution, termed ERPoT, for mobile robots operating in large-scale outdoor and challenging indoor environments, underpinned by an innovative prior polygon map. Especially, to overcome the challenge that arises as the map size grows with the expansion of the environment, the novel form of a prior map composed of multiple polygons is proposed. Benefiting from the use of polygons to concisely and accurately depict environmental occupancy, the prior polygon map achieves long-term reliable pose tracking while ensuring a compact form. More importantly, pose tracking is carried out under pure LiDAR mode, and the dense 3D point cloud is transformed into a sparse 2D scan through ground removal and obstacle selection. On this basis, a novel cost function for pose estimation through point-polygon matching is introduced, encompassing two distinct constraint forms: point-to-vertex and point-to-edge. In this study, our primary focus lies on two crucial aspects: lightweight and compact prior map construction, as well as effective and reliable robot pose tracking. Both aspects serve as the foundational pillars for future navigation across diverse mobile platforms equipped with different LiDAR sensors in varied environments. Comparative experiments based on the publicly available datasets and our self-recorded datasets are conducted, and evaluation results show the superior performance of ERPoT on reliability, prior map size, pose estimation error, and runtime over the other six approaches. The corresponding code can be accessed athttps://github.com/ghm0819/ERPoT, and the supplementary video is athttps://youtu.be/6XdcXyUrLKw.
Haiming Gao, Qibo Qiu, Hongyan Liu 0007, Dingkun Liang, Chaoqun Wang 0009, Xuebo Zhang 0003
IEEE Trans. Robotics1
2024 EMVP: Embracing Visual Foundation Model for Visual Place Recognition with Centroid-Free Probing
abstract
Visual Place Recognition (VPR) is essential for mobile robots as it enables them to retrieve images from a database closest to their current location. The progress of Visual Foundation Models (VFMs) has significantly advanced VPR by capturing representative descriptors in images. However, existing fine-tuning efforts for VFMs often overlook the crucial role of probing in effectively adapting these descriptors for improved image representation. In this paper, we propose the Centroid-Free Probing (CFP) stage, making novel use of second-order features for more effective use of descriptors from VFMs. Moreover, to control the preservation of task-specific information adaptively based on the context of the VPR, we introduce the Dynamic Power Normalization (DPN) module in both the recalibration and CFP stages, forming a novel Parameter Efficiency Fine-Tuning (PEFT) pipeline (EMVP) tailored for the VPR task. Extensive experiments demonstrate the superiority of the proposed CFP over existing probing methods. Moreover, the EMVP pipeline can further enhance fine-tuning performance in terms of accuracy and efficiency. Specifically, it achieves 93.9\%, 96.5\%, and 94.6\% Recall@1 on the MSLS Validation, Pitts250k-test, and SPED datasets, respectively, while saving 64.3\% of trainable parameters compared with the existing SOTA PEFT method.
Qibo Qiu, Haiming Gao, Honghui Yang, Haochao Ying, Wenxiao Wang 0001, Xiaofei He 0001
NeurIPS3
2024 SelFLoc: Selective feature fusion for large-scale point cloud-based place recognition
Qibo Qiu, Wenxiao Wang 0001, Haochao Ying, Dingkun Liang, Haiming Gao, Xiaofei He 0001
Knowl. Based Syst.5
2023 PriorLane: A Prior Knowledge Enhanced Lane Detection Approach Based on Transformer
abstract
Lane detection is one of the fundamental modules in self-driving. In this paper we employ a transformer-only method for lane detection, thus it could benefit from the blooming development of fully vision transformer and achieve the state-of-the-art (SOTA) performance on both CULane and TuSimple benchmarks, by fine-tuning the weight fully pre-trained on large datasets. More importantly, this paper proposes a novel and general framework called PriorLane, which is used to enhance the segmentation performance of the fully vision transformer by introducing the low-cost local prior knowledge. Specifically, PriorLane utilizes an encoder-only transformer to fuse the feature extracted by a pre-trained segmentation model with prior knowledge embeddings. Note that a Knowledge Embedding Alignment (KEA) module is adapted to enhance the fusion performance by aligning the knowledge embedding. Extensive experiments on our Zjlab dataset show that PriorLane outperforms SOTA lane detection methods by a 2.82% mIoU when prior knowledge is employed, and the code will be released at: https://github.com/vincentqqb/PriorLane.
Qibo Qiu, Haiming Gao, Wei Hua 0002, Gang Huang 0004, Xiaofei He 0001
ICRA2
2022 Bridging the Gap Between Visual Servoing and Visual SLAM: A Novel Integrated Interactive Framework
abstract
For pose stabilization task of nonholonomic mobile robots, this article proposes a novel integrated interactive framework, bridging the gap between visual servoing and simultaneous localization and mapping (SLAM). The framework consists of two cooperative components, control module for servoing task and SLAM module for feedback signals estimation. In most visual servoing methods, feedback signals for the servoing controller are estimated by means of multiple-view geometry assuming the target scene being always within the camera field of view (FOV). To handle the challenge that the target scene gets out of view during servoing process, the desired image is associated with the initial map by a two-step strategy, and an incremental map is constructed to guarantee available feedback signals estimation. In addition, on the basis of the kinematic model of the mobile robot and velocities designed by the servo controller, the predicted pose is exploited to discard moving objects in the camera FOV, thus making the proposed framework effective in dynamic scenes. Experimental results operated in different scenes without prior information demonstrate the effectiveness of the proposed approach to handle the FOV problem and dynamic scenes.Note to Practitioners—Traditional visual servoing stabilization approaches usually require that the feature points in the target scene remain within the FOV of the camera for feedback signals calculation, which is often neglected. Motivated by the requirement of continuous feedback signals to the servo controller, the SLAM technique is introduced to relax the FOV constraint during the servoing process. A novel integrated interactive framework is proposed in this article to further increase the applicability of the servoing system in practice, in which the SLAM module is also redesigned for the flexibility in dynamic scenes. The SLAM module provides feedback signals for the servo controller; meanwhile, velocities designed by the servo controller are utilized for the prediction mechanism in the SLAM module to discard features on moving objects. Experiments validate the applicability of the proposed framework in different scenarios.
Chenping Li, Xuebo Zhang 0003, Haiming Gao, Runhua Wang, Yongchun Fang
IEEE Trans Autom. Sci. Eng.3
2022 E3MoP: Efficient Motion Planning Based on Heuristic-Guided Motion Primitives Pruning and Path Optimization With Sparse-Banded Structure
abstract
To solve the autonomous navigation problem in complex environments, an efficient motion planning approach is newly presented in this paper. Considering the challenges from large-scale, partially unknown complex environments, a three-layer motion planning framework is elaborately designed, including global path planning, local path optimization, and time-optimal velocity planning. Compared with existing approaches, the novelty of this work is twofold: 1) a novel heuristic-guided pruning strategy of motion primitives is proposed and fully integrated into the state lattice-based global path planner to further improve the computational efficiency of graph search, and 2) a new soft-constrained local path optimization approach is proposed, wherein the sparse-banded system structure of the underlying optimization problem is fully exploited to efficiently solve the problem. We validate the safety, smoothness, flexibility, and efficiency of our approach in various complex simulation scenarios and challenging real-world tasks. It is shown that the computational efficiency is improved by 66.21% in the global planning stage and the motion efficiency of the robot is improved by 22.87% compared with the recent quintic Bézier curve-based state space sampling approach. We name the proposed motion planning framework E$\mathbf {^{3}} $MoP, where the number 3 not only means our approach is a three-layer framework but also means the proposed approach is efficient in three stages. Note to Practitioners—This paper is motivated by the challenges of motion planning problems of mobile robots. A three-layer motion planning framework is proposed by combining global path planning, local path optimization, and time-optimal velocity planning. For mobile robot navigation applications in semi-structured environments, optimization-based local planners are recommended. Extensive simulation and experimental results show the effectiveness of the proposed motion planning framework. However, due to the non-convexity of the path optimization formulation, the proposed local planner may get stuck in local optima. In future research, we will concentrate on extending the proposed local path optimization approach with the theory of homology classes to maintain several homotopically distinct local paths and seek global optima.
Xuebo Zhang 0003, Haiming Gao, Jing Yuan 0004, Yongchun Fang
IEEE Trans Autom. Sci. Eng.3
2019 Autonomous Indoor Exploration Via Polygon Map Construction and Graph-Based SLAM Using Directional Endpoint Features
abstract
In this paper, a novel 2-D laser-based autonomous exploration approach for mobile robots is proposed, which is based on a novel polygon map construction approach and graph-based simultaneous localization and mapping (SLAM) with directional endpoint features. This approach is composed of three modules: graph-based SLAM using directional endpoint features, polygon map construction, and exploration. Different from existing approaches in the field of 2-D SLAM, the newly proposed 2-D graph-SLAM is based on 3-D “directional endpoint” features; on this basis, a well-known data structure “circular-doubly linked list” is applied to construct a novel polygon map for navigation. Note that it is efficient for circular-doubly linked list to initialize and update the polygon map. In addition, we propose a new information entropy calculation approach to quantify the entropy of the polygon map. Then for each candidate goal, we could obtain corresponding information gain and make next decision through collision detection. Comparative experimental results with respect to the well-known Gmapping and Karto SLAM are presented to show superior performance of the proposed graph-based SLAM. The autonomous exploration experiments in the office and hallway environments show the effectiveness of the proposed approach for robotic mapping and exploration tasks.
Haiming Gao, Xuebo Zhang 0003, Jing Yuan 0004, Yongchun Fang
IEEE Trans Autom. Sci. Eng.1
1994 A Suite of Analysis Tools Based on a General Purpose Abstract Interpreter
Thomas E. Cheatham, Haiming Gao, Dan C. Stefanescu
CC2