VLDB 2026 Research / reviewers in the wild / expert
Hongbo Zhang 0004
dblp:24/3333-4
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2023
0009-0005-8753-2580ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 14 since 2021Systems, architecture and hardware · 12 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | FG-Depth: Flow-Guided Unsupervised Monocular Depth EstimationabstractThe great potential of unsupervised monocular depth estimation has been demonstrated by many works due to low annotation cost and impressive accuracy comparable to supervised methods. To further improve the performance, recent works mainly focus on designing more complex network structures and exploiting extra supervised information, e.g., semantic segmentation. These methods optimize the models by exploiting the reconstructed relationship between the target and reference images in varying degrees. However, previous methods prove that this image reconstruction optimization is prone to get trapped in local minima. In this paper, our core idea is to guide the optimization with prior knowledge from pretrained Flow-Net. And we show that the bottleneck of unsupervised monocular depth estimation can be broken with our simple but effective framework named FG-Depth. In particular, we propose (i) a flow distillation loss to replace the typical photometric loss that limits the capacity of the model and (ii) a prior flow based mask to remove invalid pixels that bring the noise in training loss. Extensive experiments demonstrate the effectiveness of each component, and our approach achieves state-of-the-art results on both KITTI and NYU-Depth-v2 datasets. Junyu Zhu, Lina Liu 0010, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 6 |
| 2023 | Self-Supervised Event-Based Monocular Depth Estimation Using Cross-Modal ConsistencyabstractAn event camera is a novel vision sensor that can capture per-pixel brightness changes and output a stream of asynchronous “events”. It has advantages over conventional cameras in those scenes with high-speed motions and challenging lighting conditions because of the high temporal resolution, high dynamic range, low bandwidth, low power consumption, and no motion blur. Therefore, several supervised monocular depth estimation from events is proposed to address scenes difficult for conventional cameras. However, depth annotation is costly and time-consuming. In this paper, to lower the annotation cost, we propose a self-supervised event-based monocular depth estimation framework named EMoDepth. EMoDepth constrains the training process using the cross-modal consistency from intensity frames that are aligned with events in the pixel coordinate. Moreover, in inference, only events are used for monocular depth prediction. Additionally, we design a multi-scale skip-connection architecture to effectively fuse features for depth estimation while maintaining high inference speed. Experiments on MVSEC and DSEC datasets demonstrate that our contributions are effective and that the accuracy can outperform existing supervised event-based and unsupervised frame-based methods. Junyu Zhu, Lina Liu 0010, Bofeng Jiang, Hongbo Zhang 0004, Wanlong Li, Yong Liu 0007 |
IROS | 5 |
| 2023 | Relative order constraint for monocular depth estimation
Chunpu Liu, Wangmeng Zuo, Guanglei Yang, Wanlong Li, Hongbo Zhang 0004, Tianyi Zang |
Appl. Intell. | 6 |
| 2023 | Hunter: Exploring High-Order Consistency for Point Cloud Registration With Severe OutliersabstractAfter decades of investigation, point cloud registration is still a challenging task in practice, especially when the correspondences are contaminated by a large number of outliers. It may result in a rapidly decreasing probability of generating a hypothesis close to the true transformation, leading to the failure of point cloud registration. To tackle this problem, we propose a transformation estimation method, named Hunter, for robust point cloud registration with severe outliers. The core of Hunter is to design a global-to-local exploration scheme to robustly find the correct correspondences. The global exploration aims to exploit guided sampling to generate promising initial alignments. To this end, a hypergraph-based consistency reasoning module is introduced to learn the high-order consistency among correct correspondences, which is able to yield a more distinct inlier cluster that facilitates the generation of all-inlier hypotheses. Moreover, we propose a preference-based local exploration module that exploits the preference information of top- k promising hypotheses to find a better transformation. This module can efficiently obtain multiple reliable transformation hypotheses by using a multi-initialization searching strategy. Finally, we present a distance-angle based hypothesis selection criterion to choose the most reliable transformation, which can avoid selecting symmetrically aligned false transformations. Experimental results on simulated, indoor, and outdoor datasets, demonstrate that Hunter can achieve significant superiority over the state-of-the-art methods, including both learning-based and traditional methods (as shown in Fig. 1). Moreover, experimental results also indicate that Hunter can achieve more stable performance compared with all other methods with severe outliers. Runzhao Yao, Shaoyi Du, Wenting Cui, Aixue Ye, Hongbo Zhang 0004, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | LTSR: Long-term Semantic Relocalization based on HD Map for Autonomous VehiclesabstractHighly accurate and robust relocalization or localization initialization ability is of great importance for autonomous vehicles (AVs). Traditional GNSS-based methods are not reliable enough in occlusion and multipath conditions. In this paper we propose a novel long-term semantic relocalization algorithm based on HD map and semantic features which are compact in representation. Semantic features appear widely on urban roads, and are robust to illumination, weather, view-point and appearance changes. Repeated structures, missed and false detections make data association (DA) highly ambiguous. To this end, a robust semantic feature matching method based on a new local semantic descriptor which encodes the spatial and normal relationship between semantic features is performed. Further, we introduce an accurate, efficient, yet simple outlier removal method which works by assessing the local and global geometric consistencies and temporal consistency of semantic matching pairs. The experimental results on our urban dataset demonstrate that our approach performs better in accuracy and robustness compared with the current state-of-the-art methods. Huayou Wang, Changliang Xue, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 6 |
| 2021 | PocoNet: SLAM-oriented 3D LiDAR Point Cloud Online Compression NetworkabstractIn this paper, we present PocoNet: Point cloud Online COmpression NETwork to address the task of SLAM-oriented compression. The aim of this task is to select a compact subset of points with high priority to maintain localization accuracy. The key insight is that points with high priority have similar geometric features in SLAM scenarios. Hence, we tackle this task as point cloud segmentation to capture complex geometric information. We calculate observation counts by matching between maps and point clouds and divide them into different priority levels. Trained by labels annotated with such observation counts, the proposed network could evaluate the point-wise priority. Experiments are conducted by integrating our compression module into an existing SLAM system to evaluate compression ratios and localization performances. Experimental results on two different datasets verify the feasibility and generalization of our approach. Jinhao Cui, Xin Kong, Xuemeng Yang, Xiangrui Zhao, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
ICRA | 9 |
| 2021 | SA-LOAM: Semantic-aided LiDAR SLAM with Loop ClosureabstractLiDAR-based SLAM system is admittedly more accurate and stable than others, while its loop closure detection is still an open issue. With the development of 3D semantic segmentation for point cloud, semantic information can be obtained conveniently and steadily, essential for high-level intelligence and conductive to SLAM. In this paper, we present a novel semantic-aided LiDAR SLAM with loop closure based on LOAM, named SA-LOAM, which leverages semantics in odometry as well as loop closure detection. Specifically, we propose a semantic-assisted ICP, including semantically matching, downsampling and plane constraint, and integrates a semantic graph-based place recognition method in our loop closure detection module. Benefitting from semantics, we can improve the localization accuracy, detect loop closures effectively, and construct a global consistent semantic map even in large-scale scenes. Extensive experiments on KITTI and Ford Campus dataset show that our system significantly improves baseline performance, has generalization ability to unseen data and achieves competitive results compared with state-of-the-art methods. Lin Li 0091, Xin Kong, Xiangrui Zhao, Wanlong Li, Hongbo Zhang 0004, Yong Liu 0007 |
ICRA | 6 |
| 2021 | IMU/Vehicle Calibration and Integrated Localization for Autonomous DrivingabstractThe localization system, which outputs vehicle position, velocity, and attitude, is one of the fundamental components in the autonomous driving vehicle. The global pose is not only used for the planning and control system, but also an important reference for the cloud source-based HD Map building and updating. The accuracy, availability, and reliability are key requirements for the localization system to ensure that the whole system runs smoothly and efficiently.IMU/Vehicle extrinsic calibration is one of the primary jobs that should be addressed. Due to the observability issue, the IMU/vehicle relative roll cannot be calibrated by the traditional maneuver-based calibration method. In this paper, we solve this issue with the proposed Multiple Orientation-based Vehicle/IMU Extrinsic Calibration (MOVIE-Cali) method, which is evaluated by Monte Carlo simulations and experiments.When the vehicle is cornering or making a U-turn, the sideslip of the tires will have negative influence on the localization system which uses Non-Holonomic Constraints (NHC)/Wheel speed sensor measurement in the model. We derive a sideslip angle model and propose an online slip parameter calibration and compensation method to improve the localization accuracy. The performance of proposed method has been evaluated by the vehicle tests. Zhenbo Liu, Leijie Wang, Hongbo Zhang 0004 |
ICRA | 4 |
| 2021 | Visual Semantic Localization based on HD Map for Autonomous Vehicles in Urban ScenariosabstractHighly accurate and robust localization ability is of great importance for autonomous vehicles (AVs) in urban scenarios. Traditional vision-based methods suffer from lost due to illumination, weather, viewing and appearance changes. In this paper we propose a novel visual semantic localization algorithm based on HD map and semantic features which are compact in representation. Semantic features are widely appeared on urban roads, and are robust to illumination, weather, viewing and appearance changes. The repeated structures, missed detections and false detections make data association (DA) highly ambiguous. To this end, a robust DA method considering local structural consistency, global pattern consistency and temporal consistency is performed. Further, we introduce a sliding window factor graph optimization framework to fuse association and odometry measurements without the requirements of high-precision absolute height information for map features.We evaluate the proposed localization framework on both simulated and real urban road. The experiments show that the proposed approach is able to achieve highly accurate localization with a mean longitudinal error of 0.43m, a mean lateral error of 0.12m and a mean yaw angle error of 0.11°. Huayou Wang, Changliang Xue, Yanxing Zhou, Hongbo Zhang 0004 |
ICRA | 5 |
| 2021 | SSC: Semantic Scan Context for Large-Scale Place RecognitionabstractPlace recognition gives a SLAM system the ability to correct cumulative errors. Unlike images that contain rich texture features, point clouds are almost pure geometric information which makes place recognition based on point clouds challenging. Existing works usually encode low-level features such as coordinate, normal, reflection intensity, etc., as local or global descriptors to represent scenes. Besides, they often ignore the translation between point clouds when matching descriptors. Different from most existing methods, we explore the use of high-level features, namely semantics, to improve the descriptor’s representation ability. Also, when matching descriptors, we try to correct the translation between point clouds to improve accuracy. Concretely, we propose a novel global descriptor, Semantic Scan Context, which explores semantic information to represent scenes more effectively. We also present a two-step global semantic ICP to obtain the 3D pose (x, y, yaw) used to align the point cloud to improve matching performance. Our experiments on the KITTI dataset show that our approach outperforms the state-of-the- art methods with a large margin. Our code is available at: https://github.com/lilin-hitcrt/SSC. Lin Li 0091, Xin Kong, Xiangrui Zhao, Tianxin Huang, Wanlong Li, Hongbo Zhang 0004, Yong Liu 0007 |
IROS | 7 |
| 2021 | BSP-MonoLoc: Basic Semantic Primitives based Monocular Localization on RoadsabstractRobust visual localization in traffic scenes is a fundamental problem for self-driving vehicles. However, it is still challenging to achieve accurate localization performance because of drastic viewpoint and illumination changes. To address the issues, we design a novel monocular localization framework based on a light-weight prior map, called BSP-MonoLoc, which leverages the 2D semantic primitives from the monocular images and the 3D basic semantic primitives from the prior map. These primitives are commonly available but lack of distinctive signature. To effectively make associations between the 2D and 3D primitives and refine the vehicle’s pose, we adopt an iterative optimization method, where an efficient hierarchical sample strategy is designed to give a good initial prediction for the associations and the pose. Experimental results on the KAIST dataset and our dataset demonstrate the proposed method can achieve high localization accuracy and run at a real-time performance. Heping Li, Changliang Xue, Hongbo Zhang 0004, Wei Gao 0014 |
IROS | 4 |
| 2021 | Semantic Segmentation-assisted Scene Completion for LiDAR Point CloudsabstractOutdoor scene completion is a challenging issue in 3D scene understanding, which plays an important role in intelligent robotics and autonomous driving. Due to the sparsity of LiDAR acquisition, it is far more complex for 3D scene completion and semantic segmentation. Since semantic features can provide constraints and semantic priors for completion tasks, the relationship between them is worth exploring. Therefore, we propose an end-to-end semantic segmentation-assisted scene completion network, including a 2D completion branch and a 3D semantic segmentation branch. Specifically, the network takes a raw point cloud as input, and merges the features from the segmentation branch into the completion branch hierarchically to provide semantic information. By adopting BEV representation and 3D sparse convolution, we can benefit from the lower operand while maintaining effective expression. Besides, the decoder of the segmentation branch is used as an auxiliary, which can be discarded in the inference stage to save computational consumption. Extensive experiments demonstrate that our method achieves competitive performance on SemanticKITTI dataset with low latency. Code and models will be released at https://github.com/jokester-zzz/SSA-SC. Xuemeng Yang, Xin Kong, Tianxin Huang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 8 |
| 2021 | Up-to-Down Network: Fusing Multi-Scale Context for 3D Semantic Scene CompletionabstractAn efficient 3D scene perception algorithm is a vital component for autonomous driving and robotics systems. In this paper, we focus on semantic scene completion, which is a task of jointly estimating the volumetric occupancy and semantic labels of objects. Since the real-world data is sparse and occluded, this is an extremely challenging task. We propose a novel framework, named Up-to-Down network (UDNet), to achieve the large-scale semantic scene completion with an encoder-decoder architecture for voxel grids. The novel up-to-down block can effectively aggregate multi-scale context information to improve labeling coherence, and the atrous spatial pyramid pooling module is leveraged to expand the receptive field while preserving detailed geometric information. Besides, the proposed multi-scale fusion mechanism efficiently aggregates global background information and improves the semantic completion accuracy. Moreover, to further satisfy the needs of different tasks, our UDNet can accomplish the multi-resolution semantic completion, achieving faster but coarser completion. Detailed experiments in the semantic scene completion benchmark of SemanticKITTI illustrate that our proposed framework surpasses the state-of-the-art methods with remarkable margins and a real-time inference speed by using only voxel grids as input. Xuemeng Yang, Tianxin Huang, Chujuan Zhang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 8 |
| 2021 | PointSiamRCNN: Target-aware Voxel-based Siamese Tracker for Point CloudsabstractCurrently, there have been many kinds of pointbased 3D trackers, while voxel-based methods are still underexplored. In this paper, we first propose a voxel-based tracker, named PointSiamRCNN, improving tracking performance by embedding target information into the search region. Our framework is composed of two parts for achieving proposal generation and proposal refinement, which fully releases the potential of the two-stage object tracking. Specifically, it takes advantage of efficient feature learning of the voxel-based Siamese network and high-quality proposal generation of the Siamese region proposal network head. In the search region, the groundtruth annotations are utilized to realize semantic segmentation, which leads to more discriminative feature learning with pointwise supervisions. Furthermore, we propose the Self and Cross Attention Module for embedding target information into the search region. Finally, the multi-scale RoI pooling module is proposed to obtain compact representations from target-aware features for proposal refinement. Exhaustive experiments on the KITTI tracking dataset demonstrate that our framework reaches the competitive performance with the state-of-the-art 3D tracking methods and achieves the state-of-the-art in terms of BEV tracking. Chujuan Zhang, Yong Liu 0007, Wanlong Li, Hongbo Zhang 0004 |
IROS | 6 |