Hui Kong 0001

dblp:94/1836-1 · DBLP profile ↗
← Back
68ranked-venue papers
15as first author
31since 2021 · last 2026
0000-0002-5303-0276ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 9 first-author · 23 since 2021Systems, architecture and hardware · 21 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Efficient Active Training for Deep LiDAR Odometry
abstract
Robust and efficient deep LiDAR odometry models are crucial for accurate localization and 3D reconstruction, but typically require extensive and diverse training data to adapt to diverse environments, leading to inefficiencies. To tackle this, we introduce an active training framework designed to selectively extract training data from diverse environments, thereby reducing the training load and enhancing model generalization. Our framework is based on two key strategies: Initial Training Set Selection (ITSS) and Active Incremental Selection (AIS). ITSS begins by breaking down motion sequences from general weather into nodes and edges for detailed trajectory analysis, prioritizing diverse sequences to form a rich initial training dataset for training the base model. For complex sequences that are difficult to analyze, especially under challenging snowy weather conditions, AIS uses scene reconstruction and prediction inconsistency to iteratively select training samples, refining the model to handle a wide range of real-world scenarios. Experiments across datasets and weather conditions validate our approach’s effectiveness. Notably, our method matches the performance of full-dataset training with just 52% of the sequence volume, demonstrating the training efficiency and robustness of our active training paradigm. By optimizing the training process, our approach sets the stage for more agile and reliable LiDAR odometry systems, capable of navigating diverse environmental conditions with greater precision.
Beibei Zhou, Zhiyuan Zhang 0004, Zhenbo Song, Jianhui Guo, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.5
2025 BHViT: Binarized Hybrid Vision Transformer
abstract
Model binarization has made significant progress in enabling real-time and energy-efficient computation for con-volutional neural networks (CNN), offering a potential solution to the deployment challenges faced by Vision Transformers (ViTs) on edge devices. However, due to the structural differences between CNN and Transformer architectures, simply applying binary CNN strategies to the ViT models will lead to a significant performance drop. To tackle this challenge, we propose BHViT, a binarization-friendly hybrid ViT architecture and its full binarization model with the guidance of three important observations. Initially, BHViT utilizes the local information interaction and hierarchical feature aggregation technique from coarse to fine levels to address redundant computations stemming from excessive tokens. Then, a novel module based on shift operations is proposed to enhance the performance of the binary Multi-Layer Perceptron (MLP) module without significantly increasing computational overhead. In addition, an innovative attention matrix binarization method based on quantization decomposition is proposed to evaluate the token’s importance in the binarized attention matrix. Finally, we propose a regularization loss to address the inadequate optimization caused by the incompatibility between the weight oscillation in the binary layers and the Adam Optimizer. Extensive experimental results demonstrate that our proposed algorithm achieves SOTA performance among binary ViT methods. The source code is released at: https://github.com/IMRL/BHViT.
Tian Gao 0004, Zhiyuan Zhang 0012, Huajun Liu, Kaijie Yin, Cheng-Zhong Xu 0001, Hui Kong 0001
CVPR7
2025 Information-Bottleneck Driven Binary Neural Network for Change Detection
abstract
In this paper, we propose Binarized Change Detection (BiCD), the first binary neural network (BNN) designed specifically for change detection. Conventional network binarization approaches, which directly quantize both weights and activations in change detection models, severely limit the network's ability to represent input data and distinguish between changed and unchanged regions. This results in significantly lower detection accuracy compared to real-valued networks. To overcome these challenges, BiCD enhances both the representational power and feature separability of BNNs, improving detection performance. Specifically, we introduce an auxiliary objective based on the Information Bottleneck (IB) principle, guiding the encoder to retain essential input information while promoting better feature discrimination. Since directly computing mutual information under the IB principle is intractable, we design a compact, learnable auxiliary module as an approximation target, leading to a simple yet effective optimization strategy that minimizes both reconstruction loss and standard change detection loss. Extensive experiments on street-view and remote sensing datasets demonstrate that BiCD establishes a new benchmark for BNN-based change detection, achieving state-of-the-art performance in this domain.
Kaijie Yin, Zhiyuan Zhang 0012, Shu Kong, Tian Gao 0004, Cheng-Zhong Xu 0001, Hui Kong 0001
ICCV6
2025 HDiffTG: A Lightweight Hybrid Diffusion-Transformer-GCN Architecture for 3D Human Pose Estimation
abstract
We propose HDiffTG, a novel 3D Human Pose Estimation (3DHPE) method that integrates Transformer, Graph Convolutional Network (GCN), and diffusion model into a unified framework. HDiffTG leverages the strengths of these techniques to significantly improve pose estimation accuracy and robustness while maintaining a lightweight design. The Transformer captures global spatiotemporal dependencies, the GCN models local skeletal structures, and the diffusion model provides step-by-step optimization for fine-tuning, achieving a complementary balance between global and local features. This integration enhances the model’s ability to handle pose estimation under occlusions and in complex scenarios. Furthermore, we introduce lightweight optimizations to the integrated model and refine the objective function design to reduce computational overhead without compromising performance. Evaluation results on the Human3.6M and MPI-INF-3DHP datasets demonstrate that HDiffTG achieves state-of-the-art (SOTA) performance on the MPI-INF-3DHP dataset while excelling in both accuracy and computational efficiency. Additionally, the model exhibits exceptional robustness in noisy and occluded environments. Source codes and models are available at https://github.com/CirceJie/HDiffTG
Yajie Fu, Chaorui Huang, Junwei Li 0009, Hui Kong 0001, Yibin Tian, Huakang Li, Zhiyuan Zhang 0004
IJCNN4
2025 TS-Diff: Two-Stage Diffusion Model for Low-Light RAW Image Enhancement
abstract
This paper presents a novel Two-Stage Diffusion Model (TS-Diff) for enhancing extremely low-light RAW images. In the pre-training stage, TS-Diff synthesizes noisy images by constructing multiple virtual cameras based on a noise space. Camera Feature Integration (CFI) modules are then designed to enable the model to learn generalizable features across diverse virtual cameras. During the aligning stage, CFIs are averaged to create a target-specific CFIT, which is fine-tuned using a small amount of real RAW data to adapt to the noise characteristics of specific cameras. A structural reparameterization technique further simplifies CFITfor efficient deployment. To address color shifts during the diffusion process, a color corrector is introduced to ensure color consistency by dynamically adjusting global color distributions. Additionally, a novel dataset, QID, is constructed, featuring quantifiable illumination levels and a wide dynamic range, providing a comprehensive benchmark for training and evaluation under extreme low-light conditions. Experimental results demonstrate that TS-Diff achieves state-of-the-art performance on multiple datasets, including QID, SID, and ELD, excelling in denoising, generalization, and color consistency across various cameras and illumination levels. These findings highlight the robustness and versatility of TS-Diff, making it a practical solution for low-light imaging applications. Source codes and models are available at https://github.com/CircccleK/TS-Diff
Zhiyuan Zhang 0004, Jiangnan Xia, Jianghan Cheng, Junwei Li 0009, Yibin Tian, Hui Kong 0001
IJCNN8
2025 PCMF2-Net: A Pyramid Cross-Modal Feature Fusion Network for Off-Road Freespace Detection
abstract
Freespace detection plays an important role in autonomous driving. In recent years, deep learning based freespace detection methods have performed well in urban scenes. However, for off-road scenes, freespace detection poses significant challenges due to the complexity of the scenes and the lack of clear edges. The existing methods have not effectively fused LiDAR data and camera images. In this paper, we propose a Pyramid Cross-Modal Feature Fusion Network (PCMF2-Net) for off-road freespace detection. The dense depth maps are concatenated with RGB images and used as input along with surface normal maps. The dual branch CNN-Transformer encoder combines convolutional neural networks and transformers to extract local and global features from RGBD images and surface normal maps, respectively. Then, in the pyramid cross-modal feature fusion module, the multi-scale and multimodal encoder features are fused in a top-down manner. In addition, we also use an edge segmentation task and a two-step training strategy to further improve performance. Experiments on the off-road freespace detection dataset (ORFD) demonstrate that the proposed PCMF2-Net achieves a competitive result of 93.9% IoU at a speed of 23 Hz.
Chunpeng Lu, Shuo Gu, Yigong Zhang, Hui Kong 0001
IROS6
2025 Pathfinder for Low-altitude Aircraft with Binary Neural Network
abstract
A prior global topological map (e.g., the OpenStreetMap, OSM) can boost the performance of autonomous mapping by a ground mobile robot. However, the prior map is usually incomplete due to lacking labeling in partial paths. To solve this problem, this paper proposes an OSM maker using airborne sensors carried by low-altitude aircraft, where the core of the OSM maker is a novel efficient pathfinder approach based on LiDAR and camera data, i.e., a binary dual-stream road segmentation model. Specifically, a multi-scale feature extraction based on the UNet architecture is implemented for images and point clouds. To reduce the effect caused by the sparsity of point cloud, an attention-guided gated block is designed to integrate image and point-cloud features. To optimize the model for edge deployment that significantly reduces storage footprint and computational demands, we propose a binarization streamline to each model component, including a variant of vision transformer (ViT) architecture as the encoder of the image branch, and new focal and perception losses to optimize the model training. The experimental results on two datasets demonstrate that our pathfinder method achieves SOTA accuracy with high efficiency in finding paths from the low-level airborne sensors, and we can create complete OSM prior maps based on the segmented road skeletons. Code and data are available at: https://github.com/IMRL/Pathfinder.
Kaijie Yin, Tian Gao 0004, Hui Kong 0001
IROS3
2025 Night-Voyager: Consistent and Efficient Nocturnal Vision-Aided State Estimation in Object Maps
abstract
Accurate and robust state estimation at nighttime is essential for autonomous robotic navigation to achieve nocturnal or round-the-clock tasks. An intuitive question arises: can low-cost standard cameras be exploited for nocturnal state estimation? Regrettably, most existing visual methods may fail under adverse illumination conditions, even with active lighting or image enhancement. A pivotal insight, however, is that streetlights in most urban scenarios act as stable and salient prior visual cues at night, reminiscent of stars in deep space aiding spacecraft voyage in interstellar navigation. Inspired by this, we propose Night-Voyager, an object-level nocturnal vision-aided state estimation framework that leverages prior object maps and keypoints for versatile localization. We also find that the primary limitation of conventional visual methods under poor lighting conditions stems from the reliance on pixel-level metrics. In contrast, metric-agnostic, nonpixel-level object detection serves as a bridge between pixel-level and object-level spaces, enabling effective propagation and utilization of object map information within the system. Night-Voyager begins with a fast initialization to solve the global localization problem. By employing an effective two-stage cross-modal data association, the system delivers globally consistent state updates using map-based observations. To address the challenge of significant uncertainties in visual observations at night, a novel matrix Lie group formulation and a feature-decoupled multistate invariant filter are introduced, ensuring consistent and efficient estimation. Through comprehensive experiments in both simulation and diverse real-world scenarios (spanning approximately 12.3 km), Night-Voyager showcases its efficacy, robustness, and efficiency, filling a critical gap in nocturnal vision-aided state estimation.
Tianxiao Gao, Mingle Zhao, Cheng-Zhong Xu 0001, Hui Kong 0001
IEEE Trans. Robotics4
2024 Night-Rider: Nocturnal Vision-aided Localization in Streetlight Maps Using Invariant Extended Kalman Filtering
abstract
Vision-aided localization for low-cost mobile robots in diverse environments has attracted widespread attention recently. Although many current systems are applicable in daytime environments, nocturnal visual localization is still an open problem owing to the lack of stable visual information. An insight from most nocturnal scenes is that the static and bright streetlights are reliable visual information for localization. Hence we propose a nocturnal vision-aided localization system in streetlight maps with a novel data association and matching scheme using object detection methods. We leverage the Invariant Extended Kalman Filter (InEKF) to fuse IMU, odometer, and camera measurements for consistent state estimation at night. Furthermore, a tracking recovery module is also designed for tracking failures. Experimental results indicate that our proposed system achieves accurate and robust localization with less than 0.2% relative error of trajectory length in four nocturnal environments.
Tianxiao Gao, Mingle Zhao, Cheng-Zhong Xu 0001, Hui Kong 0001
ICRA4
2024 Active Loop Closure for OSM-guided Robotic Mapping in Large-Scale Urban Environments
abstract
The autonomous mapping of large-scale urban scenes presents significant challenges for autonomous robots. To mitigate the challenges, global planning, such as utilizing prior GPS trajectories from OpenStreetMap (OSM), is often used to guide the autonomous navigation of robots for mapping. However, due to factors like complex terrain, unexpected body movement, and sensor noise, the uncertainty of the robot’s pose estimates inevitably increases over time, ultimately leading to the failure of robotic mapping. To address this issue, we propose a novel active loop closure procedure, enabling the robot to actively re-plan the previously planned GPS trajectory. The method can guide the robot to re-visit the previous places where the loop-closure detection can be performed to trigger the back-end optimization, effectively reducing errors and uncertainties in pose estimation. The proposed active loop closure mechanism is implemented and embedded into a real-time OSM-guided robot mapping framework. Empirical results on several large-scale outdoor scenarios demonstrate its effectiveness and promising performance.
Zezhou Sun, Mingle Zhao, Cheng-Zhong Xu 0001, Hui Kong 0001
IROS5
2024 UMAD: University of Macau Anomaly Detection Benchmark Dataset
abstract
Anomaly detection is critical in surveillance systems and patrol robots by identifying anomalous regions in images for early warning. Depending on whether reference data are utilized, anomaly detection can be categorized into anomaly detection with reference and anomaly detection without reference. Currently, anomaly detection without reference, which is closely related to out-of-distribution (OoD) object detection, struggles with learning anomalous patterns due to the difficulty of collecting sufficiently large and diverse anomaly datasets with the inherent rarity and novelty of anomalies. Alternatively, anomaly detection with reference employs the scheme of change detection to identify anomalies by comparing semantic changes between a reference image and a query one. However, there are very few ADr works due to the scarcity of public datasets in this domain. In this paper, we aim to address this gap by introducing the UMAD Benchmark Dataset. To our best knowledge, this is the first benchmark dataset designed specifically for anomaly detection with reference in robotic patrolling scenarios, e.g., where an autonomous robot is employed to detect anomalous objects by comparing a reference and a query video sequences. The reference sequences can be taken by the robot along a specified route when there are no anomalous objects in the scene. The query sequences are captured online by the robot when it is patrolling in the same scene following the same route. Our benchmark dataset is elaborated such that each query image can find a corresponding reference based on accurate robot localization along the same route in the pre-built 3D map, with which the reference and query images can be geometrically aligned using adaptive warping. Besides the proposed benchmark dataset, we evaluate the baseline models of ADr on this dataset. We hope this benchmark dataset will facilitate the advancement of ADr methods in the future. Our UMAD benchmark dataset will be publicly accessible at https://github.com/IMRL/UMAD.
Lineng Chen, Cheng-Zhong Xu 0001, Hui Kong 0001
IROS4
2024 VRSO: Visual-Centric Reconstruction for Static Object Annotation
abstract
As a part of the perception results of intelligent driving systems, static object detection (SOD) in 3D space provides crucial cues for driving environment understanding. With the rapid deployment of deep neural networks for SOD tasks, the demand for high-quality training samples soars. The traditional, also reliable, way is manual labelling over the dense LiDAR point clouds and reference images. Though most public driving datasets adopt this strategy to provide SOD ground truth (GT), it is still expensive and time-consuming in practice. This paper introduces VRSO, a visual-centric approach for static object annotation. Experiments on the Waymo Open Dataset show that the mean reprojection error from VRSO annotation is only 2.6 pixels, around four times lower than the Waymo Open Dataset labels (10.6 pixels). VRSO is distinguished in low cost, high efficiency, and high quality: (1) It recovers static objects in 3D space with only camera images as input, and (2) manual annotation is barely involved since GT for SOD tasks is generated based on an automatic reconstruction and annotation pipeline.
Chenyao Yu, Yingfeng Cai, Jiaxin Zhang 0014, Hui Kong 0001, Wei Sui
IROS4
2024 Learning to segment complex vessel-like structures with spectral transformer
Huajun Liu, Jing Yang 0053, Hui Kong 0001, Haofeng Zhang 0001
Expert Syst. Appl.4
2024 GSB: Group superposition binarization for vision transformer with limited training samples
Tian Gao 0004, Cheng-Zhong Xu 0001, Le Zhang 0001, Hui Kong 0001
Neural Networks4
2024 Semi-supervised vanishing point detection with contrastive learning
Shuo Gu, Yinbo Liu, Hui Kong 0001
Pattern Recognit.4
2024 Adaptive Fourier Convolution Network for Road Segmentation in Remote Sensing Images
abstract
Segmentation of roads in remote sensing images is a challenging task due to the inhomogeneous intensity, non-consistent contrast, and very cluttered background in remote sensing images. Recent approaches, mostly relying on convolutions or self-attention, make it difficult to extract weak and continuous road objects. Fourier neural operators provide another novel mechanism for capturing long-range and fine-grained features beyond self-attention. Based on it, we propose an adaptive Fourier convolution network (AFCNet) on the spatial-spectral domain for road segmentation in this paper. The AFCNet is built on the pipeline of the classical U-Net model and its core is the proposed Fourier neural encoder (FNE), which is built on a feed-forward layer and a flexible Fourier convolutional structure composed of Fourier-domain pooling layers, asymmetric convolutions, squeeze-excitation inspired self-attention and adaptive multiscale fusion layers. Furthermore, we combine the FNE and bottleneck in ResNet to form a hybrid global-local feature representation scheme to capture the long and weak road objects in remote sensing images. The experiments on two public datasets, the Massachusetts Roads and DeepGlobe Road Datasets, have shown that AFCNet worked with fewer parameters and outperformed most previous methods in terms of accuracy, precision, recall, and mean intersection over union (mIoU), etc.
Huajun Liu, Cailing Wang, Jinding Zhao, Suting Chen, Hui Kong 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Fourier-Deformable Convolution Network for Road Segmentation From Remote Sensing Images
abstract
Road segmentation from remote sensing images is a challenging task in capturing weak, long, and irregular road features due to the limited connectivity-preserving modeling capability. In this work, we proposed a U-shaped Fourier-deformable convolution network (FDNet) for road segmentation, which integrates the merits of deformable convolutions (DCs) and Fourier convolutions compactly. Specifically, a saliency-aware DC (SD-Conv) layer is proposed for tracing salient road features based on an iterative dynamic offset learning mechanism to grasp extremely tender and weak road objects. Meanwhile, a lightweight global feature extracting module based on spectral convolutions, namely, the adaptive Fourier convolution (AF-Conv) layer, is adopted to learn long-range dependency to extract long and continuous road structures. The proposed SD-Conv layer worked in parallel with the AF-Conv layer to construct a basic and compact block to build the U-shaped FDNet model for road segmentation. Furthermore, to maintain the continuity of road objects in complex road conditions, we introduced a topology-oriented loss function based on the Hausdorff distance (HD) on the persistence diagram (PD) of segmented results, and further combined with softDice loss components for fully supervised training. Our FDNet has been trained and evaluated on two benchmarks, and experimental results show that FDNet achieved state-of-the-art (SOTA) performance. Specifically, it achieved 80.34% on accuracy, 88.42% on precision, and 84.70% on mean intersection over union (mIoU), respectively, on the Massachusetts dataset, and achieved 99.05% on accuracy, 89.21% on precision, 88.61% on recall, and 81.37% on mIoU, respectively, on the DeepGlobe dataset, outperforming most previous methods on both datasets. Codes are available at:https://github.com/zhoucharming/FDNet.
Huajun Liu, Cailing Wang, Suting Chen, Hui Kong 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 LiDAR-SGMOS: Semantics-Guided Moving Object Segmentation with 3D LiDAR
abstract
Most of the existing moving object segmentation (MOS) methods regard MOS as an independent task, in this paper, we associate the MOS task with semantic segmentation, and propose a semantics-guided network for moving object segmentation (LiDAR-SGMOS). We first transform the range image and semantic features of the past scan into the range view of current scan based on the relative pose between scans. The residual image is obtained by calculating the normalized absolute difference between the current and transformed range images. Then, we apply a Meta-Kernel based cross scan fusion (CSF) module to adaptively fuse the range images and semantic features of current scan, the residual image and transformed features. Finally, the fused features with rich motion and semantic information are processed to obtain reliable MOS results. We also introduce a residual image augmentation method to further improve the MOS performance. Our method outperforms most LiDAR-MOS methods with only two sequential LiDAR scans as inputs on the SemanticKITTI MOS dataset.
Shuo Gu, Suling Yao, Jian Yang 0003, Cheng-Zhong Xu 0001, Hui Kong 0001
IROS5
2023 Unsupervised Cross-Spectrum Depth Estimation by Visible-Light and Thermal Cameras
abstract
Cross-spectrum depth estimation aims to provide a reliable depth map under variant-illumination conditions with a pair of dual-spectrum images. It is valuable for autonomous driving applications when vehicles are equipped with two cameras of different modalities. However, images captured by different-modality cameras can be photometrically quite different, which makes cross-spectrum depth estimation a very challenging problem. Moreover, the shortage of large-scale open-source datasets also retards further research in this field. In this paper, we propose an unsupervised visible light(VIS)-image-guided cross-spectrum (i.e., thermal and visible-light, TIR-VIS in short) depth-estimation framework. The input of the framework consists of a cross-spectrum stereo pair (one VIS image and one thermal image). First, we train a depth-estimation base network using VIS-image stereo pairs. To adapt the trained depth-estimation network to the cross-spectrum images, we propose a multi-scale feature-transfer network to transfer features from the TIR domain to the VIS domain at the feature level. Furthermore, we introduce a mechanism of cross-spectrum depth cycle-consistency to improve the depth estimation result of dual-spectrum image pairs. Meanwhile, we release to society a large cross-spectrum dataset with visible-light and thermal stereo images captured in different scenes. The experiment result shows that our method achieves better depth-estimation results than the compared existing methods. Our code and dataset are available onhttps://github.com/whitecrow1027/CrossSP_Depth.
Yubin Guo, Xinlei Qi, Jin Xie 0001, Cheng-Zhong Xu 0001, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.5
2023 CrackFormer Network for Pavement Crack Segmentation
abstract
In this paper, we rethink our earlier work on self-attention based crack segmentation, and propose an upgraded CrackFormer network (CrackFormer-II) for pavement crack segmentation, instead of only for fine-grained crack-detection tasks. This work embeds novel Transformer encoder modules into a SegNet-like encoder-decoder structure, where the basic module is composed of novel Transformer encoder blocks with effective relative positional embedding and long range interactions to extract efficient contextual information from feature-channels. Further, fusion modules of scaling-attention are proposed to integrate the results of each respective encoder and decoder block to highlight semantic features and suppress non-semantic ones. Moreover, we update the Transformer encoder blocks enhanced by the local feed-forward layer and skip-connections, and optimize the channel configurations to compress the model parameters. Compared with the original CrackFormer, the CrackFormer-II is trained and evaluated on more general crack datasets. It achieves higher accuracy than the original CrackFormer, and the state-of-the-art (SOTA) method with$6.7 \times $fewer FLOPs and$6.2 \times $fewer parameters, and its practical inference speed is comparable to most classical CNN models. The experimental results show that it achieves the F-measures on Optimal Dataset Scale (ODS) of 0.912, 0.908, 0.914 and 0.869, respectively, on the four benchmarks. Codes are available athttps://github.com/LouisNUST/CrackFormer-II.
Huajun Liu, Jing Yang 0053, Xiangyu Miao, Christoph Mertz, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.5
2023 Geometry-Aware Network for Unsupervised Learning of Monocular Camera's Ego-Motion
abstract
Deep neural networks have been shown to be effective for unsupervised monocular visual odometry that can predict the camera’s ego-motion based on an input of monocular video sequence. However, most existing unsupervised monocular methods haven’t fully exploited the extracted information from both local geometric structure and visual appearance of the scenes, resulting in degraded performance. In this paper, a novel geometry-aware network is proposed to predict the camera’s ego-motion by learning representations in both 2D and 3D space. First, to extract geometry-aware features, we design an RGB-PointCloud feature fusion module to capture information from both geometric structure and the visual appearance of the scenes by fusing local geometric features from depth-map-derived point clouds and visual features from RGB images. Furthermore, the fusion module can adaptively allocate different weights to the two types of features to emphasize important regions. Then, we devise a relevant feature filtering module to build consistency between the two views and preserve informative features with high relevance. It can capture the correlation of frame pairs in the feature-embedding space by attention mechanisms. Finally, the obtained features are fed into the pose estimator to recover the 6-DoF poses of the camera. Extensive experiments show that our method achieves promising results among the unsupervised monocular deep learning methods on the KITTI odometry and TUM-RGBD datasets.
Beibei Zhou, Jin Xie 0001, Zhong Jin, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.4
2022 A Cylindrical Convolution Network for Dense Top-View Semantic Segmentation with LiDAR Point Clouds
Shuo Gu, Cheng-Zhong Xu 0001, Hui Kong 0001
ACCV (7)4
2022 HiTPR: Hierarchical Transformer for Place Recognition in Point Cloud
abstract
Place recognition or loop closure detection is one of the core components in a full SLAM system. In this paper, aiming at strengthening the relevancy of local neighboring points and the contextual dependency among global points simultaneously, we investigate the exploitation of transformer-based network for feature extraction, and propose a Hierarchical Transformer for Place Recognition (HiTPR). The HiTPR consists of four major parts: point cell generation, short-range transformer (SRT), long-range transformer (LRT) and global descriptor aggregation. Specifically, the point cloud is initially divided into a sequence of small cells by down-sampling and nearest neighbors searching. In the SRT, we extract the local feature for each point cell. While in the LRT, we build the global dependency among all of the point cells in the whole point cloud. Experiments on several standard benchmarks demonstrate the superiority of the HiTPR in terms of average recall rate, achieving 93.71 % at top 1 % and 86.63 % at top 1 on the Oxford RobotCar dataset for example.
Zhixing Hou, Yan Yan 0002, Cheng-Zhong Xu 0001, Hui Kong 0001
ICRA4
2022 Ada-Detector: Adaptive Frontier Detector for Rapid Exploration
abstract
In this paper, we propose an efficient frontier detector method based on adaptive Rapidly-exploring Random Tree (RRT) for autonomous robot exploration. Robots can achieve real-time incremental frontier detection when they are exploring unknown environments. First, our detector adaptively adjusts the sampling space of RRT by sensing the surrounding environment structure. The adaptive sampling space can greatly improve the successful sampling rate of RRT (the ratio of the number of samples successfully added to the RRT tree to the number of sampling attempts) according to the environment structure and control the expansion bias of the RRT. Second, by generating non-uniform distributed samples, our method also solves the over-sampling problem of RRT in the sliding windows, where uniform random sampling causes over-sampling in the overlap area between two adjacent sliding windows. In this way, our detector is more inclined to sample in the latest explored area, which improves the efficiency of frontier detection and achieves incremental detection. We validated our method in three simulated benchmark scenarios. The experimental comparison shows that we reduce the frontier detection runtime by about 40% compared with the SOTA method, DSV Planner.
Zezhou Sun, Banghe Wu, Cheng-Zhong Xu 0001, Hui Kong 0001
ICRA4
2022 Learning Moving-Object Tracking with FMCW LiDAR
abstract
In this paper, we propose a learning-based moving-object tracking method utilizing the newly developed LiDAR sensor, Frequency Modulated Continuous Wave (FMCW) LiDAR. Compared with most existing commercial LiDAR sensors, FMCW LiDAR can provide additional Doppler velocity information to each 3D point of the point clouds. Benefiting from this, we can generate instance labels as ground truth in a semi-automatic manner. Given the labels, we propose a contrastive learning framework, which pulls together the features from the same instance in embedding space and pushes apart the features from different instances, to improve the tracking quality. Extensive experiments are conducted on the recorded driving data, and the results show that our method outperforms the baseline methods by a large margin.
Yi Gu 0005, Hongzhi Cheng, Kafeng Wang, Dejing Dou, Cheng-Zhong Xu 0001, Hui Kong 0001
IROS6
2022 Homography-Based Minimal-Case Relative Pose Estimation With Known Gravity Direction
abstract
In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with known gravity direction. This case is relevant to smart phones, tablets, and other camera-IMU (Inertial measurement unit) systems which have accelerometers to measure the gravity vector. We explore the rank-1 constraint on the difference between the euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and semi-calibrated (unknown focal length) problems. Based on the hidden variable technique, we convert the problems to the polynomial eigenvalue problems, and derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with the existing 6- and 7-point solvers, including results with smart phone images.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Capitalizing on RGB-FIR Hybrid Imaging for Road Detection
abstract
Traditionally, road detection approaches mostly capitalize on RGB images, 3D LiDAR point cloud or their fusion. However, RGB camera is sensitive to light conditions, while LiDAR point cloud is sparse compared with dense image pixels. In this work, a new hybrid image dataset is provided for the task of road detection based on cameras. In this dataset, the hybrid images are acquired by an optically aligned hybrid imaging device, consisting of a far-infrared (FIR) imager and an RGB camera to output pixel-wise registration of thermal and RGB frames. Then we investigate on three methods based on fully convolutional neural network (F-CNN) to demonstrate the advantages by fusing RGB-FIR images in road detection. First, a middle-fusion based model is built, where the output feature maps of encoder branches from RGB and FIR images are directly concatenated into a single-fusion branch as the decoder. Next, the originally discarded layers after fusion operation for both RGB and FIR branches are recovered as the mimic branches to imitate the distributions of the fusion outputs, which constitutes an extended cross model (ECM). Moreover, the outputs of mimic branches at different scales are also used to imitate the corresponding outputs in the fusion branch, called a hierarchical cross model (HCM). The experimental results demonstrate the effectiveness and efficiency of our fusion strategies.
Yigong Zhang, Jin Xie 0001, José M. Álvarez 0004, Cheng-Zhong Xu 0001, Jian Yang 0003, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.6
2021 Globally Optimal Relative Pose Estimation With Gravity Prior
abstract
Smartphones, tablets and camera systems used, e.g., in cars and UAVs, are typically equipped with IMUs (inertial measurement units) that can measure the gravity vector accurately. Using this additional information, the y-axes of the cameras can be aligned, reducing their relative orientation to a single degree-of-freedom. With this assumption, we propose a novel globally optimal solver, minimizing the algebraic error in the least squares sense, to estimate the relative pose in the over-determined case. Based on the epipolar constraint, we convert the optimization problem into solving two polynomials with only two unknowns. Also, a fast solver is proposed using the first-order approximation of the rotation. The proposed solvers are compared with the state-of-the-art ones on four real-world datasets with approx. 50000 image pairs in total. Moreover, we collected a dataset, by a smartphone, consisting of 10933 image pairs, gravity directions and ground truth 3D reconstructions. The source code and dataset are available at https://github.com/yaqding/opt_pose_gravity
Yaqing Ding 0001, Daniel Barath, Jian Yang 0003, Hui Kong 0001, Zuzana Kukelova
CVPR4
2021 CrackFormer: Transformer Network for Fine-Grained Crack Detection
abstract
Cracks are irregular line structures that are of interest in many computer vision applications. Crack detection (e.g., from pavement images) is a challenging task due to intensity in-homogeneity, topology complexity, low contrast and noisy background. The overall crack detection accuracy can be significantly affected by the detection performance on fine-grained cracks. In this work, we propose a Crack Transformer network (CrackFormer) for fine-grained crack detection. The CrackFormer is composed of novel attention modules in a SegNet-like encoder-decoder architecture. Specifically, it consists of novel self-attention modules with 1x1 convolutional kernels for efficient contextual information extraction across feature-channels, and efficient positional embedding to capture large receptive field contextual information for long range interactions. It also introduces new scaling-attention modules to combine outputs from the corresponding encoder and decoder blocks to suppress non-semantic features and sharpen semantic ones. The CrackFormer is trained and evaluated on three classical crack datasets. The experimental results show that the CrackFormer achieves the Optimal Dataset Scale (ODS) values of 0.871, 0.877 and 0.881, respectively, on the three datasets and outperforms the state-of-the-art methods.
Huajun Liu, Xiangyu Miao, Christoph Mertz, Cheng-Zhong Xu 0001, Hui Kong 0001
ICCV5
2021 A general elimination strategy for camera motion estimation
abstract
Camera motion estimation, such as relative pose estimation and absolute pose estimation, are fundamental problems in computer vision and robotics. To obtain the motion parameters, classical methods rely on studying the properties of the geometric matrices, e.g., rotation matrix, essential matrix, homography matrix. The well known five-point algorithm was successfully derived using the singular constraint and trace constraints on the essential matrix. However, finding all the algebraic constraints is not always trivial for some recent problems. In this paper, we propose a simple and general technique to find complete algebraic constraints so that we can derive efficient algorithms. We show that using the quaternion to formulate the rotation matrix we can eliminate any unknowns from the original equations and obtain constraints on the rest of the unknowns based on Gröbner basis. We demonstrate that this approach can be applied to almost all the camera motion estimation and show its improvement compared to the existing methods. Further more, based on this elimination technique, we exploit new constraints for the relative pose estimation with gravity prior, and derive a new globally optimal algorithm to this problem. We compare our algorithm with the state-of-the-art methods on both synthetic and real-world data, and show the benefits including accuracy and efficiency.
Yaqing Ding 0001, Yingna Su, Cheng-Zhong Xu 0001, Jian Yang 0003, Hui Kong 0001
ICRA5
2021 A Cascaded LiDAR-Camera Fusion Network for Road Detection
abstract
Most of the existing road detection methods are either single-modal based, e.g., based on LiDAR or camera, or multi-modal based with LiDAR-camera fusion. The algorithms are designed for a specific data type, and cannot cope with input data changes. In addition, the LiDAR-camera based methods can only work in day time with enough light. In this paper, we develop a novel LiDAR-camera fusion strategy, which combines the LiDAR point clouds and the camera images in a cascaded way. The proposed network has two working modes, the single-modal mode with LiDAR point clouds only and the multimodal mode with both LiDAR and camera data, so it can be used in all day scenes. The whole network consists of three parts: 1) LiDAR segmentation module, which segments road points in the LiDAR’s imagery view. 2) Sparse-to-dense module, which upsamples the sparse LiDAR feature maps to dense road detection results. 3) LiDAR-camera fusion module, which fuses the dense LiDAR feature maps with the dense camera images to obtain accurate road estimations. Experiments on the KITTI-Road dataset show that the proposed cascaded LiDAR-camera fusion network can obtain very competitive road detection performance, with a MaxF value of 96.38%, and achieve the state-of-the-art in the single-modal mode among all LiDAR-only methods.
Shuo Gu, Jian Yang 0003, Hui Kong 0001
ICRA3
2020 Hierarchical Knowledge Squeezed Adversarial Network Compression
abstract
Deep network compression has been achieved notable progress via knowledge distillation, where a teacher-student learning manner is adopted by using predetermined loss. Recently, more focuses have been transferred to employ the adversarial training to minimize the discrepancy between distributions of output from two networks. However, they always emphasize on result-oriented learning while neglecting the scheme of process-oriented learning, leading to the loss of rich information contained in the whole network pipeline. Whereas in other (non GAN-based) process-oriented methods, the knowledge have usually been transferred in a redundant manner. Observing that, the small network can not perfectly mimic a large one due to the huge gap of network scale, we propose a knowledge transfer method, involving effective intermediate supervision, under the adversarial training framework to learn the student network. Different from the other intermediate supervision methods, we design the knowledge representation in a compact form by introducing a task-driven attention mechanism. Meanwhile, to improve the representation capability of the attention-based method, a hierarchical structure is utilized so that powerful but highly squeezed knowledge is realized and the knowledge from teacher network could accommodate the size of student network. Extensive experimental results on three typical benchmark datasets, i.e., CIFAR-10, CIFAR-100, and ImageNet, demonstrate that our method achieves highly superior performances against state-of-the-art methods.
Yan Qu, Hui Kong 0001
AAAI5
2020 Minimal Solutions to Relative Pose Estimation From Two Views Sharing a Common Direction With Unknown Focal Length
abstract
We propose minimal solutions to relative pose estimation problem from two views sharing a common direction with unknown focal length. This is relevant for cameras equipped with an IMU (inertial measurement unit), e.g., smart phones, tablets. Similar to the 6-point algorithm for two cameras with unknown but equal focal lengths and 7-point algorithm for two cameras with different and unknown focal lengths, we derive new 4- and 5-point algorithms for these two cases, respectively. The proposed algorithms can cope with coplanar points, which is a degenerate configuration for these 6- and 7-point counterparts. We present a detailed analysis and comparisons with the state of the art. Experimental results on both synthetic data and real images from a smart phone demonstrate the usefulness of the proposed algorithms.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
CVPR4
2020 A two-step approach to Lidar-Camera calibration
abstract
Autonomous vehicles and robots are typically equipped with Lidar and camera. Hence, calibrating the Lidar-camera system is of extreme importance for ego-motion estimation and scene understanding. In this paper, we propose a two-step approach (coarse + fine) for the external calibration between a camera and a multiple-line Lidar. First, a new closed-form solution is proposed to obtain the initial calibration parameters. We compare our solution with the state-of-the-art SVD-based algorithm, and show the benefits of both the efficiency and stability. With the initial calibration parameters, the ICP-based calibration framework is used to register the point clouds which extracted from the camera and Lidar coordinate frames, respectively. Our method has been applied to two Lidar-camera systems: an HDL-64E Lidar-camera system, and a VLP-16 Lidar-camera system. Experimental results demonstrate that our method achieves promising performance and higher accuracy than two open-source methods.
Yingna Su, Yaqing Ding 0001, Jian Yang 0003, Hui Kong 0001
ICPR4
2020 An efficient solution to the relative pose estimation with a common direction
abstract
In this paper, we propose an efficient solution to the calibrated camera motion estimation with a common direction. This case is relevant to smart phones, tablets, and other camera-IMU (Inertial measurement unit) systems, which have accelerometers to measure the gravity direction. We can align one of the axes of the camera with this common direction so that the relative rotation between the views reduces to only 1DOF (degree of freedom). This allows us to use only three point correspondences for relative pose estimation. Unlike previous work, we derive new constraints on the simplified essential matrix using an elimination strategy based on Gröbner basis. In this case, computing the coefficients of these constraints require less computation and we only need to solve a polynomial eigenvalue problem. We show detailed analyses and comparisons against the existing 3-point algorithms, with satisfactory results obtained.
Yaqing Ding 0001, Jian Yang 0003, Hui Kong 0001
ICRA3
2020 Cascaded Non-local Neural Network for Point Cloud Semantic Segmentation
abstract
In this paper, we propose a cascaded non-local neural network for point cloud segmentation. The proposed network aims to build the long-range dependencies of point clouds for the accurate segmentation. Specifically, we develop a novel cascaded non-local module, which consists of the neighborhood-level, superpoint-level and global-level non-local blocks. First, in the neighborhood-level block, we extract the local features of the centroid points of point clouds by assigning different weights to the neighboring points. The extracted local features of the centroid points are then used to encode the superpoint-level block with the non-local operation. Finally, the global-level block aggregates the non-local features of the superpoints for semantic segmentation in an encoder-decoder framework. Benefiting from the cascaded structure, geometric structure information of different neighborhoods with the same label can be propagated. In addition, the cascaded structure can largely reduce the computational cost of the original non-local operation on point clouds. Experiments on different indoor and outdoor datasets show that our method achieves state-of-the-art performance and effectively reduces the time consumption and memory occupation.
Mingmei Cheng, Le Hui, Jin Xie 0001, Jian Yang 0003, Hui Kong 0001
IROS5
2020 Frontier Detection and Reachability Analysis for Efficient 2D Graph-SLAM Based Active Exploration
abstract
We propose an integrated approach to active exploration by exploiting the Cartographer method as the base SLAM module for submap creation and performing efficient frontier detection in the geometrically co-aligned submaps induced by graph optimization. We also carry out analysis on the reachability of frontiers and their clusters to ensure that the detected frontier can be reached by robot. Our method is tested on a mobile robot in real indoor scene to demonstrate the effectiveness and efficiency of our approach.
Zezhou Sun, Banghe Wu, Cheng-Zhong Xu 0001, Sanjay E. Sarma, Jian Yang 0003, Hui Kong 0001
IROS6
2020 LiDAR Iris for Loop-Closure Detection
abstract
In this paper, a global descriptor for a LiDAR point cloud, called LiDAR Iris, is proposed for fast and accurate loop-closure detection. A binary signature image can be obtained for each point cloud after several LoG-Gabor filtering and thresholding operations on the LiDAR-Iris image representation. Given two point clouds, their similarities can be calculated as the Hamming distance of two corresponding binary signature images extracted from the two point clouds, respectively. Our LiDAR-Iris method can achieve a pose-invariant loop-closure detection at a descriptor level with the Fourier transform of the LiDAR-Iris representation if assuming a 3D (x,y,yaw) pose space, although our method can generally be applied to a 6D pose space by re-aligning point clouds with an additional IMU sensor. Experimental results on five road-scene sequences demonstrate its excellent performance in loop-closure detection.
Ying Wang 0007, Zezhou Sun, Cheng-Zhong Xu 0001, Sanjay E. Sarma, Jian Yang 0003, Hui Kong 0001
IROS6
2019 An Efficient Solution to the Homography-Based Relative Pose Problem With a Common Reference Direction
abstract
In this paper, we propose a novel approach to two-view minimal-case relative pose problems based on homography with a common reference direction. We explore the rank-1 constraint on the difference between the Euclidean homography matrix and the corresponding rotation, and propose an efficient two-step solution for solving both the calibrated and partially calibrated (unknown focal length) problems. We derive new 3.5-point, 3.5-point, 4-point solvers for two cameras such that the two focal lengths are unknown but equal, one of them is unknown, and both are unknown and possibly different, respectively. We present detailed analyses and comparisons with existing 6 and 7-point solvers, including results with smart phone images.
Yaqing Ding 0001, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICCV4
2019 Road Detection through CRF based LiDAR-Camera Fusion
abstract
In this paper, we propose a road detection method with LiDAR-camera fusion in a novel conditional random field (CRF) framework to exploit both range and color information. In the LiDAR based part, a fast height-difference based scanning strategy is applied in the 2D LiDAR range-image domain and a dense road detection result in camera image domain can be obtained through geometric upsampling given the LiDAR-camera calibration parameters. In the camera based part, a fully convolutional network is applied in the camera image domain. Finally, we fuse the dense and binary road detection results from both LiDAR and camera in a single CRF framework. Experiments show that using a single thread of CPU, the proposed LiDAR based part can operate at a frequency of over 250Hz with sparse output in range image and 40Hz with dense result in camera image for the 64-beam Velodyne scanner. Our CRF fusion method achieves very promising road detection performance on the KITTI-Road dataset.
Shuo Gu, Yigong Zhang, Jinhui Tang 0001, Jian Yang 0003, Hui Kong 0001
ICRA5
2019 Build your own hybrid thermal/EO camera for autonomous vehicle
abstract
In this work, we propose a novel paradigm to design a hybrid thermal/EO (Electro-Optical or visible-light) camera, whose thermal and RGB frames are pixel-wisely aligned and temporally synchronized. Compared with the existing schemes, we innovate in three ways in order to make it more compact in dimension, and thus more practical and extendable for real-world applications. The first is a redesign of the structure layout of the thermal and EO cameras. The second is on obtaining a pixel-wise spatial registration of the thermal and RGB frames by a coarse mechanical adjustment and a fine alignment through a constant homography warping. The third innovation is on extending one single hybrid camera to a hybrid camera array, through which we can obtain wide-view spatially aligned thermal, RGB and disparity images simultaneously. The experimental results show that the average error of spatial-alignment of two image modalities can be less than one pixel.
Yigong Zhang, Shuo Gu, Yubin Guo, Minghao Liu 0003, Zezhou Sun, Zhixing Hou, Ying Wang 0007, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICRA12
2019 Two-View Fusion based Convolutional Neural Network for Urban Road Detection
abstract
In this paper, we propose a two-view fusion based convolutional neural network to estimate road areas in urban environments with LiDAR point clouds as input only. The proposed network takes two transformed LiDAR data representations, the LiDAR imageries and the camera-perspective maps, as inputs. It outputs pixel-wise road detection results in both the LiDAR's imagery view and the camera's perspective view simultaneously, in an end-to-end manner. To make better use of the data associations between two representations, we construct a novel mapping layer to transform features from the LiDAR's imagery view to the camera's perspective view in order to strengthen the road detection performance in the camera's perspective view. Experiments on the KITTI-Road dataset show that the proposed network can achieve the state-of-the-art performance among all LiDAR-only methods in real time.
Shuo Gu, Yigong Zhang, Jian Yang 0003, José M. Álvarez 0004, Hui Kong 0001
IROS5
2019 Histograms of the Normalized Inverse Depth and Line Scanning for Urban Road Detection
abstract
In this paper, we propose to fuse the geometric information of a 3-D LiDAR and a monocular camera to detect the urban road region ahead of an autonomous vehicle. Our method takes advantage of both the high definition of 3-D LiDAR data and the continuity of road in image representation. First, we obtain an efficient representation of LiDAR data and an organized 2-D inverse depth map, by projecting the 3-D LiDAR points onto the camera's image plane. Through the new representation, we can acquire the intermediate representations of road scenes by extracting the vertical and horizontal histograms of the normalized inverse depth. The approximate road regions can be quickly estimated with both histogram-based schemes. To accurately find the road area, we propose a row and column scanning strategy in the approximate road region to refine the detected road area. We have carried out experiments on the public KITTI-Road benchmark, and have achieved one of the best performances among the LiDAR-based road detection methods without learning procedure.
Shuo Gu, Yigong Zhang, Xia Yuan, Jian Yang 0003, Tao Wu 0001, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.6
2018 Visual Odometry for Indoor Mobile Robot by Recognizing Local Manhattan Structures
Zhixing Hou, Yaqing Ding 0001, Ying Wang 0007, Hui Kong 0001
ACCV (5)5
2018 Dijkstra Model for Stereo-Vision Based Road Detection: A Non-Parametric Method
abstract
This paper proposes a new method for detecting a road from a stereo pair of images. First, the horizon is accurately estimated by a robust, weighted-sampling RANSAC-like method in the improved v-disparity map. The vanishing point of the road region is located using both the horizon information and road flatness constraints. Then it is used as the source node of a weighted graph formed by the pixels of the left stereo-image and their adjacency relationships. The weight of each edge measures the inconsistency of adjacent pixels, and is computed using both the gray-scale and disparity information. Detecting road borders is thus reduced to finding two shortest paths from the source node to the bottom row of the image by the Dijkstra algorithm. The proposed method has been tested on 2621 image pairs of different road scenes from the KITTI dataset. Our experiments demonstrate that this training free approach detects horizon, vanishing point, and road region accurately and robustly, and compares favorably with the state of the art on the KITTI benchmark.
Yigong Zhang, Jian Yang 0003, Jean Ponce, Hui Kong 0001
ICRA4
2018 Fusion of LiDAR and Camera by Scanning in LiDAR Imagery and Image-Guided Diffusion for Urban Road Detection
abstract
This paper proposes a new method for road detection based on a 3D LiDAR and a camera. First, the original LiDAR point cloud is re-organized in an ordered way to generate a LiDAR imagery. Then the flat region is extracted from the LiDAR imagery as the candidate road region. Next, a strategy of row- and column- scanning is given in the LiDAR imagery to detect a finer road region from the candidate region. To fuse the point cloud with image information, we transform the point cloud that corresponds to the above detected road region to the image space according to the calibration parameters between the LiDAR and camera. Then we give two image-guided diffusion schemes to conduct image segmentation of road area, respectively. Our experiments demonstrate that this training free approach detects the road region fast, accurately and robustly, and compares favorably with the state-of-the-art on the KITTI benchmark.
Yigong Zhang, Shuo Gu, Jian Yang 0003, José M. Álvarez 0004, Hui Kong 0001
Intelligent Vehicles Symposium5
2018 When Dijkstra Meets Vanishing Point: A Stereo Vision Approach for Road Detection
abstract
In this paper, we propose a vanishing-point constrained Dijkstra road model for road detection in a stereo-vision paradigm. First, the stereo-camera is used to generate the u- and v-disparity maps of road image, from which the horizon can be extracted. With the horizon and ground region constraints, we can robustly locate the vanishing point of road region. Second, a weighted graph is constructed using all pixels of the image, and the detected vanishing point is treated as the source node of the graph. By computing a vanishing-point constrained Dijkstra minimum-cost map, where both disparity and gradient of gray image are used to calculate cost between two neighbor pixels, the problem of detecting road borders in image is transformed into that of finding two shortest paths that originate from the vanishing point to two pixels in the last row of image. The proposed approach has been implemented and tested over 2600 grayscale images of different road scenes in the KITTI data set. The experimental results demonstrate that this training-free approach can detect horizon, vanishing point, and road regions very accurately and robustly. It can achieve promising performance.
Yigong Zhang, Yingna Su, Jian Yang 0003, Jean Ponce, Hui Kong 0001
IEEE Trans. Image Process.5
2018 Vanishing Point Constrained Lane Detection With a Stereo Camera
abstract
In this paper, we propose a robust vanishing-point constrained lane detection method with a stereo-rig. This method can achieve promising detection performance for both straight and curved lanes without assuming any parametric lane model. First, we propose an accurate and efficient road vanishing point detection scheme based on the v-disparity and visual odometry techniques, where the v-disparity map can significantly reduce the searching space for vanishing point, and the visual odometry can benefit the vanishing point detection of both straight and curved roads. Next, we formulate the lane detection problem as a graph-search procedure, where a vanishing-point constrained Dijkstra shortest-path lane model is proposed to obtain a minimum-cost map. The two lane borders can be detected by finding two optimal paths which originate from the vanishing point to two cost-map derived terminal points, respectively. The proposed method has been tested on the KITTI and the Oxford RobotCar data sets and it works accurately and robustly on a variety of road scenes.
Yingna Su, Yigong Zhang, Jian Yang 0003, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.5
2017 Lidar-histogram for fast road and obstacle detection
abstract
Detection of traversable road regions, positive and negative obstacles, and water hazards is a fundamental task for autonomous driving vehicles, especially in off-road environment. This paper proposes an efficient method, called Lidar-histogram. It can be used to integrate the detection of traversable road regions, obstacles and water hazards into one single framework. The weak assumption of the Lidar-histogram is that a decent-sized area in front of the vehicle is flat. The Lidar-histogram is derived from an efficient organized map of Lidar point cloud, called Lidar-imagery, to index, describe and store Lidar data. The organized point-cloud map can be easily obtained by indexing the original unordered 3D point cloud to a Lidar-specific 2D coordinate system. In the Lidar-histogram representation, the 3D traversable road plane in front of vehicle can be projected as a straight line segment, and the positive and negative obstacles are projected above and below the line segment, respectively. In this way, the problem of detecting traversable road and obstacles is converted into a simple linear classification task in 2D space. Experiments have been conducted in different kinds of off-road and urban scenes, and we have obtained very promising results.
Liang Chen 0003, Jian Yang 0003, Hui Kong 0001
ICRA3
2017 Accurate and Efficient Inspection of Speckle and Scratch Defects on Surfaces of Planar Products
abstract
We propose a unified framework for detecting defects in planar industrial products or planar surfaces of nonplanar products based on a template-matching strategy. The framework includes three parts: an automatic selection of template image for a given test one, a robust geometric alignment between template and test images based on an approximate maximum clique approach, and an illumination invariant image comparison method for defect detection in the aligned images. Experimental results on challenging image datasets demonstrate the excellent performance of the proposed framework.
Hui Kong 0001, Jian Yang 0003
IEEE Trans. Ind. Informatics1
2017 On Detecting Road Regions in a Single UAV Image
abstract
Automatic detection of road regions in aerial images remains a challenging research topic. Most existing approaches work well on the requirement of users to provide some seedlike points/strokes in the road area as the initial location of road regions, or detecting particular roads such as well-paved roads or straight roads. This paper presents a fully automatic approach that can detect generic roads from a single unmanned aerial vehicles (UAV) image. The proposed method consists of two major components: automatic generation of road/nonroad seeds and seeded segmentation of road areas. To know where roads probably are (i.e., road seeds), a distinct road feature is proposed based on the stroke width transformation (SWT) of road image. To the best of our knowledge, it is the first time to introduce SWT as road features, which show the effectiveness on capturing road areas in images in our experiments. Different road features, including the SWT-based geometry information, colors, and width, are then combined to classify road candidates. Based on the candidates, a Gaussian mixture model is built to produce road seeds and background seeds. Finally, starting from these road and background seeds, a convex active contour model segmentation is proposed to extract whole road regions. Experimental results on varieties of UAV images demonstrate the effectiveness of the proposed method. Comparison with existing techniques shows the robustness and accuracy of our method to different roads.
Hailing Zhou, Hui Kong 0001, Lei Wei 0002, Douglas C. Creighton, Saeid Nahavandi
IEEE Trans. Intell. Transp. Syst.2
2016 Maximum clique based RGB-D visual odometry
abstract
In this paper, we propose a new feature-point based RGB-D visual odometry approach for estimating the relative camera motion from two consecutive frames. The approach differs from most feature-point based RGB-D visual odometry approaches in two key aspects: (1) we do not directly use point correspondences to compute relative motion, instead, we link each two distinct points to form a line segment, then utilize correspondences of the generated line segments to estimate relative motion; (2) considering the measurement noise of the RGB-D camera, we design a threshold technique to control the size of maximum clique. Several experiments on real-world dataset show that our method achieved improved accuracy when compared with other recent RGB-D based odometry methods.
Yigong Zhang, Zhixing Hou, Jian Yang 0003, Hui Kong 0001
ICPR4
2015 Efficient Road Detection and Tracking for Unmanned Aerial Vehicle
abstract
An unmanned aerial vehicle (UAV) has many applications in a variety of fields. Detection and tracking of a specific road in UAV videos play an important role in automatic UAV navigation, traffic monitoring, and ground-vehicle tracking, and also is very helpful for constructing road networks for modeling and simulation. In this paper, an efficient road detection and tracking framework in UAV videos is proposed. In particular, a graph-cut-based detection approach is given to accurately extract a specified road region during the initialization stage and in the middle of tracking process, and a fast homography-based road-tracking scheme is developed to automatically track road areas. The high efficiency of our framework is attributed to two aspects: the road detection is performed only when it is necessary and most work in locating the road is rapidly done via very fast homography-based tracking. Experiments are conducted on UAV videos of real road scenes we captured and downloaded from the Internet. The promising results indicate the effectiveness of our proposed framework, with the precision of 98.4% and processing 34 frames per second for 1046 × 595 videos on average.
Hailing Zhou, Hui Kong 0001, Lei Wei 0002, Douglas C. Creighton, Saeid Nahavandi
IEEE Trans. Intell. Transp. Syst.2
2014 Fast road detection and tracking in aerial videos
abstract
We propose a fast approach for detecting and tracking a specific road in aerial videos. It combines adaptive Gaussian Mixture Models (GMMs) to describe road colour distributions, and homography based tracking to track road geometries, where an efficient technique is developed to estimate homography transformations between two frames. Experiments are conducted on videos captured by our unmanned aerial vehicles. All the results demonstrate the effectiveness of our proposed method. We test 1755 frames from 5 videos. Our approach can achieve 0.032 seconds per frame and 2.64% segmentation error for images with 908 × 513 resolutions, on average.
Hailing Zhou, Hui Kong 0001, José M. Álvarez 0004, Douglas C. Creighton, Saeid Nahavandi
Intelligent Vehicles Symposium2
2013 A Generalized Laplacian of Gaussian Filter for Blob Detection and Its Applications
abstract
In this paper, we propose a generalized Laplacian of Gaussian (LoG) (gLoG) filter for detecting general elliptical blob structures in images. The gLoG filter can not only accurately locate the blob centers but also estimate the scales, shapes, and orientations of the detected blobs. These functions can be realized by generalizing the common 3-D LoG scale-space blob detector to a 5-D gLoG scale-space one, where the five parameters are image-domain coordinates (x, y), scales (σ(x), σ(y)), and orientation (θ), respectively. Instead of searching the local extrema of the image's 5-D gLoG scale space for locating blobs, a more feasible solution is given by locating the local maxima of an intermediate map, which is obtained by aggregating the log-scale-normalized convolution responses of each individual gLoG filter. The proposed gLoG-based blob detector is applied to both biomedical images and natural ones such as general road-scene images. For the biomedical applications on pathological and fluorescent microscopic images, the gLoG blob detector can accurately detect the centers and estimate the sizes and orientations of cell nuclei. These centers are utilized as markers for a watershed-based touching-cell splitting method to split touching nuclei and counting cells in segmentation-free images. For the application on road images, the proposed detector can produce promising estimation of texture orientations, achieving an accurate texture-based road vanishing point detection method. The implementation of our method is quite straightforward due to a very small number of tunable parameters.
Hui Kong 0001, Hatice Çinar Akakin, Sanjay E. Sarma
IEEE Trans. Cybern.1
2013 Generalizing Laplacian of Gaussian Filters for Vanishing-Point Detection
abstract
We propose a framework for road-vanishing-point detection based on a new generalized Laplacian of Gaussian (gLoG) filter. In the first part, the gLoG filter can be applied to estimate the texture orientation at each pixel of an image, and the road vanishing point can be detected based on the estimated texture orientations. However, such a texture-based road-vanishing-point detection scheme suffers from high computational complexity. In the second part, an efficient gLoG-based road-vanishing-point detection method is proposed by only using the dominant texture orientations estimated at a sparse set of salient microblob road regions, where the gLoG filter is used to detect these salient microblob areas and simultaneously estimate their dominant texture orientations. Experimental results on 1003 general road images show that the efficient gLoG-based method is significantly faster than a Gabor-filter-based method, whereas the detection accuracy is comparable. The nonefficient gLoG-based method is more accurate in detecting the vanishing point than the Gabor-based approach.
Hui Kong 0001, Sanjay E. Sarma
IEEE Trans. Intell. Transp. Syst.1
2011 Partitioning Histopathological Images: An Integrated Framework for Supervised Color-Texture Segmentation and Cell Splitting
abstract
For quantitative analysis of histopathological images, such as the lymphoma grading systems, quantification of features is usually carried out on single cells before categorizing them by classification algorithms. To this end, we propose an integrated framework consisting of a novel supervised cell-image segmentation algorithm and a new touching-cell splitting method. For the segmentation part, we segment the cell regions from the other areas by classifying the image pixels into either cell or extra-cellular category. Instead of using pixel color intensities, the color-texture extracted at the local neighborhood of each pixel is utilized as the input to our classification algorithm. The color-texture at each pixel is extracted by local Fourier transform (LFT) from a new color space, the most discriminant color space (MDC). The MDC color space is optimized to be a linear combination of the original RGB color space so that the extracted LFT texture features in the MDC color space can achieve most discrimination in terms of classification (segmentation) performance. To speed up the texture feature extraction process, we develop an efficient LFT extraction algorithm based on image shifting and image integral. For the splitting part, given a connected component of the segmentation map, we initially differentiate whether it is a touching-cell clump or a single nontouching cell. The differentiation is mainly based on the distance between the most likely radial-symmetry center and the geometrical center of the connected component. The boundaries of touching-cell clumps are smoothed out by Fourier shape descriptor before carrying out an iterative, concave-point and radial-symmetry based splitting algorithm. To test the validity, effectiveness and efficiency of the framework, it is applied to follicular lymphoma pathological images, which exhibit complex background and extracellular texture with nonuniform illumination condition. For comparison purposes, the results of the proposed segmentation algorithm are evaluated against the outputs of superpixel, graph-cut, mean-shift, and two state-of-the-art pathological image segmentation methods using ground-truth that was established by manual segmentation of cells in the original images. Our segmentation algorithm achieves better results than the other compared methods. The results of splitting are evaluated in terms of under-splitting, over-splitting, and encroachment errors. By summing up the three types of errors, we achieve a total error rate of 5.25% per image.
Hui Kong 0001, Metin Nafi Gürcan, Kamel Belkacem-Boussaid
IEEE Trans. Medical Imaging1
2010 Detecting Abandoned Objects With a Moving Camera
abstract
This paper presents a novel framework for detecting nonflat abandoned objects by matching a reference and a target video sequences. The reference video is taken by a moving camera when there is no suspicious object in the scene. The target video is taken by a camera following the same route and may contain extra objects. The objective is to find these objects. GPS information is used to roughly align the two videos and find the corresponding frame pairs. Based upon the GPS alignment, four simple but effective ideas are proposed to achieve the objective: an intersequence geometric alignment based upon homographies, which is computed by a modified RANSAC, to find all possible suspicious areas, an intrasequence geometric alignment to remove false alarms caused by high objects, a local appearance comparison between two aligned intrasequence frames to remove false alarms in flat areas, and a temporal filtering step to confirm the existence of suspicious objects. Experiments on fifteen pairs of videos show the promise of the proposed method.
Hui Kong 0001, Jean-Yves Audibert, Jean Ponce
IEEE Trans. Image Process.1
2010 General Road Detection From a Single Image
abstract
Given a single image of an arbitrary road, that may not be well-paved, or have clearly delineated edges, or some a priori known color or texture distribution, is it possible for a computer to find this road? This paper addresses this question by decomposing the road detection process into two steps: the estimation of the vanishing point associated with the main (straight) part of the road, followed by the segmentation of the corresponding road area based upon the detected vanishing point. The main technical contributions of the proposed approach are a novel adaptive soft voting scheme based upon a local voting region using high-confidence voters, whose texture orientations are computed using Gabor filters, and a new vanishing-point-constrained edge detection technique for detecting road boundaries. The proposed method has been implemented, and experiments with 1003 general road images demonstrate that it is effective at detecting road regions in challenging conditions.
Hui Kong 0001, Jean-Yves Audibert, Jean Ponce
IEEE Trans. Image Process.1
2009 Vanishing point detection for road detection
abstract
Given a single image of an arbitrary road, that may not be well-paved, or have clearly delineated edges, or some a priori known color or texture distribution, is it possible for a computer to find this road? This paper addresses this question by decomposing the road detection process into two steps: the estimation of the vanishing point associated with the main (straight) part of the road, followed by the segmentation of the corresponding road area based on the detected vanishing point. The main technical contributions of the proposed approach are a novel adaptive soft voting scheme based on variable-sized voting region using confidence-weighted Gabor filters, which compute the dominant texture orientation at each pixel, and a new vanishing-point-constrained edge detection technique for detecting road boundaries. The proposed method has been implemented, and experiments with 1003 general road images demonstrate that it is both computationally efficient and effective at detecting road regions in challenging conditions.
Hui Kong 0001, Jean-Yves Audibert, Jean Ponce
CVPR1
2006 Coupling Adaboost and Random Subspace for Diversified Fisher Linear Discriminant
abstract
Fisher linear discriminant (FLD) is a popular method for feature extraction in face recognition. However, It often suffers from the small sample size, bias and overfitting problems when dealing with the high dimensional face image data. In this paper, a framework of ensemble learning for diversified Fisher linear discriminant (EnL - DFLD) is proposed to improve the current FLD based face recognition algorithms. Firstly, the classifier ensemble in EnL - DFLD is composed of a set of diversified component FLD classifiers, which are selected intentionally by computing the diversity between the candidate component classifiers. Secondly, the candidate component classifiers are constructed by coupling the random subspace and adaboost methods, and it can also be shown that such a coupling scheme will result in more suitable component classifiers so as to increase the generalization performance of EnL - DFLD. Experiments on two common face databases verify the superiority of the proposed EnL - DFLD over the state-of-the-art algorithms in recognition accuracy
Hui Kong 0001, Eam Khwang Teoh
ICARCV1
2006 Coupling Adaboost and Random Subspace for Diversified Fisher Linear Discriminant
Hui Kong 0001
ICONIP (1)1
2005 Discriminant Low-dimensional Subspace Analysis for Face Recognition with Small Number of Training Samples
abstract
In this paper, a framework of Discriminant Low-dimensional Subspace Analysis (DLSA) method is proposed to deal with the Small Sample Size (SSS) problem in face recognition area. Firstly, it is rigorously proven that the null space of the total covariance matrix, S t, is useless for recognition. Therefore, a framework of Fisher discriminant analysis in a low-dimensional space is developed by projecting all the samples onto the range space of S t. Two algorithms are proposed in this framework, i.e., Unified Linear Discriminant Analysis (ULDA) and Modified Linear Discriminant Analysis (MLDA). The ULDA extracts discriminant information from three subspaces of this lowdimensional space. The MLDA adopts a modified Fisher criterion which can avoid the singularity problem in conventional LDA. Experimental results on a large combined database have demonstrated that the proposed ULDA and MLDA can both achieve better performance than the other state-of-the-art LDA-based algorithms in recognition accuracy. 1
Hui Kong 0001, Xuchun Li, Jian-Gang Wang 0001, Eam Khwang Teoh, Chandra Kambhamettu
BMVC1
2005 Generalized 2D Fisher Discriminant Analysis
abstract
To solve the Small Sample Size (SSS) problem, the recent linear discriminant analysis using the 2D matrix-based data representation model has demonstrated its superiority over that using the conventional vector-based data representation model in face recognition [7]. But the explicit reason why the matrix-based model is better than vectorized model has not been given until now. In this paper, a framework of Generalized 2D Fisher Discriminant Analysis (G2DFDA) is proposed. Three contributions are included in this framework: 1) the essence of these ’2D ’ methods is analyzed and their relationships with conventional ’1D ’ methods are given, 2) a Bilateral and 3) a Kernel-based 2D Fisher Discriminant Analysis methods are proposed. Extensive experiment results show its excellent performance. 1
Hui Kong 0001, Jian-Gang Wang 0001, Eam Khwang Teoh, Chandra Kambhamettu
BMVC1
2005 A Framework of 2D Fisher Discriminant Analysis: Application to Face Recognition with Small Number of Training Samples
abstract
A novel framework called 2D Fisher discriminant analysis (2D-FDA) is proposed to deal with the small sample size (SSS) problem in conventional one-dimensional linear discriminant analysis (1D-LDA). Different from the 1D-LDA based approaches, 2D-FDA is based on 2D image matrices rather than column vectors so the image matrix does not need to be transformed into a long vector before feature extraction. The advantage arising in this way is that the SSS problem does not exist any more because the between-class and within-class scatter matrices constructed in 2D-FDA are both of full-rank. This framework contains unilateral and bilateral 2D-FDA. It is applied to face recognition where only few training images exist for each subject. Both the unilateral and bilateral 2D-FDA achieve excellent performance on two public databases: ORL database and Yale face database B.
Hui Kong 0001, Lei Wang 0001, Eam Khwang Teoh, Jian-Gang Wang 0001, Ronda Venkateswarlu
CVPR (2)1
2005 Two Dimensional Fisher Discriminant Analysis: Forget About Small Sample Size Problem
abstract
This paper addresses the small sample size (SSS) problem in linear discriminant analysis (LDA) utilizing a so called 2D Fisher discriminant analysis (2D-FDA) algorithm. As opposed to traditional LDA-based approaches, 2D-FDA is based on 2D image matrices rather than 1D vectors so the image matrix does not need to be transformed into a vector before feature extraction. The between-class scatter and the within-class scatter is constructed using the original image matrices. The advantage arising in this way is that the SSS problem existing in traditional linear discriminant analysis does not occur any more. To test the performance of 2D-FDA with small number of training samples, a series of experiments are conducted on two public databases: ORL and Yale face database B. In both two trials, the 2D-FDA outperforms the other linear subspace methods when there are only very limited training images for each subject.
Hui Kong 0001, Eam Khwang Teoh, Jian-Gang Wang 0001, Ronda Venkateswarlu
ICASSP (2)1
2005 Generalized 2D principal component analysis
abstract
A two-dimensional principal component analysis (2DPCA) by J. Yang et al. (2004) was proposed and the authors have demonstrated its superiority over the conventional principal component analysis (PCA) in face recognition. But the theoretical proof why 2DPCA is better than PCA has not been given until now. In this paper, the essence of 2DPCA is analyzed and a framework of generalized 2D principal component analysis (G2DPCA) is proposed to extend the original 2DPCA in two perspectives: a bilateral-projection-based 2DPCA (B2DPCA) and a kernel-based 2DPCA (K2DPCA) schemes are introduced. Experimental results in face recognition show its excellent performance.
Hui Kong 0001, Xuchun Li, Lei Wang 0001, Earn Khwang Teoh, Ronda Venkateswarlu
IJCNN1
2005 Generalized 2D principal component analysis for face image representation and recognition
Hui Kong 0001, Lei Wang 0001, Eam Khwang Teoh, Xuchun Li, Jian-Gang Wang 0001, Ronda Venkateswarlu
Neural Networks1