Zheng Fang 0001

dblp:77/4730-1 · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
31since 2021 · last 2026
0000-0003-3887-3141ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 3 first-author · 26 since 2021Systems, architecture and hardware · 22 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
abstract
Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details, leading to semantic and geometric ambiguities, especially in complex environments. To address this limitation, we propose Ocean, an object-centric prediction framework that decomposes the scene into individual object instances to enable more accurate semantic occupancy prediction. Specifically, we first employ a lightweight segmentation model, MobileSAM, to extract instance masks from the input image. Then, we introduce a 3D Semantic Group Attention module that leverages linear attention to aggregate object-centric features in 3D space. To handle segmentation errors and missing instances, we further design a Global Similarity-Guided Attention module that leverages segmentation features for global interaction. Finally, we propose an Instance-aware Local Diffusion module that improves instance features through a generative process and subsequently refines the scene representation in the BEV space. Extensive experiments on the SemanticKITTI and SSCBench-KITTI360 benchmarks demonstrate that Ocean achieves state-of-the-art performance, with mIoU scores of 17.40 and 20.28, respectively.
Yubo Cui, Xiangru Lin, Zhiheng Li 0003, Zheng Fang 0001
AAAI5
2026 Dynamic clustering transformer for LiDAR-based 3D object detection
Yubo Cui, Zhiheng Li 0003, Zheng Fang 0001
Pattern Recognit.3
2025 LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba
abstract
Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information through the attention mechanism, enabling the 3D semantic occupancy prediction. However, these works usually face two main challenges: 1) Limited geometric information. Due to the lack of geometric information in the image itself, it is challenging to directly predict 3D space information, especially in large-scale outdoor scenes. 2) Local restricted interaction. Due to the quadratic complexity of the attention mechanism, they often use modified local attention to fuse features, resulting in a restricted fusion. To address these problems, in this paper, we propose a language-assisted 3D semantic occupancy prediction network, named LOMA. In the proposed vision-language framework, we first introduce a VL-aware Scene Generator (VSG) module to generate the 3D language feature of the scene. By leveraging the vision-language model, this module provides implicit geometric knowledge and explicit semantic information from the language. Furthermore, we present a Tri-plane Fusion Mamba (TFM) block to efficiently fuse the 3D language feature and 3D vision feature. The proposed module not only fuses the two features with global modeling but also avoids too much computation costs. Experiments on the SemanticKITTI and SSCBench-KITTI360 datasets show that our algorithm achieves new state-of-the-art performances in both geometric and semantic completion tasks. Our code will be open soon.
Yubo Cui, Zhiheng Li 0003, Jiaqiang Wang, Zheng Fang 0001
AAAI4
2025 Learning Null Geodesics for Gravitational Lensing Rendering in General Relativity
abstract
We present GravLensX, an innovative method for rendering black holes with gravitational lensing effects using neural networks. The methodology involves training neural networks to fit the spacetime around black holes and then employing these trained models to generate the path of light rays affected by gravitational lensing. This enables efficient and scalable simulations of black holes with optically thin accretion disks, significantly decreasing the time required for rendering compared to traditional methods. We validate our approach through extensive rendering of multiple black hole systems with superposed Kerr metric, demonstrating its capability to produce accurate visualizations with significantly $15\times$ reduced computational time. Our findings suggest that neural networks offer a promising alternative for rendering complex astrophysical phenomena, potentially paving a new path to astronomical visualization.
Zheng Fang 0001, Kunyi Zhang, Qiang Zhang 0029, Renjing Xu
ICCV2
2025 Optimal Brain Apoptosis
abstract
The increasing complexity and parameter count of Convolutional Neural Networks (CNNs) and Transformers pose challenges in terms of computational efficiency and resource demands. Pruning has been identified as an effective strategy to address these challenges by removing redundant elements such as neurons, channels, or connections, thereby enhancing computational efficiency without heavily compromising performance. This paper builds on the foundational work of Optimal Brain Damage (OBD) by advancing the methodology of parameter importance estimation using the Hessian matrix. Unlike previous approaches that rely on approximations, we introduce Optimal Brain Apoptosis (OBA), a novel pruning method that calculates the Hessian-vector product value directly for each parameter. By decomposing the Hessian matrix across network layers and identifying conditions under which inter-layer Hessian submatrices are non-zero, we propose a highly efficient technique for computing the second-order Taylor expansion of parameters. This approach allows for a more precise pruning process, particularly in the context of CNNs and Transformers, as validated in our experiments including VGG19, ResNet32, ResNet50, and ViT-B/16 on CIFAR10, CIFAR100 and Imagenet datasets. Our code is available at https://github.com/NEU-REAL/OBA.
Zheng Fang 0001, Delei Kong, Chenming Hu, Yuetong Fang, Renjing Xu
ICLR2
2025 CAO-RONet: A Robust 4D Radar Odometry with Exploring More Information from Low-Quality Points
abstract
Recently, 4D millimetre-wave radar exhibits more stable perception ability than LiDAR and camera under adverse conditions (e.g. rain and fog). However, low-quality radar points hinder its application, especially the odometry task that requires a dense and accurate matching. To fully explore the potential of 4D radar, we introduce a learning-based odometry framework, enabling robust ego-motion estimation from finite and uncertain geometry information. First, for sparse radar points, we propose a local completion to supplement missing structures and provide denser guideline for aligning two frames. Then, a context-aware association with a hierarchical structure flexibly matches points of different scales aided by feature similarity, and improves local matching consistency through correlation balancing. Finally, we present a window-based optimizer that uses historical priors to establish a coupling state estimation and correct errors of inter-frame matching. The superiority of our algorithm is confirmed on View-of-Delft dataset, achieving around a 50% performance improvement over previous approaches and delivering accuracy on par with LiDAR odometry. The code will be released at https://github.com/NEU-REAL/CAO-RONet.
Zhiheng Li 0003, Yubo Cui, Ningyuan Huang, Chenglin Pang, Zheng Fang 0001
ICRA5
2025 Nonlinear Motion-Guided and Spatio-Temporal Aware Network for Unsupervised Event-Based Optical Flow
abstract
Event cameras have the potential to capture continuous motion information over time and space, making them well-suited for optical flow estimation. However, most existing learning-based methods for event-based optical flow adopt frame-based techniques, ignoring the spatio-temporal characteristics of events. Additionally, these methods assume linear motion between consecutive events within the loss time window, which increases optical flow errors in long-time sequences. In this work, we observe that rich spatio-temporal information and accurate nonlinear motion between events are crucial for event-based optical flow estimation. Therefore, we propose E-NMSTFlow, a novel unsupervised event-based optical flow network focusing on long-time sequences. We propose a Spatio-Temporal Motion Feature Aware (STMFA) module and an Adaptive Motion Feature Enhancement (AMFE) module, both of which utilize rich spatio-temporal information to learn spatio-temporal data associations. Meanwhile, we propose a nonlinear motion compensation loss that utilizes the accurate nonlinear motion between events to improve the unsupervised learning of our network. Extensive experiments demonstrate the effectiveness and superiority of our method. Remarkably, our method ranks first among unsupervised learning methods on the MVSEC and DSEC-Flow datasets.
Zuntao Liu, Zheng Fang 0001
ICRA5
2025 A Coarse-to-Fine Event-based Framework for Camera Pose Relocalization with Spatio-Temporal Retrieval and Refinement Network
abstract
Most existing event-based camera pose relocalization (CPR) learning methods implicitly encode environmental information into network parameters to achieve end-to-end mapping from event stream to pose. However, these end-to-end CPR methods fail to utilize prior environmental information effectively. As the scale of the environment increases, the difficulty of this mapping relationship grows significantly, reducing the robustness of the end-to-end methods across different scenarios. To address the above issues, this paper proposes the first coarse-to-fine event-based CPR framework, which achieves a new paradigm from end-to-end pose regression network to a hierarchical approach. In the coarse localization stage, we effectively encode similarity features by incorporating the fine-grained temporal information, achieving accurate retrieval of nearby event stream. In the pose refinement stage, we present an Event Spatio-temporal Pose Refinement Network (ESPR-Net) based on the Recurrent Convolutional Neural Networks (RCNN) architecture, which is capable of learning more nu-anced spatio-temporal features to achieve accurate regression of the relative pose. Finally, we conducted a comprehensive comparison on the IJRR and M3ED dataset, achieving state-of-the-art (SOTA) performance on both. Notably, our method attains a significant 83 % performance improvement on the outdoor M3ED dataset.
Zuntao Liu, Zheng Fang 0001
ICRA5
2025 Target-Aware Viewpoint Generation for Active Robotic Exploration in Unknown Environments
abstract
When entering an unfamiliar environment, animals usually sweep off their surroundings to identify points of interest. In search and rescue robotics, autonomous exploration requires both coarse mapping of unknown areas and detailed target detection, which poses a significant challenge in balancing these tasks. To that end, we propose a target-aware robotic exploration framework that prioritizes both exploration efficiency and search completeness through three components: First, considering the computational limitations of robotic platforms, a lightweight 3D target detection method with post-fusion is introduced to detect target positions in real time. Secondly, we propose a target-aware viewpoint generation approach that integrates information gain and inspection gain to identify promising viewpoints for thorough target searches. Lastly, since a detailed examination of the environment demands numerous viewpoints, we propose a heuristic-based active exploration framework that employs a hierarchical structure to optimize exploration gain, traveling distance, and path smoothness to maximize the utility function of viewpoint sequences and ultimately find the optimal path. Extensive simulations and real-world experiments demonstrate our framework significantly enhances target search capabilities, achieving a 13 % average improvement in exploration efficiency over existing methods.
Pu Xu, Zhiheng Li 0003, Zhaoqiang Bai, Zheng Fang 0001
ICRA5
2025 RDN: An Efficient Denoising Network for 4D Radar Point Clouds
abstract
Accurate point cloud information is important for robot perception and autonomous driving. Although advanced 4D radar can provide point cloud with higher resolution than 3D radar, its data still contains a significant amount of noise due to measurement principle. To solve this issue, we propose RDN (Radar Denoising Network), a denoising network specifically designed for 4D radar. RDN includes three innovative modules: First, to overcome the noisy nature of radar points, we design a feature similarity-based farthest point sampling module (FS-FPS), which can extract representative sampling points from the noisy point cloud. Secondly, to address feature propagation issues caused by the sparse and long-range characteristics of 4D radar points, we introduce a virtual feature point prediction (VFP) module and an iterative upsampling (IUS) module. The VFP module generates virtual feature points through the network to serve as bridges for information transmission, while the IUS module uses an iterative approach to gradually refine feature propagation. The experiments on MSC-RAD4D and NTU4DRadLM datasets demonstrate the effectiveness and generalization of our method. Besides, odometry experiments prove the practical value of point cloud denoising in improving robot perception.
Ningyuan Huang, Zhiheng Li 0003, Chenglin Pang, Zheng Fang 0001
IROS4
2025 Visual Localization with Offline Google Satellite Map-Assisted for Ground Vehicles in GNSS-Denied Environment
abstract
Vehicle localization is a critical component in the planning and navigation of autonomous driving system. Generally, traditional vehicle localization methods rely on the Global Navigation Satellite System (GNSS) for self-localization. Unfortunately, GNSS can become unreliable and may fail in urban canyons, under trees, and beneath overpasses. To address this problem, we propose a visual localization framework assisted by offline Google satellite maps in GNSS-weak or GNSS-denied environments. And we introduce learning-based ground-to-satellite map feature matching method to mitigate the long-term cumulative drift of visual odometry. To reduce the negative impact of cross-view matching errors on localization accuracy, we propose a novel cross-view pose selection method to build two pose uncertainty models. Moreover, we combine the proposed method with classical SLAM methods to develop a vehicle localization framework. To verify the performance of the proposed method, we carried out the accuracy comparison experiment with state-of-the-art fusion localization methods and feature matching methods. Experimental results indicate that the proposed method achieves the best localization performance compared with the state-of-the-art methods, and our method achieves the root mean square error of 0.290m and 0.014rad in KITTI-05. The implementation code of this paper will be open-source at https://github.com/NEU-REAL/visualLocalization-with-satelliteMap.
Jibo Wang, Bairen Mao, Chenglin Pang, Shiguang Liu, Jindi Guo, Zheng Fang 0001
IROS6
2025 Robust Model-Free Path Tracking Algorithm for Hydraulic Center-Articulated Scooptrams
abstract
This paper proposes a model-free steering control method to address the path tracking challenges of Hydraulic Center-articulated Scooptrams (HCS) in narrow underground mining environments. Due to the nonlinear and time-delay characteristics of the hydraulic steering system, the HCS exhibits response lag when executing control commands. The lag time demonstrates dynamic uncertainty influenced by operating conditions, hydraulic pressure, and load variations. To address this challenge, an adaptive steering control strategy is designed. This strategy leverages the geometric relationship between the HCS and the reference path to dynamically adjust the look-ahead distance, thereby compensating for the uncertainty caused by the hydraulic system lag. Additionally, the error is mapped to the actual control input in real-time through a feedback error controller, effectively correcting control errors caused by lag without relying on a complex hydraulic system model. The proposed method was experimentally validated in a full-scale simulated mining tunnel, demonstrating considerable robustness and precise path tracking performance under uneven terrain, heavy loads, significant initial error, and bidirectional movement. This method provides a viable solution for the autonomous navigation of the HCS.
Cunguang Fang, Pu Xu, Zheng Fang 0001
IROS4
2025 Fully Asynchronous Neuromorphic Perception for Mobile Robot Dodging With Loihi Chips
abstract
Sparse and asynchronous sensing and processing in natural organisms lead to ultra low-latency and energy-efficient perception. Inspired by these characteristics, event cameras have attracted extensive attention from academia and industry. However, the mainstream event stream processing paradigms (e.g., event frames, 3D voxels) encounter issues such as feature loss, event stacking, and high computational burden. These problems deviate from the intended purpose of event cameras. To address these issues, we propose a fully asynchronous neuromorphic paradigm that integrates event cameras, spiking networks, and neuromorphic processors (Intel Loihi). This paradigm can faithfully process each event asynchronously as it arrives, mimicking the spike-driven signal processing in biological brains. We first propose a Key-Event-Point module to extract the key event stream from the raw event stream, effectively addressing the issue of limited transmission bandwidth when processing events on neuromorphic processors. Then, we propose the complete pipeline for the dodging network based on spiking neural networks to achieve offline training and online inference. Finally, we compare the proposed paradigm with event frames and 3D voxels processing paradigms in detail on the real mobile robot dodging task. Experimental results show that our scheme exhibits better robustness than image-like methods with different time windows and light conditions. Additionally, the energy consumption per inference of our scheme on the embedded Loihi processor is only 4.30% of that of the event spike tensor method on NVIDIA Jetson Orin NX with energy-saving mode, and 1.64% of that of the event frame method on the same neuromorphic processor. To the best of our knowledge, this is the first time that a fully asynchronous neuromorphic paradigm has been implemented for solving sequential tasks on a real mobile robot.Note to Practitioners—As a neuromorphic visual sensor, the event camera offers a novel approach to achieving robust, low-power perception in robotics, owing to its high temporal resolution, wide dynamic range, and low information redundancy. However, the sparse and asynchronous nature of event streams presents a significant challenge for data processing. The mainstream approach preprocesses event streams into various representations (e.g., event frames, 3D voxels) before performing subsequent operations, leading to issues such as feature loss, event stacking, and high computational burden. In this paper, we propose a fully neuromorphic system that leverages the asynchronous characteristics of event streams and spiking neural networks to enable asynchronous processing of events. Our asynchronous processing of event streams enhances robustness across different time windows and lighting conditions while significantly reducing power consumption during inference. Our fully asynchronous neuromorphic system has been validated on a real mobile robot and is expected to advance the robot perception system towards biological perception.
Delei Kong, Chenming Hu, Zheng Fang 0001
IEEE Trans Autom. Sci. Eng.4
2025 Coupling and Decoupling: Towards Temporal Feedback for 3D Object Detection
abstract
3D object detection has garnered significant attention within the academic community, primarily due to its broad utility in domains such as autonomous driving and robotics. Prior research efforts have predominantly concentrated on leveraging temporal contextual information embedded within sequential data to enhance the current feature representations. However, a notable limitation of these endeavors lies in their inadequate treatment of the inherent noise present within historical sequences, thereby constraining the efficiency of fusion methods. In this paper, we propose a new temporal feedback network, named TFNet, to model and correct the temporal noise by designing acoupling-decouplingmechanism. Central to our approach are two distinct modules: (i) Foreground Feature Enhancement, which amplifies sparse instance details across temporal frames, thereby furnishing essential local information priors for subsequent fusion; and (ii) Coupling-Decoupling Feature Interaction, designed to first aggregate temporal contextual information and then disentangle fusion features into frame-specific representations. Leveraging a feedback strategy, this module can adaptively enhance useful information and eliminate noise within individual frame features. Empirical evaluations conducted on the nuScenes benchmark demonstrate the effectiveness of TFNet, achieving the new state-of-the-art performance without any bells and whistles.
Yubo Cui, Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Zhiheng Li 0003, Zheng Fang 0001
IEEE Trans. Multim.6
2024 EventRPG: Event Data Augmentation with Relevance Propagation Guidance
abstract
Event camera, a novel bio-inspired vision sensor, has drawn a lot of attention for its low latency, low power consumption, and high dynamic range. Currently, overfitting remains a critical problem in event-based classification tasks for Spiking Neural Network (SNN) due to its relatively weak spatial representation capability. Data augmentation is a simple but efficient method to alleviate overfitting and improve the generalization ability of neural networks, and saliency-based augmentation methods are proven to be effective in the image processing field. However, there is no approach available for extracting saliency maps from SNNs. Therefore, for the first time, we present Spiking Layer-Time-wise Relevance Propagation rule (SLTRP) and Spiking Layer-wise Relevance Propagation rule (SLRP) in order for SNN to generate stable and accurate CAMs and saliency maps. Based on this, we propose EventRPG, which leverages relevance propagation on the spiking neural network for more efficient augmentation. Our proposed method has been evaluated on several SNN structures, achieving state-of-the-art performance in object recognition tasks including N-Caltech101, CIFAR10-DVS, with accuracies of 85.62% and 85.55%, as well as action recognition task SL-Animals with an accuracy of 91.59%. Our code is available at https://github.com/myuansun/EventRPG.
Donghao Zhang 0004, ZongYuan Ge, Jia Li 0057, Zheng Fang 0001, Renjing Xu
ICLR6
2024 SeqTrack3D: Exploring Sequence Information for Robust 3D Point Cloud Tracking
abstract
3D single object tracking (SOT) is an important and challenging task for the autonomous driving and mobile robotics. Most existing methods perform tracking between two consecutive frames while ignoring the motion patterns of the target over a series of frames, which would cause performance degradation in the scenes with sparse points. To break through this limitation, we introduce "Sequence-to-Sequence" tracking paradigm and a tracker named SeqTrack3D to capture target motion across continuous frames. Unlike previous methods that primarily adopted three strategies: matching two consecutive point clouds, predicting relative motion, or utilizing sequential point clouds to address feature degradation, our SeqTrack3D combines both historical point clouds and bounding box sequences. This novel method ensures robust tracking by leveraging location priors from historical boxes, even in scenes with sparse points. Extensive experiments conducted on large-scale datasets show that SeqTrack3D achieves new state-of-the-art performances, improving by 6.00% on NuScenes and 14.13% on Waymo dataset. The code will be made public at https://github.com/aron-lin/seqtrack3d.
Zhiheng Li 0003, Yubo Cui, Zheng Fang 0001
ICRA4
2024 Observation Time Difference: an Online Dynamic Objects Removal Method for Ground Vehicles
abstract
In the process of urban environment mapping, the sequential accumulations of dynamic objects will leave a large number of traces in the map. These traces will usually have bad influences on the localization accuracy and navigation performance of the robot. Therefore, dynamic objects removal plays an important role for creating clean map. However, conventional dynamic objects removal methods usually run offline. That is, the map is reprocessed after it is constructed, which undoubtedly increases additional time costs. To tackle the problem, this paper proposes a novel method for online dynamic objects removal for ground vehicles. According to the observation time difference between the object and the ground where it is located, dynamic objects are classified into two types: suddenly appear and suddenly disappear. For these two kinds of dynamic objects, we propose downward retrieval and upward retrieval methods to eliminate them respectively. We validate our method on SemanticKITTI dataset and author-collected dataset with highly dynamic objects. Compared with other state-of-the-art methods, our method is more efficient and robust, and reduces the running time per frame by more than 60% on average. Our method will be open-sourced on GitHub1.
Rongguang Wu, Chenglin Pang, Xuankang Wu, Zheng Fang 0001
ICRA4
2024 FlowTrack: Point-level Flow Network for 3D Single Object Tracking
abstract
3D single object tracking (SOT) is a crucial task in fields of mobile robotics and autonomous driving. Traditional motion-based approaches achieve target tracking by estimating the relative movement of target between two consecutive frames. However, they usually overlook local motion information of the target and fail to exploit historical frame information effectively. To overcome the above limitations, we propose a point-level flow method with multi-frame information for 3D SOT task, called FlowTrack. Specifically, by estimating the flow for each point in the target, our method could capture the local motion details of target, thereby improving the tracking performance. Meanwhile, to handle scenes with sparse points, we present a learnable target feature as the bridge to efficiently integrate target information from past frames. Moreover, we design a Instance Flow Head to transform dense point-level flow into instance-level motion, effectively aggregating local motion information to obtain global target motion. Finally, our method achieves competitive performance with improvements of 5.9% on the KITTI and 2.9% on the NuScenes, compared to the next best method.
Yubo Cui, Zhiheng Li 0003, Zheng Fang 0001
IROS4
2024 ASML-VDIO: Visual-Depth-Inertial Odometry using Selected Accurate and Stable Multi-Modal Landmarks in Structural Environments
abstract
In complex indoor structural scenes such as shopping centers and malls, camera pose estimation using pure point features is easy to fail due to the difficulty in extracting sufficient and stable point features from weak textures or dynamic environments. Recent works have attempted to address these challenges by introducing line features. However, the addition of line features increases the number of parameters and landmarks for BA (Bundle Adjustment), leading to efficiency reduction. This is a common issue in multi-modal SLAM (Simultaneous Localization And Mapping). To address this issue, this paper proposes a novel visual-depth-inertial odometry (ASML-VDIO) framework by combining RGB-D and IMU sensors. To improve the efficiency of BA, the proposed landmark classification method classifies 3D landmarks into accurate landmarks and other landmarks based on spatial consistency verification and depth range limitation. Then, accurate landmarks are fixed, and only other landmarks are optimized in the optimization of BA. Furthermore, to remove line features extracted from dynamic objects (pedestrian, shopping-car, etc), we propose a dynamic line removal method that combines geometric constraints and motion constraints of line features. Finally, the method is evaluated on public and author-collected datasets, showing competitive accuracy and robustness in complex indoor structural scenes while 71% speedup on optimization thread with same constraints.
Xingjian Luo, Chenglin Pang, Xuankang Wu, Zheng Fang 0001
IROS4
2024 EverySync: An Open Hardware Time Synchronization Sensor Suite for Common Sensors in SLAM
abstract
Multi-sensor fusion systems have been widely applied in various fields, including mobile robot, simultaneous localization and mapping (SLAM), and autonomous driving. For a tightly coupled multi-sensor fusion system, strict time synchronization between sensors will improve the accuracy of the system. However, there is currently a lack of open-source and general-purpose hardware synchronization systems for Cameras, IMUs, LiDARs, GNSS/RTK in the academic community. Therefore, we propose EverySync, an open hardware time synchronization system to address this gap. The synchronization accuracy of the system was evaluated through multiple experiments, achieving an accuracy of less than 1 ms. And, real-world experiments proved that hardware time synchronization improves the accuracy of the SLAM system. This open-source system is available on GitHub.
Xuankang Wu, Haoxiang Sun, Rongguang Wu, Zheng Fang 0001
IROS4
2024 Trans-Rotor: An Active Omnidirectional Aerial-Ground Vehicle With Differential Gear Joint Transformation Mechanism
abstract
Aerial-ground vehicles have shown great potential in various fields due to their superior mobility and outstanding endurance. However, most of morphing aerial-ground vehicles consider little about controllability and traversability in ground mode. We present a novel aerial-ground vehicle called TransRotor. By proposing a differential gear joint, we equip TransRotor with omnidirectional mobility in both air and ground mode. Besides, using a four-wheel-steering model in ground mode provides better traversability and ground flexibility. Moreover, we design mid-mode transformation for Trans-Rotor, which provides smooth and rapid mode switching. In this work, we firstly propose a novel design of an aerial-ground vehicle. Then, we propose a decoupled controller considering the four-wheel-steer model to achieve autonomous navigation of the vehicle. Comprehensive experiments and a benchmark comparison are carried out to validate the outstanding performance of the proposed system, where the system shows ground flexibility and saves energy up to more than 95%.
Xuankang Wu, Haoxiang Sun, Tong Xiao 0001, Yanzhang Pan, Zheng Fang 0001
IROS5
2024 PARE: A Plane-Assisted Autonomous Robot Exploration Framework in Unknown and Uneven Terrain
abstract
Identifying traversable areas is a critical task for unmanned vehicles exploring safely through unstructured environments. In practice, the ambiguity in perceiving terrain traversability usually brings great challenges for autonomous exploration in unknown and uneven terrain, which often leads to conservative strategies or potential risk of vehicle damage, resulting in many unexplored areas in the environment. To that end, this paper proposes a plane-assisted autonomous robot exploration framework (PARE) to achieve maximum volume and safe autonomous exploration. The process is carried out by a three-step dual-layer framework: constructing a local tree using Plane-Assisted RRT* (PA-RRT*), calculating exploration gain based on terrain information, and maintaining a global search graph. Firstly, the planar feature metrics (flatness, sparsity, elevation variation, slope and slope variation) are introduced to determine the terrain traversability. Secondly, to completely explore the rugged environment, we propose a dual-layer exploration framework comprising local and global strategies. A local planner based on PA-RRT* is proposed to find the best path by evaluating the planar information and the volumetric gain within the local exploration tree. Meanwhile, a global planner constructed by graph is proposed to record unexplored nodes with high exploration gain from the local tree to ensure a high level of exploration volume. Extensive simulation and real-world experiments demonstrate that our method significantly outperforms existing frameworks, with an average improvement of more than 12% in exploration volume.
Pu Xu, Zhaoqiang Bai, Zheng Fang 0001
IROS4
2024 Geometry-aided Underwater 3D Mapping Using Side-scan Sonar
abstract
In recent years, the interest in underwater exploration with Autonomous Underwater Vehicles (AUVs) equipped with side-scan sonars (SSS) has grown considerably. However, state-of-the-art SSS Simultaneous Localization and Mapping (SLAM) systems encounter challenges in data association across large viewpoint changes. Additionally, these systems assume that the seabed is a flat surface, leading to significant mapping error in uneven underwater terrains. To address these challenges, we propose a framework that leverages the side-scan sonar geometry to facilitate data association and improve mapping accuracy. The framework begins with a preprocessing module that extracts feature points and provides initial estimates of the elevation angles of the landmarks. Then, a non-consecutive data association module applies epipolar line search to establish correspondences between the current and historical frames. Finally, the mapping module uses side-scan sonar bundle adjustment to recover the positions of the landmarks. The proposed method is evaluated using an underwater terraced fields dataset. Our method achieves over 90% matching rate and reduces the average mapping error from 3.799 to 0.134.
Yiqiao Yang, Chenglin Pang, Chengdong Wu 0001, Zheng Fang 0001
IROS4
2024 OTD: An Online Dynamic Traces Removal Method Based on Observation Time Difference
abstract
Three-dimensional point cloud map plays an important role in 3-D reconstruction, autonomous robot navigation, autonomous driving, and environmental monitoring. Nowadays, 3-D point cloud map could be obtained through frame-by-frame accumulation of LiDAR point cloud using SLAM technology. However, during this process, the movements of dynamic objects in the environment will leave a large number of traces on the point cloud map, causing difficulties in the subsequent use of the map, such as city model construction and robot autonomous navigation. Therefore, dynamic traces removal is crucial for building clean static maps. However, existing methods for dynamic traces removal are mostly offline, which inevitably incurs additional time consumption. To address this problem, this article proposes an online dynamic traces removal method. We take voxels as the smallest unit for dynamic traces removal, and voxels containing dynamic traces are called dynamic voxels, otherwise they are called static voxels. Our method is based on the assumption that static voxels always appear and disappear simultaneously with the ground below them. Therefore, we call voxel that appears later than the ground as suddenly appear dynamic voxel, and voxel that disappears earlier than the ground as suddenly disappear dynamic voxel. We call this method of judging dynamic voxels as observation time difference, and propose downward retrieval and upward retrieval methods to remove these two types of dynamic voxels, respectively. We tested our proposed method on SemanticKITTI, UrbanLoco, and author-collected datasets. Experimental results show that our method is more accurate and robust than existing online dynamic traces removal methods. And compared with other methods, our method shortens the time of processing each frame of point cloud by more than 60%. Our method is open-sourced on GitHub:https://github.com/RongguangWu/OTD.
Rongguang Wu, Zheng Fang 0001, Chenglin Pang, Xuankang Wu
IEEE Trans. Geosci. Remote. Sens.2
2024 Intersection Is Also Needed: A Novel LiDAR-Based Road Intersection Dataset and Detection Method
abstract
3D object detection is crucial for autonomous driving. However, most existing methods focus on the foreground objects, such as vehicles and pedestrians, while ignoring some important background objects for traffic scene understanding, especially road intersections. Moreover, existing datasets (e.g., KITTI, Waymo) do not provide the labels for intersections, and the evaluation metric is also unsuitable for intersection detection. To address the above issues, we first present a LiDAR-based intersection dataset on the basis of KITTI dataset, calledKITTI-Intersection Dataset. The new dataset includes 4718 frames with 5178 instances belonging to Forkroad and Crossroad, respectively. To weaken the impact of uncertain intersection size on the performance evaluation, we introduce CEIOU instead of IOU as a new evaluation metric. Then, we proposeMInsectDetandMMInsectDet, two LiDAR-based detection methods, to solve the intersection detection problem. We start with a lightweight BEV backbone to alleviate the influence of numerous dynamic foreground objects at the intersection and obtain discriminative features. After that, to obtain more abundant and complete intersection features, we propose a Multi-Representation Backbone that integrates the BEV and voxel features to achieve better detection performance. Furthermore, in order to better adapt to various appearances and sizes of intersection, we propose a Class-Aware MultiHead, which classifies and regresses different categories with specific head. Finally, we evaluate our MInsectDet and MMInsectDet methods on the proposed KITTI-Intersection Dataset with the state-of-the-art foreground 3D detection methods. The results show that MMInsectDet achieves the best performance, and MInsectDet ranks second but could run at 65.0 FPS.
Zhiheng Li 0003, Yubo Cui, Zheng Fang 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Real-Time 3D Single Object Tracking With Transformer
abstract
LiDAR-based 3D single object tracking is a challenging issue in robotics and autonomous driving. Currently, existing approaches usually suffer from the problem that objects at long distance often have very sparse or partially-occluded point clouds, which makes the features extracted by the model ambiguous. Ambiguous features will make it hard to locate the target object and finally lead to bad tracking results. To solve this problem, we utilize the powerful Transformer architecture and propose aPoint-Track-Transformer (PTT)module for point cloud-based 3D single object tracking task. Specifically, PTT module generates fine-tuned attention features by computing attention weights, which guides the tracker focusing on the important features of the target and improves the tracking ability in complex scenarios. To evaluate our PTT module, we embed PTT into the dominant method and construct a novel 3D SOT tracker named PTT-Net. In PTT-Net, we embed PTT into the voting stage and proposal generation stage, respectively. PTT module in the voting stage could model the interactions among point patches, which learns context-dependent features. Meanwhile, PTT module in the proposal generation stage could capture the contextual information between object and background. We evaluate our PTT-Net on KITTI and NuScenes datasets. Experimental results demonstrate the effectiveness of PTT module and the superiority of PTT-Net, which surpasses the baseline by a noticeable margin,$\sim$10% in the Car category. Meanwhile, our method also has a significant performance improvement in sparse scenarios. In general, the combination of transformer and tracking pipeline enables our PTT-Net to achieve state-of-the-art performance on both two datasets. Additionally, PTT-Net could run in real-time at 40FPS on NVIDIA 1080Ti GPU. Our code is open-sourced for the research community athttps://github.com/shanjiayao/PTT.
Jiayao Shan, Sifan Zhou, Yubo Cui, Zheng Fang 0001
IEEE Trans. Multim.4
2022 Sparse point-voxel aggregation network for efficient point cloud semantic segmentation
abstract
Abstract Effective and efficient semantic segmentation of 3D point cloud data is important for many tasks. Many methods for point cloud semantic segmentation rely on computationally expensive sampling and grouping layers to process irregular points, while others convert irregular points into regular volumetric grids and process them with a 3D U‐Net‐based semantic segmentation network. However, most of these methods suffer from high computational costs and cannot be applied to the real‐time processing of large‐scale point clouds. To address these issues, we propose a computationally efficient point‐voxel‐based network architecture named Sparse Point‐Voxel Aggregation Network (SPVAN) for point cloud semantic segmentation. It consists of an encoding layer that consists of sparse convolution and MLP layers and a new decoding layer called Point Feature Aggregation Layer (PFAL) that is only composed of feature interpolation and MLP layers. Compared with recent popular point‐voxel‐based methods with the U‐Net‐based network, our method does not need 3D convolution networks in the decoding layer and thus achieves a higher speed. Experimental results on the large‐scale SemanticKITTI dataset show that our method gets a good balance between the efficiency and the performance. Moreover, our method achieves on‐par or better performance than previous methods for semantic segmentation on the challenging S3DIS dataset.
Zheng Fang 0001, Binyu Xiong
IET Comput. Vis.1
2022 RGB-D SLAM in Dynamic Environments Using Point Correlations
abstract
In this paper, a simultaneous localization and mapping (SLAM) method that eliminates the influence of moving objects in dynamic environments is proposed. This method utilizes the correlation between map points to separate points that are part of the static scene and points that are part of different moving objects into different groups. A sparse graph is first created using Delaunay triangulation from all map points. In this graph, the vertices represent map points, and each edge represents the correlation between adjacent points. If the relative position between two points remains consistent over time, there is correlation between them, and they are considered to be moving together rigidly. If not, they are considered to have no correlation and to be in separate groups. After the edges between the uncorrelated points are removed during point-correlation optimization, the remaining graph separates the map points of the moving objects from the map points of the static scene. The largest group is assumed to be the group of reliable static map points. Finally, motion estimation is performed using only these points. The proposed method was implemented for RGB-D sensors, evaluated with a public RGB-D benchmark, and tested in several additional challenging environments. The experimental results demonstrate that robust and accurate performance can be achieved by the proposed SLAM method in both slightly and highly dynamic environments. Compared with other state-of-the-art methods, the proposed method can provide competitive accuracy with good real-time performance.
Weichen Dai 0001, Yu Zhang 0018, Ping Li 0017, Zheng Fang 0001, Sebastian A. Scherer
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 3D Object Tracking with Transformer
Yubo Cui, Zheng Fang 0001, Jiayao Shan, Zuoxu Gu, Sifan Zhou
BMVC2
2021 Vanishing Point Aided LiDAR-Visual-Inertial Estimator
abstract
In this paper, we propose a vanishing point aided LiDAR-Visual-Inertial estimator to achieve real-time, low-drift and robust pose estimation. The proposed method is mainly composed of 3 sequential modules, namely IMU-aided vanishing point (VP) detection module, voxel-map based feature depth association module, and visual inertial fixed-lag smoother module. The IMU-aided VP detection module will detect feature points, line segments and vanishing points to establish robust correspondences in successive frames. In particular, we propose to use 1-line RANSAC method to provide stable VP hypotheses and polar grid to accelerate vanishing point hypothesis validation. After that, we propose a novel voxel-map based feature depth association method, to retrieve depth and assign depth to visual feature efficiently. Finally, the visual inertial fixed-lag smoother module is proposed to jointly minimize error terms. Experiments show that our method outperforms the state-of-the-art visual-inertial odometry and LiDAR-visual estimator in both indoor and outdoor environments.
Zheng Fang 0001, Shibo Zhao, Yongnan Chen, Shan An
ICRA2
2021 PTT: Point-Track-Transformer Module for 3D Single Object Tracking in Point Clouds
abstract
3D single object tracking is a key issue for robotics. In this paper, we propose a transformer module called Point-Track-Transformer (PTT) for point cloud-based 3D single object tracking. PTT module contains three blocks for feature embedding, position encoding, and self-attention feature computation. Feature embedding aims to place features closer in the embedding space if they have similar semantic information. Position encoding is used to encode coordinates of point clouds into high dimension distinguishable features. Self-attention generates refined attention features by computing attention weights. Besides, we embed the PTT module into the open-source state-of-the-art method P2B to construct PTT-Net. Experiments on the KITTI dataset reveal that our PTT-Net surpasses the state-of-the-art by a noticeable margin $\left( {\sim 10\% } \right)$. Additionally, PTT-Net could achieve real-time performance (~40FPS) on NVIDIA 1080Ti GPU. Our code is open-sourced for the robotics community at https://github.com/shanjiayao/PTT.
Jiayao Shan, Sifan Zhou, Zheng Fang 0001, Yubo Cui
IROS3
2020 Neural Coding Strategies for Event-Based Vision Data
abstract
Neural coding schemes are powerful tools used within neuroscience. This paper introduces three different neural coding scheme formations for event-based vision data which are designed to emulate the neural behaviour exhibited by neurons under stimuli. Presented are phase-of-firing and two sparse neural coding schemes. It is determined that machine learning approaches, i.e. Convolutional Neural Network combined with a Stacked Autoencoder network, produce powerful descriptors of the patterns within events. These coding schemes are deployed in an existing action recognition template and evaluated using two popular event-based data sets.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001
ICASSP5
2020 Post-Stimulus Time-Dependent Event Descriptor
abstract
Event-based image processing is a relatively new domain in the field of computer vision. Much research has been carried out on adapting event-based data to comply with established techniques from frame-based computer vision. On the contrary, this paper presents a descriptor which is designed specifically for direct use with event-based data and therefore can be considered to be a pure event-based vision descriptor as it only uses events emitted from event-based vision devices without transforming the data to accommodate frame-based vision techniques. This novel descriptor is known as the Post-stimulus Time-dependent Event Descriptor (P-TED). P-TED is comprised of two features extracted from event data which describe motion and the underlying pattern of transmission respectively. Furthermore a framework is presented which leverages the P-TED descriptor to classify motions within event data. This framework is compared against another state-of-the-art event-based vision descriptor as well as an established frame-based approach.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001
ICIP5
2020 Reducing-Over-Time Tree for Event-based Data
abstract
This paper presents a novel Reducing-Over-Time (ROT) binary tree structure for event-based vision data and subtypes of the tree structure. A framework is presented using ROT, that takes advantage of the self-balancing and self-pruning nature of the tree structure to extract spatial-temporal information. The ROT framework is paired with an established motion classification technique and performance is evaluated against other state-of-the-art techniques using four datasets. Additionally, the ROT framework as a processing platform is compared with other event-based vision processing platforms in terms of memory usage and is found to be one of the most memory efficient platforms available.
Shane Harrigan, Sonya A. Coleman, Dermot Kerr, Yogarajah Pratheepan, Zheng Fang 0001, Chengdong Wu 0001
ICPR5
2020 TP-TIO: A Robust Thermal-Inertial Odometry with Deep ThermalPoint
abstract
To achieve robust motion estimation in visually degraded environments, thermal odometry has been an attraction in the robotics community. However, most thermal odometry methods are purely based on classical feature extractors, which is difficult to establish robust correspondences in successive frames due to sudden photometric changes and large thermal noise. To solve this problem, we propose ThermalPoint, a lightweight feature detection network specifically tailored for producing keypoints on thermal images, providing notable anti-noise improvements compared with other state-of-the-art methods. After that, we combine ThermalPoint with a novel radiometric feature tracking method, which directly makes use of full radiometric data and establishes reliable correspondences between sequential frames. Finally, taking advantage of an optimization-based visual-inertial framework, a deep feature-based thermal-inertial odometry (TP-TIO) framework is proposed and evaluated thoroughly in various visually degraded environments. Experiments show that our method outperforms state-of-the-art visual and laser odometry methods in smoke-filled environments and achieves competitive accuracy in normal environments.
Shibo Zhao, Zheng Fang 0001, Sebastian A. Scherer
IROS4
2019 A Robust Laser-Inertial Odometry and Mapping Method for Large-Scale Highway Environments
abstract
In this paper, we propose a novel laser-inertial odometry and mapping method to achieve real-time, low-drift and robust pose estimation in large-scale highway environments. The proposed method is mainly composed of four sequential modules, namely scan pre-processing module, dynamic object detection module, laser-inertial odometry module and laser mapping module. Scan pre-processing module uses inertial measurements to compensate the motion distortion of each laser scan. Then, the dynamic object detection module is used to detect and remove dynamic objects from each laser scan by applying CNN segmentation network. After obtaining the undistorted point cloud without moving objects, the laser-inertial odometry module uses an Error State Kalman Filter to fuse the data of laser and IMU and output the coarse pose estimation at high frequency. Finally, the laser mapping module performs a fine processing step and the “Frame-to-Model'' scan matching strategy is used to create a static global map. We compare the performance of our method with two state-of-the-art methods, LOAM and SuMa, using KITTI dataset and real highway scene dataset. Experiment results show that our method performs better than the state-of-the-art methods in real highway environments and achieves competitive accuracy on the KITTI dataset.
Shibo Zhao, Zheng Fang 0001, HaoLai Li, Sebastian A. Scherer
IROS2
2018 Feature Regions Segmentation Based RGB-D Visual Odometry in Dynamic Environment
abstract
A novel RGB-D visual odometry method for dynamic environment is proposed. Majority of visual odometry systems can only work in static environments, which limits their applications in real world. In order to improve the accuracy and robustness of visual odometry in dynamic environment, a Feature Regions Segmentation algorithm is proposed to resist the disturbance caused by the moving objects. The matched features are divided into different regions to separate the moving objects from the static background. The features in the largest region which belong to the static background are used to estimate the camera pose finally. The effectiveness of our visual odometry method is verified in a dynamic environment of our lab. Furthermore, an exhaustive experimental evaluation is conducted on benchmark datasets including static environments and dynamic environments compared with the state-of-art visual odometry systems. The accuracy comparison results show that the proposed algorithm outperforms those systems in large scale dynamic environments. Our method tracks the camera movement correctly while others failed. In addition, our method can give the same good performances in static environment. Experiments demonstrate that the proposed RGB-D visual odometry can obtain accurate and robust estimation results in dynamic environments.
Yu Zhang 0018, Weichen Dai 0001, Ping Li 0017, Zheng Fang 0001
IECON5
2015 Real-time onboard 6DoF localization of an indoor MAV in degraded visual environments using a RGB-D camera
abstract
Real-time and reliable localization is a prerequisite for autonomously performing high-level tasks with micro aerial vehicles(MAVs). Nowadays, most existing methods use vision system for 6DoF pose estimation, which can not work in degraded visual environments. This paper presents an onboard 6DoF pose estimation method for an indoor MAV in challenging GPS-denied degraded visual environments by using a RGB-D camera. In our system, depth images are mainly used for odometry estimation and localization. First, a fast and robust relative pose estimation (6DoF Odometry) method is proposed, which uses the range rate constraint equation and photometric error metric to get the frame-to-frame transform. Then, an absolute pose estimation (6DoF Localization) method is proposed to locate the MAV in a given 3D global map by using a particle filter. The whole localization system can run in real-time on an embedded computer with low CPU usage. We demonstrate the effectiveness of our system in extensive real environments on a customized MAV platform. The experimental results show that our localization system can robustly and accurately locate the robot in various practical challenging environments.
Zheng Fang 0001, Sebastian A. Scherer
ICRA1
2014 Experimental study of odometry estimation methods using RGB-D cameras
abstract
Lightweight RGB-D cameras that can provide rich 2D visual and 3D point cloud information are well suited to the motion estimation of indoor micro aerial vehicles (MAVs). In recent years, several RGB-D visual odometry methods which process data from the sensor in different ways have been proposed. However, it is unclear which methods are preferable for online odometry estimation on a computation-limited, fast moving MAV in practical indoor environments. This paper presents a detailed analysis and comparison of several state-of-the-art real-time odometry estimation methods in a variety of challenging scenarios, with a special emphasis on the trade-off among accuracy, robustness and computation speed. An experimental comparison is conducted using public available benchmark datasets and author-collected datasets including long corridors, illumination changing environments and fast motion scenarios. Experimental results present both quantitative and qualitative differences among these methods and provide some guidelines on choosing the “right” algorithm for an indoor MAV according to the quality of the RGB-D data and environment characteristics.
Zheng Fang 0001, Sebastian A. Scherer
IROS1