VLDB 2026 Research / reviewers in the wild / expert
Rui Fan 0001
dblp:03/1805-1
· DBLP profile ↗
61ranked-venue papers
13as first author
48since 2021 · last 2026
0000-0003-2593-6596ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 5 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 15 since 2021Systems, architecture and hardware · 15 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Ultrasound-based Reliable Disease Diagnosis Using Causal InferenceabstractAligning the decision-making process of deep learning models with that of experienced sonographers is essential for ultrasound-based reliable disease diagnosis. Although existing methods have made significant progress in this aspect, their alignments are primarily associational rather than causal, leading to pseudo-correlations between features and diagnostic results. Such a biased diagnosis blindly models the sonographer's diagnostic skills and attention to specific patterns, which we argue hardly produces an AI diagnoser that is comparable to human experts. To address this issue, we propose a causality-based diagnostic framework to align the model's diagnostic behaviors with those of experts. Specifically, by delving into both conspicuous and inconspicuous confounders within the ultrasound images, the back-door and front-door adjustment causal learning modules are proposed to promote unbiased learning by mitigating potential pseudo-correlations. In addition, we integrate causal inference into a well-designed dual-branch model with feature interaction bridges for compatibility with multimodal ultrasound inputs. To fully evaluate our method, we conduct comparative studies on different diseases and ultrasound modalities. In particular, we publish a carefully constructed multimodal ultrasound dataset for breast lesion diagnosis and segmentation. Sufficient comparative and ablation studies on this dataset emphasize that our method outperforms state-of-the-art methods. Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Ping Zhong 0002, Yixiong Liang, Rui Fan 0001, Jianxin Wang 0001 |
AAAI | 6 |
| 2026 | Edge Self-Adversarial Augmentation Enhances Graph Contrastive Learning Against Neighborhood InconsistencyabstractRecent studies have shown that unsupervised graph contrastive learning (GCL) is vulnerable to adversarial attacks. Automatic adversarial augmentation techniques are proposed to improve both the effectiveness and robustness of GCL. Existing methods typically regard unsupervised contrastive loss as the adversarial goal, essentially aiming to maximize inter-view instance-wise discrepancies between adversarial and original views. However, such attacks overlook intra-view neighborhood inconsistency, which hinders the robustness of GCL models against local neighborhood noises, resulting in performance degradation on low-homophily graphs. To tackle this issue, we propose a novel adversarial contrastive paradigm, named Edge self-aDversarial Augmentation for Graph Contrastive Learning (EDA-GCL). We theoretically establish that the adversarial objective of the intra-view neighborhood is equivalent to maximizing the discrepancy between bidirectional edge features. Hence, we build our adversarial framework based on edge self-adversarial learning. It generates pairwise adversarial augmentations from the original view by learning distinct neighborhood connectivity structures. The learned pairwise adversarial views are utilized for GCL model training in the minimization stage. Notably, this edge-level adversarial approach reduces the computational complexity to the level of the edge number. Experiments on various graph tasks and complex noise scenarios demonstrate the superiority and robustness of our EDA-GCL. Chunchun Chen, Chenrun Wang, Yiwei Fu, Xin Sun 0003, Rui Fan 0001, Wei Ye 0001 |
AAAI | 8 |
| 2026 | WP-CrackNet: A collaborative adversarial learning framework for end-to-end weakly-supervised road crack detection
Nachuan Ma, Zhengfei Song, Chengxi Zhang, Rui Fan 0001, Lihua Xie 0001 |
Neurocomputing | 6 |
| 2026 | FDPMambaFuse: A frequency-domain and parallel Mamba-based model for multimodal medical image fusion
Baiya Li, Boheng Zhang, Yang Liu 0106, Haorui Huang, Cailing Lin, Rui Fan 0001, Mingjian Sun |
Neural Networks | 6 |
| 2026 | Enhancing graph learning interpretability through modulating cluster information flow
Wei Ye 0001, Rui Fan 0001, Jungong Han |
Pattern Recognit. | 4 |
| 2026 | Integrating Disparity Confidence Estimation Into Relative Depth Prior-Guided Unsupervised Stereo MatchingabstractUnsupervised stereo matching has garnered significant attention for its independence from costly disparity annotations. Typical unsupervised methods rely on the multi-view consistency assumption for training networks, which suffer considerably from stereo matching ambiguities, such as repetitive patterns and texture-less regions. A feasible solution lies in transferring 3D geometric knowledge from a relative depth map to the stereo matching networks. However, existing knowledge transfer methods learn depth ranking information from randomly built sparse correspondences, which makes inefficient utilization of 3D geometric knowledge and introduces noise from mistaken disparity estimates. This work proposes a novel unsupervised learning framework to address these challenges, which comprises a plug-and-play disparity confidence estimation algorithm and two depth prior-guided loss functions. Specifically, the local coherence consistency between neighboring disparities and their corresponding relative depths is first checked to obtain disparity confidence. Afterwards, quasi-dense correspondences are built using only confident disparity estimates to facilitate efficient depth ranking learning. Finally, a dual disparity smoothness loss is proposed to boost stereo matching performance at disparity discontinuities. Experimental results demonstrate that our method achieves state-of-the-art stereo matching accuracy on the KITTI Stereo benchmarks among all unsupervised stereo matching methods. Mingjian Sun, Cairong Zhao, Hanli Wang, Alexander V. Dvorkovich, Rui Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Treasure Hunting: Embodied Contrastive Learning-Enhanced Coarse-to-Fine Object Seeking With Explorer and Discriminator CooperationabstractObject navigation (ObjcetNav), which enables an agent to seek any instance of an object category, has shown great advances. However, current agents are built upon occlusion-prone visual observations or compressed 2-D maps, which hinder their embodied perception of 3-D scene geometry. Furthermore, existing methods usually decouple ObjectNav into the exploration and exploitation subtasks, easily leading to ambiguous object localization and blind exploration. To address these issues, we first propose an embodied contrastive learning (ECL) method with geometric consistency (GC) and behavioral awareness (BA), which motivates agents to encode 3-D scene layouts and semantic cues actively. The BA is modeled by predicting navigational actions based on multiframe visual images, as behaviors causing differences between adjacent visual sensations are crucial for learning correlations among continuous visions. The GC is modeled by aligning the behavior-aware visual stimulus with 3-D semantic shapes through unsupervised contrastive learning. Then, based on the above ECL pretraining, a coarse-to-fine ObjectNav policy with explorer and discriminator cooperation is proposed, inspired by the treasure-hunting mindset. Concretely, the explorer is designed to adaptively switch the action spaces, thereby switching the global and local exploration thoughts according to the accumulated scene priors. The discriminator is designed to discriminate the target's authenticity using behavior-aware visual features and geometric invariance priors, which permits mimicking the human behavior of "approaching to confirm" when distinguishing objects from a distance. As expected, our ECL method performs well on object detection (ObjDet) and instance segmentation (InstSeg) tasks. Our ECL-enhanced ObjectNav strategy outperforms state-of-the-art (SOTA) methods on Matterport3D (MP3D), Gibson, and HM3D datasets. Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Ping Zhong 0002, Rui Fan 0001, Jianxin Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | EventPillars: Pillar-based Efficient Representations for Event DataabstractEvent Cameras offer appealing advantages, including power efficiency and ultra-low latency, driving forward advancements in edge applications. In order to leverage mature frame-based algorithms, most approaches typically compute dense, image-like representations from sparse, asynchronous events. However, they are often unable to capture comprehensive information or are computationally intensive, which hinders the edge deployment of event-based vision. Meanwhile, pillar-based paradigms have been proven to be efficient and well established for dense representations of sparse data. Hence, from a novel pillar-based perspective, we present EventPillars, an efficient, comprehensive framework for dense event representations. To summarize, it (i) incorporates the Temporal Event Range to describe an intact temporal distribution, (ii) Activates the Event Polarities to explicitly record the scene dynamics, (iii) enhances the target awareness by a spatial attention prior from Normalized Event Density, (iv) can be plug-and-played into different downstream tasks. Extensive experiments show that our EventPillars records a new state-of-the-art precision on object recognition and detection datasets with surprisingly 9.2× and 4.5× lower computation and storage consumption. This brings a new insight into dense event representations and is promising to boost the edge deployment of event-based vision. Rui Fan 0001, Weidong Hao, Juntao Guan, Lai Rui, Lin Gu 0003, Fanhong Zeng, Zhangming Zhu |
AAAI | 1 |
| 2025 | ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy PredictionabstractInferring the 3D structure of a scene from a single image is an ill-posed and challenging problem in the field of vision-centric autonomous driving. Existing methods usually employ neural radiance fields to produce voxelized 3D occupancy, lacking instance-level semantic reasoning and temporal photometric consistency. In this paper, we propose ViPOcc, which leverages the visual priors from vision foundation models (VFMs) for fine-grained 3D occupancy prediction. Unlike previous works that solely employ volume rendering for RGB and depth image reconstruction, we introduce a metric depth estimation branch, in which an inverse depth alignment module is proposed to bridge the domain gap in depth distribution between VFM predictions and the ground truth. The recovered metric depth is then utilized in temporal photometric alignment and spatial geometric alignment to ensure accurate and consistent 3D occupancy prediction. Additionally, we also propose a semantic-guided non-overlapping Gaussian mixture sampler for efficient, instance-aware ray sampling, which addresses the redundant and imbalanced sampling issue that still exists in previous state-of-the-art methods. Extensive experiments demonstrate the superior performance of ViPOcc in both 3D occupancy prediction and depth estimation tasks on diverse public datasets. Xijing Zhang, Tanghui Li, Yanting Zhang 0001, Rui Fan 0001 |
AAAI | 6 |
| 2025 | An Efficient Hybrid Vision Transformer for Tinyml Applications
Fanhong Zeng, Huanan Li, Juntao Guan, Rui Fan 0001 |
ICCV | 4 |
| 2025 | RLCNet: A Novel Deep Feature-Matching-Based Method for Online Target-Free Radar-LiDAR CalibrationabstractWhile millimeter-wave radars are widely used in robotics and autonomous driving, extrinsic calibration with other sensors remains challenging due to the sparsity and uncertainty of radar point clouds. In this paper, we propose a novel deep feature-matching-based online extrinsic calibration approach for a 4D millimeter-wave radar and 3D LiDAR system. We formulate the calibration problem as a crossmodal point cloud registration task, initiating with keypointlevel matching followed by dense matching refinement. Efficient yet powerful neural networks are employed to extract prior keypoint matches, which are then expanded to surrounding regions, establishing dense point correspondences. Our approach effectively leverages the majority of the information from millimeter-wave radar, mitigating the impact of radar point cloud sparsity. We evaluate our approach on two datasets, and experimental results demonstrate that it outperforms state-of-the-art baseline methods and achieves an average improvement of 66.96% in calibration success rate, while reducing translational error and rotational error by 23.84% and 30.31%, respectively. Our implementation will be made open-source at https://github.com/nubot-nudt/RLCNet. Kai Luan, Chenghao Shi, Xieyuanli Chen, Rui Fan 0001, Zhiqiang Zheng 0002, Huimin Lu 0002 |
ICRA | 4 |
| 2025 | KDMOS:Knowledge Distillation for Motion SegmentationabstractMotion Object Segmentation (MOS) is crucial for autonomous driving, as it enhances localization, path planning, map construction, scene flow estimation, and future state prediction. While existing methods achieve strong performance, balancing accuracy and real-time inference remains a challenge. To address this, we propose a logits-based knowledge distillation framework for MOS, aiming to improve accuracy while maintaining real-time efficiency. Specifically, we adopt a Bird’s Eye View (BEV) projection-based model as the student and a non-projection model as the teacher. To handle the severe imbalance between moving and non-moving classes, we decouple them and apply tailored distillation strategies, allowing the teacher model to better learn key motion-related features. This approach significantly reduces false positives and false negatives. Additionally, we introduce dynamic upsampling, optimize the network architecture, and achieve a 7.69% reduction in parameter count, mitigating overfitting. Our method achieves a notable IoU of 78.8% on the hidden test set of the SemanticKITTI-MOS dataset and delivers competitive results on the Apollo dataset. The KDMOS implementation is available at https://github.com/SCNU-RISLAB/KDMOS. Chunyu Cao, Jintao Cheng, Linfan Zhan, Rui Fan 0001, Zhijian He |
IROS | 5 |
| 2025 | PanopticSplatting: End-to-End Panoptic Gaussian SplattingabstractOpen-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the dependence of hand-designed components. To streamline the pipeline and achieve global optimization, we propose PanopticSplatting, an end-to-end system for open-vocabulary panoptic reconstruction. Our method introduces query-guided Gaussian segmentation with local cross attention, lifting 2D instance masks without cross-frame association in an end-to-end way. The local cross attention within view frustum effectively reduces the training memory, making our model more accessible to large scenes with more Gaussians and objects. In addition, to address the challenge of noisy labels in 2D pseudo masks, we propose label blending to promote consistent 3D segmentation with less noisy floaters, as well as label warping on 2D predictions which enhances multi-view coherence and segmentation accuracy. Our method demonstrates strong performances in 3D scene panoptic reconstruction on the ScanNet-V2 and ScanNet++ datasets, compared with both NeRF-based and Gaussian-based panoptic reconstruction methods. Moreover, PanopticSplatting can be easily generalized to numerous variants of Gaussian splatting, and we demonstrate its robustness on different Gaussian base models. Changjian Jiang, Sitong Mao, Shunbo Zhou, Rui Fan 0001, Rong Xiong, Yue Wang 0020 |
IROS | 6 |
| 2025 | Preference-driven Knowledge Distillation for Few-shot Node ClassificationabstractGraph neural networks (GNNs) can efficiently process text-attributed graphs (TAGs) due to their message-passing mechanisms, but their training heavily relies on the human-annotated labels. Moreover, the complex and diverse local topologies of nodes of real-world TAGs make it challenging for a single mechanism to handle. Large language models (LLMs) perform well in zero-/few-shot learning on TAGs but suffer from a scalability challenge. Therefore, we propose a preference-driven knowledge distillation (PKD) framework to synergize the complementary strengths of LLMs and various GNNs for few-shot node classification. Specifically, we develop a GNN-preference-driven node selector that effectively promotes prediction distillation from LLMs to teacher GNNs. To further tackle nodes' intricate local topologies, we develop a node-preference-driven GNN selector that identifies the most suitable teacher GNN for each node, thereby facilitating tailored knowledge distillation from teacher GNNs to the student GNN. Extensive experiments validate the efficacy of our proposed framework in few-shot node classification on real-world TAGs.
Our code can be available at <https://github.com/GEEX-Weixing/PKD>. Chunchun Chen, Rui Fan 0001, Xiaofeng Cao 0002, Sourav Medya, Wei Ye 0001 |
NeurIPS | 3 |
| 2025 | DepthMatch: Semi-Supervised RGB-D Scene Parsing Through Depth-Guided RegularizationabstractRGB-D scene parsing methods effectively capture both semantic and geometric features of the environment, demonstrating great potential under challenging conditions such as extreme weather and low lighting. However, existing RGB-D scene parsing methods predominantly rely on supervised training strategies, which require a large amount of manually annotated pixel-level labels that are both time-consuming and costly. To overcome these limitations, we introduce DepthMatch, a semi-supervised learning framework that is specifically designed for RGB-D scene parsing. To make full use of unlabeled data, we propose complementary patch mix-up augmentation to explore the latent relationships between texture and spatial features in RGB-D image pairs. We also design a lightweight spatial prior injector to replace traditional complex fusion modules, improving the efficiency of heterogeneous feature fusion. Furthermore, we introduce depth-guided boundary loss to enhance the model's boundary prediction capabilities. Experimental results demonstrate that DepthMatch exhibits high applicability in both indoor and outdoor scenes, achieving state-of-the-art results on the NYUv2 dataset and ranking first on the KITTI Semantics benchmark. Our source code will be publicly available at mias.group/DepthMatch. Jiahang Li 0001, Sergey Vityazev, Alexander V. Dvorkovich, Rui Fan 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | MMFSeg: Multi-Structure Multi-Feature Fusion for Segmentation of Road Potholes
Yanning Guo, Rui Fan 0001, Yuxiang Sun 0002 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature MatchingabstractPoint-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in the modality gap between unstructured point clouds and structured images, especially under sparse single-frame LiDAR settings. Existing methods typically extract features separately from point clouds and images, then rely on hand-crafted or learned matching strategies. This separate encoding fails to bridge the modality gap effectively, and more critically, these methods struggle with the sparsity and noise of single-frame LiDAR, often requiring point cloud accumulation or additional priors to improve reliability. Inspired by recent progress in detector-free matching paradigms, we revisit the projection-based approach and introduce the detector-free framework for direct point-pixel matching between LiDAR and camera views. To further enhance matching reliability, we introduce a repeatability scoring mechanism that acts as a soft visibility prior. This guides the network to suppress unreliable matches in regions with low intensity variation, improving robustness under sparse input. Extensive experiments on KITTI, nuScenes, and MIAS-LCEC-TF70 benchmarks demonstrate that our method achieves state-of-the-art performance, outperforming prior approaches on nuScenes (even those relying on accumulated point clouds), despite using only single-frame LiDAR. Yanting Zhang 0001, Fangjun Ding, Shen Cai, Yanchao Dong, Rui Fan 0001 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Environment-Driven Online LiDAR-Camera Extrinsic CalibrationabstractLiDAR-camera extrinsic calibration (LCEC) is crucial for multi-modal data fusion in autonomous robotic systems. Existing methods, whether target-based or target-free, typically rely on customized calibration targets or fixed scene types, which limit their applicability in real-world scenarios. To address these challenges, we present EdO-LCEC, the first environment-driven online calibration approach. Unlike traditional target-free methods, EdO-LCEC employs a generalizable scene discriminator to estimate the feature density of the application environment. Guided by this feature density, EdO-LCEC extracts LiDAR intensity and depth features from varying perspectives to achieve higher calibration accuracy. To overcome the challenges of cross-modal feature matching between LiDAR and camera, we introduce dual-path correspondence matching (DPCM), which leverages both structural and textural consistency for reliable 3D-2D correspondences. Furthermore, we formulate the calibration process as a joint optimization problem that integrates global constraints across multiple views and scenes, thereby enhancing overall accuracy. Extensive experiments on real-world datasets demonstrate that EdO-LCEC outperforms state-of-the-art methods, particularly in scenarios involving sparse point clouds or partially overlapping sensor views. Hongbo Zhao 0009, Ping Zhong 0002, Xiao-Hu Zhou, Wei Ye 0001, Rui Fan 0001 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Real-Time Metric-Semantic Mapping for Autonomous Navigation in Outdoor EnvironmentsabstractThe creation of a metric-semantic map, which encodes human-prior knowledge, represents a high-level abstraction of environments. However, constructing such a map poses challenges related to the fusion of multi-modal sensor data, the attainment of real-time mapping performance, and the preservation of structural and semantic information consistency. In this paper, we introduce an online metric-semantic mapping system that utilizes LiDAR-Visual-Inertial sensing to generate a global metric-semantic mesh map of large-scale outdoor environments. Leveraging GPU acceleration, our mapping process achieves exceptional speed, with frame processing taking less than$7ms$, regardless of scenario scale. Furthermore, we seamlessly integrate the resultant map into a real-world navigation system, enabling metric-semantic-based terrain assessment and autonomous point-to-point navigation within a campus environment. Through extensive experiments conducted on both publicly available and self-collected datasets comprising 24 sequences, we demonstrate the effectiveness of our mapping and navigation methodologies. Note to Practitioners—This paper tackles the challenge of autonomous navigation for mobile robots in complex, unstructured environments with rich semantic elements. Traditional navigation relies on geometric analysis and manual annotations, struggling to differentiate similar structures like roads and sidewalks. We propose an online mapping system that creates a global metric-semantic mesh map for large-scale outdoor environments, utilizing GPU acceleration for speed and overcoming the limitations of existing real-time semantic mapping methods, which are generally confined to indoor settings. Our map integrates into a real-world navigation system, proven effective in localization and terrain assessment through experiments with both public and proprietary datasets. Future work will focus on integrating kernel-based methods to improve the map’s semantic accuracy. Jianhao Jiao, Ruoyu Geng, Yuanhang Li, Ren Xin, Jin Wu 0002, Lujia Wang 0001, Ming Liu 0001, Rui Fan 0001, Dimitrios Kanoulas |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2025 | TiCoSS: Tightening the Coupling Between Semantic Segmentation and Stereo Matching Within a Joint Learning FrameworkabstractSemantic segmentation and stereo matching, respectively analogous to the ventral and dorsal streams in our human brain, are two key components of autonomous driving perception systems. Addressing these two tasks with separate networks is no longer the mainstream direction in developing computer vision algorithms, particularly with the recent advances in large vision models and embodied artificial intelligence. The trend is shifting towards combining them within a joint learning framework, especially emphasizing feature sharing between the two tasks. The major contributions of this study lie in comprehensively tightening the coupling between semantic segmentation and stereo matching. Specifically, this study makes three key contributions: (1) a tightly coupled, gated feature fusion strategy, (2) a hierarchical deep supervision strategy, and (3) a coupling tightening loss function. The combined use of these technical contributions results in TiCoSS, a state-of-the-art joint learning framework that simultaneously tackles semantic segmentation and stereo matching. Through extensive experiments on the KITTI, vKITTI2, and Cityscapes datasets, along with both qualitative and quantitative analyses, we validate the effectiveness of our developed strategies and loss function. Our approach demonstrates superior performance compared to prior arts, with a notable increase in mean intersection over union by over 9%. Guanfeng Tang, Jiahang Li 0001, Ping Zhong 0002, Wei Ye 0001, Xieyuanli Chen, Huimin Lu 0002, Rui Fan 0001 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Three-Filters-to-Normal+: Revisiting Discontinuity Discrimination in Depth-to-Normal TranslationabstractThis article introduces three-filters-to-normal$+$(3F2N$+$), an extension of our previous work three-filters-to-normal (3F2N), with a specific focus on incorporating discontinuity discrimination capability into surface normal estimators (SNEs). 3F2N$+$achieves this capability by utilizing a novel discontinuity discrimination module (DDM), which combines depth curvature minimization and correlation coefficient maximization through conditional random fields (CRFs). To evaluate the robustness of SNEs on noisy data, we create a large-scale synthetic surface normal (SSN) dataset containing 20 scenarios (ten indoor scenarios and ten outdoor scenarios with and without random Gaussian noise added to depth images). Extensive experiments demonstrate that 3F2N$+$achieves greater performance than all other geometry-based surface normal estimators, with average angular errors of 7.85$^\circ$, 8.95$^\circ$, 9.25$^\circ$, and 11.98$^\circ$on the clean-indoor, clean-outdoor, noisy-indoor, and noisy-outdoor datasets, respectively. We conduct three additional experiments to demonstrate the effectiveness of incorporating our proposed 3F2N$+$into downstream robot perception tasks, including freespace detection, 6D object pose estimation, and point cloud completion. Our source code and datasets are publicly available at https://mias.group/3F2Nplus.Note to Practitioners—The primary motivation behind this work arises from the need to develop a high-performing surface normal estimator for practical robotics and computer vision applications. While geometry-based surface normal estimators have been widely used in these domains, the existing solutions focus merely on discontinuity discrimination. To tackle this problem, this article introduces a plug-and-play module that leverages both depth curvature and correlation coefficient to quantify discontinuity levels, thereby optimizing surface normal estimation, particularly near or on discontinuous regions. Moreover, this article also introduces a large-scale public dataset with random noise added to depth images, providing a more realistic and robust platform for algorithm evaluation within this research community. Extensive experimental results demonstrate that our method outperforms other state-of-the-art algorithms. Jingwei Yang 0002, Bohuan Xue, Deming Wang, Rui Fan 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | LIX: Implicitly Infusing Spatial Geometric Prior Knowledge Into Visual Semantic Segmentation for Autonomous DrivingabstractDespite the impressive performance achieved by data-fusion networks with duplex encoders for visual semantic segmentation, they become ineffective when spatial geometric data are not available. Implicitly infusing the spatial geometric prior knowledge acquired by a data-fusion teacher network into a single-modal student network is a practical, albeit less explored research avenue. This article delves into this topic and resorts to knowledge distillation approaches to address this problem. We introduce the Learning to Infuse "X" (LIX) framework, with novel contributions in both logit distillation and feature distillation aspects. We present a mathematical proof that underscores the limitation of using a single, fixed weight in decoupled knowledge distillation and introduce a logit-wise dynamic weight controller as a solution to this issue. Furthermore, we develop an adaptively-recalibrated feature distillation algorithm, including two novel techniques: feature recalibration via kernel regression and feature consistency quantification via centered kernel alignment. Extensive experiments conducted with intermediate-fusion and late-fusion networks across various public datasets provide both quantitative and qualitative evaluations, demonstrating the superior performance of our LIX framework when compared to other state-of-the-art approaches. Source code is available at https://mias.group/LIX. Sicen Guo, Ziwei Long, Ioannis Pitas, Rui Fan 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | These Maps Are Made by Propagation: Adapting Deep Stereo Networks to Road Scenarios With Decisive Disparity DiffusionabstractStereo matching has emerged as a cost-effective solution for road surface 3D reconstruction, garnering significant attention towards improving both computational efficiency and accuracy. This article introduces decisive disparity diffusion (D3Stereo), marking the first exploration of dense deep feature matching that adapts pre-trained deep convolutional neural networks (DCNNs) to previously unseen road scenarios. A pyramid of cost volumes is initially created using various levels of learned representations. Subsequently, a novel recursive bilateral filtering algorithm is employed to aggregate these costs. A key innovation of D3Stereo lies in its alternating decisive disparity diffusion strategy, wherein intra-scale diffusion is employed to complete sparse disparity images, while inter-scale inheritance provides valuable prior information for higher resolutions. Extensive experiments conducted on our created UDTIRI-Stereo and Stereo-Road datasets underscore the effectiveness of D3Stereo strategy in adapting pre-trained DCNNs and its superior performance compared to all other explicit programming-based algorithms designed specifically for road surface 3D reconstruction. Additional experiments conducted on the Middlebury dataset with backbone DCNNs pre-trained on the ImageNet database further validate the versatility of D3Stereo strategy in tackling general stereo matching problems. Our source code and supplementary material are publicly available at https://mias.group/D3-Stereo. Yikang Zhang 0001, Ioannis Pitas, Rui Fan 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | DCPI-Depth: Explicitly Infusing Dense Correspondence Prior to Unsupervised Monocular Depth EstimationabstractThere has been a recent surge of interest in learning to perceive depth from monocular videos in an unsupervised fashion. A key challenge in this field is achieving robust and accurate depth estimation in regions with weak textures or where dynamic objects are present. This study makes three major contributions by delving deeply into dense correspondence priors to provide existing frameworks with explicit geometric constraints. The first novel contribution is a contextual-geometric depth consistency loss, which employs depth maps triangulated from dense correspondences based on estimated ego-motion to guide the learning of depth perception from contextual information, since explicitly triangulated depth maps capture accurate relative distances among pixels. The second novel contribution arises from the observation that there exists an explicit, deducible relationship between optical flow divergence and depth gradient. A differential property correlation loss is therefore designed to refine depth estimation with a specific emphasis on local variations. The third novel contribution is a bidirectional stream co-adjustment strategy that enhances the interaction between rigid and optical flows, encouraging the former towards more accurate correspondence and making the latter more adaptable across various scenarios under the static scene hypotheses. DCPI-Depth, a framework that incorporates all these innovative components and couples two bidirectional and collaborative streams, achieves state-of-the-art performance and generalizability across multiple public datasets, outperforming all existing prior arts. Specifically, it demonstrates accurate depth estimation in texture-less and dynamic regions, and shows more reasonable smoothness. Our source code is publicly available at https://mias.group/DCPI-Depth. Mengtan Zhang, Rui Fan 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Zone-YOLO: Vision-Language Object Detection Using Zone PromptabstractObject detection in complex traffic scenarios is crucial for Intelligent Transportation Systems (ITS). At present, most real-time traffic object detection methods primarily rely on YOLO-style vision-only detectors, limiting their potential for further improvement. Vision-Language Object Detection (VLOD) has made promising progress currently, yet its adoption in the realm of ITS remains limited. Previous VLOD methods utilize text features in the classification task, without fully exploring their impact on the regression process for object localization. Besides, existing multi-modal fusion approaches fail to fuse text features with multi-scale image features at corresponding scales, which is detrimental to the representation capability of the model. In this work, we dive into the limitations above and introduce Zone-YOLO to improve the VLOD to a new level. Specifically, we propose Scale-Aware Modal Fusion (SAMF) to fully exploit the text and image features and learn to fuse the multi-modal representations seamlessly at different scales with channel- and modal-wise enhancement. Moreover, we present a novel Zone Prompt learning method to introduce text features into regression process and capture the zone-class-entity triple co-occurrence, which significantly improves the localization performance of the model. Extensive experiments show that Zone-YOLO outperforms the comparative methods by a considerable margin, achieving 55.1 AP, 72.1 AP50 and 71.2 APL on COCO. The competitive results on BDD100K and VisDrone2019 further demonstrate the superiority of Zone-YOLO on efficient traffic object detection. Jiaxiong Yang, Ning Jia 0003, Rui Fan 0001, Yougang Sun |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | MF-MOS: A Motion-Focused Model for Moving Object SegmentationabstractMoving object segmentation (MOS) provides a reliable solution for detecting traffic participants and thus is of great interest in the autonomous driving field. Dynamic capture is always critical in the MOS problem. Previous methods capture motion features from the range images directly. Differently, we argue that the residual maps provide greater potential for motion information, while range images contain rich semantic guidance. Based on this intuition, we propose MF-MOS, a novel motion-focused model with a dual-branch structure for LiDAR moving object segmentation. Novelly, we decouple the spatial-temporal information by capturing the motion from residual maps and generating semantic features from range images, which are used as movable object guidance for the motion branch. Our straightforward yet distinctive solution can make the most use of both range images and residual maps, thus greatly improving the performance of the LiDAR-based MOS task. Remarkably, our MF-MOS achieved a leading IoU of 76.7% on the MOS leaderboard of the SemanticKITTI dataset upon submission, demonstrating the current state-of-the-art performance. The implementation of our MF-MOS has been released at https://github.com/SCNU-RISLAB/MF-MOS. Jintao Cheng, Kang Zeng, Zhuoxu Huang, Jin Wu 0002, Chengxi Zhang, Xieyuanli Chen, Rui Fan 0001 |
ICRA | 8 |
| 2024 | Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor DistillationabstractCorrespondence matching plays a crucial role in numerous robotics applications. In comparison to conventional hand-crafted methods and recent data-driven approaches, there is significant interest in plug-and-play algorithms that make full use of pre-trained backbone networks for multi-scale feature extraction and leverage hierarchical refinement strategies to generate matched correspondences. The primary focus of this paper is to address the limitations of deep feature matching (DFM), a state-of-the-art (SoTA) plug-and-play correspondence matching approach. First, we eliminate the pre-defined threshold employed in the hierarchical refinement process of DFM by leveraging a more flexible nearest neighbor search strategy, thereby preventing the exclusion of repetitive yet valid matches during the early stages. Our second technical contribution is the integration of a patch descriptor, which extends the applicability of DFM to accommodate a wide range of backbone networks pre-trained across diverse computer vision tasks, including image classification, semantic segmentation, and stereo matching. Taking into account the practical applicability of our method in real-world robotics applications, we also propose a novel patch descriptor distillation strategy to further reduce the computational complexity of correspondence matching. Extensive experiments conducted on three public datasets demonstrate the superior performance of our proposed method. Specifically, it achieves an overall performance in terms of mean matching accuracy of 0.68, 0.92, and 0.95 with respect to the tolerances of 1, 3, and 5 pixels, respectively, on the HPatches dataset, outperforming all other SoTA algorithms. Our source code, demo video, and supplement are publicly available at mias.group/GCM. Ziwei Long, Yanting Zhang 0001, Jin Wu 0002, Zhijun Fang 0001, Rui Fan 0001 |
ICRA | 6 |
| 2024 | SG-RoadSeg: End-to-End Collision-Free Space Detection Sharing Encoder Representations Jointly Learned via Unsupervised Deep StereoabstractCollision-free space detection is of utmost importance for autonomous robot perception and navigation. State-of-the-art (SoTA) approaches generally extract features from RGB images and an additional source or modality of 3-D information, such as depth or disparity images, using a pair of independent encoders. The extracted features are subsequently fused and decoded to yield semantic predictions of collision-free spaces. Such feature-fusion approaches become infeasible in scenarios, where the sensor for 3-D information acquisition is unavailable, or just when multi-sensor calibration falls short of the necessary precision. To overcome these limitations, this paper introduces a novel end-to-end collision-free space detection network, referred to as SG-RoadSeg, built upon our previous work SNE-RoadSeg. A key contribution of this paper is a strategy for sharing encoder representations that are co-learned through both semantic segmentation and unsupervised stereo matching tasks, enabling the features extracted from RGB images to contain both semantic and spatial geometric information. The unsupervised deep stereo serves as an auxiliary functionality, capable of generating accurate disparity maps that can be used by other perception tasks that require depth-related data. Comprehensive experimental results on the KITTI road and semantics datasets validate the effectiveness of our proposed architecture and encoder representation sharing strategy. SG-RoadSeg also demonstrates superior performance than other SoTA collision-free space detection approaches. Our source code, demo video, and supplement are publicly available at mias.group/SG-RoadSeg. Wei Ye 0001, Rui Fan 0001 |
ICRA | 7 |
| 2024 | Dive Deeper into Rectifying Homography for Stereo Camera Online Self-CalibrationabstractAccurate estimation of stereo camera extrinsic parameters is crucial to guarantee the performance of stereo matching algorithms. In prior arts, the online self-calibration of stereo cameras has commonly been formulated as a specialized visual odometry problem, without taking into account the principles of stereo rectification. In this paper, we first delve deeply into the concept of rectifying homography, which serves as the cornerstone for the development of our novel stereo camera online self-calibration algorithm, for cases where only a single pair of images is available. Furthermore, we introduce a simple yet effective solution for global optimum extrinsic parameter estimation in the presence of stereo video sequences. Additionally, we emphasize the impracticality of using three Euler angles and three components in the translation vectors for performance quantification. Instead, we introduce four new evaluation metrics to quantify the robustness and accuracy of extrinsic parameter estimation, applicable to both single-pair and multi-pair cases. Extensive experiments conducted across indoor and outdoor environments using various experimental setups validate the effectiveness of our proposed algorithm. The comprehensive evaluation results demonstrate its superior performance in comparison to the baseline algorithm. Our source code, demo video, and supplement are publicly available at mias.group/StereoCalibrator. Hongbo Zhao 0009, Yikang Zhang 0001, Rui Fan 0001 |
ICRA | 4 |
| 2024 | Rotation-equivariant correspondence matching based on a dual-activation mixer
Shuai Su, Ronghao Dang, Rui Fan 0001 |
Neurocomputing | 3 |
| 2024 | UDTIRI: An Online Open-Source Intelligent Road Inspection Benchmark SuiteabstractIn the emerging field of urban digital twins (UDTs), there are extensive and captivating opportunities for leveraging cutting-edge deep learning techniques. Particularly within the specialized area of intelligent road inspection (IRI), a noticeable gap exists, underscored by the current dearth of dedicated research efforts and the lack of large-scale well-annotated datasets. To foster advancements in this burgeoning field, we have launched an online open-source benchmark suite, referred to as UDTIRI. Along with this article, we introduce the road pothole detection task, the first online competition published within this benchmark suite. This task provides a well-annotated dataset, comprising 1,000 RGB images and their pixel/instance-level ground-truth annotations, captured in diverse real-world scenarios under different illumination and weather conditions. Our benchmark provides a systematic and thorough evaluation of state-of-the-art object detection, semantic segmentation, and instance segmentation networks, developed based on either convolutional neural networks or Transformers. We anticipate that our benchmark suite will serve as a catalyst for the integration of advanced UDT techniques into IRI. By providing algorithms with a more comprehensive understanding of diverse road conditions, we seek to unlock their untapped potential and foster innovation in this critical domain. Sicen Guo, Jiahang Li 0001, Dacheng Zhou, Denghuang Zhang, Shuai Su, Xingyi Zhu, Rui Fan 0001 |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2024 | UP-CrackNet: Unsupervised Pixel-Wise Road Crack Detection via Adversarial Image RestorationabstractOver the past decade, automated methods have been developed to detect cracks more efficiently, accurately, and objectively, with the ultimate goal of replacing conventional manual visual inspection techniques. Among these methods, semantic segmentation algorithms have demonstrated promising results in pixel-wise crack detection tasks. However, training such networks requires a large amount of human-annotated datasets with pixel-level annotations, which is a highly labor-intensive and time-consuming process. Moreover, supervised learning-based methods often struggle with poor generalizability in unseen datasets. Therefore, we propose an unsupervised pixel-wise road crack detection network, known as UP-CrackNet. Our approach first generates multi-scale square masks and randomly selects them to corrupt undamaged road images by removing certain regions. Subsequently, a generative adversarial network is trained to restore the corrupted regions by leveraging the semantic context learned from surrounding uncorrupted regions. During the testing phase, an error map is generated by calculating the difference between the input and restored images, which allows for pixel-wise crack detection. Our comprehensive experimental results demonstrate that UP-CrackNet outperforms other general-purpose unsupervised anomaly detection algorithms, and exhibits satisfactory performance and superior generalizability when compared with state-of-the-art supervised crack segmentation algorithms. Our source code is publicly available at mias.group/UP-CrackNet. Nachuan Ma, Rui Fan 0001, Lihua Xie 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | D2NT: A High-Performing Depth-to-Normal TranslatorabstractSurface normal holds significant importance in visual environmental perception, serving as a source of rich geometric information. However, the state-of-the-art (SoTA) surface normal estimators (SNEs) generally suffer from an unsatisfactory trade-off between efficiency and accuracy. To resolve this dilemma, this paper first presents a superfast depth-to-normal translator (D2NT), which can directly translate depth images into surface normal maps without calculating 3D coordinates. We then propose a discontinuity-aware gradient (DAG) filter, which adaptively generates gradient convolution kernels to improve depth gradient estimation. Finally, we propose a surface normal refinement module that can easily be integrated into any depth-to-normal SNEs, substantially improving the surface normal estimation accuracy. Our proposed algorithm demonstrates the best accuracy among all other existing real-time SNEs and achieves the SoTA trade-off between efficiency and accuracy. Bohuan Xue, Ming Liu 0001, Rui Fan 0001 |
ICRA | 5 |
| 2023 | Transparent Objects: A Corner Case in Stereo MatchingabstractStereo matching is a common technique used in 3D perception, but transparent objects such as reflective and penetrable glass pose a challenge as their disparities are often estimated inaccurately. In this paper, we propose transparency-aware stereo (TA-Stereo), an effective solution to tackle this issue. TA-Stereo first utilizes a semantic segmentation or salient object detection network to identify transparent objects, and then homogenizes them to enable stereo matching algorithms to handle them as non-transparent objects. To validate the effectiveness of our proposed TA-Stereo strategy, we collect 260 images containing transparent objects from the KITTI Stereo 2012 and 2015 datasets and manually label pixel-level ground truth. We evaluate our strategy with six deep stereo networks and two types of transparent object detection methods. Our experiments demonstrate that TA-Stereo significantly improves the disparity accuracy of transparent objects. Our project webpage can be accessed at mias.group/TA-Stereo. Shuai Su, Rui Fan 0001 |
ICRA | 4 |
| 2023 | E3CM: Epipolar-constrained cascade correspondence matching
Chenbo Zhou, Shuai Su, Rui Fan 0001 |
Neurocomputing | 4 |
| 2023 | One-Vote Veto: Semi-Supervised Learning for Low-Shot Glaucoma DiagnosisabstractConvolutional neural networks (CNNs) are a promising technique for automated glaucoma diagnosis from images of the fundus, and these images are routinely acquired as part of an ophthalmic exam. Nevertheless, CNNs typically require a large amount of well-labeled data for training, which may not be available in many biomedical image classification applications, especially when diseases are rare and where labeling by experts is costly. This article makes two contributions to address this issue: 1) It extends the conventional Siamese network and introduces a training method for low-shot learning when labeled data are limited and imbalanced, and 2) it introduces a novel semi-supervised learning strategy that uses additional unlabeled training data to achieve greater accuracy. Our proposed multi-task Siamese network (MTSN) can employ any backbone CNN, and we demonstrate with four backbone CNNs that its accuracy with limited training data approaches the accuracy of backbone CNNs trained with a dataset that is 50 times larger. We also introduce One-Vote Veto (OVV) self-training, a semi-supervised learning strategy that is designed specifically for MTSNs. By taking both self-predictions and contrastive predictions of the unlabeled training data into account, OVV self-training provides additional pseudo labels for fine-tuning a pre-trained MTSN. Using a large (imbalanced) dataset with 66,715 fundus photographs acquired over 15 years, extensive experimental results demonstrate the effectiveness of low-shot learning with MTSN and semi-supervised learning with OVV self-training. Three additional, smaller clinical datasets of fundus images acquired under different conditions (cameras, instruments, locations, populations) are used to demonstrate the generalizability of the proposed methods. Rui Fan 0001, Christopher Bowd, Nicole Brye, Mark Christopher, Robert N. Weinreb, David J. Kriegman, Linda M. Zangwill |
IEEE Trans. Medical Imaging | 1 |
| 2022 | SDA-SNE: Spatial Discontinuity-Aware Surface Normal Estimation via Multi-Directional Dynamic ProgrammingabstractThe state-of-the-art (SoTA) surface normal estimators (SNEs) generally translate depth images into surface normal maps in an end-to-end fashion. Although such SNEs have greatly minimized the trade-off between efficiency and accuracy, their performance on spatial discontinuities, e.g., edges and ridges, is still unsatisfactory. To address this issue, this paper first introduces a novel multi-directional dynamic programming strategy to adaptively determine inliers (co-planar 3D points) by minimizing a (path) smoothness energy. The depth gradients can then be refined iteratively using a novel recursive polynomial interpolation algorithm, which helps yield more reasonable surface normals. Our introduced spatial discontinuity-aware (SDA) depth gradient refinement strategy is compatible with any depth-to-normal SNEs. Our proposed SDA-SNE achieves much greater performance than all other SoTA approaches, especially near/on spatial discontinuities. We further evaluate the performance of SDA-SNE with respect to different iterations, and the results suggest that it converges fast after only a few iterations. This ensures its high efficiency in various robotics and computer vision applications requiring real-time performance. Additional experiments on the datasets with different extents of random noise further validate our SDA-SNE's robustness and environmental adaptability. Our source code, demo video, and supplementary material are publicly available at mias.group/SDA-SNE. Nan Ming, Rui Fan 0001 |
3DV | 3 |
| 2022 | Urban Digital Twins for Intelligent Road InspectionabstractUrban digital twin (UDT) technologies offer new opportunities for intelligent road inspection (IRI). This paper first reviews the state-of-the-art algorithms used in the two key components of UDT-based IRI systems: (1) multi-temporal, multi-dimension, multi-score, and heterogeneous road data acquisition, and (2) road distress detection. This paper then summarizes the UDTIRI competition, organized in conjunction with IEEE Bigdata 2022. More details on our competition are available at sites.google.com/view/udtiri-workshop/bigdata-2022. Rui Fan 0001, Yikang Zhang 0001, Sicen Guo, Jiahang Li 0001, Shuai Su, Yanting Zhang 0001, Wenshuo Wang 0001, Yu Jiang 0003, Mohammud Junaid Bocus, Xingyi Zhu |
IEEE Big Data | 1 |
| 2022 | UnDAF: A General Unsupervised Domain Adaptation Framework for Disparity or Optical Flow EstimationabstractDisparity and optical flow estimation are respectively 1D and 2D dense correspondence matching (DCM) tasks in nature. Unsupervised domain adaptation (UDA) is crucial for their success in new and unseen scenarios, enabling networks to draw inferences across different domains without manually-labeled ground truth. In this paper, we propose a general UDA framework (UnDAF) for disparity or optical flow estimation. Unlike existing approaches based on adversarial learning that suffers from pixel distortion and dense correspondence mismatch after domain alignment, our UnDAF adopts a straightforward but effective coarse-to-fine strategy, where a co-teaching strategy (two networks evolve by complementing each other) refines DCM estimations after Fourier transform initializes domain alignment. The simplicity of our approach makes it extremely easy to guide adaptation across different domains, or more practically, from synthetic to real-world domains. Extensive experiments carried out on the KITTI and MPI Sintel benchmarks demonstrate the accuracy and robustness of our UnDAF, advancing all other state-of-the-art UDA approaches for disparity or optical flow estimation. Our project page is available at https://sites.google.com/view/undaf. Hengli Wang, Rui Fan 0001, Peide Cai, Ming Liu 0001, Lujia Wang 0001 |
ICRA | 2 |
| 2022 | Rethinking Road Surface 3-D Reconstruction and Pothole Detection: From Perspective Transformation to Disparity Map SegmentationabstractPotholes are one of the most common forms of road damage, which can severely affect driving comfort, road safety, and vehicle condition. Pothole detection is typically performed by either structural engineers or certified inspectors. However, this task is not only hazardous for the personnel but also extremely time consuming. This article presents an efficient pothole detection algorithm based on road disparity map estimation and segmentation. We first incorporate the stereo rig roll angle into shifting distance calculation to generalize perspective transformation. The road disparities are then efficiently estimated using semiglobal matching. A disparity map transformation algorithm is then performed to better distinguish the damaged road areas. Subsequently, we utilize simple linear iterative clustering to group the transformed disparities into a collection of superpixels. The potholes are finally detected by finding the superpixels, whose intensities are lower than an adaptively determined threshold. The proposed algorithm is implemented on an NVIDIA RTX 2080 Ti GPU in CUDA. The experimental results demonstrate that our proposed road pothole detection algorithm achieves state-of-the-art accuracy and efficiency. Rui Fan 0001, Umar Özgünalp, Yuan Wang 0015, Ming Liu 0001, Ioannis Pitas |
IEEE Trans. Cybern. | 1 |
| 2022 | Loop-Box: Multiagent Direct SLAM Triggered by Single Loop Closure for Large-Scale MappingabstractIn this article, we present a multiagent framework for real-time large-scale 3-D reconstruction applications. In SLAM, researchers usually build and update a 3-D map after applying nonlinear pose graph optimization techniques. Moreover, many multiagent systems are prevalently using odometry information from additional sensors. These methods generally involve extensive computer vision algorithms and are tightly coupled with various sensors. We develop a generic method for the key challenging scenarios in multiagent 3-D mapping based on different camera systems. The proposed framework performs actively in terms of localizing each agent after the first loop closure between them. It is shown that the proposed system only uses monocular cameras to yield real-time multiagent large-scale localization and 3-D global mapping. Based on the initial matching, our system can calculate the optimal scale difference between multiple 3-D maps and then estimate an accurate relative pose transformation for large-scale global mapping. M. Usman Maqbool Bhutta, Manohar Kuse, Rui Fan 0001, Ming Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Dynamic Fusion Module Evolves Drivable Area and Road Anomaly Detection: A Benchmark and AlgorithmsabstractJoint detection of drivable areas and road anomalies is very important for mobile robots. Recently, many semantic segmentation approaches based on convolutional neural networks (CNNs) have been proposed for pixelwise drivable area and road anomaly detection. In addition, some benchmark datasets, such as KITTI and Cityscapes, have been widely used. However, the existing benchmarks are mostly designed for self-driving cars. There lacks a benchmark for ground mobile robots, such as robotic wheelchairs. Therefore, in this article, we first build a drivable area and road anomaly detection benchmark for ground mobile robots, evaluating existing state-of-the-art (SOTA) single-modal and data-fusion semantic segmentation CNNs using six modalities of visual features. Furthermore, we propose a novel module, referred to as the dynamic fusion module (DFM), which can be easily deployed in existing data-fusion networks to fuse different types of visual features effectively and efficiently. The experimental results show that the transformed disparity image is the most informative visual feature and the proposed DFM-RTFNet outperforms the SOTAs. In addition, our DFM-RTFNet achieves competitive performance on the KITTI road benchmark. Hengli Wang, Rui Fan 0001, Yuxiang Sun 0002, Ming Liu 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | SCV-Stereo: Learning Stereo Matching From a Sparse Cost VolumeabstractConvolutional neural network (CNN)-based stereo matching approaches generally require a dense cost volume (DCV) for disparity estimation. However, generating such cost volumes is computationally-intensive and memory-consuming, hindering CNN training and inference efficiency. To address this problem, we propose SCV-Stereo, a novel CNN architecture, capable of learning dense stereo matching from sparse cost volume (SCV) representations. Our inspiration is derived from the fact that DCV representations are somewhat redundant and can be replaced with SCV representations. Benefiting from these SCV representations, our SCV-Stereo can update disparity estimations in an iterative fashion for accurate and efficient stereo matching. Extensive experiments carried out on the KITTI Stereo benchmarks demonstrate that our SCV-Stereo can significantly minimize the trade-off between accuracy and efficiency for stereo matching. Our project page is https://sites.google.com/view/scv-stereo. Hengli Wang, Rui Fan 0001, Ming Liu 0001 |
ICIP | 2 |
| 2021 | Co-Teaching: an Ark to Unsupervised Stereo MatchingabstractStereo matching is a key component of autonomous driving perception. Recent unsupervised stereo matching approaches have received adequate attention due to their advantage of not requiring disparity ground truth. These approaches, however, perform poorly near occlusions. To overcome this drawback, in this paper, we propose CoT-Stereo, a novel unsupervised stereo matching approach. Specifically, we adopt a co-teaching framework where two networks interactively teach each other about the occlusions in an unsupervised fashion, which greatly improves the robustness of unsupervised stereo matching. Extensive experiments on the KITTI Stereo benchmarks demonstrate the superior performance of CoT-Stereo over all other state-of-the-art unsupervised stereo matching approaches in terms of both accuracy and speed. Our project webpage is https://sites.google.com/view/cot-stereo. Hengli Wang, Rui Fan 0001, Ming Liu 0001 |
ICIP | 2 |
| 2021 | S2P2: Self-Supervised Goal-Directed Path Planning Using RGB-D Data for Robotic WheelchairsabstractPath planning is a fundamental capability for autonomous navigation of robotic wheelchairs. With the impressive development of deep-learning technologies, imitation learning-based path planning approaches have achieved effective results in recent years. However, the disadvantages of these approaches are twofold: 1) they may need extensive time and labor to record expert demonstrations as training data; and 2) existing approaches could only receive high-level commands, such as turning left/right. These commands could be less sufficient for the navigation of mobile robots (e.g., robotic wheelchairs), which usually require exact poses of goals. We contribute a solution to this problem by proposing S2P2, a self-supervised goal-directed path planning approach. Specifically, we develop a pipeline to automatically generate planned path labels given as input RGB-D images and poses of goals. Then, we present a best-fit regression plane loss to train our data-driven path planning model based on the generated labels. Our S2P2 does not need pre-built maps, but it can be integrated into existing map-based navigation systems through our framework. Experimental results show that our S2P2 outperforms traditional path planning algorithms, and increases the robustness of existing map-based navigation systems. Our project page is available at https://sites.google.com/view/s2p2. Hengli Wang, Yuxiang Sun 0002, Rui Fan 0001, Ming Liu 0001 |
ICRA | 3 |
| 2021 | SNE-RoadSeg+: Rethinking Depth-Normal Translation and Deep Supervision for Freespace DetectionabstractFreespace detection is a fundamental component of autonomous driving perception. Recently, deep convolutional neural networks (DCNNs) have achieved impressive performance for this task. In particular, SNE-RoadSeg, our previously proposed method based on a surface normal estimator (SNE) and a data-fusion DCNN (RoadSeg), has achieved impressive performance in freespace detection. However, SNE-RoadSeg is computationally intensive, and it is difficult to execute in real time. To address this problem, we introduce SNE-RoadSeg+, an upgraded version of SNE-RoadSeg. SNE-RoadSeg+ consists of 1) SNE+, a module for more accurate surface normal estimation, and 2) RoadSeg+, a data-fusion DCNN that can greatly minimize the trade-off between accuracy and efficiency with the use of deep supervision. Extensive experimental results have demonstrated the effectiveness of our SNE+ for surface normal estimation and the superior performance of our SNE-RoadSeg+ over all other freespace detection approaches. Specifically, our SNE-RoadSeg+ runs in real time, and meanwhile, achieves the state-of-the-art performance on the KITTI road benchmark. Our project page is at https://www.sne-roadseg.site/sne-roadseg-plus. Hengli Wang, Rui Fan 0001, Peide Cai, Ming Liu 0001 |
IROS | 2 |
| 2021 | Agile reactive navigation for a non-holonomic mobile robot using a pixel processor arrayabstractAbstract This paper presents an agile reactive navigation strategy for driving a non‐holonomic ground vehicle around a pre‐set course of gates in a cluttered environment using a low‐cost processor array sensor. This enables machine vision tasks to be performed directly upon the sensor's image plane, rather than using a separate general‐purpose computer. The authors demonstrate a small ground vehicle running through or avoiding multiple gates at high speed using minimal computational resources. To achieve this, target tracking algorithms are developed for the Pixel Processing Array and captured images are then processed directly on the vision sensor acquiring target information for controlling the ground vehicle. The algorithm can run at up to 2000 fps outdoors and 200 fps at indoor illumination levels. Conducting image processing at the sensor level avoids the bottleneck of image transfer encountered in conventional sensors. The real‐time performance of on‐board image processing and robustness is validated through experiments. Experimental results demonstrate the algorithm's ability to enable a ground vehicle to navigate at an average speed of 2.20 m/s for passing through multiple gates and 3.88 m/s for a ‘slalom’ task in an environment featuring significant visual clutter. Laurie Bose, Colin Greatwood, Jianing Chen 0005, Rui Fan 0001, Tom Richardson 0002, Stephen J. Carey, Piotr Dudek, Walterio W. Mayol-Cuevas |
IET Image Process. | 5 |
| 2021 | Graph Attention Layer Evolves Semantic Segmentation for Road Pothole Detection: A Benchmark and AlgorithmsabstractExisting road pothole detection approaches can be classified as computer vision-based or machine learning-based. The former approaches typically employ 2D image analysis/ understanding or 3D point cloud modeling and segmentation algorithms to detect (i.e., recognize and localize) road potholes from vision sensor data, e.g., RGB images and/or depth/disparity images. The latter approaches generally address road pothole detection using convolutional neural networks (CNNs) in an end-to-end manner. However, road potholes are not necessarily ubiquitous and it is challenging to prepare a large well-annotated dataset for CNN training. In this regard, while computer vision-based methods were the mainstream research trend in the past decade, machine learning-based methods were merely discussed. Recently, we published the first stereo vision-based road pothole detection dataset and a novel disparity transformation algorithm, whereby the damaged and undamaged road areas can be highly distinguished. However, there are no benchmarks currently available for state-of-the-art (SoTA) CNNs trained using either disparity images or transformed disparity images. Therefore, in this paper, we first discuss the SoTA CNNs designed for semantic segmentation and evaluate their performance for road pothole detection with extensive experiments. Additionally, inspired by graph neural network (GNN), we propose a novel CNN layer, referred to as graph attention layer (GAL), which can be easily deployed in any existing CNN to optimize image feature representations for semantic segmentation. Our experiments compare GAL-DeepLabv3+, our best-performing implementation, with nine SoTA CNNs on three modalities of training data: RGB images, disparity images, and transformed disparity images. The experimental results suggest that our proposed GAL-DeepLabv3+ achieves the best overall pothole detection accuracy on all training data modalities. The source code, dataset, and benchmark are publicly available at mias.group/GAL-Pothole-Detection. Rui Fan 0001, Hengli Wang, Yuan Wang 0015, Ming Liu 0001, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 2020 | SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection
Rui Fan 0001, Hengli Wang, Peide Cai, Ming Liu 0001 |
ECCV (30) | 1 |
| 2020 | Applying Surface Normal Information in Drivable Area and Road Anomaly Detection for Ground Mobile RobotsabstractThe joint detection of drivable areas and road anomalies is a crucial task for ground mobile robots. In recent years, many impressive semantic segmentation networks, which can be used for pixel-level drivable area and road anomaly detection, have been developed. However, the detection accuracy still needs improvement. Therefore, we develop a novel module named the Normal Inference Module (NIM), which can generate surface normal information from dense depth images with high accuracy and efficiency. Our NIM can be deployed in existing convolutional neural networks (CNNs) to refine the segmentation performance. To evaluate the effectiveness and robustness of our NIM, we embed it in twelve state-of-the-art CNNs. The experimental results illustrate that our NIM can greatly improve the performance of the CNNs for drivable area and road anomaly detection. Furthermore, our proposed NIM-RTFNet ranks 8th on the KITTI road benchmark and exhibits a real-time inference speed. Hengli Wang, Rui Fan 0001, Yuxiang Sun 0002, Ming Liu 0001 |
IROS | 2 |
| 2020 | Pothole Detection Based on Disparity Transformation and Road Surface ModelingabstractPothole detection is one of the most important tasks for road maintenance. Computer vision approaches are generally based on either 2D road image analysis or 3D road surface modeling. However, these two categories are always used independently. Furthermore, the pothole detection accuracy is still far from satisfactory. Therefore, in this paper, we present a robust pothole detection algorithm that is both accurate and computationally efficient. A dense disparity map is first transformed to better distinguish between damaged and undamaged road areas. To achieve greater disparity transformation efficiency, golden section search and dynamic programming are utilized to estimate the transformation parameters. Otsu's thresholding method is then used to extract potential undamaged road areas from the transformed disparity map. The disparities in the extracted areas are modeled by a quadratic surface using least squares fitting. To improve disparity map modeling robustness, the surface normal is also integrated into the surface modeling process. Furthermore, random sample consensus is utilized to reduce the effects caused by outliers. By comparing the difference between the actual and modeled disparity maps, the potholes can be detected accurately. Finally, the point clouds of the detected potholes are extracted from the reconstructed 3D road surface. The experimental results show that the successful detection accuracy of the proposed system is around 98.7% and the overall pixel-level accuracy is approximately 99.6%. Rui Fan 0001, Umar Özgünalp, Brett Hosking, Ming Liu 0001, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 2020 | Corrections to "Pothole Detection Based on Disparity Transformation and Road Surface Modeling"abstractUnfortunately, we made two minor mistakes in the above paper. First of all, the first graph on row (c) inFig. 11was same as the third graph on row (c) inFig. 11. Secondly, “precision” and “recall” inTable IIIneed to be switched. The correct figure and table have no influence on the discussion and conclusions in the above paper, and they are given here. Rui Fan 0001, Umar Özgünalp, Brett Hosking, Ming Liu 0001, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 2020 | Road Damage Detection Based on Unsupervised Disparity Map SegmentationabstractThis article presents a novel road damage detection algorithm based on unsupervised disparity map segmentation. Firstly, a disparity map is transformed by minimizing an energy function with respect to stereo rig roll angle and road disparity projection model. Instead of solving this energy minimization problem using non-linear optimization techniques, we directly find its numerical solution. The transformed disparity map is then segmented using Otus's thresholding method, and the damaged road areas can be extracted. The proposed algorithm requires no parameters when detecting road damage. The experimental results illustrate that our proposed algorithm performs both accurately and efficiently. The pixel-level road damage detection accuracy is approximately 97.56%. The source code is publicly available at: https://github.com/ruirangerfan/unsupervised_disparity_map_segmentation.git. Rui Fan 0001, Ming Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Using DP Towards A Shortest Path Problem-Related ApplicationabstractThe detection of curved lanes is still challenging for autonomous driving systems. Although current cutting-edge approaches have performed well in real applications, most of them are based on strict model assumptions. Similar to other visual recognition tasks, lane detection can be formulated as a two-dimensional graph searching problem, which can be solved by finding several optimal paths along with line segments and boundaries. In this paper, we present a directed graph model, in which dynamic programming is used to deal with a specific shortest path problem. This model is particularly suitable to represent objects with long continuous shape structure, e.g., lanes and roads. We apply the designed model and proposed an algorithm for detecting lanes by formulating it as the shortest path problem. To evaluate the performance of our proposed algorithm, we tested five sequences (including 1573 frames) from the KITTI database. The results showed that our method achieves an average successful detection precision of 97.5%. Jianhao Jiao, Rui Fan 0001, Ming Liu 0001 |
ICRA | 2 |
| 2019 | Real-Time Binocular Vision Implementation on an SoC TMS320C6678 DSP
Rui Fan 0001, Sicheng Duanmu, Yilong Zhu, Jianhao Jiao, Mohammud Junaid Bocus, Yang Yu 0028, Lujia Wang 0001, Ming Liu 0001 |
ICVS | 1 |
| 2019 | Automatic Calibration of Multiple 3D LiDARs in Urban EnvironmentsabstractMultiple LiDARs have progressively emerged on autonomous vehicles for rendering a rich view and dense measurements. However, the lack of precise calibration negatively affects their potential applications. In this paper, we propose a novel system that enables automatic multi-LiDAR calibration method without any calibration target, prior environment information, and manual initialization. Our approach starts with a hand-eye calibration by aligning the motion of each sensor. The initial results are then refined by an appearance-based method by minimizing a cost function constructed by point-plane distance. Experimental results on simulated and real-world data demonstrate the reliability and accuracy of our calibration approach. The proposed approach can calibrate a multi-LiDAR system with the rotation and translation errors less than 0. 04rad and 0. 1m respectively for a mobile platform. Jianhao Jiao, Yang Yu 0028, Qinghai Liao, Haoyang Ye, Rui Fan 0001, Ming Liu 0001 |
IROS | 5 |
| 2019 | Road Crack Detection Using Deep Convolutional Neural Network and Adaptive ThresholdingabstractCrack is one of the most common road distresses which may pose road safety hazards. Generally, crack detection is performed by either certified inspectors or structural engineers. This task is, however, time-consuming, subjective and labor-intensive. In this paper, a novel road crack detection algorithm which is based on deep learning and adaptive image segmentation is proposed. Firstly, a deep convolutional neural network is trained to determine whether an image contains cracks or not. The images containing cracks are then smoothed using bilateral filtering, which greatly minimizes the number of noisy pixels. Finally, cracks are extracted from the road surface using an adaptive thresholding method. The experimental results illustrate that our network can classify images with an accuracy of 99.92%, and the cracks can be successfully extracted from the images using our proposed thresholding algorithm. Rui Fan 0001, Mohammud Junaid Bocus, Yilong Zhu, Jianhao Jiao, Fulong Ma, Ming Liu 0001 |
IV | 1 |
| 2019 | A Novel Dual-Lidar Calibration Algorithm Using Planar SurfacesabstractMultiple lidars are used on mobile vehicles for rendering a broad view to enhance the performance of perception systems. However, precise calibration of multiple lidars is challenging since the feature correspondences in scan points are sparse for providing enough constraints. To address this problem, existing methods require fixed calibration targets in scenes or rely exclusively on additional sensors. In this paper, we present a novel method that enables automatic lidar calibration without these restrictions. Three linearly independent planar surfaces appearing in surroundings is utilized to find correspondences. Two components are developed to ensure the extrinsic parameters to be found: a closed-form solver for initialization and an optimizer for refinement by minimizing a nonlinear cost function. Simulation and experimental results demonstrate the accuracy of our calibration approach with the rotation and translation errors smaller than 0.05rad and 0.1m respectively. Jianhao Jiao, Qinghai Liao, Yilong Zhu, Tianyu Liu 0008, Yang Yu 0028, Rui Fan 0001, Lujia Wang 0001, Ming Liu 0001 |
IV | 6 |
| 2018 | A novel disparity transformation algorithm for road segmentation
Rui Fan 0001, Mohammud Junaid Bocus, Naim Dahnoun |
Inf. Process. Lett. | 1 |
| 2018 | Road Surface 3D Reconstruction Based on Dense Subpixel Disparity Map EstimationabstractVarious 3D reconstruction methods have enabled civil engineers to detect damage on a road surface. To achieve the millimetre accuracy required for road condition assessment, a disparity map with subpixel resolution needs to be used. However, none of the existing stereo matching algorithms are specially suitable for the reconstruction of the road surface. Hence in this paper, we propose a novel dense subpixel disparity estimation algorithm with high computational efficiency and robustness. This is achieved by first transforming the perspective view of the target frame into the reference view, which not only increases the accuracy of the block matching for the road surface but also improves the processing speed. The disparities are then estimated iteratively using our previously published algorithm where the search range is propagated from three estimated neighbouring disparities. Since the search range is obtained from the previous iteration, errors may occur when the propagated search range is not sufficient. Therefore, a correlation maxima verification is performed to rectify this issue, and the subpixel resolution is achieved by conducting a parabola interpolation enhancement. Furthermore, a novel disparity global refinement approach developed from the Markov Random Fields and Fast Bilateral Stereo is introduced to further improve the accuracy of the estimated disparity map, where disparities are updated iteratively by minimising the energy function that is related to their interpolated correlation polynomials. The algorithm is implemented in C language with a near real-time performance. The experimental results illustrate that the absolute error of the reconstruction varies from 0.1 mm to 3 mm. Rui Fan 0001, Xiao Ai, Naim Dahnoun |
IEEE Trans. Image Process. | 1 |
| 2017 | Multiple Lane Detection Algorithm Based on Novel Dense Vanishing Point EstimationabstractThe detection of multiple curved lane markings is still a challenge for advanced driver assistance systems today, due to interference such as road markings and shadows cast by roadside structures and vehicles. The vanishing point Vpcontains the global information of the road image. Hence, Vp-based lane detection algorithms are quite insensitive to interference. When curved lanes are assumed, Vpshifts with respect to the rows of the image. In this paper, a Vpfor each individual row of the image is estimated by first extracting a Vpy(vertical position of the Vp) for each individual row of the image from the v-disparity. Then, based on the estimated Vpy's, a 2-D Vpx(horizontal position of the Vp) accumulator is efficiently formed. Thus, by globally optimizing this 2-D Vpxaccumulator, globally optimum Vps for the road image are extracted. Then, estimated Vps are utilized for multiple curved lane marking detection on nonflat road surfaces. The resultant system achieves a detection rate of 99% in 1862 frames of six stereo vision test sequences. Umar Özgünalp, Rui Fan 0001, Xiao Ai, Naim Dahnoun |
IEEE Trans. Intell. Transp. Syst. | 2 |