Fan Lu 0001

dblp:35/6336-1 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-4932-3850ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Post-Refining LiDAR Point Cloud Registration via Diffusion-Based Correspondence Refinement
abstract
LiDAR point cloud registration is a fundamental problem in 3D computer vision. Recent learning-based methods have significantly improved the robustness and accuracy of LiDAR point cloud registration, while most of these methods are designed for global registration. In this paper, we focus on the post-refinement problem of LiDAR point cloud registration, which has been largely overlooked in previous learning-based approaches. To address the correspondence error between LiDAR point clouds, we formulate the transformation refinement problem as a correspondence optimization problem and propose DCR, a diffusion-based correspondence refinement model. DCR is built upon denoising diffusion models and recovers accurate correspondences from noisy initial correspondences estimated by arbitrary global registration methods. To achieve precise correspondence refinement, we employ a residual-based modeling scheme to constrain the output distribution and design a conditional generation paradigm to control the randomness of diffusion models. We combine DCR with both learning-based and traditional global registration methods and perform experiments on three large-scale outdoor LiDAR point cloud datasets to verify the performance. Extensive experimental results demonstrate DCR’s versatility and effectiveness in improving the accuracy of LiDAR point cloud registration.
Fan Lu 0001, Tianhang Wang, Bin Li 0087, Alois C. Knoll, Zhijun Li 0001, Guang Chen 0001
IEEE Trans Autom. Sci. Eng.1
2026 I2EKD: Efficient and Versatile Image-to-Event Knowledge Distillation
abstract
Recently, general-purpose features for event camera data have become increasingly important in advancing event-based vision applications. Current methods typically adopt pre-training paradigms, yielding promising performance. However, the limited data and sparse spatial information of events hinder effective use of pretraining for rich semantic learning. In this paper, we tackle semantic scarcity by transferring knowledge from large pre-trained image models, without increasing event training data. Concretely, we propose a novel image-to-event knowledge distillation method named I2EKD. Acknowledging that different backbones suit different applications, we fix the teacher and keep the student architecture flexible. To improve versatility, we equip I2EKD with two model-agnostic objectives at the logit and feature levels. Additionally, without task-specific objectives or labels, I2EKD avoids re-distillation and transfers well to downstream applications. Furthermore, leveraging DINOv2 as the teacher, whose feature distribution is built from billions of data, the student can swiftly mimic the superior distribution in a data-efficient manner. Compared with the SOTA pre-training method, I2EKD generates outperforming or comparable features with 1/15 training cost (1/10 data × 2/3 epochs). Extensive experiments on different vision tasks (object recognition, semantic segmentation, and monocular depth) verify the effectiveness of our method. Notably, I2EKD achieves top-1 object recognition accuracy of 70.72%, leading the pre-training SOTA by 5.89%.
Hu Cao, Sanqing Qu, Fan Lu 0001, Yan Zhong 0001, Zhichao Lu, Luziwei Leng, Guang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 An Online-Training-Free Adaptor for Open Heterogeneous Collaborative Perception via Diffusion Model
abstract
Collaborative perception seeks to mitigate the limitations of single-vehicle perception, such as occlusions, by facilitating communication and information sharing among connected vehicles. However, most existing works assume a homogeneous scenario where all vehicles share identity sensor types and perception model architectures. In contrast, real-world systems often involve heterogeneous agents with diverse sensor configurations and independently developed models. In such settings, directly exchanging features without proper alignment can significantly degrade performance and hinder effective collaboration. While some methods have been proposed to address heterogeneity, they typically require retraining or access to internal model parameters, making them impractical for scalable deployment. To address these challenges, we propose DiffAlign, a plug-and-play adapter that enables feature alignment across heterogeneous agents in a training-free and model-agnostic manner. DiffAlign treats received BEV features as noisy latent representations and progressively refines them through a pretrained diffusion process. This alignment strategy does not require access to model internals or any retraining, which makes it both scalable and privacy-preserving while supporting diverse sensor modalities and perception backbones. Extensive experiments on simulated OPV2V and real-world V2V4Real datasets demonstrate that DiffAlign consistently improves detection performance in heterogeneous settings, improving CoBEVT by 132.01% and 91.95%, respectively. Our method provides a practical path toward scalable, generalizable, and deployment-ready collaborative perception.
Tianhang Wang, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Hu Cao, Alois C. Knoll, Guang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse Corruptions
abstract
Collaborative perception enhances single-vehicle perception by integrating sensory data from multiple connected vehicles. However, existing studies often assume ideal conditions, overlooking resilience to real-world challenges, such as adverse weather and sensor malfunctions, which is critical for safe deployment. To address this gap, we introduce RCP-Bench, the first comprehensive benchmark designed to evaluate the robustness of collaborative detection models under a wide range of real-world corruptions. RCP-Bench includes three new datasets (i.e., OPV2V-C, V2XSet-C, and DAIR-V2X-C) that simulate six collaborative cases and 14 types of camera corruption resulting from external environmental factors, sensor failures, and temporal misalignments. Extensive experiments on 10 leading collaborative perception models reveal that, while these models perform well under ideal conditions, they are significantly affected by corruptions. To improve robustness, we propose two simple yet effective strategies, RCP-Drop and RCP-Mix, based on training regularization and feature augmentation. Additionally, we identify several critical factors influencing robustness, such as backbone architecture, camera number, feature fusion methods, and the number of connected vehicles. We hope that RCP-Bench, along with these strategies and insights, will stimulate future research toward developing more robust collaborative perception models. Our benchmark toolkit is available at https://github.com/LuckyDush/RCP-Bench.
Shihang Du, Sanqing Qu, Tianhang Wang, Yunwei Zhu, Fan Lu 0001, Guang Chen 0001
CVPR7
2025 Points, Images and Texts: Boosting Point Cloud Completion with Multi-Modal Features
abstract
Point cloud completion is crucial for reconstructing accurate shapes in many 3D visual applications. Recent approaches incorporate images into the completion pipeline, introducing geometric clues and global constraints. However, their fusion processes often fail to reconstruct detailed parts and maintain global consistency simultaneously. Except for images, text is another important clue for recognizing the target's characteristics. Thus, in this work, we propose to combine multiple modalities including points, images and texts for point cloud completion. Specifically, inspired by recently pre-trained large language models, we generate the description texts for images by Visual Question Answering (VQA) models and introduce Visual-Textual Embedding (VTE) models to extract joint features of image-text pairs. Furthermore, we describe the edge geometric patterns by multi-scale edge convolution to guide the refinement of shapes in local areas. Then we adopt cross attention mechanism to effectively fuse multi-modal features and refine the coarse shape. Extensive experiments on commonly used benchmarks demonstrate our method's superior performance over previous uni-modal and cross-modal methods.
ChengKai Xia, Fan Lu 0001, Bin Li 0087, Alois C. Knoll, Guang Chen 0001
ICRA2
2025 Event-Aided Progressive Neural Radiance Fields Reconstruction Under Challenging Illumination
abstract
Scene reconstruction and novel view synthesis are extensively utilized in domains such as autonomous driving, facilitating the creation of more comprehensive datasets and accelerating algorithm development. Current techniques have achieved promising results in image-based synthesis. Nonethe-less, they encounter difficulties including limited dynamic range, motion blur, and suboptimal performance under challenging illumination. Furthermore, these techniques depend significantly on precise camera poses, which are difficult to acquire in practical situations. Event cameras, characterized by their high dynamic range and temporal resolution, demonstrate advantages in challenging illumination and rapid motion scenarios and have been utilized in several autonomous driving applications. This study is the first to combine image and event camera for neural radiance field (NeRF) reconstruction in driving scenarios. The proposed method leverages the unique properties of event camera and adopts a progressive reconstruction strategy to jointly optimize camera poses during training, reducing reliance on pose precision and enabling more robust and accurate scene reconstruction under challenging illumination. Experimental results demonstrate that the proposed approach attains enhanced performance in novel view synthesis, exhibiting a 3.5% increase in PSNR and a 2% increase in SSIM on the DSEC dataset compared to baseline method.
Zongtao Bu, Fan Lu 0001, Sanqing Qu, Bin Li 0087, Guang Chen 0001
IV2
2025 Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition
abstract
Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread availability of RGB cameras and the robustness of point clouds in providing accurate spatial geometry and distance information. However, current methods rely on intermediate modalities that capture either the vertical or horizontal field of view, limiting their ability to fully exploit the complementary information from both sensors. In this work, we propose an innovative initial retrieval + re-rank method that effectively combines information from range (or RGB) images and Bird's Eye View (BEV) images. Our approach relies solely on a computationally efficient global descriptor similarity search process to achieve re-ranking. Additionally, we introduce a novel similarity label supervision technique to maximize the utility of limited training data. Specifically, we employ points average distance to approximate appearance similarity and incorporate an adaptive margin, based on similarity differences, into the vanilla triplet loss. Experimental results on the KITTI dataset demonstrate that our method significantly outperforms state-of-the-art approaches.
Jianyi Peng, Fan Lu 0001, Bin Li 0087, Sanqing Qu, Guang Chen 0001
IV2
2025 OOD-Barrier: Build a Middle-Barrier for Open-Set Single-Image Test Time Adaptation via Vision Language Models
abstract
In real-world environments, a well-designed model must be capable of handling dynamically evolving distributions, where both in-distribution (ID) and out-of-distribution (OOD) samples appear unpredictably and individually, making real-time adaptation particularly challenging. While open-set test-time adaptation has demonstrated effectiveness in adjusting to distribution shifts, existing methods often rely on batch processing and struggle to manage single-sample data stream in open-set environments. To address this limitation, we propose Open-IRT, a novel open-set Intermediate-Representation-based Test-time adaptation framework tailored for single-image test-time adaptation with vision-language models. Open-IRT comprises two key modules designed for dynamic, single-sample adaptation in open-set scenarios. The first is Polarity-aware Prompt-based OOD Filter module, which fully constructs the ID-OOD distribution, considering both the absolute semantic alignment and relative semantic polarity. The second module, Intermediate Domain-based Test-time Adaptation module, constructs an intermediate domain and indirectly decomposes the ID-OOD distributional discrepancy to refine the separation boundary during the test-time. Extensive experiments on a range of domain adaptation benchmarks demonstrate the superiority of Open-IRT. Compared to previous state-of-the-art methods, it achieves significant improvements on representative benchmarks, such as CIFAR-100C and SVHN — with gains of +8.45\% in accuracy, -10.80\% in FPR95, and +11.04\% in AUROC.
Boyang Peng, Sanqing Qu, Tianpei Zou, Fan Lu 0001, Siheng Chen, Yong Wu 0007, Guang Chen 0001
NeurIPS4
2025 Multimodal LiDAR-Camera Novel View Synthesis with Unified Pose-free Neural Fields
abstract
Pose-free Neural Radiance Field (NeRF) aims at novel view synthesis (NVS) without relying on accurate poses, exhibiting significant practical value. Image and LiDAR point cloud are two pivotal modalities in autonomous driving scenarios. While demonstrating impressive performance, single-modality pose-free NeRFs often suffer from local optima due to the limited geometric information provided by dense image textures or the sparse, textureless nature of point clouds. Although prior methods have explored the complementary strengths of both modalities, they have only leveraged inherently sparse point clouds for discrete, non-pixel-wise depth supervision, and are limited to NVS of images. As a result, a Multimodal Unified Pose-free framework remains notably absent. In light of this, we propose MUP, a pose-free framework for LiDAR-Camera joint NVS in large-scale scenes. This unified framework enables continuous depth supervision for image reconstruction using LiDAR-Fields rather than discrete point clouds. By leveraging multimodal inputs, pose optimization receives gradients from the rendering loss of point cloud geometry and image texture, thereby alleviating the issue of local optima commonly encountered in single-modality pose-free tasks. Moreover, to further guide pose optimization of NeRF, we propose a multimodal geometric optimizer that leverages geometric relations from point clouds and photometric regularization from adjacent image frames. Besides, to alleviate the domain gap between modalities, we propose a multimodal-specific coarse-to-fine training approach for unified, compact reconstruction. Extensive experiments on KITTI-360 and NuScenes datasets demonstrate MUP's superiority in accomplishing geometry-aware, modality-consistent, and pose-free 3D reconstruction.
Weiyi Xue, Fan Lu 0001, Yunwei Zhu, Zehan Zheng, Sanqing Qu, Jiangtong Li, Haiyun Wei, Guang Chen 0001
NeurIPS2
2025 CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement Learning
abstract
Constrained hybrid-action reinforcement learning (RL) promises to learn a safe policy within a parameterized action space, which is particularly valuable for safety-critical applications involving discrete-continuous hybrid action spaces. However, existing hybrid-action RL algorithms primarily focus on reward maximization, which faces significant challenges for tasks involving both cost constraints and hybrid action spaces. In this work, we propose a novel Constrained Hybrid-action Policy Optimization algorithm (CHPO) to address the problems of constrained hybrid-action RL. Concretely, we rethink the limitations of hybrid-action RL in handling safe tasks with parameterized action spaces and reframe the objective of constrained hybrid-action RL by introducing the concept of Constrained Parameterized-action Markov Decision Process (CPMDP). Subsequently, we present a constrained hybrid-action policy optimization algorithm to confront the constrained hybrid-action problems and conduct theoretical analyses demonstrating that the CHPO converges to the optimal solution while satisfying safety constraints. Finally, extensive experiments demonstrate that the CHPO achieves competitive performance across multiple experimental tasks.
Ao Zhou 0005, Jiayi Guan, Li Shen 0008, Fan Lu 0001, Sanqing Qu, Junqiao Zhao, Guang Chen 0001
NeurIPS4
2025 Deep reinforcement learning as an interaction agent to steer fragment-based 3D molecular generation for protein pockets
abstract
Designing high-affinity molecules for protein targets (especially novel protein families) is a crucial yet challenging task in drug discovery. Recently, there has been tremendous progress in structure-based 3D molecular generative models that incorporate structural information of protein pockets. However, the capacity for molecular representation learning and the generalization for capturing interaction patterns need substantial further developments. Here, we propose AMG, a framework that leverages deep reinforcement learning as a pocket-ligand interaction agent (IA) to gradually steer fragment-based 3D molecular generation targeting protein pockets. AMG is trained using a two-stage strategy to capture interaction features and explicitly optimize the IA. The framework also introduces a pair of separate encoders for pockets and ligands, coupled with a dedicated pre-training strategy. This enables AMG to enhance its generalization ability by leveraging a vast repository of undocked pockets and molecules, thus mitigating the constraints posed by the limited quantity and quality of available datasets. Extensive evaluations demonstrate that AMG significantly outperforms five state-of-the-art baselines in affinity performance while maintaining proper drug-likeness properties. Furthermore, visual analysis confirms the superiority of AMG at capturing 3D molecular geometrical features and interaction patterns within pocket-ligand complexes, indicating its considerable promise for various structure-based downstream tasks.
Sanqing Qu, Fan Lu 0001, Zhixin Tian, Alois C. Knoll, Guang Chen 0001, Shaorong Gao, Yanping Zhang 0007
Briefings Bioinform.4
2025 MHCPP: A Motion-Based Historical Enhancement Collaborative Perception and Prediction Framework
Shihang Du, Sanqing Qu, Tianhang Wang, Fan Lu 0001, Guang Chen 0001, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.5
2024 LiDAR4D: Dynamic Neural Fields for Novel Space-Time View LiDAR Synthesis
abstract
Although neural radiance fields (NeRFs) have achieved triumphs in image novel view synthesis (NVS), LiDAR NVS remains largely unexplored. Previous LiDAR NVS methods employ a simple shift from image NVS methods while ignoring the dynamic nature and the large-scale reconstruction problem of LiDAR point clouds. In light of this, we propose LiDAR4D, a differentiable LiDAR-only frameworkfor novel space-time LiDAR view synthesis. In consideration of the sparsity and large-scale characteristics, we design a 4D hybrid representation combined with multi-planar and grid features to achieve effective reconstruction in a coarse-to-fine manner. Furthermore, we introduce geometric constraints derivedfrom point clouds to improve temporal consistency. For the realistic synthesis of LiDAR point clouds, we incorporate the global optimization of ray-drop prob-ability to preserve cross-region patterns. Extensive experiments on KITTI-360 and NuScenes datasets demonstrate the superiority of our method in accomplishing geometry-aware and time-consistent dynamic reconstruction. Codes are available at https://github.com/ispc-labILiDAR4D.
Zehan Zheng, Fan Lu 0001, Weiyi Xue, Guang Chen 0001, Changjun Jiang 0002
CVPR2
2024 PCDepth: Pattern-based Complementary Learning for Monocular Depth Estimation by Best of Both Worlds
abstract
Event cameras can record scene dynamics with high temporal resolution, providing rich scene details for monocular depth estimation (MDE) even at low-level illumination. Therefore, existing complementary learning approaches for MDE fuse intensity information from images and scene details from event data for better scene understanding. However, most methods directly fuse two modalities at pixel level, ignoring that the attractive complementarity mainly impacts high-level patterns that only occupy a few pixels. For example, event data is likely to complement contours of scene objects. In this paper, we discretize the scene into a set of high-level patterns to explore the complementarity and propose a Pattern-based Complementary learning architecture for monocular Depth estimation (PCDepth). Concretely, PCDepth comprises two primary components: a complementary visual representation learning module for discretizing the scene into high-level patterns and integrating complementary patterns across modalities and a refined depth estimator aimed at scene reconstruction and depth prediction while maintaining an efficiency-accuracy balance. Through pattern-based complementary learning, PCDepth fully exploits two modalities and achieves more accurate predictions than existing methods, especially in challenging nighttime scenarios. Extensive experiments on MVSEC and DSEC datasets verify the effectiveness and superiority of our PCDepth. Remarkably, compared with state-of-the-art, PCDepth achieves a 37.9% improvement in accuracy in MVSEC nighttime scenarios.
Sanqing Qu, Fan Lu 0001, Zongtao Bu, Florian Röhrbein, Alois C. Knoll, Guang Chen 0001
IROS3
2024 RCDN: Towards Robust Camera-Insensitivity Collaborative Perception via Dynamic Feature-based 3D Neural Modeling
abstract
Collaborative perception is dedicated to tackling the constraints of single-agent perception, such as occlusions, based on the multiple agents' multi-view sensor inputs. However, most existing works assume an ideal condition that all agents' multi-view cameras are continuously available. In reality, cameras may be highly noisy, obscured or even failed during the collaboration. In this work, we introduce a new robust camera-insensitivity problem: how to overcome the issues caused by the failed camera perspectives, while stabilizing high collaborative performance with low calibration cost? To address above problems, we propose RCDN, a Robust Camera-insensitivity collaborative perception with a novel Dynamic feature-based 3D Neural modeling mechanism. The key intuition of RCDN is to construct collaborative neural rendering field representations to recover failed perceptual messages sent by multiple agents. To better model collaborative neural rendering field, RCDN first establishes a geometry BEV feature based time-invariant static field with other agents via fast hash grid modeling. Based on the static background field, the proposed time-varying dynamic field can model corresponding motion vector for foregrounds with appropriate positions. To validate RCDN, we create OPV2V-N, a new large-scale dataset with manual labelling under different camera failed scenarios. Extensive experiments conducted on OPV2V-N show that RCDN can be ported to other baselines and improve their robustness in extreme camera-insensitivity setting. Our code and datasets will be available soon.
Tianhang Wang, Fan Lu 0001, Zehan Zheng, Zhijun Li 0001, Guang Chen 0001, Changjun Jiang 0002
NeurIPS2
2024 GeoNLF: Geometry guided Pose-Free Neural LiDAR Fields
abstract
Although recent efforts have extended Neural Radiance Field (NeRF) into LiDAR point cloud synthesis, the majority of existing works exhibit a strong dependence on precomputed poses. However, point cloud registration methods struggle to achieve precise global pose estimation, whereas previous pose-free NeRFs overlook geometric consistency in global reconstruction. In light of this, we explore the geometric insights of point clouds, which provide explicit registration priors for reconstruction. Based on this, we propose Geometry guided Neural LiDAR Fields (GeoNLF), a hybrid framework performing alternately global neural reconstruction and pure geometric pose optimization. Furthermore, NeRFs tend to overfit individual frames and easily get stuck in local minima under sparse-view inputs. To tackle this issue, we develop a selective-reweighting strategy and introduce geometric constraints for robust optimization. Extensive experiments on NuScenes and KITTI-360 datasets demonstrate the superiority of GeoNLF in both novel view synthesis and multi-view registration of low-frequency large-scale point clouds.
Weiyi Xue, Zehan Zheng, Fan Lu 0001, Haiyun Wei, Guang Chen 0001, Changjun Jiang 0002
NeurIPS3
2024 HDMNet: A Hierarchical Matching Network with Double Attention for Large-scale Outdoor LiDAR Point Cloud Registration
abstract
Outdoor LiDAR point clouds are typically large-scale and complexly distributed. To achieve efficient and accurate registration, emphasizing the similarity among local regions and prioritizing global local-to-local matching is of utmost importance, subsequent to which accuracy can be enhanced through cost-effective fine registration. In this paper, a novel hierarchical neural network with double attention named HDMNet is proposed for large-scale outdoor LiDAR point cloud registration. Specifically, A novel feature consistency enhanced double-soft matching network is introduced to achieve two-stage matching with high flexibility while enlarging the receptive field with high efficiency in a patch-to-patch manner, which significantly improves the registration performance. Moreover, in order to further utilize the sparse matching information from deeper layer, we develop a novel trainable embedding mask to incorporate the confidence scores of correspondences obtained from pose estimation of deeper layer, eliminating additional computations. The high-confidence keypoints in the sparser point cloud of the deeper layer correspond to a high-confidence spatial neighborhood region in shallower layer, which will receive more attention, while the features of non-key regions will be masked. Extensive experiments are conducted on two large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HDMNet.
Weiyi Xue, Fan Lu 0001, Guang Chen 0001
WACV2
2023 NeuralPCI: Spatio-Temporal Neural Field for 3D Point Cloud Multi-Frame Non-Linear Interpolation
abstract
In recent years, there has been a significant increase in focus on the interpolation task of computer vision. Despite the tremendous advancement of video interpolation, point cloud interpolation remains insufficiently explored. Meanwhile, the existence of numerous nonlinear large motions in real-world scenarios makes the point cloud interpolation task more challenging. In light of these issues, we present NeuralPCI: an end-to-end 4D spatiotemporal Neural field for 3D Point Cloud Interpolation, which implicitly integrates multi-frame information to handle nonlinear large motions for both indoor and outdoor scenarios. Furthermore, we construct a new multi-frame point cloud interpolation dataset called NL-Drive for large nonlinear motions in autonomous driving scenes to better demonstrate the superiority of our method. Ultimately, NeuralPCI achieves state-of-the-art performance on both DHB (Dynamic Human Bodies) and NL-Drive datasets. Beyond the interpolation task, our method can be naturally extended to point cloud extrapolation, morphing, and auto-labeling, which indicates its substantial potential in other domains. Codes are available at https://github.com/ispc-lab/NeuralPCI.
Zehan Zheng, Danni Wu, Ruisi Lu, Fan Lu 0001, Guang Chen 0001, Changjun Jiang 0002
CVPR4
2023 Urban Radiance Field Representation with Deformable Neural Mesh Primitives
abstract
Neural Radiance Fields (NeRFs) have achieved great success in the past few years. However, most current methods still require intensive resources due to ray marching-based rendering. To construct urban-level radiance fields efficiently, we design Deformable Neural Mesh Primitive (DNMP), and propose to parameterize the entire scene with such primitives. The DNMP is a flexible and compact neural variant of classic mesh representation, which enjoys both the efficiency of rasterization-based rendering and the powerful neural representation capability for photo-realistic image synthesis. Specifically, a DNMP consists of a set of connected deformable mesh vertices with paired vertex features to parameterize the geometry and radiance information of a local area. To constrain the degree of freedom for optimization and lower the storage budgets, we enforce the shape of each primitive to be decoded from a relatively low-dimensional latent space. The rendering colors are decoded from the vertex features (interpolated with rasterization) by a view-dependent MLP. The DNMP provides a new paradigm for urban-level scene representation with appealing properties: (1) High-quality rendering. Our method achieves leading performance for novel view synthesis in urban scenarios. (2) Low computational costs. Our representation enables fast rendering (2.07 ms/1K pixels) and low peak memory usage (110 MB/1K pixels). We also present a lightweight version that can run 33× faster than vanilla NeRFs, which is comparable to the highly-optimized Instant-NGP. Project page: https://dnmp.github.io/.
Fan Lu 0001, Guang Chen 0001, Hongsheng Li 0001, Kwan-Yee Lin, Changjun Jiang 0002
ICCV1
2023 Sparse-to-Dense Matching Network for Large-Scale LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Previous learning-based methods for LiDAR point cloud registration can be categorized into two schemes: dense-to-dense matching methods and sparse-to-sparse matching methods. However, for large-scale outdoor LiDAR point clouds, solving dense point correspondences is time-consuming, whereas sparse keypoint matching easily suffers from keypoint detection error. In this paper, we propose SDMNet, a novel Sparse-to-Dense Matching Network for large-scale outdoor LiDAR point cloud registration. Specifically, SDMNet performs registration in two sequential stages: sparse matching stage and local-dense matching stage. In the sparse matching stage, we sample a set of sparse points from the source point cloud and then match them to the dense target point cloud using a spatial consistency enhanced soft matching network and a robust outlier rejection module. Furthermore, a novel neighborhood matching module is developed to incorporate local neighborhood consensus, significantly improving performance. The local-dense matching stage is followed for fine-grained performance, where dense correspondences are efficiently obtained by performing point matching in local spatial neighborhoods of high-confidence sparse correspondences. Extensive experiments on three large-scale outdoor LiDAR point cloud datasets demonstrate that the proposed SDMNet achieves state-of-the-art performance with high efficiency.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Yibing Zhan, Zhijun Li 0001, Dacheng Tao, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 HRegNet: A Hierarchical Network for Efficient and Accurate Outdoor LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Outdoor LiDAR point clouds are typically large-scale and complexly distributed, which makes the registration challenging. In this paper, we propose an efficient hierarchical network named HRegNet for large-scale outdoor LiDAR point cloud registration. Instead of using all points in the point clouds, HRegNet performs registration on hierarchically extracted keypoints and descriptors. The overall framework combines the reliable features in deeper layer and the precise position information in shallower layers to achieve robust and precise registration. We present a correspondence network to generate correct and accurate keypoints correspondences. Moreover, bilateral consensus and neighborhood consensus are introduced for keypoints matching, and novel similarity features are designed to incorporate them into the correspondence network, which significantly improves the registration performance. In addition, we design a consistency propagation strategy to effectively incorporate spatial consistency into the registration pipeline. The whole network is also highly efficient since only a small number of keypoints are used for registration. Extensive experiments are conducted on three large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HRegNet. The source code of the proposed HRegNet is available at https://github.com/ispc-lab/HRegNet2.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Sanqing Qu, Rongqi Gu, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 RPP-Net: Rigid Constrained Point Cloud Prediction Network
abstract
Forecasting the future environment is an essential and fundamental capability of autonomous driving systems. In contrast to the widely studied video prediction, only a few literatures have explored LiDAR point cloud prediction. To achieve future point cloud generation, most existing methods are based on free-form 3D scene flow prediction. However, these simple scene flow prediction-based methods may cause distortions since the motions of 3D scenes can be seen as a combination of rigid (static background) and flow (dynamic foreground) motions. To address this issue, we propose a simple but effective rigid constrained point cloud prediction network named RPP-Net. The key component of the proposed method is the hybrid motion decoder, which generates motion masks and combines flow motions and rigid motions to produce hybrid motions without additional annotations and sensors. Besides, to reduce running time, we provide a hierarchical cell, which uses a feature encoder to extract deep features and an RPP-RNN module to capture temporal correlations across frames. To evaluate the effectiveness of our RPP-Net, we have conducted extensive experiments on both KITTI dataset and Argoverse dataset and the results show that our RPP-Net significantly outperforms existing methods and achieves a new state-of-the-art.
Tianpei Zou, Guang Chen 0001, Fan Lu 0001, Zhijun Li 0001, Sanqing Qu, Alois C. Knoll, Changjun Jiang 0002
IEEE Trans. Intell. Transp. Syst.3
2022 MoNet: Motion-Based Point Cloud Prediction Network
abstract
Predicting the future can significantly improve the safety of intelligent vehicles, which is a key component in autonomous driving. 3D point clouds can accurately model 3D information of surrounding environment and are crucial for intelligent vehicles to perceive the scene. Therefore, prediction of 3D point clouds has great significance for intelligent vehicles, which can be utilized for numerous further applications. However, due to point clouds are unordered and unstructured, point cloud prediction is challenging and has not been deeply explored in current literature. In this paper, we propose a novel motion-based neural network named MoNet. The key idea of the proposed MoNet is to integrate motion features between two consecutive point clouds into the prediction pipeline. The introduction of motion features enables the model to more accurately capture the variations of motion information across frames and thus make better predictions for future motion. In addition, content features are introduced to model the spatial content of individual point clouds. A recurrent neural network named MotionRNN is proposed to capture the temporal correlations of both features. Moreover, an attention-based motion align module is proposed to address the problem of missing motion features in the inference pipeline. Extensive experiments on two large-scale outdoor LiDAR point cloud datasets demonstrate the performance of the proposed MoNet. Moreover, we perform experiments on applications using the predicted point clouds and the results indicate the great application potential of the proposed method.
Fan Lu 0001, Guang Chen 0001, Zhijun Li 0001, Yinlong Liu, Sanqing Qu, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.1
2021 PointINet: Point Cloud Frame Interpolation Network
abstract
LiDAR point cloud streams are usually sparse in time dimension, which is limited by hardware performance. Generally, the frame rates of mechanical LiDAR sensors are 10 to 20 Hz, which is much lower than other commonly used sensors like cameras. To overcome the temporal limitations of LiDAR sensors, a novel task named Point Cloud Frame Interpolation is studied in this paper. Given two consecutive point cloud frames, Point Cloud Frame Interpolation aims to generate intermediate frame(s) between them. To achieve that, we propose a novel framework, namely Point Cloud Frame Interpolation Network (PointINet). Based on the proposed method, the low frame rate point cloud streams can be upsampled to higher frame rates. We start by estimating bi-directional 3D scene flow between the two point clouds and then warp them to the given time step based on the 3D scene flow. To fuse the two warped frames and generate intermediate point cloud(s), we propose a novel learning-based points fusion module, which simultaneously takes two warped point clouds into consideration. We design both quantitative and qualitative experiments to evaluate the performance of the point cloud frame interpolation method and extensive experiments on two large scale outdoor LiDAR datasets demonstrate the effectiveness of the proposed PointINet. Our code is available at https://github.com/ispc-lab/PointINet.git.
Fan Lu 0001, Guang Chen 0001, Sanqing Qu, Zhijun Li 0001, Yinlong Liu, Alois C. Knoll
AAAI1
2021 HRegNet: A Hierarchical Network for Large-scale Outdoor LiDAR Point Cloud Registration
abstract
Point cloud registration is a fundamental problem in 3D computer vision. Outdoor LiDAR point clouds are typically large-scale and complexly distributed, which makes the registration challenging. In this paper, we propose an efficient hierarchical network named HRegNet for large-scale out-door LiDAR point cloud registration. Instead of using all points in the point clouds, HRegNet performs registration on hierarchically extracted keypoints and descriptors. The overall framework combines the reliable features in deeper layer and the precise position information in shallower layers to achieve robust and precise registration. We present a correspondence network to generate correct and accurate keypoints correspondences. Moreover, bilateral consensus and neighborhood consensus are introduced for keypoints matching and novel similarity features are designed to in-corporate them into the correspondence network, which significantly improves the registration performance. Besides, the whole network is also highly efficient since only a small number of keypoints are used for registration. Extensive experiments are conducted on two large-scale outdoor LiDAR point cloud datasets to demonstrate the high accuracy and efficiency of the proposed HRegNet. The project website is https://ispc-group.github.io/hregnet.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Sanqing Qu, Rongqi Gu
ICCV1
2020 RSKDD-Net: Random Sample-based Keypoint Detector and Descriptor
abstract
Keypoint detector and descriptor are two main components of point cloud registration. Previous learning-based keypoint detectors rely on saliency estimation for each point or farthest point sample (FPS) for candidate points selection, which are inefficient and not applicable in large scale scenes. This paper proposes Random Sample-based Keypoint Detector and Descriptor Network (RSKDD-Net) for large scale point cloud registration. The key idea is using random sampling to efficiently select candidate points and using a learning-based method to jointly generate keypoints and corresponding descriptors. To tackle the information loss of random sampling, we exploit a novel random dilation cluster strategy to enlarge the receptive field of each sampled point and an attention mechanism to aggregate the positions and features of neighbor points. Furthermore, we propose a matching loss to train the descriptor in a weakly supervised manner. Extensive experiments on two large scale outdoor LiDAR datasets show that the proposed RSKDD-Net achieves state-of-the-art performance with more than 15 times faster than existing methods. Our code is available at https://github.com/ispc-lab/RSKDD-Net.
Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Zhongnan Qu, Alois C. Knoll
NeurIPS1