Pei An

dblp:154/3152 · DBLP profile ↗
← Back
32ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0002-3645-8465ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries
abstract
Semantic occupancy has emerged as a powerful representation in world models for its ability to capture rich spatial semantics. However, most existing occupancy world models rely on static and fixed embeddings or grids, which inherently limit the flexibility of perception. Moreover, their ``in-place classification" over grids exhibits a potential misalignment with the dynamic and continuous nature of real scenarios. In this paper, we propose SparseWorld, a novel 4D occupancy world model that is flexible, adaptive, and efficient, powered by sparse and dynamic queries. We propose a Range-Adaptive Perception module, in which learnable queries are modulated by the ego vehicle states and enriched with temporal-spatial associations to enable extended-range perception. To effectively capture the dynamics of the scene, we design a State-Conditioned Forecasting module, which replaces classification-based forecasting with regression-guided formulation, precisely aligning the dynamic queries with the continuity of the 4D environment. In addition, We specifically devise a Temporal-Aware Self-Scheduling training strategy to enable smooth and efficient training. Extensive experiments demonstrate that SparseWorld achieves state-of-the-art performance across perception, forecasting, and planning tasks. Comprehensive visualizations and ablation studies further validate the advantages of SparseWorld in terms of flexibility, adaptability, and efficiency.
Chenxu Dang, Jason Bao, Pei An, Xinyue Tang, An Pan, Jie Ma 0003, Bingchuan Sun
AAAI4
2026 Density-aware few-parametric networks for robust few-shot point cloud semantic segmentation
Yudong Liang, Pei An, Qiong Liu 0001, You Yang 0002
Neurocomputing2
2026 Rethinking the refinement stage of 3D object detection: A multi-task learning perspective with Mixture-of-Experts
Bingqian Wu, Pei An, Siwen Quan, Qiao Wu, Chu'ai Zhang, Jiaqi Yang 0002
J. Vis. Commun. Image Represent.2
2026 HGS-OCC: Real-time 3D occupancy prediction via hybrid depth optimization with multi-representation geometry-semantic integration
Zaipeng Duan, Xuzhong Hu, Pei An, Chenxu Dang, Jie Ma 0003
Pattern Recognit.3
2026 FSF-Net: Enhance 4D occupancy forecasting with coarse BEV scene flow for autonomous driving
Erxin Guo, Pei An, You Yang 0002, Qiong Liu 0001, Anan Liu
Pattern Recognit.2
2026 Dual-domain homogeneous fusion with cross-modal mamba and progressive decoder for 3D object detection
Xuzhong Hu, Zaipeng Duan, Pei An, Ziwen Xu, Jie Ma 0003
Pattern Recognit.3
2026 DFS-Net: A Dense Focal Stack Image Generation Network From Misaligned Multi-Focus Images
abstract
Dense focal stack images inherently encode depth cues and are crucial for various 3D vision applications. However, existing generation methods are susceptible to misalignment and introduce a domain gap between synthetic and real-world data due to off-axis aberrations. To address these challenges, we introduce DFS-Net, an aberration-aware dense focal stack image generation network. DFS-Net consists of two core modules: all-in-focus image synthesis and aberration-aware point spread function (PSF) generation. The all-in-focus image synthesis is achieved through a densely connected fusion network based on multi-scale focus migration and focus property detection. This fusion network can effectively fuse misaligned multi-focus images into an all-in-focus image. The aberration-aware PSF generation is realized through a multi-layer perceptron (MLP) network. Supervised by ray-tracing-based PSFs, the MLP network can generate spatially varying PSFs for arbitrary spatial positions and focus distances. By selecting a set of focus distances, the generated PSF maps are locally convolved with the all-in-focus image to produce an aberration-aware dense focal stack. We conduct extensive comparative experiments on all-in-focus image fusion and focal stack generation against state-of-the-art methods. The experimental results demonstrate that DFS-Net can synthesize all-in-focus images with high subjective and objective quality, as well as generate dense focal stacks that closely approximate ray-tracing results. In addition, we conduct comparative experiments on the depth-from-focus and salient object detection tasks using the generated focal stacks. The experimental results demonstrate that our DFS-Net can significantly enhance the performance of existing depth-from-focus and salient object detection models. The code and dataset will be publicly available at https://github.com/North-Li/DFS-Net.
Zhilong Li, Pei An, You Yang 0002, Qiong Liu 0001, Dan Song 0006, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2025 FASTer: Focal token Acquiring-and-Scaling Transformer for Long-term 3D Objection Detection
abstract
Recent top-performing temporal 3D detectors based on Lidars have increasingly adopted region-based paradigms. They first generate coarse proposals, followed by encoding and fusing regional features. However, indiscriminate sampling and fusion often overlook the varying contributions of individual points and lead to exponentially increased complexity as the number of input frames grows. Moreover, arbitrary result-level concatenation limits the global information extraction. In this paper, we propose a Focal Token Acquring-and-Scaling Transformer (FASTer), which dynamically selects focal tokens and condenses token sequences in an adaptive and lightweight manner. Empha-sizing the contribution of individual tokens, we propose a simple but effective Adaptive Scaling mechanism to capture geometric contexts while sifting out focal points. Adaptively storing and processing only focal points in historical frames dramatically reduces the overall complexity. Furthermore, a novel Grouped Hierarchical Fusion strategy is proposed, progressively performing sequence scaling and Intra-Group Fusion operations to facilitate the exchange of global spatial and temporal information. Experiments on the Waymo Open Dataset demonstrate that our FASTer significantly outperforms other state-of-the-art detectors in both performance and efficiency while also exhibiting improved flexibility and robustness. The code is available at https://github.com/MSunDYY/FASTer.git.
Chenxu Dang, Zaipeng Duan, Pei An, Xuzhong Hu, Jie Ma 0003
CVPR3
2025 SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction
abstract
Multimodal 3D occupancy prediction has garnered significant attention for its potential in autonomous driving. However, most existing approaches are single-modality: camera-based methods lack depth information, while LiDAR-based methods struggle with occlusions. Current lightweight methods primarily rely on the Lift-Splat-Shoot (LSS) pipeline, which suffers from inaccurate depth estimation and fails to fully exploit the geometric and semantic information of 3D LiDAR points. Therefore, we propose a novel multimodal occupancy prediction network called SDG-OCC, which incorporates a joint semantic and depth-guided view transformation coupled with a fusion-to-occupancy-driven active distillation. The enhanced view transformation constructs accurate depth distributions by integrating pixel semantics and co-point depth through diffusion and bilinear discretization. The fusion-to-occupancy-driven active distillation extracts rich semantic information from multimodal data and selectively transfers knowledge to image features based on LiDAR-identified regions. Finally, for optimal performance, we introduce SDG-Fusion, which uses fusion alone, and SDG-KL, which integrates both fusion and distillation for faster inference. Our method achieves state-of-the-art (SOTA) performance with real-time processing on the Occ3D-nuScenes dataset and shows comparable performance on the more challenging SurroundOcc-nuScenes dataset, demonstrating its effectiveness and robustness. The code will be released at https://github.com/DzpLab/SDGOCC.
Zaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An, Junfeng Ding, Jie Zhan, YunBiao Xu, Jie Ma 0003
CVPR4
2025 Unlocking Generalization Power in LiDAR Point Cloud Registration
abstract
In real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety in autonomous driving and other LiDAR-based applications. However, current methods fall short in achieving this level of generalization. To address these limitations, we propose UGP, a pruned framework designed to enhance generalization power for LiDAR point cloud registration. The core insight in UGP is the elimination of cross-attention mechanisms to improve generalization, allowing the network to concentrate on intra-frame feature extraction. Additionally, we introduce a progressive self-attention module to reduce ambiguity in large-scale scenes and integrate Bird’s Eye View (BEV) features to incorporate semantic information about scene elements. Together, these enhancements significantly boost the network’s generalization performance. We validated our approach through various generalization experiments in multiple outdoor scenes. In cross-distance generalization experiments on KITTI and nuScenes, UGP achieved state-of-the-art mean Registration Recall rates of 94.5% and 91.4%, respectively. In cross-dataset generalization from nuScenes to KITTI, UGP achieved a state-of-the-art mean Registration Recall of 90.9%. Code will be available at https://github.com/peakpang/UGP
Zhenxuan Zeng, Qiao Wu, Xiyu Zhang 0001, Lin Wu 0001, Pei An, Jiaqi Yang 0002, Peng Wang 0015
CVPR5
2025 MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnP
abstract
Image-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. The differential perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing the projective constraints on 2D-3D correspondences. However, differential PnP is highly sensitive to noise and outliers in the predicted correspondences. This issue hinders the effectiveness of correspondence learning. Inspired by the robustness of blind PnP against noise and outliers in correspondences, we propose an approximated blind PnP based correspondence learning approach. To mitigate the high computational cost of blind PnP, we simplify blind PnP to an amenable task of minimizing Chamfer distance between learned 2D and 3D keypoints, called MinCD-PnP. To effectively solve MinCD-PnP, we design a lightweight multi-task learning module, named as MinCD-Net, which can be easily integrated into the existing I2P registration architectures. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected datasets demonstrate that MinCD-Net outperforms state-of-the-art methods and achieves a higher inlier ratio (IR) and registration recall (RR) in both cross-scene and cross-dataset settings.
Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Liangliang Nan
ICCV1
2025 Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology Relationship
abstract
Image-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios.
Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Jie Ma 0003, Liangliang Nan
IJCAI1
2025 Op-Occ-Net: Neural Ordinary Differential Equation Based 4D Occupancy Forecasting for Autonomous Vehicles
Junfeng Ding, Erxin Guo, Pei An, Jie Ma 0003, Ruixiang Zhao
PRCV (11)3
2025 DSConv: Fine-Grained Dynamic Sequence Convolution for 3D Understanding
Pei An, Chenxu Dang, Zaipeng Duan, Jie Ma 0003
PRCV (10)2
2025 Enhance Image-to-Point-Cloud Registration with Beltrami Flow
Pei An, You Yang 0002, Jiaqi Yang 0002, Muyao Peng, Qiong Liu 0001, Liangliang Nan
Int. J. Comput. Vis.1
2025 Lidar-camera range-view fusion for 3D object detection in autonomous driving
Xuzhong Hu, Zaipeng Duan, Pei An, Jun Zhang 0062, Jie Ma 0003
Multim. Syst.3
2024 3D Single-Object Tracking in Point Clouds with High Temporal Variation
Qiao Wu, Kun Sun 0002, Pei An, Mathieu Salzmann, Yanning Zhang 0001, Jiaqi Yang 0002
ECCV (7)3
2024 Low-Rank Completion Based Normal Guided Lidar Point Cloud Up-Sampling
abstract
Commercial inexpensive LiDAR sensor generally suffers low vertical resolutions, whose point cloud is sparse and may not be able to satisfy future metaverse applications. LiDAR point cloud up-sampling is a task to increase the vertical resolution while preserving the structural details. Scene representation is the central pillar of point cloud up-sampling. However, the sparsity of point cloud hinders the extraction of scene representation. In this paper, we find that low-rank representation can describe the primary scene structure approximately, and convert up-sampling as low-rank tensor completion problem. To decrease problem complexity, we leverage range view projection to convert the problem as low-rank depth completion, and present a low-rank normal guided up-sampling approach. It uses normal as guidance to smooth range depth. Extensive experiments show that our method outperforms current methods. In 2× up-sampling task, it achieves as low as 41cm of mean absolute error (MAE), which is 282% and 32% smaller than interpolation and traditional matrix completion methods, respectively. Hence, we believe the proposed method benefits to the field of metaverse.
Pei An, You Yang 0002, Jie Ma 0003
ICASSP1
2024 OL-Reg: Registration of Image and Sparse LiDAR Point Cloud With Object-Level Dense Correspondences
abstract
Image and point cloud registration (2D-3D registration) is an essential prerequisite for multi-modal feature fusion. However, due to the significant feature difference of point cloud and image, it is challenging to establish 2D-3D correspondences. Targeting for the background of autonomous driving, we propose 2D-3D registration method with object-level correspondence (OL-Reg) in this paper. Object-level correspondence consists of object bounding box and object contour in 2D image and 3D space. The first step is to match 2D-3D objects. Due to sensor pose and field of view (FoV) difference, object shape and occlusion is different in image and point cloud, causing the difficulty of object matching. To solve this issue, we represent object as 3D bounding box, and design 2D-3D object matching with 3D box projection (Box-Proj) constraint. It aligns object 3D bounding box in image and point cloud. After that, the next step is to build 2D-3D correspondence from the matched objects. To extract correspondence from object with irregular shape, we notice the distance constraint of object surface and rays back-projected from object contour, and present projection based iterative closest point (Proj-ICP). Towards the stability of Proj-ICP, object-level regularization term is designed. Experiment is conducted in KITTI object and odometry dataset. With the pre-trained 3D object detector, results suggest that OL-Reg has the better performance than current approaches in tasks of re-localization and extrinsic calibration. And source code will be released soon1.
Pei An, Xuzhong Hu, Junfeng Ding, Jun Zhang 0062, Jie Ma 0003, You Yang 0002, Qiong Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Survey of Extrinsic Calibration on LiDAR-Camera System for Intelligent Vehicle: Challenges, Approaches, and Trends
abstract
A system with light detection and ranging (LiDAR) and camera (named as LiDAR-camera system) plays the essential role in intelligent vehicle (IV), for it provides 3D spatial and 2D texture features for 3D scene understanding. To leverage LiDAR point cloud and image, extrinsic calibration is a crucial technique, for it can align 2D pixel and 3D point in the pixel-level accuracy. With the rapid development of IV, calibration demand is shifted from offline to online, from the specific scenes to the open scenes. It brings new challenge to the calibration task. Although numbers of approaches have been proposed in the last decade, there lacks an in-depth summary about this topic. Thus, we conduct a survey of extrinsic calibration. Theoretically, the key of calibration is to build correspondence from LiDAR point cloud and optical image. From the viewpoint of correspondence, we attempt to divide the mainstream approaches into explicit and implicit correspondence based methods. After that, we summarize both the strength and weakness of the current works, provide the methods comparison, and list the open-source implementations. Finally, we analyze the tendency of calibration approach, discuss the remained problems in this field. We believe that this survey benefits to the community of autonomous driving.
Pei An, Junfeng Ding, Siwen Quan, Jiaqi Yang 0002, You Yang 0002, Qiong Liu 0001, Jie Ma 0003
IEEE Trans. Intell. Transp. Syst.1
2024 SP-Det: Leveraging Saliency Prediction for Voxel-Based 3D Object Detection in Sparse Point Cloud
abstract
Voxel is one of the common structural representation of 3D point cloud. Due to the sparsity of point cloud generated by light detection and ranging (LiDAR), there is the extreme imbalance in the foreground and background voxels. It decreases the accuracy of 3D object detection, has the negative effect on intelligent driving safety. To overcome this problem, we present a saliency prediction based 3D object detector SP-Det in this article. Although foreground voxels have the sufficient feature of object, it is difficult to localize the foreground region from voxel space with the larger background region. We design an auxiliary learning task, saliency prediction (SP). It benefits 3D detector in identifying the foreground region. SP task uses label diffusion to alleviate the label imbalance. It reduces the learning difficulty of saliency in voxel and bird's eye view (BEV) spaces. After that, to strengthen feature interaction from the sparse foreground region, we design saliency fusion (SF) module to fuse the learning result in SP task. It utilizes voxel and BEV saliency maps as progressive attention to resist the redundant feature from background region. To aggregate more foreground feature inside 3D and BEV region of interest (RoI), we design hybrid grid maps based RoI pooling (Hybrid-RoI pooling). Experiments are conducted in STF dataset. The adverse weather enlarges the sparsity of LiDAR point cloud, increasing the difficulty of object detection. SP-Det identifies and leverages foreground region, and achieves the performance better than the current methods. Hence, we believe that SP-Det benefits to LiDAR based 3D scene understanding in the adverse weather.
Pei An, Yucong Duan, Yuliang Huang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001
IEEE Trans. Multim.1
2024 ESC-Net: Alleviating Triple Sparsity on 3D LiDAR Point Clouds for Extreme Sparse Scene Completion
abstract
3D scene completion (SC) has made progress in the last three years. From the application of mobile robot system, SC should support the downstream task (i.e. mapping or perception), instead of only predicting the completed scenes. However, as the low-cost few-beam LiDAR is widely applied in mobile robot, gap between SC and downstream tasks is large. To generate the high quality completion result, the bottleneck lies in the triple sparsity of input, ground truth (GT) occupancy, and GT foreground. To deal with the triple sparsity, we present an extreme sparse scene completion network (ESC-Net). At first, input sparsity hides most of the spatial information of the scene. A feature completion (FC) decoder is designed to mine the spatial feature using feature-level completion. Then, GT occupancy sparsity hinders representation learning of the real scene with continuous surfaces. A multi-view multi-task attention (MMA) loss is presented to recover the high-quality object boundaries via correcting occupancy and semantic labels of regions from 3D and bird's eye view (BEV) spaces. After that, GT foreground sparsity is the imbalance of foreground and background GT labels. It causes the inaccuracy of local 3D object completion. A combination network (ESC-Net-D) is presented to recover 3D structural details of both foreground and background. Experiment is conducted on KITTI and SemanticPOSS datasets. It shows that ESC-Net has the performance higher than current methods not only on completion task, but also on the downstream tasks (i.e. 3D registration, 3D object detection). Hence, we believe that ESC-Net benefits to the community of mobile robot. Source code is released soon.
Pei An, Siwen Quan, Junfeng Ding, Jie Ma 0003, You Yang 0002, Qiong Liu 0001
IEEE Trans. Multim.1
2023 ProUDA: Progressive unsupervised data augmentation for semi-Supervised 3D object detection on point cloud
Pei An, Junxiong Liang, Tao Ma 0004, Yanfei Chen, Liheng Wang, Jie Ma 0003
Pattern Recognit. Lett.1
2023 RS-Aug: Improve 3D Object Detection on LiDAR With Realistic Simulator Based Data Augmentation
abstract
Light detection and ranging (LiDAR) is an essential sensor for three dimensional (3D) object detection via generating 3D point cloud of the surroundings, and it has been widely used in the various visual applications, especially autonomous driving. However, limited numbers of labeled LiDAR datasets brutally restrain the development of 3D object detector, and this situation breeds an urgent demand on data augmentation in this field. By far, most of the traditional methods reuse the labeled samples, while those unlabeled are hastily untaken. Motivated by this, we propose aRealisticSimulator based data augmentation (RS-Aug). It aims to construct augmented real scenes to enrich the diversity of training dataset. To train 3D object detector in a supervised learning way, the first step of RS-Aug is auto-annotation. Time-continuous LiDAR frames are used to construct the dense scene, which is beneficial to annotation and the subsequent rendering augmentation. However, 3D points with incorrect semantic labels are naturally gathered during multi-view reconstruction, causing the negative effect on auto-annotation. We propose an algorithm of cluster guided$k$-nearest neighbor (c-$k$NN). It emphasizes on de-nosing semantic labels of clustered points using distance and intensity constraints. Then, the next step of RS-Aug is rendering augmentation on the real scene. To enhance the rendering quality using collision and distance constraints with the less computation complexity, we propose a scheme of heuristic search (HS) based object insertion. It estimates the proper position of the inserted object from 2D bird’s eye view (BEV). Experiments demonstrate the de-noising accuracy of c-$k$NN, rendering quality of HS based object insertion, and improvement of RS-Aug on object detection.
Pei An, Junxiong Liang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Distribution Aware VoteNet for 3D Object Detection
abstract
Occlusion is common in the actual 3D scenes, causing the boundary ambiguity of the targeted object. This uncertainty brings difficulty for labeling and learning. Current 3D detectors predict the bounding box directly, regarding it as Dirac delta distribution. However, it does not fully consider such ambiguity. To deal with it, distribution learning is used to efficiently represent the boundary ambiguity. In this paper, we revise the common regression method by predicting the distribution of the 3D box and then present a distribution-aware regression (DAR) module for box refinement and localization quality estimation. It contains scale adaptive (SA) encoder and joint localization quality estimator (JLQE). With the adaptive receptive field, SA encoder refines discriminative features for precise distribution learning. JLQE provides a reliable location score by further leveraging the distribution statistics, correlating with the localization quality of the targeted object. Combining DAR module and the baseline VoteNet, we propose a novel 3D detector called DAVNet. Extensive experiments on both ScanNet V2 and SUN RGB-D datasets demonstrate that the proposed DAVNet achieves significant improvement and outperforms state-of-the-art 3D detectors.
Junxiong Liang, Pei An, Jie Ma 0003
AAAI2
2022 Deep structural information fusion for 3D object detection on LiDAR-camera system
Pei An, Junxiong Liang, Bin Fang 0007, Jie Ma 0003
Comput. Vis. Image Underst.1
2022 Lambertian Model-Based Normal Guided Depth Completion for LiDAR-Camera System
abstract
Depth completion is an essential task for the dense scene reconstruction on light detection and ranging (LiDAR)-camera system. Learning-based method achieves precise depth completion results on specific data sets. However, for the general outdoor scenes with insufficient labeled data sets, an efficient nonlearning method is still required. In this letter, from the geometrical constraint between depth and normal, a novel nonlearning normal guided depth completion method is proposed. For the objects in the outdoor scene, local brightness normal (LBN) constraint is derived from the Lambertian model. It is used to recover dense normal from RGB image and sparse normal. After that, we present a pipeline for depth completion with the guidance of dense normal. Extensive experiments on the KITTI depth completion data set demonstrate that our method achieves smaller root mean squared error (RMSE) than current nonlearning methods.
Pei An, Wenxing Fu, Yingshuo Gao, Jie Ma 0003, Jun Zhang 0062, Bin Fang 0007
IEEE Geosci. Remote. Sens. Lett.1
2022 NCFT: Automatic Matching of Multimodal Image Based on Nonlinear Consistent Feature Transform
abstract
Automatic matching of multimodal images remains a critical challenging task in many remote sensing and computer vision applications. Due to significant nonlinear radiation distortions (NRDs) between multimodal images, it is difficult for many traditional feature matching methods which are sensitive to NRD to achieve satisfactory matching performance. To cope with this problem, this letter proposes a novel feature matching method named nonlinear consistent feature transform (NCFT) that is robust to large NRD. There are three main contributions to NCFT. First, we propose a new consistent feature map instead of image intensity for feature point detection and description, the consistent feature map encodes the structure information and provides a rich and robust feature. Second, we propose a mean-residual maximum index map (MR-MIM) for feature description and the MR-MIM is constructed from the Log-Gabor convolution sequence on the consistent feature map. Finally, the structure descriptors are built according to the MR-MIM, and multimodal image matching is achieved by computing the correspondence. The extensive experimental results demonstrate that NCFT can effectively overcome the problem of NRD, NCFT outperforms other state-of-the-art methods and improves the matching accuracy and robustness on different multimodal image datasets.
Yucong Duan, Bin Fang 0007, Pei An, Jie Ma 0003
IEEE Geosci. Remote. Sens. Lett.5
2021 Feature Interactive Representation for Point Cloud Registration
abstract
Point cloud registration is the process of using the common structures in two point clouds to splice them together. To find out these common structures and make these structures match more accurately, interacting information of the source and target point clouds is essential. However, limited attention has been paid to explicitly model such feature interaction. To this end, we propose a Feature Interactive Representation learning Network (FIRE-Net), which can explore feature interaction among the source and target point clouds from different levels. Specifically, we first introduce a Combined Feature Encoder (CFE) based on feature interaction intra point cloud. The CFE extracts interactive features intra each point cloud and combines them to enhance the ability of the network to describe the local geometric structure. Then, we propose a feature interaction mechanism inter point clouds which includes a Local Interaction Unit (LIU) and a Global Interaction Unit (GIU). The former is used to interact information between point pairs across two point clouds, thus the point features in one point cloud and its similar point features in another point cloud can be aware of each other. The latter is applied to change the per-point features depending on the global cross information of two point clouds, thus one point cloud has the global perception of another. Extensive experiments on partially overlapping point cloud registration show that our method achieves state-of-the-art performance.
Bingli Wu, Jie Ma 0003, Gaojie Chen 0003, Pei An
ICCV4
2021 Straight Sampling Network for Point Cloud Learning
abstract
Sampling operation is a bottleneck of the hierarchical point cloud learning. Existing learnable sampling methods generate a “soft” virtual subset in the training phase, thus distorting the original underlying shape and losing 3D geometric information. In this paper, we propose a novel end-to-end discrete sampling method, named Straight Sampling, to output a “hard” authentic subset with the assistance of Straight Through Estimator. Equipped with Straight Sampling, a hierarchical architecture is developed to learn an effective representation. By grouping and pooling the sampled points in 3D Euclidean space, the network benefits from semantic features as well as 3D geometric information to achieve state-of-the-art performance.
Gaojie Chen 0003, Jie Ma 0003, Pei An
ICIP4
2020 Novel calibration method for camera array in spherical arrangement
Pei An, Qiong Liu 0001, Firas Abedi, You Yang 0002
Signal Process. Image Commun.1
2014 The Stochastic Loss of Spikes in Spiking Neural P Systems: Design and Implementation of Reliable Arithmetic Circuits
abstract
Spiking neural P systems (in short, SN P systems) have been introduced as computing devices inspired by the structure and functioning of neural cells. The presence of unreliable components in SN P systems can be considered in many different aspects.
Matteo Cavaliere, Pei An, Sarma B. K. Vrudhula, Yu Cao 0001
Fundam. Informaticae3