Jianwei Wan

dblp:05/8763 · DBLP profile ↗
← Back
34ranked-venue papers
0as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Occupancy Prediction Guided Point Cloud Geometry Compression
abstract
Efficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signal-to- noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based models struggle with limited receptive fields to capture long-range dependencies, while Transformer-built architectures always neglect fine-grained details due to their reliance on global self-attention. This paper presents a Transformer-efficient occupancy prediction Network, termed TopNet, to overcome these challenges by developing several novel components designed to enhance both global context modeling and local structure preservation: Locally-enhanced Context Encoding (LeCE) for improving local structural awareness and enhancing the translation-invariance of the octree nodes, Adaptive-Length Sliding Window Attention (AL-SWA) for capturing both global and local dependencies while adaptively adjusting attention weights based on the input window length, Spatial-Gated-enhanced Channel Mixer (SG-CM) for efficient feature aggregation from ancestors and siblings, and Latent-guided Node Occupancy Predictor (LNOP) for improving prediction accuracy of spatially adjacent octree nodes in local context. Comprehensive experiments across three large-scale outdoor sparse LiDAR datasets, including SemanticKITTI, nuScenes, and LiDAR-CS, as well as two indoor dense human body datasets, including 8iVFB and MVUB, and one indoor dense scenario dataset, ScanNet, demonstrate that our TopNet achieves state-of-the-art compression performance with fewer parameters.
Yifan Zhang 0030, Ting Liu 0017, Xinpu Liu, Ke Xu 0013, Jianwei Wan, Yulan Guo, Hanyun Wang
IEEE Trans. Circuits Syst. Video Technol.6
2026 OctGLP-Net: Learning Octree-Structured Context Entropy Model With Global-Local Perception for Point Cloud Geometry Compression
abstract
The insufficient exploitation of spatial correlations among octree node context and feature interactions across spatial and channel dimensions limits the reconstruction performance of current point cloud geometry compression (PCGC) models. To solve these issues, this paper presents an octree-structured context entropy model OctGLP-Net with global-local perception, which mainly consists of a Local Spatial Perception (LocSP) module, a Scaled-cosine Attention based Spatial Interaction (SASI) module, and a Locally-enhanced Feed-forward Spatial and Channel Interaction (LFSCI) module. First, we propose the LocSP to extract local context information from octree nodes. Then, we introduce the SASI to fully exploit the spatial correlations among nodes. Finally, to effectively interact local features across spatial and channel dimensions, we devise a LFSCI network by employing depth-wise and point-wise convolutions to realize fine-grained local feature extraction from octree nodes. Experimental results on both sparse LiDAR and dense object benchmark datasets demonstrate that our method achieves state-of-the-art lossy/lossless compression performance. Our method obtains higher reconstruction quality (D1/D2 PSNR) and smaller chamfer distance (CD) at similar bits per point (BPP) on the SemanticKITTI, nuScenes, and LiDAR-CS datasets, and lower bitrate on the Owlii, 8iVFB and MVUB datasets. Significantly, the proposed OctGLP-Net also exhibits strong generalization abilities when applied to unseen nuScenes and LiDAR-CS datasets. In addition, downstream object detection task on the nuScenes dataset with different compression precisions further demonstrate the superiority and robustness of our method.
Ke Xu 0013, Xinpu Liu, Jianwei Wan, Yulan Guo, Hanyun Wang
IEEE Trans. Intell. Transp. Syst.4
2025 TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry Compression
abstract
Efficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signalto-noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based models struggle with limited receptive fields to capture long-range dependencies, while Transformer-built architectures always neglect fine-grained details due to their reliance on global selfattention. In this paper, we propose a Transformer-efficient occupancy prediction Network, termed TopNet, to overcome these challenges by developing several novel components: Locally-enhanced Context Encoding (LeCE) for enhancing the translation-invariance of the octree nodes, Adaptive-Length Sliding Window Attention (ALSWA) for capturing both global and local dependencies while adaptively adjusting attention weights based on the input window length, Spatial-Gated-enhanced Channel Mixer (SG-CM) for efficient feature aggregation from ancestors and siblings, and Latent-guided Node Occupancy Predictor (LNOP) for improving prediction accuracy of spatially adjacent octree nodes. Comprehensive experiments across both indoor and outdoor point cloud datasets demonstrate that our TopNet achieves state-ofthe-art performance with fewer parameters, further advancing the reduction-efficiency boundaries of PCGC. The code is available at https://github.com/xinjiewang1995/TopNet.
Yifan Zhang 0030, Ting Liu 0017, Xinpu Liu, Ke Xu 0013, Jianwei Wan, Yulan Guo, Hanyun Wang
CVPR6
2025 Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models
abstract
Exploratory data analysis (EDA), coupled with SQL, is essential for data analysts involved in data exploration and analysis. However, data analysts often encounter two primary challenges: (1) the need to craft SQL queries skillfully and (2) the requirement to generate suitable visualization types that enhance the interpretation of query results. Due to its significance, substantial research efforts have been made to explore different approaches to address these challenges, including leveraging large language models (LLMs). However, existing methods fail to meet real-world data exploration requirements primarily due to (1) complex database schema, (2) unclear user intent, (3) limited cross-domain generalization capability, and (4) insufficient end-to-end text-to-visualization capability. This paper presents TiInsight, an automated SQL-based cross-domain exploratory data analysis system. First, we propose a hierarchical data context (i.e., HDC), which leverages LLMs to summarize the contexts related to the database schema, which is crucial for open-world EDA systems to generalize across data domains. Second, the EDA system is divided into four components (i.e., stages): HDC generation, question clarification and decomposition, text-to-SQL generation (i.e., TiSQL), and data visualization (i.e., TiChart). Finally, we implemented an end-to-end EDA system with a user-friendly GUI in the production environment at PingCAP. We have also open-sourced all APIs of TiInsight to facilitate research within the EDA community. Through extensive evaluations by a real-world user study, we demonstrate that TiInsight offers remarkable performance compared to human experts. Additionally, TiSQL achieves an execution accuracy of 86.3% on the Spider dataset when using GPT-4. It also attains an execution accuracy of 60.98% on the Bird test dataset.
Jun-Peng Zhu, Boyan Niu, Peng Cai 0001, Zheming Ni, Jianwei Wan, Kai Xu 0003, Xuan Zhou 0001, Guanglei Bao
Proc. VLDB Endow.5
2025 GCFI-Net: Global-Local Cross-Spatial-Channel Feature Interaction Network for Point Cloud Geometry Compression
abstract
Efficiently compressing large-scale point cloud data under limited bandwidth and computing resource conditions has become a critical issue to be addressed in mobile computing platforms. Although the octree structure can efficiently represent large-scale and complex point clouds, existing octree-based Point Cloud Geometry Compression (PCGC) approaches typically focus on exploiting either spatial or channel features individually, neglecting the interaction across spatial-channel dimensions. In addition, current approaches are also limited to small-scale point clouds due to reliance on global Transformer or local convolutional neural network (CNN). To solve these issues, we introduce GCFI-Net, a global-local cross-spatial-channel feature interaction network for predicting the occupancy probability distribution of each octree node in this paper. In the GCFI-Net, we propose a Multiscale Convolutional Fusion-based Spatial Interaction (MCFSI) module to capture global context and model spatial interactions, and a Global-Local Cross-Channel Interaction (GLCCI) module with dual pathways to integrate global and local cross-channel information. Additionally, we propose a Multiscale-enhanced Spatial and Channel Interaction (MSCI) module to aggregate features from ancestor and sibling nodes, which further enhances the octree node representation ability. Extensive experiments on large-scale sparse LiDAR and dense human body point clouds demonstrate that the proposed GCFI-Net achieves superior compression performance with fewer parameters compared to state-of-the-art PCGC methods.
Yifan Zhang 0030, Xinpu Liu, Ke Xu 0013, Jianwei Wan, Yulan Guo, Hanyun Wang
IEEE Trans. Mob. Comput.5
2025 DuInNet: Dual-Modality Feature Interaction for Point Cloud Completion
abstract
To further promote the development of multimodal point cloud completion, we contribute a large-scale multimodal point cloud completion benchmark ModelNet-MPC with richer shape categories and more diverse test data, which contains nearly 400,000 pairs of high-quality point clouds and rendered images of 40 categories. Besides the fully supervised point cloud completion task, two additional tasks including denoising completion and zero-shot learning completion are proposed in ModelNet-MPC, to simulate real-world scenarios and verify the robustness to noise and the transfer ability across categories of current methods. Meanwhile, considering that existing multimodal completion pipelines usually adopt a unidirectional fusion mechanism and ignore the shape prior contained in the image modality, we propose a Dual-Modality Feature Interaction Network (DuInNet) in this paper. DuInNet iteratively interacts features between point clouds and images to learn both geometric and texture characteristics of shapes with the dual feature interactor. To adapt to specific tasks such as fully supervised, denoising, and zero-shot learning point cloud completions, an adaptive point generator is proposed to generate complete point clouds in blocks with different weights for these two modalities. Extensive experiments on the ShapeNet-ViPC and ModelNet-MPC benchmarks demonstrate that DuInNet exhibits superiority, robustness and transfer ability in all completion tasks over state-of-the-art methods. The code and dataset will be available athttps://github.com/xinpuliu/DuInNet.
Xinpu Liu, Baolin Hou, Hanyun Wang, Ke Xu 0013, Jianwei Wan, Yulan Guo
IEEE Trans. Multim.5
2024 Chat2Query: A Zero-Shot Automatic Exploratory Data Analysis System with Large Language Models
abstract
Data analysts often encounter two primary challenges while conducting exploratory data analysis by SQL: (1) the need to skillfully craft SQL queries, and (2) the requirement to generate suitable visualizations that enhance the interpretation of query results. The emergence of large language models (LLMs) has inaugurated a paradigm shift in text-to-SQL and data-to-chart. This paper presents Chat2Query, an LLM -empowered zero-shot automatic exploration data analysis system. Firstly, Chat2Query provides a user-friendly interface that allows users to employ natural languages to interact with the database directly. Secondly, Chat2Query offers an LLM -empowered text-to-SQL generator, SQL rewriter, SQL formatter, and data-to-chart generator. Thirdly, Chat2Query is uniquely distinguished by its underlying incorporation of the TiDB Serverless, fostering superior elasticity and scalability. This strategic integration empowers Chat2Query with the capability to seamlessly adapt to change workloads, aligning with the evolving demands of the user. We have implemented and deployed Chat2Query in the production environment, and demonstrate its usability and efficiency in three representative real-world scenarios.
Jun-Peng Zhu, Boyan Niu, Zheming Ni, Jianwei Wan
ICDE7
2024 Information geometry based extreme low-bit neural network for point cloud
Yanxin Ma, Ke Xu 0013, Jianwei Wan
Pattern Recognit.4
2024 Terahertz ISAR Imaging With Nonrigid Vibration Compensation and Sidelobe Suppression of Space Targets
abstract
Terahertz (THz) inverse synthetic aperture radar (ISAR) imaging of space targets has the advantage of high resolution and all-day capability for spatial situational awareness (SSA). However, compared with microwave ISAR imaging, microvibrations of the radar platform and target can cause image defocusing and quality degradation in the THz band. Commonly, the existing algorithms considered vibration as simple harmonic motion in a rigid body and estimated the vibration parameters for phase error compensation. Nevertheless, these methods are dependent on the prominent points and cannot focus on the complex nonrigid vibration of the space target. Aiming at these problems, an ISAR imaging framework that can be leveraged for complex nonrigid vibration compensation is proposed. First, a compressed sensing (CS) model is established to compensate for the vibration phase error in ISAR imaging. Then, the alternating direction method of multipliers (ADMMs) is employed to optimize the CS model. Subsequently, considering that strong scattering points of space targets will cover up the weak scattering information in ISAR images, a sidelobe suppression constraint is appended. Thus, the weak scattering information is retained while removing the sidelobe and noise. It is worth mentioning that the ISAR imaging algorithm is derived to a 2-D matrix form to reduce memory usage and improve computational efficiency. Finally, simulated and measured data are utilized to verify the superiority of the proposed algorithm and prove the potential of ISAR imaging for practical space targets.
Zhian Yuan, Xu Chen 0055, Bin Deng 0002, Jianwei Wan, Hongqiang Wang 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 OctPCGC-Net: Learning Octree-Structured Context Entropy Model for Point Cloud Geometry Compression
Hanyun Wang, Ke Xu 0013, Jianwei Wan, Yulan Guo
PRCV (2)4
2023 V2P-SSD: Single-Stage 3-D Object Detection With Voxel-to-Point Transformation
abstract
We study the problem of efficient object detection in 3-D point clouds with the voxel-point framework. Considering a large number of redundant and dense proposals are usually generated for small-sized objects during inference in voxel-based single-stage detectors, the existing detectors usually introduce extra subnetworks to filter and further refine the redundancy proposals. Albeit feasible, the computational and memory cost also increase during inference. In this letter, we introduce a novel voxel-to-point 3-D detector, termed V2P-SSD, which is a novel and lightweight pipeline that jointly integrates the voxel backbone and point head together in a single-stage framework. Different from dense predictions in feature maps, voxels related to objects in our framework are sampled with a fixed number and then transformed into points. Consequently, the point head is used to dynamically generate object proposals. Our voxel-to-point detection paradigm demonstrates a significant precision improvement on small-sized objects without introducing extra memory footprints. Extensive experiments conducted on KITTI and ONCE benchmarks validate the superiority of our method.
Yifan Zhang 0030, Qingyong Hu, Ke Xu 0013, Jianwei Wan, Yulan Guo
IEEE Geosci. Remote. Sens. Lett.4
2023 A gradient optimization and manifold preserving based binary neural network for point cloud
Ke Xu 0013, Yanxin Ma, Jianwei Wan
Pattern Recognit.4
2022 Not All Points Are Equal: Learning Highly Efficient Point-based Detectors for 3D LiDAR Point Clouds
abstract
We study the problem of efficient object detection of 3D LiDAR point clouds. To reduce the memory and computational cost, existing point-based pipelines usually adopt task-agnostic random sampling or farthest point sampling to progressively downsample input point clouds, despite the fact that not all points are equally important to the task of object detection. In particular, the foreground points are inherently more important than background points for object detectors. Motivated by this, we propose a highly-efficient single-stage point-based 3D detector in this paper, termed IA-SSD. The key of our approach is to exploit two learnable, task-oriented, instance-aware downsampling strategies to hierarchically select the foreground points belonging to objects of interest. Additionally, we also introduce a contextual centroid perception module to further estimate precise instance centers. Finally, we build our IA-SSD following the encoder-only architecture for efficiency. Extensive experiments conducted on several large-scale detection benchmarks demonstrate the competitive performance of our IA-SSD. Thanks to the low memory footprint and a high degree of parallelism, it achieves a superior speed of 80+ frames-per-second on the KITTI dataset with a single RTX2080Ti GPU. The code is available at https://github.com/yifanzhang713/IA-SSD.
Yifan Zhang 0030, Qingyong Hu, Guoquan Xu, Yanxin Ma, Jianwei Wan, Yulan Guo
CVPR5
2022 Adaptive Channel Encoding Transformer for Point Cloud Analysis
Guoquan Xu, Hezhi Cao, Yifan Zhang 0030, Yanxin Ma, Jianwei Wan, Ke Xu 0013
ICANN (3)5
2022 Point cloud completion by dynamic transformer with adaptive neighbourhood feature fusion
abstract
Abstract Point cloud completion aims to reconstruct detailed structures from partial observations. However, previous methods often suffered from inaccurate neighbourhood feature extraction and rough reconstruction of complex structures, which is difficult to complete detailed semantic shapes. To solve this problem, we present a method that applies dynamic transformers with adaptive neighbourhood feature fusion operations to resume complete point clouds. Firstly, we propose an adaptive neighbourhood feature extraction module, which contains a learnable global neighbourhood selection strategy and a traditional local k‐nearest neighbour strategy to dynamically select neighbourhood points according to the shape of different objects. Secondly, we observed that traditional point generation methods based on folding‐series operations limit their capacity of generating complex and faithful shapes. Inspired by cell division, we regard the process of reconstructing as point splitting and propose a genetic hierarchical point generation module, which means current points can inherit shape features from previous points and generate more detailed structures of their own. Extensive experiments of point cloud completion are carried out on PCN and Completion3D datasets to verify the effectiveness of our method. The average category chamfer distances of our method are 7.17(×10 −3 ) in PCN and 7.96(×10 −4 ) in Completion3D, which has better completion performance than other methods.
Xinpu Liu, Guoquan Xu, Ke Xu 0013, Jianwei Wan, Yanxin Ma
IET Comput. Vis.4
2022 AGFA-Net: Adaptive Global Feature Augmentation Network for Point Cloud Completion
abstract
Completing shapes of point clouds from partial scans is a fundamental problem for 3-D vision and remote sensing. However, recent methods mainly relied on K-nearest neighbors (KNN) operations to extract local features of point clouds, which are susceptible to outliers and have limited ability to capture features from long-range context information. In this letter, we propose a new framework with an encoder–decoder architecture, named adaptive global feature augmentation network (AGFA-Net) for point cloud completion. The network mainly consists of spatial and channel attention blocks. Spatial attention blocks are used to replace KNN operations and aggregate global features adaptively by calculating per-point attention values, and channel attention blocks are used to augment useful features of geometric details. Meanwhile, several skip connections are added between different attention blocks to selectively convey geometric features from local regions of partial point clouds to the completion process. Experimental results and analyses demonstrate that our method can generate finer shapes of point clouds and outperforms other state-of-the-art methods under widely used benchmark point completion network (PCN) dataset and several terrestrial laser scanning (TLS) data.
Xinpu Liu, Yanxin Ma, Ke Xu 0013, Jianwei Wan, Yulan Guo
IEEE Geosci. Remote. Sens. Lett.4
2017 PolSAR image segmentation based on hierarchical region merging and segment refinement with WMRF model
abstract
In this paper, a superpixel-based segmentation method is proposed for PolSAR images by utilizing hierarchical region merging and segment refinement. The loss of the energy function, which determines the consistency of two adjacent regions from the statistical aspect, is applied to guide the merging procedure. In addition to the edge penalty term, the homogeneity measurement is also employed to prevent merging the regions that are from different land covers or objects. Based on the merged segments, the segment refinement is applied to further improve the segmentation accuracy by iteratively relabeling the edge pixels. It uses a maximum a posterior (MAP) criterion using the statistical distribution of the pixels and the Markov random field (MRF) model. The performance of the proposed method is validated on an experimental PolSAR dataset from the ESAR system.
Wei Wang 0099, Qinglin Zhai, Yifang Ban, Jun Zhang 0044, Jianwei Wan
IGARSS5
2017 An Improved Azimuth Reconstruction Method for Multichannel SAR Using Vandermonde Matrix
abstract
To overcome the contradiction between wide swath and high resolution in synthetic aperture radar systems, a multichannel azimuth reconstruction method is investigated to unambiguously recover the Doppler spectrum. The proposed method is derived from the least squares principle by exploiting a Vandermonde component of the system matrix. The Vandermonde matrix is Doppler independent and data independent. Reconstruction filter weightings can be easily achieved, and performance, including signal-to-noise ratio (SNR) and azimuth ambiguity-to-signal ratio, can be explicitly expressed. By well-conditioning the Vandermonde matrix and coherent processing of all channels, the proposed method improves the reconstruction performance. In simulated reconstruction, compared with the conventional matrix inversion method, the SNR increases by approximately 30 dB.
Jianwei Wan, Qin Xin 0004, Mi He, Yongjian Nian
IEEE Geosci. Remote. Sens. Lett.2
2016 Multichannel azimuth reconstruction of high resolution wide swath SAR via vandermonde matrix
abstract
To overcome the contradiction between wide unambiguous swath and high azimuth resolution in synthetic aperture radar (SAR) system, a multi-channel reconstruction method is investigated to unambiguously recover the azimuth signal even for nonuniform spatial samplings. This reconstruction method is derived from matrix inversion method under least square principle. By taking the system additive noise into consideration, the ill-condition problem of the conventional matrix inversion method is overcome. By decomposing the system matrix into a Vandermonde matrix and a diagonal matrix the reconstruction filters can be easily achieved by calculating the inversion of the Vandermonde matrix which is independent of the Doppler frequency.
Jianwei Wan, Qin Xin 0004
IGARSS2
2016 A Comprehensive Performance Evaluation of 3D Local Feature Descriptors
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan, Ngai Ming Kwok
Int. J. Comput. Vis.5
2016 Integrating Contextual Information With H/̄α Decomposition for PolSAR Data Classification
abstract
The use of contextual information is beneficial to improve both the accuracy and reliability of image classification. Based on the robust fuzzy${c}$-means (RFCM) clustering method and an adaptive Markov random field model, this letter proposes a contextual${H}/{\bar {\alpha }}$classifier for polarimetric synthetic aperture radar images. At each iterative step of RFCM clustering, the prior probability extracted from the local neighborhood is combined with the fuzzy membership derived from inherent polarimetric characteristics, thus the enhanced fuzzy membership is more reliable. In addition, an adaptive smoothing factor is proposed for use during contextual information retrieval, which can prevent oversmoothing and preserve the local spatial details. The experimental results implemented using AIRSAR and ESAR L-band data validate the efficacy of the proposed method. Compared with the iterated Wishart classifier and fuzzy${H}/{\bar {\alpha }}$classifier, the proposed method significantly improves the classification accuracy, with less noise and increased preservation of details.
Wei Wang 0099, Deliang Xiang, Jun Zhang 0044, Jianwei Wan
IEEE Geosci. Remote. Sens. Lett.4
2015 Efficient detection of ground moving targets in FMCW SAR by focusing
abstract
Frequency-Modulated Continuous-Wave Synthetic Aperture Radar (FMCW SAR) is a compact promising remote imaging sensor. Providing FMCW SAR system with simultaneous GMTI application is an appealing and difficult work. To discriminate the target optimally, the total moving target range walk must be taken in account, especially considering its high resolution property. In this paper a concept of relative motion is extended to FMCW SAR. A novel moving target signal model is proposed. By searching the relative velocity, the moving target can be focused accurately like a fixed target in squint SAR mode. An optimal detection and parameter estimation scheme is ultimately introduced.
Qin Xin 0004, Jianwei Wan
IGARSS3
2015 A novel local surface feature for 3D object recognition under clutter and occlusion
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, Min Lu 0001
Inf. Sci.4
2015 Lossless and near-lossless compression of hyperspectral images based on distributed source coding
Yongjian Nian, Mi He, Jianwei Wan
J. Vis. Commun. Image Represent.3
2014 Performance Evaluation of 3D Local Feature Descriptors
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan, Jun Zhang 0044
ACCV (2)5
2014 Automatic markerless registration of mobile LiDAR point-clouds
abstract
Point-cloud registration plays a significant role in the area of mobile LiDAR data processing. This paper proposes an automatic markerless registration algorithm for lidar point-clouds. It first introduces a local feature for point-cloud representation. The feature is invariant to rotations and translations of a point-cloud. It then presents a point-cloud registration method using geometric consistency check and the Iterative Closest Points (ICP) algorithm. Comparative experiments were performed on a publicly available dataset. Experimental results show that our algorithm is very accurate and outperforms the spin image and SHOT based algorithms.
Min Lu 0001, Yulan Guo, Jun Zhang 0044, Jianwei Wan, Jonathan Li 0001
IGARSS4
2014 3D Object Recognition in Cluttered Scenes with Local Surface Features: A Survey
abstract
3D object recognition in cluttered scenes is a rapidly growing research area. Based on the used types of features, 3D object recognition methods can broadly be divided into two categories-global or local feature based methods. Intensive research has been done on local surface feature based methods as they are more robust to occlusion and clutter which are frequently present in a real-world scene. This paper presents a comprehensive survey of existing local surface feature based 3D object recognition methods. These methods generally comprise three phases: 3D keypoint detection, local surface feature description, and surface matching. This paper covers an extensive literature survey of each phase of the process. It also enlists a number of popular and contemporary databases together with their relevant attributes.
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Min Lu 0001, Jianwei Wan
IEEE Trans. Pattern Anal. Mach. Intell.5
2014 An Accurate and Robust Range Image Registration Algorithm for 3D Object Modeling
abstract
Range image registration is a fundamental research topic for 3D object modeling and recognition. In this paper, we propose an accurate and robust algorithm for pairwise and multi-view range image registration. We first extract a set of Rotational Projection Statistics (RoPS) features from a pair of range images, and perform feature matching between them. The two range images are then registered using a transformation estimation method and a variant of the Iterative Closest Point (ICP) algorithm. Based on the pairwise registration algorithm, we propose a shape growing based multi-view registration algorithm. The seed shape is initialized with a selected range image and then sequentially updated by performing pairwise registration between itself and the input range images. All input range images are iteratively registered during the shape growing process. Extensive experiments were conducted to test the performance of our algorithm. The proposed pairwise registration algorithm is accurate, and robust to small overlaps, noise and varying mesh resolutions. The proposed multi-view registration algorithm is also very accurate. Rigorous comparisons with the state-of-the-art show the superiority of our algorithm.
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, Min Lu 0001
IEEE Trans. Multim.4
2013 3D free form object recognition using rotational projection statistics
abstract
Recognizing 3D objects in the presence of clutter and occlusion is a challenging task. This paper presents a 3D free form object recognition system based on a novel local surface feature descriptor. For a randomly selected feature point, a local reference frame (LRF) is defined by calculating the eigenvectors of the covariance matrix of a local surface, and a feature descriptor called rotational projection statistics (RoPS) is constructed by calculating the statistics of the point distribution on 2D planes defined from the LRF. It finally proposes a 3D object recognition algorithm based on RoPS features. Candidate models and transformation hypotheses are generated by matching the scene features against the model features in the library, these hypotheses are then tested and verified by aligning the model to the scene. Comparative experiments were performed on two publicly available datasets and an overall recognition rate of 98.8% was achieved. Experimental results show that our method is robust to noise, mesh resolution variations and occlusion.
Yulan Guo, Mohammed Bennamoun, Ferdous Sohel, Jianwei Wan, Min Lu 0001
WACV4
2013 Rotational Projection Statistics for 3D Local Surface Description and Object Recognition
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Min Lu 0001, Jianwei Wan
Int. J. Comput. Vis.5
2012 Near lossless compression of hyperspectral images based on distributed source coding
Yongjian Nian, Jianwei Wan
Sci. China Inf. Sci.2
2010 Singular point detection using Discrete Hodge Helmholtz Decomposition in fingerprint images
abstract
Identification of singular points in the ridge structure is an important problem in fingerprint matching. This paper presents a fingerprint singular point detection method that is applicable to various fingerprint images regardless of their resolutions. Using the Discrete Hodge Helmholtz Decomposition (DHHD) method, potential structures of singular points can be extracted. We also calculate Poincare Index (PI) of the image. Then we combine DHHD and PI to detect singular points. The result of the experiment on the public databases with different images demonstrates that the proposed method has high accuracy in locating singular points.
Hengzhen Gao, Mrinal Mandal 0001, Gencheng Guo, Jianwei Wan
ICASSP4
2008 Cooperative localization method for multi-robot based on PF-EKF
Jianwei Wan, Yun-Hui Liu 0001, JinXin Shao
Sci. China Ser. F Inf. Sci.2
2006 Neural network-aided adaptive unscented Kalman filter for nonlinear state estimation
abstract
The extended Kalman filter (EKF) is well known as a state estimation method for a nonlinear system and has been used to train a multilayered neural network (MNN) by augmenting the state with unknown connecting weights. However, EKF has the inherent drawbacks such as instability due to linearization and costly calculation of Jacobian matrices, and its performance degrades greatly, especially when the nonlinearity is severe. In this letter, first a more robust learning algorithm for an MNN-based on unscented Kalman filter (UKF) is derived. Since it gives a more accurate estimate of the linkweights, the convergence performance is improved. The algorithm is then extended further to develop a NN-aided UKF for nonlinear state estimation. The NN in this algorithm is used to approximate the uncertainty of the system model due to mismodeling, extreme nonlinearities, etc. The UKF is used for both NN online training and state estimation simultaneously. Simulation results show that the new algorithm is very effective and is closer to optimal fashion in nonlinear filtering compared with traditional methods.
Ronghui Zhan, Jianwei Wan
IEEE Signal Process. Lett.2