VLDB 2026 Research / reviewers in the wild / expert
Zhiyuan Zhang 0004
dblp:72/1760-4
· DBLP profile ↗
27ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0003-3945-5638ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Active Training for Deep LiDAR OdometryabstractRobust and efficient deep LiDAR odometry models are crucial for accurate localization and 3D reconstruction, but typically require extensive and diverse training data to adapt to diverse environments, leading to inefficiencies. To tackle this, we introduce an active training framework designed to selectively extract training data from diverse environments, thereby reducing the training load and enhancing model generalization. Our framework is based on two key strategies: Initial Training Set Selection (ITSS) and Active Incremental Selection (AIS). ITSS begins by breaking down motion sequences from general weather into nodes and edges for detailed trajectory analysis, prioritizing diverse sequences to form a rich initial training dataset for training the base model. For complex sequences that are difficult to analyze, especially under challenging snowy weather conditions, AIS uses scene reconstruction and prediction inconsistency to iteratively select training samples, refining the model to handle a wide range of real-world scenarios. Experiments across datasets and weather conditions validate our approach’s effectiveness. Notably, our method matches the performance of full-dataset training with just 52% of the sequence volume, demonstrating the training efficiency and robustness of our active training paradigm. By optimizing the training process, our approach sets the stage for more agile and reliable LiDAR odometry systems, capable of navigating diverse environmental conditions with greater precision. Beibei Zhou, Zhiyuan Zhang 0004, Zhenbo Song, Jianhui Guo, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | WA-FDNet: A Unified Weight Adaptation Network for Multimodal Image Fusion and Object Detection
Yanyin Guo, Junwei Li 0009, Zhiyuan Zhang 0004 |
CGI (2) | 4 |
| 2025 | Bridging the Modality Gap: Advancing Multimodal Human Pose Estimation with Modality-Adaptive Pose Estimator and Novel Benchmark Datasets
Jiangnan Xia, Zhiyuan Zhang 0004, Yanyin Guo, Jianghan Cheng, Junwei Li 0009 |
CVM (3) | 2 |
| 2025 | FCAD: Feature-Coupled Anisotropic Diffusion for Continuous Graph LearningabstractIn this work, we propose a novel continuous graph neural network called FCAD (Feature-Coupled Anisotropic Diffusion) for the task of node classification on graphs. Our approach is motivated by the success of feature-coupled anisotropic diffusion PDEs in multivalued image restoration. Our method introduces a total variation regularization-inspired anisotropic term to control diffusion between nodes and incorporates a learnable parameterization for feature coupling during the diffusion process. Our model performs competitively against several GNN baselines for both heterophilous and homophilous graphs, demonstrating notable benefits for heterophilous graphs due to the learnable feature coupling. Amitoz Azad, Zhiyuan Zhang 0004 |
ECAI | 2 |
| 2025 | Improving Multimodal Human Pose Estimation by Adversarial Modality Enhancement†abstractHuman pose estimation in computer vision predominantly focuses on the visible modality, with limited research on the infrared modality. No existing methods demonstrate robust performance across both modalities, missing their complementary strengths. This gap arises from the lack of a multimodal benchmark and the difficulty of developing robust multimodal capabilities. To address this, we introduce MMPD, a novel visible-infrared multimodal pose benchmark with high-quality annotations for both modalities. Leveraging MMPD, we expose the limitations of state-of-the-art methods due to modality variance. To overcome this challenge, we propose a novel method-agnostic scheme called AMMPE. By employing the Modality Adversarial Enhancement Stage and Modality Interaction Stage, AMMPE easily incorporates multimodal information without additional pose annotations and enhances effective modality interaction. Extensive experiments demonstrate that AMMPE improves performance in both visible and infrared modalities, achieving excellent modality robustness. The code and benchmark is avaible at: https://github.com/ICANDOALLTHINGSSS/Adversarial-Multi-Modality-Pose-Estimation Jiangnan Xia, Yanyin Guo, Jianghan Cheng, Junwei Li 0009, Zhiyuan Zhang 0004 |
ICASSP | 7 |
| 2025 | A Dental Periapical X-ray Images Segmentation Network Based on Pixel-Wise Contrastive Learning with Dual Attention MechanismsabstractPeriapical radiographs tend to have poor quality due to factors like acquisition techniques, equipment limitations, and patient differences. These factors lead to discrepancies in images, making precise segmentation challenging. However, accurate segmentation is crucial in dental practices. To address this issue, deep learning can be employed to improve segmentation accuracy and efficiency, thereby providing more reliable support for clinical diagnosis. In this study, we propose a deep learning network based on an encoder-decoder architecture, which integrates a dual attention mechanism and pixel-wise contrastive learning to address the tooth segmentation problem. We design the Dual Attention Contrast (DAC) module, which enhances feature maps through joint spatial and channel attention before utilizing optimized multi-scale features for both segmentation prediction and pixel-wise contrastive learning. This module, implemented with a multi-level deployment strategy, strengthens the network's ability to extract discriminative features from the multi-scale anatomical structures in periapical radiographs while constructing a globally structured feature space across datasets. The dual attention mechanism improves the recognition of key local features through spatial-channel collaborative calibration, while pixel-wise contrastive learning explicitly constrains the topological structure of the feature space, effectively mitigating the impact of inter-image variations on segmentation accuracy. Experimental results show that the proposed model outperforms current mainstream models across various evaluation metrics in periapical radiograph segmentation tasks. Yibin Tian, Zhiyuan Zhang 0004, Xueyang Zhang, Shan LianLei |
IJCNN | 4 |
| 2025 | HDiffTG: A Lightweight Hybrid Diffusion-Transformer-GCN Architecture for 3D Human Pose EstimationabstractWe propose HDiffTG, a novel 3D Human Pose Estimation (3DHPE) method that integrates Transformer, Graph Convolutional Network (GCN), and diffusion model into a unified framework. HDiffTG leverages the strengths of these techniques to significantly improve pose estimation accuracy and robustness while maintaining a lightweight design. The Transformer captures global spatiotemporal dependencies, the GCN models local skeletal structures, and the diffusion model provides step-by-step optimization for fine-tuning, achieving a complementary balance between global and local features. This integration enhances the model’s ability to handle pose estimation under occlusions and in complex scenarios. Furthermore, we introduce lightweight optimizations to the integrated model and refine the objective function design to reduce computational overhead without compromising performance. Evaluation results on the Human3.6M and MPI-INF-3DHP datasets demonstrate that HDiffTG achieves state-of-the-art (SOTA) performance on the MPI-INF-3DHP dataset while excelling in both accuracy and computational efficiency. Additionally, the model exhibits exceptional robustness in noisy and occluded environments. Source codes and models are available at https://github.com/CirceJie/HDiffTG Yajie Fu, Chaorui Huang, Junwei Li 0009, Hui Kong 0001, Yibin Tian, Huakang Li, Zhiyuan Zhang 0004 |
IJCNN | 7 |
| 2025 | TS-Diff: Two-Stage Diffusion Model for Low-Light RAW Image EnhancementabstractThis paper presents a novel Two-Stage Diffusion Model (TS-Diff) for enhancing extremely low-light RAW images. In the pre-training stage, TS-Diff synthesizes noisy images by constructing multiple virtual cameras based on a noise space. Camera Feature Integration (CFI) modules are then designed to enable the model to learn generalizable features across diverse virtual cameras. During the aligning stage, CFIs are averaged to create a target-specific CFIT, which is fine-tuned using a small amount of real RAW data to adapt to the noise characteristics of specific cameras. A structural reparameterization technique further simplifies CFITfor efficient deployment. To address color shifts during the diffusion process, a color corrector is introduced to ensure color consistency by dynamically adjusting global color distributions. Additionally, a novel dataset, QID, is constructed, featuring quantifiable illumination levels and a wide dynamic range, providing a comprehensive benchmark for training and evaluation under extreme low-light conditions. Experimental results demonstrate that TS-Diff achieves state-of-the-art performance on multiple datasets, including QID, SID, and ELD, excelling in denoising, generalization, and color consistency across various cameras and illumination levels. These findings highlight the robustness and versatility of TS-Diff, making it a practical solution for low-light imaging applications. Source codes and models are available at https://github.com/CircccleK/TS-Diff Zhiyuan Zhang 0004, Jiangnan Xia, Jianghan Cheng, Junwei Li 0009, Yibin Tian, Hui Kong 0001 |
IJCNN | 2 |
| 2025 | LSFDNet: A Single-Stage Fusion and Detection Network for Ships Using SWIR and LWIR
Yanyin Guo, Runxuan An, Junwei Li 0009, Zhiyuan Zhang 0004 |
ACM Multimedia | 4 |
| 2025 | CSSA-Fusion: Channel Selective and Spatial Alignment Infrared-Visible Image FusionabstractInfrared-visible image fusion aims to integrate complementary information from two modalities to generate images with enriched semantic content. However, existing methods often neglect two critical aspects: the design of a local–global feature enhancement architecture and spatial alignment. To address these challenges, we propose Channel Selective and Spatial Alignment Fusion (CSSA-Fusion), a novel framework composed of two synergistic modules. The first is a selective channel and redundancy suppression module, which introduces a dual-branch selective channel attention mechanism to jointly capture local saliency and global channel importance for enhanced feature representation, and an informativeness–redundancy separation strategy to suppress redundant information while preserving discriminative features. The second is a directional feature processing module, consisting of a mechanism that decouples and recombines modality-specific and common representations to mitigate mutual interference, and a spatial alignment module that performs geometric alignment via horizontal and vertical coordinate decomposition to correct spatial discrepancies between modalities. Extensive experiments on benchmark datasets demonstrate that CSSA-Fusion consistently outperforms state-of-the-art deep learning methods on multiple quality metrics. The fused images exhibit superior visual quality with well-preserved textures and enhanced semantic details. Zhongrui Xiao, Zhiyuan Zhang 0004, Yibin Tian |
MMAsia | 5 |
| 2025 | VDGPG: A Virtual Data-Guided Prompt Generation Framework for Incremental Learning with Application to Wafer Defect DetectionabstractAlthough convolutional neural networks have been widely used for wafer defect detection in semiconductor manufacturing, they typically rely on static offline datasets to train models. These models show strong reliability when detecting known defect types, but struggle with unknown ones, posing challenges in model adaptation and leading to high maintenance costs. Incremental Learning (IL) offers a solution that allows models to continuously adapt to new types of defects without accessing full historical data. This paper introduces a Virtual Data-Guided Prompt Generation (VDGPG) framework, a novel IL approach for wafer defect detection that integrates prompt-guided learning and dual-branch virtual data generation. Specifically, VDGPG assigns task-specific prompt vectors to individual attention heads, using a channel attention gating mechanism and similarity computation to select prompt vectors from a prompt pool. This enables the model to focus more effectively on the relevant features for each category of defects. The dual-branch virtual data generation module generates diverse virtual samples, with a special emphasis on contour edge generation, which guides the model to learn features of potential new categories proactively. Experiments using the public WM-811K dataset demonstrate that VDGPG achieves significant performance improvements in wafer defect detection over existing IL methods. Bingwen Liu, Yibin Tian, Shanglei Chai, Zhiyuan Zhang 0004 |
SMC | 4 |
| 2025 | Detection of Incomplete Root Canal Obturations in Dental X-ray Images via Spatial-Semantic Attention and Dynamic Feature CalibrationabstractTo address the challenges of low resolution, loss of small target features, and interference from complex anatomical structures in detecting incomplete root canal obturations in dental periapical radiographs, this article proposes an improved YOLOv8 model. First, we design a Convolution module with Space-to-Depth Transformation (SDT-Conv) that preserves feature map resolution through spatial depth-wise separable convolutions, effectively mitigating loss of small targets caused by downsampling operations. Second, we construct a Dynamic Iterative Token Aggregator (DITA) architecture that enhances global feature representation through hyper-token spatial aggregation and semantic correlation, while employing a spatial-semantic dual-stream attention mechanism to strengthen multiscale feature fusion capabilities, thereby providing richer feature information for the entire network. Finally, we embed an Efficient Multiscale Attention (EMA) dynamic calibration mechanism in the detection head, which optimizes feature responses through cross-channel weight adaptation, enabling the model to precisely localize small object boundaries. The experimental results demonstrate that the improved model achieves 81.5% mAP@50 on the validation set, representing a 12.8% improvement over YOLOv8n. It effectively overcomes the challenges posed by variations in obturation materials, dental structure occlusions, and low-contrast interference. Zhiqi Ren, Shanglei Chai, Zhiyuan Zhang 0004, Xueyang Zhang, Yibin Tian |
SMC | 3 |
| 2025 | YOLOv8-CTCD: An Improved YOLOv8 for Cherry Tomato Cluster Detection in Robotic HarvestingabstractCherry tomato harvesting is generally performed manually. Robotic harvesting is gaining increasing interest from both academia and industry. This paper proposes a cherry tomato cluster detection algorithm based on YOLOv8, named YOLOv8-CTCD. First, the YOLOv8 input channels are adjusted to enable 4-channel RGB-D images as input. Subsequently, a CARAFE-M module is designed to replace the upsampling method in YOLOv8n. It maintains a lightweight architecture while achieving a larger receptive field, allowing effective aggregation of contextual information. In addition, it assigns greater weight to more important features. Moreover, a C2f-MLCA module is introduced into YOLOv8, which integrates information from feature maps at different levels and enhances the network’s capability of feature extraction. It also integrates the SPPELAN module to strengthen its feature fusion capability. YOLOv8-CTCD has been evaluated using a private cherry tomato dataset obtained from a greenhouse farm. The experimental results show that it achieves an mAP@50 of 93.8% and an mAP@50:90 of 68%, which represents improvements of 2.1% and 3% over YOLOv8n, respectively. Shanglei Chai, Zhiyuan Zhang 0004, Yibin Tian |
SMC | 3 |
| 2024 | Test-Time Augmentation for 3D Point Cloud Classification and SegmentationabstractData augmentation is a powerful technique to enhance the performance of a deep learning task but has received less attention in 3D deep learning. It is well known that when 3D shapes are sparsely represented with low point density, the performance of the downstream tasks drops significantly. This work explores test-time augmentation (TTA) for 3D point clouds. We are inspired by the recent revolution of learning implicit representation and point cloud upsampling, which can produce high-quality 3D surface reconstruction and proximity-to-surface, respectively. Our idea is to leverage the implicit field reconstruction or point cloud upsampling techniques as a systematic way to augment point cloud data. Mainly, we test both strategies by sampling points from the reconstructed results and using the sampled point cloud as test-time augmented data. We show that both strategies are effective in improving accuracy. We observed that point cloud upsampling for test-time augmentation can lead to more significant performance improvement on downstream tasks such as object classification and segmentation on the ModelNet40, ShapeNet, ScanObjectNN, and SemanticKITTI datasets, especially for sparse point clouds. Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang 0004, Binh-Son Hua, Sai-Kit Yeung |
3DV | 3 |
| 2024 | RISurConv: Rotation Invariant Surface Attention-Augmented Convolutions for 3D Point Cloud Classification and Segmentation
Zhiyuan Zhang 0004, Licheng Yang 0003, Zhiyu Xiang |
ECCV (28) | 1 |
| 2024 | Teeth Segmentation from Bite-Wing X-Ray Images by Integrating Nested Dual UNet with Swin TransformersabstractIn medical practice, the precision of image segmentation is crucial for diagnosis and treatment evaluations. Specifically, in dentistry, accurate teeth segmentation from bite-wing images is important for automatic and objective evaluations of root canal treatments. This study introduces$\text{Swin}-\mathrm{U}^{2} \text{Net}$, a model merging the nested dual UNet with residual U-block and Swin Transformers. It combines the local feature extraction capability of the former and the global attention and context understanding of the latter. It has been evaluated for tooth root segmentation using 500 bite-wing dental x-ray images obtained from a root canal treatment clinic. It achieved the best segmentation outcome in terms of Intersection over Union (IOU) and the third best result in terms of Dice Similarity Coefficient (DSC) with the second least amount of network parameters among six UNet-like models, thus it is effective and efficient. Yibin Tian, Zhiyuan Zhang 0004, Xueyang Zhang, Bingran Du |
SMC | 3 |
| 2024 | An adaptive network fusing light detection and ranging height-sliced bird's-eye view and vision for place recognition
Zuo Jiang, Yibin Ye, Junwei Li 0009, Zhiyuan Zhang 0004 |
Eng. Appl. Artif. Intell. | 7 |
| 2022 | CVFNet: Real-time 3D Object Detection by Learning Cross View FeaturesabstractIn recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve time-consuming operations such as 3D convolutions on voxels or ball query among points, making the resulting network inappropriate for time critical applications. On the other hand, 2D view-based methods feature high computing efficiency while usually obtaining inferior performance than the voxel or point based methods. In this work, we present a real-time view-based single stage 3D object detector, namely CVFNet to fulfill this task. To strengthen the cross-view feature learning under the condition of demanding efficiency, our framework extracts the features of different views and fuses them in an efficient progressive way. We first propose a novel Point-Range feature fusion module that deeply integrates point and range view features in multiple stages. Then, a special Slice Pillar is designed to well maintain the 3D geometry when transforming the obtained deep point-view features into bird's eye view. To better balance the ratio of samples, a sparse pillar detection head is presented to focus the detection on the nonempty grids. We conduct experiments on the popular KITTI and NuScenes benchmark, and state-of-the-art performances are achieved in terms of both accuracy and speed. Jiaqi Gu 0004, Zhiyu Xiang, Tingming Bai, Lingxuan Wang, Xijun Zhao, Zhiyuan Zhang 0004 |
IROS | 7 |
| 2022 | RIConv++: Effective Rotation Invariant Convolutions for 3D Point Clouds Deep Learning
Zhiyuan Zhang 0004, Binh-Son Hua, Sai-Kit Yeung |
Int. J. Comput. Vis. | 1 |
| 2020 | Global Context Aware Convolutions for 3D Point Cloud UnderstandingabstractRecent advances in deep learning for 3D point clouds have shown great promises in scene understanding tasks thanks to the introduction of convolution operators to consume 3D point clouds directly in a neural network. Point cloud data, however, could have arbitrary rotations, especially those acquired from 3D scanning. Recent works show that it is possible to design point cloud convolutions with rotation invariance property, but such methods generally do not perform as well as translation-invariant only convolution. We found that a key reason is that compared to point coordinates, rotation-invariant features consumed by point cloud convolution are not as distinctive. To address this problem, we propose a novel convolution operator that enhances feature distinction by integrating global context information from the input point cloud to the convolution. To this end, a globally weighted local reference frame is constructed in each point neighborhood in which the local point set is decomposed into bins. Anchor points are generated in each bin to represent global shape features. A convolution can then be performed to transform the points and anchor features into final rotation-invariant features. We conduct several experiments on point cloud classification, part segmentation, shape retrieval, and normals estimation to evaluate our convolution, which achieves state-of-the-art accuracy under challenging rotations. Zhiyuan Zhang 0004, Binh-Son Hua, Yibin Tian, Sai-Kit Yeung |
3DV | 1 |
| 2020 | 3D Dental Biometrics: Automatic Pose-invariant Dental Arch Extraction and MatchingabstractA novel automatic pose-invariant dental arch extraction and matching framework is developed for 3D dental identification using laser-scanned dental plasters. In our previous attempt [1]-[5], 3D point-based algorithms have been developed and they have shown a few advantages over existing 2D dental identifications. This study is a continuous effort in developing arch-based algorithms to extract and match dental arch feature in an automatic and pose-invariant way. As best as we know, this is the first attempt at automatic dental arch extraction and matching for 3D dental identification. A Radial Ray Algorithm (RRA) is proposed by projecting dental arch shape from 3D to 2D. This algorithm is fully automatic and fast. Preliminary identification result is obtained by matching 11 postmortem (PM) samples against 200 ante-mortem (AM) samples. 72.7% samples achieved top 5% accuracy. 90.9% samples achieved top 10% accuracy and all 11 samples (100%) achieved top 15.5% accuracy out of the 200-rank list. In addition, the time for identifying a single subject from 200 subjects has been significantly reduced from 45 minutes to 5 minutes by matching the extracted 2D dental arch. Although the extracted 2D arch feature is not as accurate and discriminative as the full 3D arch, it may serve as an important filter feature to improve the identification speed in future investigations. Zhiyuan Zhang 0004 |
ICPR | 2 |
| 2019 | Rotation Invariant Convolutions for 3D Point Clouds Deep LearningabstractRecent progresses in 3D deep learning has shown that it is possible to design special convolution operators to consume point cloud data. However, a typical drawback is that rotation invariance is often not guaranteed, resulting in networks that generalizes poorly to arbitrary rotations. In this paper, we introduce a novel convolution operator for point clouds that achieves rotation invariance. Our core idea is to use low-level rotation invariant geometric features such as distances and angles to design a convolution operator for point cloud learning. The well-known point ordering problem is also addressed by a binning approach seamlessly built into the convolution. This convolution operator then serves as the basic building block of a neural network that is robust to point clouds under 6-DoF transformations such as translation and rotation. Our experiment shows that our method performs with high accuracy in common scene understanding tasks such as object classification and segmentation. Compared to previous and concurrent works, most importantly, our method is able to generalize and achieve consistent results across different scenarios in which training and testing can contain arbitrary rotations. Our implementation is publicly available at our project page. Zhiyuan Zhang 0004, Binh-Son Hua, David W. Rosen, Sai-Kit Yeung |
3DV | 1 |
| 2019 | ShellNet: Efficient Point Cloud Convolutional Neural Networks Using Concentric Shells StatisticsabstractDeep learning with 3D data has progressed significantly since the introduction of convolutional neural networks that can handle point order ambiguity in point cloud data. While being able to achieve good accuracies in various scene understanding tasks, previous methods often have low training speed and complex network architecture. In this paper, we address these problems by proposing an efficient end-to-end permutation invariant convolution for point cloud deep learning. Our simple yet effective convolution operator named ShellConv uses statistics from concentric spherical shells to define representative features and resolve the point order ambiguity, allowing traditional convolution to perform on such features. Based on ShellConv we further build an efficient neural network named ShellNet to directly consume the point clouds with larger receptive fields while maintaining less layers. We demonstrate the efficacy of ShellNet by producing state-of-the-art results on object classification, object part segmentation, and semantic scene segmentation while keeping the network very fast to train. Zhiyuan Zhang 0004, Binh-Son Hua, Sai-Kit Yeung |
ICCV | 1 |
| 2016 | Efficient 3D dental identification via signed feature histogram and learning keypoint detection
Zhiyuan Zhang 0004, Sim Heng Ong, Kelvin Weng Chiong Foong |
Pattern Recognit. | 1 |
| 2013 | Symmetry Robust Descriptor for Non-Rigid Surface MatchingabstractAbstract In this paper, we propose a novel shape descriptor that is robust in differentiating intrinsic symmetric points on geometric surfaces. Our motivation is that even the state‐of‐theart shape descriptors and non‐rigid surface matching algorithms suffer from symmetry flips. They cannot differentiate surface points that are symmetric or near symmetric. Hence a left hand of one human model may be matched to a right hand of another. Our Symmetry Robust Descriptor (SRD) is based on a signed angle field, which can be calculated from the gradient fields of the harmonic fields of two point pairs. Experiments show that the proposed shape descriptor SRD results in much less symmetry flips compared to alternative methods. We further incorporate SRD into a stand‐alone algorithm to minimize symmetry flips in finding sparse shape correspondences. SRD can also be used to augment other modern non‐rigid shape matching algorithms with ease to alleviate symmetry confusions. Zhiyuan Zhang 0004, KangKang Yin, Kelvin Weng Chiong Foong |
Comput. Graph. Forum | 1 |
| 2012 | Improved spin images for 3D surface matching using signed anglesabstractDespite the popularity of spin images in surface matching and registration, disadvantages such as noise sensitivity and low discriminative ability still hindered their usefulness in real applications. In this paper, a novel approach was proposed for improving the spin images. The proposed method modified the standard spin images by using angle information between the normals of reference point and neighboring points. This information largely increased the robustness to noise without losing the intrinsic advantages of spin images. Moreover, signs were defined to incorporate the directions of angles which were shown to be able to further improve the descriptive power. Experiments were also conducted to show the outperformance of improved spin images under different levels of noise, and good agreements were obtained by comparing with the standard spin images and a recent popular 3D descriptor. Zhiyuan Zhang 0004, Sim Heng Ong, Kelvin Weng Chiong Foong |
ICIP | 1 |
| 2009 | Multi-view Ear Recognition Based on Moving Least Square Pose Interpolation
Heng Liu 0003, David Zhang 0001, Zhiyuan Zhang 0004 |
ICIC (2) | 3 |