VLDB 2026 Research / reviewers in the wild / expert
Zhenhua Guo 0001
dblp:41/294-1
· DBLP profile ↗
92ranked-venue papers
16as first author
31since 2021 · last 2026
0000-0002-8201-0864ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 8 first-author · 14 since 2021Security and privacy · 8 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFPA: Dual-level Feature Perturbation Augmentation for Whole Slide Image Analysis
Chaojun Zhang, Zhenhua Guo 0001, Xiaoshuang Shi |
ICIC (7) | 3 |
| 2025 | Reti-Diff: Illumination Degradation Image Restoration with Retinex-based Latent Diffusion ModelabstractIllumination degradation image restoration (IDIR) techniques aim to improve the visibility of degraded images and mitigate the adverse effects of deteriorated illumination. Among these algorithms, diffusion-based models (DM) have shown promising performance but are often burdened by heavy computational demands and pixel misalignment issues when predicting the image-level distribution. To tackle these problems, we propose to leverage DM within a compact latent space to generate concise guidance priors and introduce a novel solution called Reti-Diff for the IDIR task. Specifically, Reti-Diff comprises two significant components: the Retinex-based latent DM (RLDM) and the Retinex-guided transformer (RGformer). RLDM is designed to acquire Retinex knowledge, extracting reflectance and illumination priors to facilitate detailed reconstruction and illumination correction. RGformer subsequently utilizes these compact priors to guide the decomposition of image features into their respective reflectance and illumination components. Following this, RGformer further enhances and consolidates these decomposed features, resulting in the production of refined images with consistent content and robustness to handle complex degradation scenarios. Extensive experiments demonstrate that Reti-Diff outperforms existing methods on three IDIR tasks, as well as downstream applications. Chunming He, Chengyu Fang 0001, Yulun Zhang 0001, Longxiang Tang, Jinfa Huang, Kai Li 0012, Zhenhua Guo 0001, Xiu Li 0001, Sina Farsiu |
ICLR | 7 |
| 2025 | Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous DrivingabstractWith the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road scene segmentation approaches are predominantly trained on closed-set data, resulting in insufficient detection capabilities for out-of-distribution (OOD) objects. To overcome this limitation, road anomaly detection methods have been proposed. However, existing methods primarily depend on image inpainting and OOD distribution detection techniques, facing two critical issues: (1) inadequate consideration of the objectiveness attributes of anomalous regions, causing incomplete segmentation when anomalous objects share similarities with known classes, and (2) insufficient attention to environmental constraints, leading to the detection of anomalies irrelevant to autonomous driving tasks. In this paper, we propose a novel framework termed Segmenting Objectiveness and Task-Awareness (SOTA) for autonomous driving scenes. Specifically, SOTA enhances the segmentation of objectiveness through a Semantic Fusion Block (SFB) and filters anomalies irrelevant to road navigation tasks using a Scene-understanding Guided Prompt-Context Adaptor (SG-PCA). Extensive empirical evaluations on multiple benchmark datasets, including Fishyscapes Lost and Found, Segment-Me-If-You-Can, and RoadAnomaly, demonstrate that the proposed SOTA consistently improves OOD detection performance across diverse detectors, achieving robust and accurate segmentation outcomes. Mi Zheng, Guanglei Yang, Zitong Huang, Zhenhua Guo 0001, Kevin Han, Wangmeng Zuo |
ACM Multimedia | 4 |
| 2025 | Diffusion Models in Low-Level Vision: A SurveyabstractDeep generative models have gained considerable attention in low-level vision tasks due to their powerful generative capabilities. Among these, diffusion model-based approaches, which employ a forward diffusion process to degrade an image and a reverse denoising process for image generation, have become particularly prominent for producing high-quality, diverse samples with intricate texture details. Despite their widespread success in low-level vision, there remains a lack of a comprehensive, insightful survey that synthesizes and organizes the advances in diffusion model-based techniques. To address this gap, this paper presents the first comprehensive review focused on denoising diffusion models applied to low-level vision tasks, covering both theoretical and practical contributions. We outline three general diffusion modeling frameworks and explore their connections with other popular deep generative models, establishing a solid theoretical foundation for subsequent analysis. We then categorize diffusion models used in low-level vision tasks from multiple perspectives, considering both the underlying framework and the target application. Beyond natural image processing, we also summarize diffusion models applied to other low-level vision domains, including medical imaging, remote sensing, and video processing. Additionally, we provide an overview of widely used benchmarks and evaluation metrics in low-level vision tasks. Our review includes an extensive evaluation of diffusion model-based techniques across six representative tasks, with both quantitative and qualitative analysis. Finally, we highlight the limitations of current diffusion models and propose four promising directions for future research. This comprehensive review aims to foster a deeper understanding of the role of denoising diffusion models in low-level vision. Chunming He, Yuqi Shen, Chengyu Fang 0001, Fengyang Xiao, Longxiang Tang, Yulun Zhang 0001, Wangmeng Zuo, Zhenhua Guo 0001, Xiu Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | UniM2AE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
Jian Zou 0005, Guanglei Yang, Zhenhua Guo 0001, Tao Luo 0014, Chun-Mei Feng 0001, Wangmeng Zuo |
ECCV (22) | 4 |
| 2024 | Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged ObjectsabstractCamouflaged object detection (COD) is the challenging task of identifying camouflaged objects visually blended into surroundings. Albeit achieving remarkable success, existing COD detectors still struggle to obtain precise results in some challenging cases. To handle this problem, we draw inspiration from the prey-vs-predator game that leads preys to develop better camouflage and predators to acquire more acute vision systems and develop algorithms from both the prey side and the predator side. On the prey side, we propose an adversarial training framework, Camouflageator, which introduces an auxiliary generator to generate more camouflaged objects that are harder for a COD method to detect. Camouflageator trains the generator and detector in an adversarial way such that the enhanced auxiliary generator helps produce a stronger detector. On the predator side, we introduce a novel COD method, called Internal Coherence and Edge Guidance (ICEG), which introduces a camouflaged feature coherence module to excavate the internal coherence of camouflaged objects, striving to obtain more complete segmentation results. Additionally, ICEG proposes a novel edge-guided separated calibration module to remove false predictions to avoid obtaining ambiguous boundaries. Extensive experiments show that ICEG outperforms existing COD detectors and Camouflageator is flexible to improve various COD detectors, including ICEG, which brings state-of-the-art COD performance. Chunming He, Kai Li 0012, Yachao Zhang 0001, Yulun Zhang 0001, Chenyu You, Zhenhua Guo 0001, Xiu Li 0001, Martin Danelljan, Fisher Yu 0001 |
ICLR | 6 |
| 2024 | An Adaptive Region-Based Transformer for Nonrigid Medical Image Registration With a Self-Constructing Latent GraphabstractNonrigid registration of medical images is formulated usually as an optimization problem with the aim of seeking out the deformation field between a referential-moving image pair. During the past several years, advances have been achieved in the convolutional neural network (CNN)-based registration of images, whose performance was superior to most conventional methods. More lately, the long-range spatial correlations in images have been learned by incorporating an attention-based model into the transformer network. However, medical images often contain plural regions with structures that vary in size. The majority of the CNN- and transformer-based approaches adopt embedding of patches that are identical in size, disallowing representation of the inter-regional structural disparities within an image. Besides, it probably leads to the structural and semantical inconsistencies of objects as well. To address this issue, we put forward an innovative module called region-based structural relevance embedding (RSRE), which allows adaptive embedding of an image into unequally-sized structural regions based on the similarity of self-constructing latent graph instead of utilizing patches that are identical in size. Additionally, a transformer is integrated with the proposed module to serve as an adaptive region-based transformer (ART) for registering medical images nonrigidly. As demonstrated by the experimental outcomes, our ART is superior to the advanced nonrigid registration approaches in performance, whose Dice score is 0.734 on the LPBA40 dataset with 0.318% foldings for deformation field, and is 0.873 on the ADNI dataset with 0.331% foldings. Xiu Li 0001, Zhenhua Guo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Camouflaged Object Detection with Feature Decomposition and Edge ReconstructionabstractCamouflaged object detection (COD) aims to address the tough issue of identifying camouflaged objects visually blended into the surrounding backgrounds. COD is a challenging task due to the intrinsic similarity of camouflaged objects with the background, as well as their ambiguous boundaries. Existing approaches to this problem have developed various techniques to mimic the human visual system. Albeit effective in many cases, these methods still struggle when camouflaged objects are so deceptive to the vision system. In this paper, we propose the FEature Decomposition and Edge Reconstruction (FEDER) model for COD. The FEDER model addresses the intrinsic similarity of foreground and background by decomposing the features into different frequency bands using learnable wavelets. It then focuses on the most informative bands to mine subtle cues that differentiate foreground and background. To achieve this, a frequency attention module and a guidance-based feature aggregation module are developed. To combat the ambiguous boundary problem, we propose to learn an auxiliary edge reconstruction task alongside the COD task. We design an ordinary differential equation-inspired edge reconstruction module that generates exact edges. By learning the auxiliary task in conjunction with the COD task, the FEDER model can generate precise prediction maps with accurate object boundaries. Experiments show that our FEDER model significantly outperforms state-of-the-art methods with cheaper computational and memory costs. The code will be available at https://github.com/ChunmingHe/FEDER. Chunming He, Kai Li 0012, Yachao Zhang 0001, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001 |
CVPR | 6 |
| 2023 | A Two-Branch Network for Video Anomaly Detection with Spatio-Temporal Feature LearningabstractVideo anomaly detection is very challenging, as most anomalies are rare and inconclusive. Previous weakly supervised learning approaches utilize the classifier trained with video-level labels to locate anomalous clips from the video. However, the anomalous clips often contain both anomalies and numerous irrelevant background behaviors, increasing the difficulty of localization. In this work, we propose a two-branch network to obtain the global and each local object’s action information of the clip respectively, where the local objects are extracted by a pre-trained object detector. This local-cum-global perception highlights the anomalous features from the background noise. We further propose a spatio-temporal relationship network, which is based on the attention mechanism to model the spatial relations of different objects and the temporal correlations among different clips to efficiently capture the spatio-temporal distribution of anomalies in the video. Extensive experiments on two benchmarks show that our method achieves significant performance gains. Guoqiu Li, Shengjie Chen, Yujiu Yang 0001, Zhenhua Guo 0001 |
ICASSP | 4 |
| 2023 | Degradation-Resistant Unfolding Network for Heterogeneous Image FusionabstractHeterogeneous image fusion (HIF) techniques aim to enhance image quality by merging complementary information from images captured by different sensors. Among these algorithms, deep unfolding network (DUN)-based methods achieve promising performance but still suffer from two issues: they lack a degradation-resistant-oriented fusion model and struggle to adequately consider the structural properties of DUNs, making them vulnerable to degradation scenarios. In this paper, we propose a Degradation-Resistant Unfolding Network (DeRUN) for the HIF task to generate high-quality fused images even in degradation scenarios. Specifically, we introduce a novel HIF model for degradation resistance and derive its optimization procedures. Then, we incorporate the optimization unfolding process into the proposed DeRUN for end-to-end training. To ensure the robustness and efficiency of DeRUN, we employ a joint constraint strategy and a lightweight partial weight sharing module. To train DeRUN, we further propose a gradient direction-based entropy loss with powerful texture representation capacity. Extensive experiments show that DeRUN significantly outperforms existing methods on four HIF tasks, as well as downstream applications, with cheaper computational and memory costs. Chunming He, Kai Li 0012, Guoxia Xu, Yulun Zhang 0001, Runze Hu, Zhenhua Guo 0001, Xiu Li 0001 |
ICCV | 6 |
| 2023 | Weakly-Supervised Concealed Object Segmentation with SAM-based Pseudo Labeling and Multi-scale Feature GroupingabstractWeakly-Supervised Concealed Object Segmentation (WSCOS) aims to segment objects well blended with surrounding environments using sparsely-annotated data for model training. It remains a challenging task since (1) it is hard to distinguish concealed objects from the background due to the intrinsic similarity and (2) the sparsely-annotated training data only provide weak supervision for model learning. In this paper, we propose a new WSCOS method to address these two challenges. To tackle the intrinsic similarity challenge, we design a multi-scale feature grouping module that first groups features at different granularities and then aggregates these grouping results. By grouping similar features together, it encourages segmentation coherence, helping obtain complete segmentation results for both single and multiple-object images. For the weak supervision challenge, we utilize the recently-proposed vision foundation model, ``Segment Anything Model (SAM)'', and use the provided sparse annotations as prompts to generate segmentation masks, which are used to train the model. To alleviate the impact of low-quality segmentation masks, we further propose a series of strategies, including multi-augmentation result ensemble, entropy-based pixel-level weighting, and entropy-based image-level selection. These strategies help provide more reliable supervision to train the segmentation model. We verify the effectiveness of our method on various WSCOS tasks, and experiments demonstrate that our method achieves state-of-the-art performance on these tasks. Chunming He, Kai Li 0012, Yachao Zhang 0001, Guoxia Xu, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001 |
NeurIPS | 7 |
| 2023 | Benchmarking Large Language Models on CMExam - A comprehensive Chinese Medical Exam DatasetabstractRecent advancements in large language models (LLMs) have transformed the field of question answering (QA). However, evaluating LLMs in the medical field is challenging due to the lack of standardized and comprehensive datasets. To address this gap, we introduce CMExam, sourced from the Chinese National Medical Licensing Examination. CMExam consists of 60K+ multiple-choice questions for standardized and objective evaluations, as well as solution explanations for model reasoning evaluation in an open-ended manner. For in-depth analyses of LLMs, we invited medical professionals to label five additional question-wise annotations, including disease groups, clinical departments, medical disciplines, areas of competency, and question difficulty levels. Alongside the dataset, we further conducted thorough experiments with representative LLMs and QA algorithms on CMExam. The results show that GPT-4 had the best accuracy of 61.6% and a weighted F1 score of 0.617. These results highlight a great disparity when compared to human accuracy, which stood at 71.6%. For explanation tasks, while LLMs could generate relevant reasoning and demonstrate improved performance after finetuning, they fall short of a desired standard, indicating ample room for improvement. To the best of our knowledge, CMExam is the first Chinese medical exam dataset to provide comprehensive medical annotations. The experiments and findings of LLM evaluation also provide valuable insights into the challenges and potential solutions in developing Chinese medical QA systems and LLM evaluation pipelines. Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Helin Wang, Chenyu You, Zhenhua Guo 0001, Lei Zhu 0017, Michael Lingzhi Li |
NeurIPS | 9 |
| 2023 | Reconstruct face from features based on genetic algorithm using GAN generator as a distribution constraint
Xingbo Dong, Zhihui Miao, Zhe Jin 0001, Zhenhua Guo 0001, Andrew Beng Jin Teoh |
Comput. Secur. | 6 |
| 2023 | Inducing semantic hierarchy structure in empirical risk minimization with optimal transport measures
Wanqing Xie, Yubin Ge, Site Li, Zhenhua Guo 0001, Xiaofeng Liu 0001 |
Neurocomputing | 5 |
| 2023 | Self-paced resistance learning against overfitting on noisy labels
Xiaoshuang Shi, Zhenhua Guo 0001, Kang Li 0004, Yun Liang 0012, Xiaofeng Zhu 0001 |
Pattern Recognit. | 2 |
| 2023 | A simple and effective patch-Based method for frame-level face anti-spoofing
Shengjie Chen, Yujiu Yang 0001, Zhenhua Guo 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | DVS-Voltmeter: Stochastic Process-Based Event Simulator for Dynamic Vision Sensors
Songnan Lin, Zhenhua Guo 0001, Bihan Wen |
ECCV (7) | 3 |
| 2022 | AACP: Model Compression by Accurate and Automatic Channel PruningabstractChannel pruning is formulated as a neural architecture search (NAS) problem recently, which achieves impressive performance in model compression. However, prior arts only considered one kind of constraint (FLOPs, inference latency or model size) when pruning a neural network. This will lead to an unbalanced-pruning problem, where the FLOPs of a pruned network are under budget but the inference latency and model size are still unaffordable. Another challenge is that the supernet training process of typical NAS-based channel pruning methods is computationally expensive. To address these problems, we propose a novel Accurate and Automatic Channel Pruning (AACP) method. Firstly, we impose multiple constraints on channel pruning to address the unbalanced-pruning problem. To solve this complicated multi-objective problem, AACP proposes Improved Differential Evolution (IDE) algorithm which is more effective than typical evolutionary algorithms in searching for optimal architectures. Secondly, AACP proposes a Pruned Structure Accuracy Estimator (PSAE) which can estimate the performance of sub-networks without training a supernet and speeds up the performance estimation process. Our method achieves state-of-the-art performance on several benchmarks. On CIFAR10, our method reduces 65% FLOPs of ResNet110 with an improvement of 0.26% top-1 accuracy. On ImageNet, we reduce 42% FLOPs of ResNet50 with a small loss of 0.06% top-1 accuracy. The code is available at https://github.com/linlb11/AACP. Lanbo Lin, Shengjie Chen, Yujiu Yang 0001, Zhenhua Guo 0001 |
ICPR | 4 |
| 2022 | Augmenting Anchors by the Detector ItselfabstractUsually, it is difficult to determine the scale and aspect ratio of anchors for anchor-based object detection methods. Current state-of-the-art object detectors either determine anchor parameters according to objects' shape and scale in a dataset, or avoid this problem by utilizing anchor-free methods, however, the former scheme is dataset-specific and the latter methods could not get better performance than the former ones. In this paper, we propose a novel anchor augmentation method named AADI, which means Augmenting Anchors by the Detector Itself. AADI is not an anchor-free method, instead, it can convert the scale and aspect ratio of anchors from a continuous space to a discrete space, which greatly alleviates the problem of anchors' designation. Furthermore, AADI is a learning-based anchor augmentation method, but it does not add any parameters or hyper-parameters, which is beneficial for research and downstream tasks. Extensive experiments on COCO dataset demonstrate the effectiveness of AADI, specifically, AADI achieves significant performance boosts on many state-of-the-art object detectors (eg. at least +2.4 box AP on Faster R-CNN, +2.2 box AP on Mask R-CNN, and +0.9 box AP on Cascade Mask R-CNN). We hope that this simple and cost-efficient method can be widely used in object detection. Code and models are available at https://github.com/WanXiaopei/aadi. Xiaopei Wan, Guoqiu Li, Yujiu Yang 0001, Zhenhua Guo 0001 |
IJCAI | 4 |
| 2022 | A robust context attention network for human hand detection
Zhihuai Xie, Wentian Zhao, Zhenhua Guo 0001 |
Expert Syst. Appl. | 4 |
| 2022 | A novel 2D contactless fingerprint matching method
Hao Gui, Yujiu Yang 0001, Zhenhua Guo 0001 |
Neurocomputing | 5 |
| 2022 | Query2Set: Single-to-Multiple Partial Fingerprint Recognition Based on Attention MechanismabstractCurrently, fingerprint authentication systems in mobile devices, which have limited-size fingerprint sensors, are mainly based on partial fingerprint matching algorithms. To cover all areas of the finger, the system usually collects multiple partially overlapping partial fingerprints during the enrollment. Existing recognition methods either perform score-level fusion after single-to-single matching, or perform single-to-single matching after image-level mosaicking. However, these two-stage methods have the risk of discarding some real information or introducing some fake information. In this paper, we define this “query2set” task and propose a novel single-to-multiple partial fingerprint recognition method based on atttention mechanism. Our end-to-end deep model can adaptively extract and fuse appropriate features from a set of fingerprints for matching based on the input query fingerprint. Experiments indicate that our method outperforms several state-of-the-art single-to-single approaches and provides a new insight of fingerprint recognition on mobile devices. Shengjie Chen, Zhenhua Guo 0001, Xiu Li 0001, Dongliang Yang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | BioCanCrypto: An LDPC Coded Bio-Cryptosystem on Fingerprint Cancellable TemplateabstractBiometrics as a means of personal authentication has demonstrated strong viability in the past decade. However, directly deriving a unique cryptographic key from biometric data is a non-trivial task due to the fact that biometric data is usually noisy and presents large intra-class variations. Moreover, biometric data is permanently associated with the user, which leads to security and privacy issues. Cancellable biometrics and bio-cryptosystem are two main branches to address those issues, yet both approaches fall short in terms of accuracy performance, security, and privacy. In this paper, we propose a Bio-Crypto system on fingerprint Cancellable template (Bio-CanCrypto), which bridges cancellable biometrics and bio-cryptosystem to achieve a middle-ground for alleviating the limitations of both. Specifically, a cancellable transformation is applied on a fixed-length fingerprint feature vector to generate cancellable templates. Next, an LDPC coding mechanism is introduced into a reusable fuzzy extractor scheme and used to extract the stable cryptographic key from the generated cancellable templates. The proposed system can achieve both cancellability and reusability in one scheme. Experiments are conducted on a public fingerprint dataset, i.e., FVC2002. The results demonstrate that the proposed LDPC coded reusable fuzzy extractor is effective and promising. Xingbo Dong, Zhe Jin 0001, Leshan Zhao, Zhenhua Guo 0001 |
IJCB | 4 |
| 2021 | Adversarial Unsupervised Domain Adaptation with Conditional and Label Shift: Infer, Align and IterateabstractIn this work, we propose an adversarial unsupervised domain adaptation (UDA) method under inherent conditional and label shifts, in which we aim to align the distributions w.r.t. both p(x|y) and p(y). Since labels are inaccessible in a target domain, conventional adversarial UDA methods assume that p(y) is invariant across domains and rely on aligning p(x) as an alternative to the p(x|y) alignment. To address this, we provide a thorough theoretical and empirical analysis of the conventional adversarial UDA methods under both conditional and label shifts, and propose a novel and practical alternative optimization scheme for adversarial UDA. Specifically, we infer the marginal p(y) and align p(x|y) iteratively at the training stage, and precisely align the posterior p(y|x) at the testing stage. Our experimental results demonstrate its effectiveness on both classification and segmentation UDA and partial UDA. Xiaofeng Liu 0001, Zhenhua Guo 0001, Site Li, Fangxu Xing, Jane You, C.-C. Jay Kuo, Georges El Fakhri, Jonghye Woo |
ICCV | 2 |
| 2021 | Weakly Supervised Fingerprint Pore Extraction With Convolutional Neural NetworkabstractFingerprint recognition has been used for person identification for centuries, and fingerprint features are divided into three levels. The level 3 feature is the fingerprint pore, which can be used to improve the performance of the automatic fingerprint recognition performance and to prevent spoofing in high-resolution fingerprints. Therefore, the accurate extraction of fingerprint pores is quite important. With the development of convolutional neural networks (CNNs), researchers have made great progress in fingerprint feature extraction. However, these supervised-based methods require manually labelled pores to train the network, and labelling pores is very tedious and time consuming because there are hundreds of pores in one fingerprint. In this paper, we design a weakly supervised pore extraction method that avoids manual label processing and trains the network with a noisy label. This method can achieve results comparable with a supervised CNN-based method. Rongxiao Tang, Shuang Sun 0004, Feng Liu 0013, Zhenhua Guo 0001 |
ICIP | 4 |
| 2021 | A Robust Method with DropBlock for Face Anti-SpoofingabstractFace anti-spoofing is an important component for reliable face recognition. Previous deep learning methods usually exploit a large network structure and are prone to be overfitting when the face anti-spoofing database is a small scale. And in some occasions, multi-frame information or other auxiliary information is not available. In order to address these issues, we propose an end-to-end deep learning approach with DropBlock layer to distinguish between a fake face and a genuine face, which makes use of frame-only RGB image. In this paper, we evaluate the proposed approach with different settings on three popular databases, i.e., REPLAY-ATTACK, CASIA-FASD, and CASIA-SURF. The proposed approach can achieve competitive results on these databases. The experimental results show that our model is more robust and can learn more generalized features for face anti-spoofing. Zhuo Zhou, Zhenhua Guo 0001 |
IJCNN | 3 |
| 2021 | An Adaptive Iterative Inpainting Method with More Information ExplorationabstractThe CNN-based image inpainting methods have achieved promising performance because of its outstanding semantic understanding and reasoning potentialities. However, previous works could not get satisfied results in some situations because information is not fully explored. In this paper, we propose a new method by combining three innovative ideas. First, to increase the diversity of the semantic information obtained by the network in image synthesis, we propose a multiple hidden space perceptual (MHSP) loss, which extracts high-level features from multiple pre-trained autoencoders. Second, we adopt an adaptive iterative reasoning (AIR) stategy to reduce the calculations under small-hole circumstances while ensuring the performance in large-hole circumstances. Third, we find that color inconsistencies occasionally occurred in the final image merging process, so we add a novel interval maximum saturation (IMS) loss to the final loss function. Experiments on the benchmark datasets show our method performs favorably against state-of-the-art approaches. Code is made publicly available at: https://github.com/IC-LAB/adaptive_iterative_inpainting. Shengjie Chen, Zhenhua Guo 0001, Bo Yuan 0003 |
ACM Multimedia | 2 |
| 2021 | Surface and Internal Fingerprint Reconstruction From Optical Coherence Tomography Through Convolutional Neural NetworkabstractOptical coherence tomography (OCT), as a non-destructive and high-resolution fingerprint acquisition technology, is robust against poor skin conditions and resistant to spoof attacks. It measures fingertip information on and beneath skin as 3D volume data, containing the surface fingerprint, internal fingerprint and sweat glands. Various methods have been proposed to extract internal fingerprints, which ignore the inter-slice dependence and often require manually selected parameters. In this article, a modified U-Net that combines residual learning, bidirectional convolutional long short-term memory and hybrid dilated convolution (denoted as BCL-U Net) for OCT volume data segmentation and two fingerprint reconstruction approaches are proposed. To the best of our knowledge, it is the first time that simultaneous and automatic extraction is performed for surface fingerprint, internal fingerprint and sweat gland. The proposed BCL-U Net utilizes the spatial dependence in OCT volume data and deals with segmentation of objects with diverse sizes to achieve accurate extraction. Comparisons have been performed to demonstrate the advantages of the proposed method. A thorough evaluation of the recognition abilities of internal and surface fingerprints is conducted using a dataset significantly larger than previous studies. Four databases containing internal and surface fingerprints are generated from 1572 OCT volume data by the proposed method. The internal fingerprint matching experiment has achieved a lowest equal error rate (EER) of 0.95%. Mixed internal and surface fingerprint matching experiment is also performed and achieves an EER of 3.67%, verifying the consistency of the internal and surface fingerprints. The matching experiments for fingers under poor skin conditions show a 2.47% EER of internal fingerprints that is much lower than that of surface fingerprints, which proves the advantage of internal fingerprints and indicates the potential of the internal fingerprints to supplement or replace the surface fingerprints for some specific applications. Baojin Ding, Haixia Wang 0002, Peng Chen 0008, Yilong Zhang 0001, Zhenhua Guo 0001, Jianjiang Feng, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Loss-Based Attention for Interpreting Image-Level Prediction of Convolutional Neural NetworksabstractAlthough deep neural networks have achieved great success on numerous large-scale tasks, poor interpretability is still a notorious obstacle for practical applications. In this paper, we propose a novel and general attention mechanism, loss-based attention, upon which we modify deep neural networks to mine significant image patches for explaining which parts determine the image decision-making. This is inspired by the fact that some patches contain significant objects or their parts for image-level decision. Unlike previous attention mechanisms that adopt different layers and parameters to learn weights and image prediction, the proposed loss-based attention mechanism mines significant patches by utilizing the same parameters to learn patch weights and logits (class vectors), and image prediction simultaneously, so as to connect the attention mechanism with the loss function for boosting the patch precision and recall. Additionally, different from previous popular networks that utilize max-pooling or stride operations in convolutional layers without considering the spatial relationship of features, the modified deep architectures first remove them to preserve the spatial relationship of image patches and greatly reduce their dependencies, and then add two convolutional or capsule layers to extract their features. With the learned patch weights, the image-level decision of the modified deep architectures is the weighted sum on patches. Extensive experiments on large-scale benchmark databases demonstrate that the proposed architectures can obtain better or competitive performance to state-of-the-art baseline networks with better interpretability. The source codes are available on: https://github.com/xsshi2015/Loss-based-Attention-for-Interpreting-Image-level-Prediction-of-Convolutional-Neural-Networks. Xiaoshuang Shi, Fuyong Xing, Kaidi Xu, Pingjun Chen, Yun Liang 0012, Zhiyong Lu, Zhenhua Guo 0001 |
IEEE Trans. Image Process. | 7 |
| 2021 | A Scalable Optimization Mechanism for Pairwise Based Discrete HashingabstractMaintaining the pairwise relationship among originally high-dimensional data into a low-dimensional binary space is a popular strategy to learn binary codes. One simple and intuitive method is to utilize two identical code matrices produced by hash functions to approximate a pairwise real label matrix. However, the resulting quartic problem in term of hash functions is difficult to directly solve due to the non-convex and non-smooth nature of the objective. In this paper, unlike previous optimization methods using various relaxation strategies, we aim to directly solve the original quartic problem using a novel alternative optimization mechanism to linearize the quartic problem by introducing a linear regression model. Additionally, we find that gradually learning each batch of binary codes in a sequential mode, i.e. batch by batch, is greatly beneficial to the convergence of binary code learning. Based on this significant discovery and the proposed strategy, we introduce a scalable symmetric discrete hashing algorithm that gradually and smoothly updates each batch of binary codes. To further improve the smoothness, we also propose a greedy symmetric discrete hashing algorithm to update each bit of batch binary codes. Moreover, we extend the proposed optimization mechanism to solve the non-convex optimization problems for binary code learning in many other pairwise based hashing algorithms. Extensive experiments on benchmark single-label and multi-label databases demonstrate the superior performance of the proposed mechanism over recent state-of-the-art methods on two kinds of retrieval tasks: similarity and ranking order. The source codes are available on https://github.com/xsshi2015/Scalable-Pairwise-based-Discrete-Hashing. Xiaoshuang Shi, Fuyong Xing, Zizhao Zhang 0002, Manish Sapkota, Zhenhua Guo 0001, Lin Yang 0002 |
IEEE Trans. Image Process. | 5 |
| 2021 | A Novel Multicamera System for High-Speed Touchless Palm RecognitionabstractPalm-related biometrics have been widely studied for a long time, as the palm contains many distinctive patterns. However, most of the existing systems are designed to work within an ideal environment, such as in front of a unicolor background or in a large enclosure. Those preconditions can avoid influences of ambient light and hand distance change, but at the same time, they also limit the applications of palm recognition. In the work reported in this paper, we designed a novel red-green-blue and depth-based four-camera system that can capture the palm-related images separately in real time. The techniques of region-of-interest (ROI) location, ROI alignment, and light-source intensity optimization were studied. The ROI location method is modified to increase the robustness of hand gesture variation. Based on the depth information, we proposed the coordinate mapping and inclination rectification methods to obtain aligned ROI pairs. Using this device, we collected a video-based multimodal palm image database. After the parameter optimization and information fusion, the equal-error-rate of our approach on this database is lower than 0.47%. The recognition rate obtained from the support-vector-machine-based fusion is higher than 99.8%. The experimental results prove that the proposed system achieves advantages of anti-spoofing, high speed, high accuracy, and small size. David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001, Nan Luo |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | A Joint Super-Resolution and Deformable Registration Network for 3D Brain Images
Zhenhua Guo 0001 |
ICPR | 2 |
| 2020 | Face Anti-Spoofing Using Spatial Pyramid PoolingabstractFace recognition system is vulnerable to many kinds of presentation attacks, so how to effectively detect whether the image is from the real face is particularly important. At present, many deep learning-based anti-spoofing methods have been proposed. But these approaches have some limitations, for example, global average pooling (GAP) easily loses local information of faces, single-scale features easily ignore information differences in different scales, while a complex network is prone to be overfitting. In this paper, we propose a face anti-spoofing approach using spatial pyramid pooling (SPP). Firstly, we use ResNet-18 with a small amount of parameter as the basic model to avoid overfitting. Further, we use spatial pyramid pooling module in the single model to enhance local features while fusing multi-scale information. The effectiveness of the proposed method is evaluated on three databases, CASIA-FASD, Replay-Attack and CASIA-SURF. The experimental results show that the proposed approach can achieve state-of-the-art performance. Zhuo Zhou, Zhenhua Guo 0001 |
ICPR | 3 |
| 2020 | A Multi-Task Neural Network for Action Recognition with 3D Key-PointsabstractAction recognition and 3D human pose estimation are fundamental problems in computer vision and closely related areas. In this work, we propose a multi-task neural network for action recognition and 3D human pose estimation. Results of previous methods are usually error-prone especially when tested against the images taken in-the-wild, leading error results in action recognition. To solve this problem, we propose a principled approach to generate high quality 3D pose ground truth given any in-the-wild image with a person inside. We achieve this by first devising a novel stereo inspired neural network to directly map any 2D pose to high quality 3D counterpart. Based on the high-quality 3D labels, we carefully design the multi-task framework for action recognition and 3D human pose estimation. The proposed architecture can utilize shallow, deep features of images, and in-the-wild 3D human key-points to guide a more precise result. High quality 3D key-points can fully reflect morphological features of motions, thus boost the performance on action recognition. Experimental results demonstrate that 3D pose estimation leads to significantly higher performance on action recognition than separated learning. We also evaluate the generalization ability of our method both quantitatively and qualitatively. The proposed architecture performs favorably against the baseline 3D pose estimation methods. In addition, the reported results on Penn Action and NTU datasets demonstrate the effectiveness of our method on the action recognition task. Rongxiao Tang, Zhenhua Guo 0001 |
ICPR | 3 |
| 2020 | Anchor-Based Self-Ensembling for Semi-Supervised Deep Pairwise Hashing
Xiaoshuang Shi, Zhenhua Guo 0001, Fuyong Xing, Yun Liang 0012, Lin Yang 0002 |
Int. J. Comput. Vis. | 2 |
| 2020 | Pre-registration of translated/distorted fingerprints based on correlation and the orientation field
Zhenhua Guo 0001, Jane You |
Inf. Sci. | 2 |
| 2020 | Dependency-Aware Attention Control for Image Set-Based Face RecognitionabstractThis paper considers the problem of image set-based face verification and identification. Unlike traditional single sample (an image or a video) setting, this situation assumes the availability of a set of heterogeneous collection of orderless images and videos. The samples can be taken at different check points, different identity documents $etc$ . The importance of each image is usually considered either equal or based on a quality assessment of that image independent of other images and/or videos in that image set. How to model the relationship of orderless images within a set remains a challenge. We address this problem by formulating it as a Markov Decision Process (MDP) in a latent space. Specifically, we first propose a dependency-aware attention control (DAC) network, which uses actor-critic reinforcement learning for attention decision of each image to exploit the correlations among the unordered images. An off-policy experience replay is introduced to speed up the learning process. Moreover, the DAC is combined with a temporal model for videos using divide and conquer strategies. We also introduce a pose-guided representation (PGR) scheme that can further boost the performance at extreme poses. We propose a parameter-free PGR without the need for training as well as a novel metric learning-based PGR for pose alignment without the need for pose detection in testing stage. Extensive evaluations on IJB-A/B/C, YTF, Celebrity-1000 datasets demonstrate that our method outperforms many state-of-art approaches on the set-based as well as video-based face recognition databases. Xiaofeng Liu 0001, Zhenhua Guo 0001, Jane You, B. V. K. Vijaya Kumar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Permutation-Invariant Feature Restructuring for Correlation-Aware Image Set-Based RecognitionabstractWe consider the problem of comparing the similarity of image sets with variable-quantity, quality and un-ordered heterogeneous images. We use feature restructuring to exploit the correlations of both inner&inter-set images. Specifically, the residual self-attention can effectively restructure the features using the other features within a set to emphasize the discriminative images and eliminate the redundancy. Then, a sparse/collaborative learning-based dependency-guided representation scheme reconstructs the probe features conditional to the gallery features in order to adaptively align the two sets. This enables our framework to be compatible with both verification and open-set identification. We show that the parametric self-attention network and non-parametric dictionary learning can be trained end-to-end by a unified alternative optimization scheme, and that the full framework is permutation-invariant. In the numerical experiments we conducted, our method achieves top performance on competitive image set/video-based face recognition and person re-identification benchmarks. Xiaofeng Liu 0001, Zhenhua Guo 0001, Site Li, Ping Jia, Lingsheng Kong, Jane You, B. V. K. Vijaya Kumar |
ICCV | 2 |
| 2019 | Low-resolution palmprint image denoising by generative adversarial networks
Shengjie Chen, Zhenhua Guo 0001, Yushen Zuo |
Neurocomputing | 3 |
| 2019 | Structured orthogonal matching pursuit for feature selection
Xiaoshuang Shi, Fuyong Xing, Zhenhua Guo 0001, Hai Su, Fujun Liu, Lin Yang 0002 |
Neurocomputing | 3 |
| 2019 | Binary Filter for Fast Vessel Pattern Extraction
Shuang Sun 0004, Shidong Li, Zhenhua Guo 0001 |
Neural Process. Lett. | 3 |
| 2019 | A non-rigid registration method with application to distorted fingerprint matching
Zhenhua Guo 0001, Jane You |
Pattern Recognit. | 2 |
| 2019 | Joint learning for voice based disease detection
Kebin Wu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001 |
Pattern Recognit. | 4 |
| 2019 | Non-rigid medical image registration using image field in Demons algorithm
Zhenhua Guo 0001, Jane You |
Pattern Recognit. Lett. | 2 |
| 2019 | Similarity mapping for robust face recognition via a single training sample per person
Qin Li 0001, Xiaoshuang Shi, Zhenhua Guo 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | ECG-based personal recognition using a convolutional neural network
Zhibo Xiao, Zhenhua Guo 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | Learning Discriminant Direction Binary Palmprint DescriptorabstractPalmprint directions have been proved to be one of the most effective features for palmprint recognition. However, most existing direction-based palmprint descriptors are hand-craft designed and require strong prior knowledge. In this paper, we propose a discriminant direction binary code (DDBC) learning method for palmprint recognition. Specifically, for each palmprint image, we first calculate the convolutions of the direction-based templates and palmprint and form the informative convolution difference vectors by computing the convolution difference between the neighboring directions. Then, we propose a simple yet effective model to learn feature mapping functions that can project these convolution difference vectors into DDBCs. For all training samples: (1) the variance of the learned binary codes is maximized; (2) the intra-class distance of the binary codes is minimized; and (3) the inter-class distance of the binary codes is maximized. Finally, we cluster the block-wise histograms of DDBC forming the discriminant direction binary palmprint descriptor for palmprint recognition. The experimental results on four challenging contactless palmprint databases clearly demonstrate the effectiveness of the proposed method. Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Zhenhua Guo 0001, Jie Wen 0001, Wei Jia 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Robust Face Detector with Fully Convolutional Networks
Yingcheng Su, Xiaopei Wan, Zhenhua Guo 0001 |
PRCV (3) | 3 |
| 2018 | Palmprint gender classification by convolutional neural networkabstractPalmprint gender classification can revolutionise the performance of authentication systems, reduce searching space and speed up matching rate. However, to the best of their knowledge, there is no literature addressing this issue. The authors design a new convolutional neural network (CNN) structure, fine‐tuning Visual Geometry Group Network, up to 19 layers to achieve a 20‐layer network, for palmprint gender classification. Experimental results show that the proposed structure could achieve good performance for gender classification. They also investigate palmprint images with 15 different kinds of spectra. They empirically find that a palmprint image acquired by the Blue spectrum could achieve 89.2% correct classification and could be considered as a suitable spectrum for gender classification. The neural network is able to classify a 224 × 224 × 3‐pixel palmprint image in <23 ms, verifying that the proposed CNN is an effective real‐time solution. Zhihuai Xie, Zhenhua Guo 0001, Chengshan Qian |
IET Comput. Vis. | 2 |
| 2018 | Robust principal component analysis via optimal mean by joint ℓ2, 1 and Schatten p-norms minimization
Xiaoshuang Shi, Feiping Nie 0001, Zhihui Lai 0001, Zhenhua Guo 0001 |
Neurocomputing | 4 |
| 2018 | Learning acoustic features to detect Parkinson's disease
Kebin Wu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001 |
Neurocomputing | 4 |
| 2018 | Self-learning for face clustering
Xiaoshuang Shi, Zhenhua Guo 0001, Fuyong Xing, Jinzheng Cai, Lin Yang 0002 |
Pattern Recognit. | 2 |
| 2017 | Fingerprint pose estimation based on faster R-CNNabstractFingerprint pose estimation is one of the bottlenecks of indexing in large scale database. The existing methods of pose estimation are based on manually appointed features (e.g. special points, ridges, orientation filed). In this paper, we propose a method based on deep learning to achieve accurate pose estimation. Faster R-CNN is adopted to detect the center point and rough direction, followed by intra-class and inter-class combination to calculate the precise direction. Extensive experiments on NIST-14 show that (1) the predicted poses are close to manual annotations even when the fingerprints are incomplete or noisy, (2) the estimated poses for matching fingerprint pairs are very consistent and (3) by registering fingerprints using the estimated pose, the accuracy of a state-of-the-art fingerprint indexing system is further improved. Jiahong Ouyang, Jianjiang Feng, Jiwen Lu, Zhenhua Guo 0001, Jie Zhou 0001 |
IJCB | 4 |
| 2017 | License Plate Detection Using Deep Cascaded Convolutional Neural Networks in Complex Scenes
Zhenhua Guo 0001 |
ICONIP (2) | 3 |
| 2017 | Offline Signature Verification Using Local Features and Decision TreesabstractThe most difficult problem of offline signature verification (SV) is that a signature is merely a static image missing a lot of the dynamic information associated with it. In this paper, three separate pseudo-dynamic features based on the gray level: gradient based local binary pattern (GLBP), statistical features of gray level co-occurrence matrix (SGLCM), simplified histogram of oriented gradients (SHOG) are proposed for writer-independent offline SV. These gray-level features can convey both texture information and the relative structural relationship of signature strokes. In addition, our experiments prove that the proposed features contain complementary information. Using random forests (RFs) as classifier, a fusion of the proposed features could achieve 7.42% and 0.08% average error rate (AER) for GPDS-253 and CEDAR datasets, respectively, which show the effectiveness of the proposed system. The implication of this paper is that part dynamic information could be extracted from a static gray level image. Zhenhua Guo 0001, Zhenyin Fan, Youbin Chen |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2017 | Active learning via local structure reconstruction
Qin Li 0001, Xiaoshuang Shi, Linfei Zhou, Zhifeng Bao, Zhenhua Guo 0001 |
Pattern Recognit. Lett. | 5 |
| 2017 | Color-Guided Depth Recovery via Joint Local Structural and Nonlocal Low-Rank RegularizationabstractHigh-quality depth recovery from RGB-D data has received increasingly more attention in recent years due to their wide applications from depth-based image rendering to three-dimensional imaging and video. Sharp contrast between high-quality color images and low-quality depth maps presents severe challenges to the development of color-guided depth recovery techniques. Previous works have emphasized either locally varying characteristics of color-depth dependence or nonlocal similarities around the discontinuities of the scene geometry. Therefore, it is desirable to exploit both local and nonlocal structural constraints for optimizing the performance of color-guided depth recovery. In this work, we propose a unified variational approach via joint local and nonlocal regularization. The local regularization term consists of two complementary parts-one characterizing the color-depth dependence in the gradient domain and the other in the spatial domain; nonlocal regularization involves a low-rank constraint suitable for large-scale depth discontinuities. Extensive experimental results are reported to show that our approach outperforms several existing state-of-the-art depth recovery methods on both synthetic and real-world data sets. Weisheng Dong, Guangming Shi, Xin Li 0005, Kefan Peng, Jinjian Wu, Zhenhua Guo 0001 |
IEEE Trans. Multim. | 6 |
| 2017 | Door Knob Hand Recognition SystemabstractBiometric applications have been used globally in everyday life. However, conventional biometrics is created and optimized for high-security scenarios. Being used in daily life by ordinary untrained people is a new challenge. Facing this challenge, designing a biometric system with prior constraints of ergonomics, we propose ergonomic biometrics design model, which attains the physiological factors, the psychological factors, and the conventional security characteristics. With this model, a novel hand-based biometric system, door knob hand recognition system (DKHRS), is proposed. DKHRS has the identical appearance of a conventional door knob, which is an optimum solution in both physiological factors and psychological factors. In this system, a hand image is captured by door knob imaging scheme, which is a tailored omnivision imaging structure and is optimized for this predetermined door knob appearance. Then features are extracted by local Gabor binary pattern histogram sequence method and classified by projective dictionary pair learning. In the experiment on a large data set including 12 000 images from 200 people, the proposed system achieves competitive recognition performance comparing with conventional biometrics like face and fingerprint recognition systems, with an equal error rate of 0.091%. This paper shows that a biometric system could be built with a reliable recognition performance under the ergonomic constraints. Xiaofeng Qu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Large Margin Coupled Mapping for Low Resolution Face Recognition
Zhenhua Guo 0001, Xiu Li 0001, Youbin Chen |
PRICAI | 2 |
| 2016 | Face recognition using part-based dense sampling local features
Zhenhua Guo 0001, Youbin Chen |
Neurocomputing | 3 |
| 2016 | Two-Dimensional Whitening Reconstruction for Enhancing Robustness of Principal Component AnalysisabstractPrincipal component analysis (PCA) is widely applied in various areas, one of the typical applications is in face. Many versions of PCA have been developed for face recognition. However, most of these approaches are sensitive to grossly corrupted entries in a 2D matrix representing a face image. In this paper, we try to reduce the influence of grosses like variations in lighting, facial expressions and occlusions to improve the robustness of PCA. In order to achieve this goal, we present a simple but effective unsupervised preprocessing method, two-dimensional whitening reconstruction (TWR), which includes two stages: 1) A whitening process on a 2D face image matrix rather than a concatenated 1D vector; 2) 2D face image matrix reconstruction. TWR reduces the pixel redundancy of the internal image, meanwhile maintains important intrinsic features. In this way, negative effects introduced by gross-like variations are greatly reduced. Furthermore, the face image with TWR preprocessing could be approximate to a Gaussian signal, on which PCA is more effective. Experiments on benchmark face databases demonstrate that the proposed method could significantly improve the robustness of PCA methods on classification and clustering, especially for the faces with severe illumination changes. Xiaoshuang Shi, Zhenhua Guo 0001, Feiping Nie 0001, Lin Yang 0002, Jane You, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Dynamic background estimation and complementary learning for pixel-wise foreground/background segmentation
Weifeng Ge, Zhenhua Guo 0001, Yuhan Dong, Youbin Chen |
Pattern Recognit. | 2 |
| 2016 | Robust Texture Image Representation by Scale Selective Local Binary PatternsabstractLocal binary pattern (LBP) has successfully been used in computer vision and pattern recognition applications, such as texture recognition. It could effectively address grayscale and rotation variation. However, it failed to get desirable performance for texture classification with scale transformation. In this paper, a new method based on dominant LBP in scale space is proposed to address scale variation for texture classification. First, a scale space of a texture image is derived by a Gaussian filter. Then, a histogram of pre-learned dominant LBPs is built for each image in the scale space. Finally, for each pattern, the maximal frequency among different scales is considered as the scale invariant feature. Extensive experiments on five public texture databases (University of Illinois at Urbana-Champaign, Columbia Utrecht Database, Kungliga Tekniska Högskolan-Textures under varying Illumination, Pose and Scale, University of Maryland, and Amsterdam Library of Textures) validate the efficiency of the proposed feature extraction scheme. Coupled with the nearest subspace classifier, the proposed method could yield competitive results, which are 99.36%, 99.51%, 99.39%, 99.46%, and 99.71% for UIUC, CUReT, KTH-TIPS, UMD, and ALOT, respectively. Meanwhile, the proposed method inherits simple and efficient merits of LBP, for example, it could extract scale-robust feature for a 200×200 image within 0.24 s, which is applicable for many real-time applications. Zhenhua Guo 0001, Xingzheng Wang, Jie Zhou 0001, Jane You |
IEEE Trans. Image Process. | 1 |
| 2015 | Within-class penalty based multi-class support vector machineabstractSupport vector machine (SVM) is a widely used maximum margin classifier, but the classification performance is largely affected by outliers. In this paper, we propose a novel multi-class SVM method to reduce the influence of outliers on the classification performance. Our proposed method includes an efficient optimization model via considering the within-class scatter and an optimization way. Specifically, the method is based on one assumption that penalizing the within-class scatter can reduce the number of misclassified outliers near the decision boundary, because data points of each class could be compacted by the within-class penalty. Experiments on benchmark databases demonstrate the effectiveness of the assumption and the proposed method. Xiaoshuang Shi, Zhenhua Guo 0001, Yujiu Yang 0001, Lin Yang 0002 |
ICIP | 2 |
| 2015 | Online personal verification by palmvein image through palmprint-like and palmvein information
Qin Li 0001, Xiu Li 0001, Zhenhua Guo 0001, Jane You |
Neurocomputing | 3 |
| 2015 | A Framework of Joint Graph Embedding and Sparse Regression for Dimensionality ReductionabstractOver the past few decades, a large number of algorithms have been developed for dimensionality reduction. Despite the different motivations of these algorithms, they can be interpreted by a common framework known as graph embedding. In order to explore the significant features of data, some sparse regression algorithms have been proposed based on graph embedding. However, the problem is that these algorithms include two separate steps: (1) embedding learning and (2) sparse regression. Thus their performance is largely determined by the effectiveness of the constructed graph. In this paper, we present a framework by combining the objective functions of graph embedding and sparse regression so that embedding learning and sparse regression can be jointly implemented and optimized, instead of simply using the graph spectral for sparse regression. By the proposed framework, supervised, semisupervised, and unsupervised learning algorithms could be unified. Furthermore, we analyze two situations of the optimization problem for the proposed framework. By adopting an ℓ2,1-norm regularization for the proposed framework, it can perform feature selection and subspace learning simultaneously. Experiments on seven standard databases demonstrate that joint graph embedding and sparse regression method can significantly improve the recognition performance and consistently outperform the sparse regression method. Xiaoshuang Shi, Zhenhua Guo 0001, Zhihui Lai 0001, Yujiu Yang 0001, Zhifeng Bao, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Background Subtraction with Dynamic Noise Sampling and Complementary LearningabstractBackground subtraction is a popular technique used in accurate foreground extraction with a stationary background. Since most outdoor surveillance videos are taken in complex environments, their "stationary" backgrounds change in some unknown patterns, which make the perfect foreground extraction very difficult. Based on visual background extractor (ViBe) scheme, in this paper we propose a new background subtraction algorithm which includes two innovative mechanisms and several other improved technique tricks. The paper inherits and develops background modeling based on pixel sample values, and use dynamic noise sampling and complementary learning to overcome the pixel-wise background model's intrinsic shortcomings. Besides, the algorithm works on the quantitative analysis without any estimation of the probability density function (pdf). Hence, it takes relatively low computational cost. Extensive experiments on a popular public dataset show that the proposed method has much better precision than ViBe, and could get the best precision and the highest average ranking compared with 27 state-of-the-art algorithms presented on the change detection website. Weifeng Ge, Yuhan Dong, Zhenhua Guo 0001, Youbin Chen |
ICPR | 3 |
| 2014 | A tentative comparison on CDN and NDNabstractWith the pretty prompt growth in Internet content, future Internet is emerging as the main usage shifting from traditional host-to-host model to content dissemination model, e.g. video makes up more than half of Internet traffic. ISPs, content providers and other third parties have widely deployed content delivery networks (CDNs) to support digital content distribution. Though CDN is an ad-hoc solution to the content dissemination problem, there are still big challenges, such as complicated control plane. By contrast, as a wholly new designed network architecture, named data networking (NDN) incorporates content delivery function in its network layer, its stateful routing and forwarding plane can effectively detect and adapt to the dynamic and ever-changing Internet. In this paper, we try to explore the similarities and differences between CDN and NDN. Hence, we evaluate the distribution efficiency, network security and protocol overhead between CDN and NDN. Especially in the implementation phase, we conduct their testbeds separately with the same topology to derive their performance of content delivery. Finally, summarizing our main results, we gather that: 1) NDN has its own advantage on lots of aspects, including security, scalability and quality of service (QoS); 2) NDN make full use of surrounding resources and is more adaptive to the dynamic and ever-changing Internet; 3) though CDN is a commercial and mature architecture, in some scenarios, NDN can perform better than CDN under the same topology and caching storage. In a word, NDN is practical to play an even greater role in the evolution of the Internet based on the massive distribution and retrieval in the future. Ge Ma, Zhen Chen 0001, Zhenhua Guo 0001, Yixin Jiang, Xiaobin Guo |
SMC | 4 |
| 2014 | BreadZip: a combination of network traffic data and bitmap index encoding algorithmabstractNowadays, rapid evolution of computers and mobile devices has caused the explosive increase in network traffic. So it becomes more and more necessary to archive network traffic for analyzing network events and a lot of emerging applications. Compression is fundamental for traffic archival solution to save the storage space, and indexing is effective to accelerate search queries for archive of traffic data. In this paper, we propose BreadZip (blocks row-reordering and adaptive index zip), a combination of initial traffic data and index compression. BreadZip has three main advantages. 1) to improve compressing efficiency and reduce memory footprint, traffic data is reordered in sequence and divided into fixed-size blocks; 2) to accelerate queries, an improved bitmap indexes with smaller volume than traditional will be introduced; 3) to save space, both traffic blocks and bitmap indexes are compressed in different simple run-length encoding methods respectively. Finally, our empirical results on network traffic from CAIDA (Cooperative Association for Internet Data Analysis) show that our solution can significantly reduce the volume of traffic data, while simultaneously preserving the ability to perform selectively queries with response times in seconds. Ge Ma, Zhenhua Guo 0001, Xiu Li 0001, Zhen Chen 0001, Yixin Jiang, Xiaobin Guo |
SMC | 2 |
| 2014 | Face recognition by sparse discriminant analysis via joint L2, 1-norm minimization
Xiaoshuang Shi, Yujiu Yang 0001, Zhenhua Guo 0001, Zhihui Lai 0001 |
Pattern Recognit. | 3 |
| 2013 | Facial image medical analysis system using quantitative chromatic feature
Xingzheng Wang, Bob Zhang 0001, Zhenhua Guo 0001, David Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2013 | Is local dominant orientation necessary for the classification of rotation invariant texture?
Zhenhua Guo 0001, Qin Li 0001, Lin Zhang 0014, Jane You, David Zhang 0001, Wenhuang Liu |
Neurocomputing | 1 |
| 2013 | Distal-Interphalangeal-Crease-Based User Authentication SystemabstractTouchless-based fingerprint recognition technology is thought to be an alternative to touch-based systems to solve problems of hygienic, latent fingerprints, and maintenance. However, there are few studies about touchless fingerprint recognition systems due to the lack of a large database and the intrinsic drawback of low ridge-valley contrast of touchless fingerprint images. This paper proposes an end-to-end solution for user authentication systems based on touchless fingerprint images in which a multiview strategy is adopted to collect images and the robust fingerprint feature of touchless image is extracted for matching with high recognition accuracy. More specifically, a touchless multiview fingerprint capture device is designed to generate three views of raw images followed by preprocessing steps including region of interest (ROI) extraction and image correction. The distal interphalangeal crease (DIP)-based feature is then extracted and matched to recognize the human's identity in which part selection is introduced to improve matching efficiency. Experiments are conducted on two sessions of touchless multiview fingerprint image database with 541 fingers acquired about two weeks apart. An EER of ~ 1.7% can be achieved by using the proposed DIP-based feature, which is much better than touchless fingerprint recognition by using scale invariant feature transformation (SIFT) and minutiae features. The given fusion results show that it is effective to combine the DIP-based feature, minutiae, and SIFT feature for touchless fingerprint recognition systems. The EER is as low as ~ 0.5%. Feng Liu 0013, David Zhang 0001, Zhenhua Guo 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Local directional derivative pattern for rotation invariant texture classification
Zhenhua Guo 0001, Qin Li 0001, Jane You, David Zhang 0001, Wenhuang Liu |
Neural Comput. Appl. | 1 |
| 2012 | Phase congruency induced local features for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Zhenhua Guo 0001 |
Pattern Recognit. | 4 |
| 2012 | Feature Band Selection for Online Multispectral Palmprint RecognitionabstractA palmprint is a unique and reliable biometric feature with high usability. In the past decades, many palmprint recognition systems have been successfully developed. However, most of the previous work used the white light as the illumination source, and the recognition accuracy and anti-spoof capability is limited. Recently, multispectral imaging has attracted considerable research attention as it can acquire more discriminative information in a short time. One crucial step in developing online multispectral palmprint systems is how to determine the optimal number of spectral bands and select the most representative bands to build the system. This paper presents a study on feature band selection by analyzing hyperspectral palmprint data (520-1050 nm). Our experimental results showed that three spectral bands could provide most of the discriminate information of a palmprint. This finding could be used as the guidance for designing new online multispectral palmprint systems. Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wenhuang Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Texture Image Classification Using Complex Texton
Zhenhua Guo 0001, Qin Li 0001, Lin Zhang 0014, Jane You, Wenhuang Liu |
ICIC (2) | 1 |
| 2011 | Online joint palmprint and palmvein verification
David Zhang 0001, Zhenhua Guo 0001, Guangming Lu 0002, Lei Zhang 0006, Wangmeng Zuo |
Expert Syst. Appl. | 2 |
| 2011 | Empirical study of light source selection for palmprint recognition
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wangmeng Zuo, Guangming Lu 0002 |
Pattern Recognit. Lett. | 1 |
| 2010 | The multiscale competitive code via sparse representation for palmprint verificationabstractPalm lines are the most important features for palmprint recognition. They are best considered as typical multiscale features, where the principal lines can be represented at a larger scale while the wrinkles at a smaller scale. Motivated by the success of coding-based palmprint recognition methods, this paper investigates a compact representation of multiscale palm line orientation features, and proposes a novel method called the sparse multiscale competitive code (SMCC). The SMCC method first defines a filter bank of second derivatives of Gaussians with different orientations and scales, and then uses the l1-norm sparse coding to obtain a robust estimation of the multiscale orientation field. Finally, a generalized competitive code is used to encode the dominant orientation. Experimental results show that the SMCC achieves higher verification accuracy than state-of-the-art palmprint recognition methods, yet uses a smaller template size than other multiscale methods. Wangmeng Zuo, Zhouchen Lin, Zhenhua Guo 0001, David Zhang 0001 |
CVPR | 3 |
| 2010 | Hierarchical multiscale LBP for face and palmprint recognitionabstractLocal binary pattern (LBP), fast and simple for implementation, has shown its superiority in face and palmprint recognition. To extract representative features, “uniform” LBP was proposed and its effectiveness has been validated. However, all “non-uniform” patterns are clustered into one pattern, so a lot of useful information is lost. In this study, the authors propose to build a hierarchical multiscale LBP histogram for an image. The useful information of “non-uniform” patterns at large scale is dug out from its counterpart of small scale. The main advantage of the proposed scheme is that it can fully utilize LBP information while it does not need any training step, which may be sensitive to training samples. Experiments on one public face database and one palmprint database show the effectiveness of the proposed method. Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001, Xuanqin Mou |
ICIP | 1 |
| 2010 | Rotation invariant texture classification using adaptive LBP with directional statistical featuresabstractLocal Binary Pattern (LBP) has been widely used in texture classification because of its simplicity and computational efficiency. Traditional LBP codes the sign of the local difference and uses the histogram of the binary code to model the given image. However, the directional statistical information is ignored in LBP. In this paper, some directional statistical features, specifically the mean and standard deviation of the local absolute difference are extracted and used to improve the LBP classification efficiency. In addition, the least square estimation is used to adaptively minimize the local difference for more stable directional statistical features, and we call this scheme the adaptive LBP (ALBP). By coupling the directional statistical features with ALBP, a new rotation invariant texture classification method is presented. Experiments on a large texture database show that the proposed texture feature extraction and classification scheme could significantly improve the classification accuracy of LBP. Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001 |
ICIP | 1 |
| 2010 | Monogenic-LBP: A new approach for rotation invariant texture classificationabstractAnalysis of two-dimensional textures has many potential applications in computer vision. In this paper, we investigate the problem of rotation invariant texture classification, and propose a novel texture feature extractor, namely Monogenic-LBP (M-LBP). M-LBP integrates the traditional Local Binary Pattern (LBP) operator with the other two rotation invariant measures: the local phase and the local surface type computed by the 1st-order and 2nd-order Riesz transforms, respectively. The classification is based on the image's histogram of M-LBP responses. Extensive experiments conducted on the CUReT database demonstrate the overall superiority of M-LBP over the other state-of-the-art methods evaluated. Lin Zhang 0014, Lei Zhang 0006, Zhenhua Guo 0001, David Zhang 0001 |
ICIP | 3 |
| 2010 | Feature Band Selection for Multispectral Palmprint RecognitionabstractPalm print is a unique and reliable biometric characteristic with high usability. Many palm print recognition algorithms and systems have been successfully developed in the past decades. Most of the previous works use the white light sources for illumination. Recently, it has been attracting much research attention on developing new biometric systems with both high accuracy and high anti-spoof capability. Multispectral palm print imaging and recognition can be a potential solution to such systems because it can acquire more discriminative information for personal identity recognition. One crucial step in developing such systems is how to determine the minimal number of spectral bands and select the most representative bands to build the multispectral imaging system. This paper presents preliminary studies on feature band selection by analyzing hyper spectral palm print data (420nm~1100nm). Our experiments showed that 2 spectral bands at 700nm and 960nm could provide most discriminate information of palm print. This finding could be used as the guidance for designing multispectral palm print systems in the future. Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001 |
ICPR | 1 |
| 2010 | A unified distance measurement for orientation coding in palmprint verification
Zhenhua Guo 0001, Wangmeng Zuo, Lei Zhang 0006, David Zhang 0001 |
Neurocomputing | 1 |
| 2010 | Rotation invariant texture classification using LBP variance (LBPV) with global matching
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001 |
Pattern Recognit. | 1 |
| 2010 | A Completed Modeling of Local Binary Pattern Operator for Texture ClassificationabstractIn this correspondence, a completed modeling of the local binary pattern (LBP) operator is proposed and an associated completed LBP (CLBP) scheme is developed for texture classification. A local region is represented by its center pixel and a local difference sign-magnitude transform (LDSMT). The center pixels represent the image gray level and they are converted into a binary code, namely CLBP-Center (CLBP_C), by global thresholding. LDSMT decomposes the image local differences into two complementary components: the signs and the magnitudes, and two operators, namely CLBP-Sign (CLBP_S) and CLBP-Magnitude (CLBP_M), are proposed to code them. The traditional LBP is equivalent to the CLBP_S part of CLBP, and we show that CLBP_S preserves more information of the local structure than CLBP_M, which explains why the simple LBP operator can extract the texture features reasonably well. By combining CLBP_S, CLBP_M, and CLBP_C features into joint or hybrid distributions, significant improvement can be made for rotation invariant texture classification. Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2009 | Is White Light the Best Illumination for Palmprint Recognition?
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006 |
CAIP | 1 |
| 2009 | Rotation Invariant Texture Classification Using Binary Filter Response Pattern (BFRP)
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001 |
CAIP | 1 |
| 2009 | Palmprint verification using consistent orientation codingabstractDeveloping accurate and robust palmprint verification algorithms is one of the key issues in automatic palmprint recognition systems. Recently, orientation based coding algorithms, such as Competitive Code (CompCode) and Orthogonal Line Ordinal Features (OLOF), have been proposed and have been attracting much research attention. Such algorithms could achieve high accuracy with high feature matching speed for real time implementation. By investigating the relationship between these two different coding schemes, we propose in this paper a feature-level fusion scheme for palmprint verification. Only the stable features which are consistent between the two codes are extracted for matching. The experimental results on the public palmprint database show that the proposed fusion code could achieve at least 14% EER (Equal Error Rate) reduction compared with either of the original codes. Zhenhua Guo 0001, Wangmeng Zuo, Lei Zhang 0006, David Zhang 0001 |
ICIP | 1 |
| 2009 | Palmprint verification using binary orientation co-occurrence vector
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wangmeng Zuo |
Pattern Recognit. Lett. | 1 |
| 2007 | Palmprint Verification using Complex Wavelet TransformabstractPalmprint is a unique and reliable biometric characteristic with high usability. With the increasing demand of automatic palmprint authentication systems, the development of accurate and robust palmprint verification algorithms has been attracting a lot of interests. The relative translation, rotation and distortion between two palmprint images will introduce much error in palmprint matching. However, an accurate registration of palmprint images is too time-consuming. In this paper, we propose a modified complex wavelet structural similarity index (CW-SSIM) to compute the matching score and hence identify the input palmprint. Since CW-SSIM is robust to translation, small rotation and distortion, a fast rough alignment of palmprint images is sufficient. CW-SSIM is also insensitive to luminance and contrast changes. Our experimental results show that the proposed scheme outperforms the state-of-the-art methods by achieving a higher genuine acceptance rate and a lower false acceptance rate simultaneously. Lei Zhang 0006, Zhenhua Guo 0001, Zhou Wang 0001, David Zhang 0001 |
ICIP (2) | 2 |