David Zhang 0001

dblp:z/DavidZhang · also David Dapeng Zhang · DBLP profile ↗
← Back
543ranked-venue papers
17as first author
112since 2021 · last 2026
0000-0002-5027-5286ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 297 · 9 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 182 · 1 first-author · 36 since 2021Human-computer interaction and ubiquitous computing · 46 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 18 · 1 first-author · 3 since 2021Security and privacy · 12 · 4 since 2021Systems, architecture and hardware · 4 · 1 first-authorComputer networks · 3Software engineering, systems software and programming languages · 2Theory of computation · 2 · 1 first-author
YearPublicationVenuePosition
2026 Robust stochastic configuration networks ensemble with fuzzy granular computing
Chenglong Zhang 0001, Zihao Liao, Shicheng Dai, David Zhang 0001, Haiwei Hou
Expert Syst. Appl.5
2026 Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm
Fuxiang Huang, Xiaowei Fu, Shiyu Ye, Wen Li 0001, Xinbo Gao 0001, David Zhang 0001, Lei Zhang 0038
Int. J. Comput. Vis.7
2026 Weakly Supervised Salient Object Detection with Text Supervision
Zhihao Wu 0002, Jie Wen 0001, LinLin Shen, Xiaopeng Fan 0001, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
Int. J. Comput. Vis.7
2026 Multimodal artificial intelligence for disease diagnosis: Advances, applications, and challenges
Shaozhe Wang, Fan Zhang 0070, Yu Liu 0023, Huafeng Li 0001, Junyu Dong, David Zhang 0001
Pattern Recognit.8
2026 Exploiting Hu invariant moments and deep features for image retrieval
Guanghai Liu 0001, David Zhang 0001
Pattern Recognit.3
2026 Pseudocylindrical Convolutions for Learned Omnidirectional Image Compression
abstract
Equirectangular projection (ERP) is a convenient form to store omnidirectional images, but it is neither equal-area nor conformal, creating challenges for subsequent visual communication. When used for image compression, ERP amplifies sampling density and deforms objects near the poles, hindering perceptually optimal bit allocation. Here, we present one of the earliest endeavors to apply deep neural networks to omnidirectional image compression. We first propose parametric pseudocylindrical representations that generalize common pseudocylindrical map projections. A tractable greedy algorithm is introduced to identify (sub-)optimal representation configurations, guided by a proxy objective for rate-distortion performance. We then develop pseudocylindrical convolutions, which can be efficiently implemented by standard convolutions with “pseudocylindrical padding.” To demonstrate the utility of the proposed pseudocylindrical representations and convolutions, we implement an end-to-end omnidirectional image compression method, consisting of an analysis transform, a uniform quantizer, a synthesis transform, and an entropy model. Experiments show that our optimized method achieves consistently better rate-distortion performance compared to the state-of-the-art.
Mu Li 0005, Kede Ma, Jinxing Li 0003, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 LearnMat: Semantic-Aware Self-Supervision Fine-Grained Visual Recognition
abstract
Self-supervised learning has shown potential in fine-grained visual recognition (FGVR). However, existing self-supervised learning methods are often susceptible to irrelevant patterns during training and lack the ability to capture the critical subtle differences in FGVR, leading to suboptimal performance. Moreover, existing approaches focus primarily on uni-modal visual concepts. Despite the emergence of powerful vision-language models (VLMs) in various high-level vision tasks, their potential in self-supervised FGVR remains largely unexplored. To this end, we propose a novel self-supervised learning (LearnMat) framework, that effectively filters out irrelevant feature interference and extracts more important and subtle discriminative features during training. Specifically, LearnMat consists of two key modules: the semantic awareness module (SAM) and the insight extraction module (IEM). In the SAM, we introduce a novel vision-language-grounded semantic distillation strategy using a corpus of generic, category-agnostic textual attributes, that injects explicit semantic constraints into self-supervised training and improves robustness to background interference. Complementarily, the IEM exploits gradient-based signals from the input image to highlight subtle differences and localize key discriminative regions, mitigating inter-class similarity and intra-class variation, and enhancing fine-grained discrimination. Extensive experiments across multiple popular FGVR datasets show that LearnMat significantly outperforms recent state-of-the-art methods, highlighting its marked effectiveness. Our code is avaliable at https://github.com/Heng-CHY/LearnMat.
ShuaiHeng Li, Fan Zhang 0070, Yangyang Shu, Guanbin Li, Junyu Dong, Lingqiao Liu, David Zhang 0001
IEEE Trans. Image Process.8
2026 Diagnosing and Improving Vector-Quantization-Based Blind Image Restoration
abstract
Vector-Quantization (VQ) based discrete generative models are widely used to learn powerful high-quality (HQ) priors for blind image restoration (BIR). In this paper, we diagnose the side-effects of discrete VQ process essential to VQ-based BIR methods: 1) confining the representation capacity of HQ codebook, 2) being error-prone for code index prediction on low-quality (LQ) images, and 3) under-valuing the importance of input LQ image. These motivate us to learn continuous feature representation of HQ codebook for better restoration performance than using discrete VQ process. To further improve the restoration fidelity, we propose a new Self-in-Cross-Attention (SinCA) module to augment the HQ codebook with the feature of input LQ image, and perform cross-attention between LQ feature and input-augmented codebook. By this way, our SinCA leverages the input LQ image to enhance the representation of codebook for restoration fidelity. Experiments on four typical VQ-based BIR methods demonstrate that, by replacing the VQ process with a transformer using our SinCA, they achieve better quantitative and qualitative performance on blind image super-resolution and blind face restoration. The code and pre-trained models are publicly released at https://github.com/lhy-85/SinCA.
Zengyou Wang, Xiantong Zhen, Ran Gu, David Zhang 0001, Jun Xu 0019
IEEE Trans. Image Process.6
2026 Soft Supervision-Guided Spatial-Temporal Refinement Network for Video-Based Visible-Infrared Person Re-Identification
abstract
Thanks to automatic switch between visible and infrared modes, person re-identification (Re-ID) in 24-hour has been possible through cross-modal retrieval. Instead of exploiting still images, video-based cross-modal person Re-ID is studied in this paper. Specifically, a large-scale dataset 'HITSZ-PVCM' is first collected, consisting of as many as 1,681 identities and 839,632 frames. Generally, videos contain much richer pedestrian appearances. However, most existing works only generate temporal representations by whole frames, inevitably losing fine-grained details. Furthermore, training a network by metric losses (e.g., center loss) is a common strategy, while such point-to-point constraints are too strong and limit model generalization due to existing diversity among intra-class samples. Here, we propose a Soft Supervision guided Spatial-Temporal Refinement (S3TR) network to tackle these problems. Specifically, S3TR refines each frame guided by a coarse temporal feature, so that more discriminative features are extracted and transformed to a sequential representation. Followed by a global-local mutual learning module, the modality gap is then erased without losing fine-grained details. Furthermore, we propose a novel soft-clustering center loss to measure intra-/inter-class similarity/dissimilarity in a group-to-group way, efficiently improving model generalization. To the best of our knowledge, HITSZ-PVCM is the largest dataset and S3TR achieves superior performances compared with state-of-the-arts.
Jinxing Li 0003, Chuhao Zhou, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.8
2026 A Cosine Network for Image Super-Resolution
abstract
Deep convolutional neural networks can use hierarchical information to progressively extract structural information to recover high-quality images. However, preserving the effectiveness of the obtained structural information is important in image super-resolution. In this paper, we propose a cosine network for image super-resolution (CSRNet) by improving a network architecture and optimizing the training strategy. To extract complementary homologous structural information, odd and even heterogeneous blocks are designed to enlarge the architectural differences and improve the performance of image super-resolution. Combining linear and non-linear structural information can overcome the drawback of homologous information and enhance the robustness of the obtained structural information in image super-resolution. Taking into account the local minimum of gradient descent, a cosine annealing mechanism is used to optimize the training procedure by performing warm restarts and adjusting the learning rate. Experimental results illustrate that the proposed CSRNet is competitive with state-of-the-art methods in image super-resolution.
Chunwei Tian, Bob Zhang 0001, Zhiwu Li 0001, C. L. Philip Chen, David Zhang 0001
IEEE Trans. Image Process.6
2026 Progressive Fusion of Multi-Scale Mamba Context and Local Detail Priors for Infrared Small Target Detection
abstract
Infrared Small Target Detection (IRSTD) requires strong target-level detection capability, which depends on effective modeling of long-range global dependencies. This demand has driven the transition from CNN-based approaches to Transformer-based architectures. Although Transformers improve global context modeling, their high computational cost limits practical deployment. Recent advances in Mamba enable efficient long-range dependency modeling with reduced complexity, offering a promising alternative that alleviates the efficiency limitations of Transformers while preserving target-level detection performance. However, Mamba is not inherently tailored for IRSTD, as it lacks explicit mechanisms for capturing fine-grained local details and modeling background variations across multiple spatial scales. To address these limitations, we propose MCFNet, an encoder-decoder framework that integrates Mamba to enhance target-level detection performance with moderate computational cost. MCFNet introduces a Detail-Capturable Convolution Block to strengthen local detail perception and a Multi-scale Contextual Mamba Block to improve background modeling across different scales. While the resulting dual-branch design enhances both global semantics and local details, it also introduces challenges in feature fusion. To this end, a Feature Fusion Decoding Module is further proposed to enable effective collaboration between global and local representations. Extensive experiments on multiple public IRSTD benchmark datasets demonstrate that MCFNet consistently outperforms existing methods in both pixel-level and target-level metrics, achieving higher detection accuracy with reduced false alarms. The code of our model is available at: https://github.com/Fihven/MCFNet.
Xiangjun Zhu, Fei-wei Qin, Changmiao Wang, Jin Fan 0003, Fei Lin 0006, Jing Bai 0004, Chenglong Zhang 0001, David Zhang 0001
IEEE Trans. Image Process.8
2026 Data-Driven Study on Why Wrist Pulse Can Assess Health Conditions
abstract
Pulse diagnosis (PD) is a traditional diagnostic technique in which physicians palpate the wrist pulse at the radial artery to evaluate an individual's health status. However, the scientific validity of PD has been questioned due to the lack of quantitative and objective evidence. In particular, variations in wrist pulse waves across different health conditions remain underexplored, and the correlations between pulse wave characteristics and physiological states are not yet well understood. To fill these gaps, we design a data-driven analysis architecture that integrates a quantitative analysis method, medical knowledge, and association establishment. The quantitative analysis method systematically quantifies the pulse waves, presenting the pulse characteristics, typical pulse wave, feature distributions, and recognition performance. Experiments on 900 samples demonstrate the differences in pulse waveforms from various diseases. Further, feature visualization and classification present distinct variations in pulse across different health conditions. By integrating the quantitative results with medical knowledge, we establish the correlations between pulse waves and health status, interpreting how physiological changes affect pulse waveform morphology. This study is the first to quantitatively explore why wrist pulse can evaluate health conditions, thereby addressing longstanding skepticism surrounding PD. The findings of this paper significantly enhance PD's credibility.
Chaoxun Guo, Shicheng Dai, Chenglong Zhang 0001, Baoyuan Wu, David Zhang 0001
IEEE J. Biomed. Health Informatics6
2025 Leukocyte classification using relative-relationship-guided contrastive learning
Qinghua Lin, Jiawei Wu 0001, Taotao Lai, Rongteng Wu, David Zhang 0001
Expert Syst. Appl.6
2025 Weakly Supervised Salient Object Detection With Oversize Bounding Box
Zhihao Wu 0002, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
Int. J. Comput. Vis.4
2025 Crucial and irreplaceable 3D features for facial beauty analysis
Yahan Sun, Tianhao Peng 0003, David Zhang 0001
Knowl. Based Syst.4
2025 Meta-distribution-based ensemble sampler for imbalanced semi-supervised learning
Zhihan Ning, Chaoxun Guo, David Zhang 0001
Pattern Recognit.3
2025 Multi-Modal Cross-Subject Emotion Feature Alignment and Recognition With EEG and Eye Movements
abstract
Multi-modal emotion recognition has attracted much attention in human-computer interaction, because it provides complementary information for the recognition model. However, the distribution drift among subjects and the heterogeneity of different modalities pose challenges to multi-modal emotion recognition, thereby limiting its practical application. Most of the current multi-modal emotion recognition methods are difficult to suppress above uncertainties in fusion. In this paper, we propose a cross-subject multi-modal emotion recognition framework, which jointly learns subject-independent representation and common feature between EEG and eye movements. First, we design the dynamic adversarial domain adaptation for cross-subject distribution alignment, dynamically selecting source domains in training. Second, we simultaneously capture intra-modal and inter-modal emotion-related features by both self-attention and cross-attention mechanisms, thus obtaining the robust and complementary representation of emotional information. Then, two contrastive loss functions are imposed on above network to further reduce inter-modal heterogeneity, and mine higher-order semantic similarity between synchronously collected multi-modal data. Finally, we used the output of the softmax layer as the predicted value. The experimental results on several multi-modal emotion datasets with EEG and eye movements demonstrate that our method is significantly superior to the state-of-the-art emotion recognition approaches.
Qi Zhu 0001, Lunke Fei, Chuhang Zheng, Wei Shao 0005, David Zhang 0001, Daoqiang Zhang
IEEE Trans. Affect. Comput.6
2025 SIAVC: Semi-Supervised Framework for Industrial Accident Video Classification
abstract
Semi-supervised learning suffers from the imbalance of labeled and unlabeled training data in the video surveillance scenario. In this paper, we propose a new semi-supervised learning method called SIAVC for industrial accident video classification. Specifically, we design a video augmentation module called the Super Augmentation Block (SAB). SAB adds Gaussian noise and randomly masks video frames according to historical loss on the unlabeled data for model optimization. Then, we propose a Video Cross-set Augmentation Module (VCAM) to generate diverse pseudo-label samples from the high-confidence unlabeled samples, which alleviates the mismatch of sampling experience and provides high-quality training data. Additionally, we construct a new industrial accident surveillance video dataset with frame-level annotation, namely ECA9, to evaluate our proposed method. Compared with the state-of-the-art semi-supervised learning based methods, SIAVC demonstrates outstanding video classification performance, achieving 88.76% and 89.13% accuracy on ECA9 and Fire Detection datasets, respectively. The source code and the constructed dataset ECA9 will be released inhttps://github.com/AlchemyEmperor/SIAVC.
Qinghua Lin, Haoyi Fan, Tiesong Zhao, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Hierarchical Fuzzy Stochastic Configuration Network Based on Canonical Correlation Analysis for Multiview Pulse Signal Fusion
abstract
Pulse-based diagnostic (PBD) is a crucial noninvasive approach for disease diagnosis which acquires multi-view pulse signals through various sensors for disease classification. Takagi-Sugeno-Kang (TSK) fuzzy systems and multi-view methodologies are extensively employed for multi-view pulse signal fusion in PBD systems. However, problem still exists in integrating multi-view learning mechanism and classical fuzzy systems. The most critical challenges involve effectively utilizing multi-view information while improving computational efficiency and ensuring universal approximation property. Consequently, we proposed a hierarchical fuzzy stochastic configuration network based on canonical correlation analysis (HFSCN-CCA) method to enhance the performance of TSK fuzzy systems on multi-view datasets. Specially, we replaced the deep canonical correlation analysis (DCCA) framework with data-independent incremental hierarchical stochastic configuration network based on canonical correlation analysis (HSCN-CCA) method to increase the computational efficiency and ensure universal approximation property while maximizing the correlation. Additionally, we optimized the traditional TSK fuzzy system with a novel joint membership function to capture sample-specific multi-view information for higher representation capability. Furthermore, we conducted multiple experiments on diverse disease datasets to validate the performance of HFSCN-CCA and successfully demonstrated the complementarity of the two signals in multiple disease classification tasks.
Shicheng Dai, Chenglong Zhang 0001, Chaoxun Guo, Baoyuan Wu, David Zhang 0001
IEEE Trans. Fuzzy Syst.5
2025 Deep Ensemble Stochastic Configuration Network via Graph Intuitionistic Fuzzy for Depression Recognition
abstract
Depression is an affective disorder that poses a serious threat to both mental and physical health. Utilizing fuzzy -based neural network models for the identification and screening of depression can facilitate early intervention and treatment. Intuitionistic fuzzy stochastic configuration networks (IFSCNs) utilize a cost-sensitive learning framework, which enhances generalization performance for solving binary classification problems. However, IFSCNs ignore the impact of the relative neighborhood density of imbalanced samples with outliers. To learn more discriminant information from class imbalance depression recognition task, in this paper, we propose a novel deep ensemble stochastic configuration network via graph intuitionistic fuzzy, termed as DeSCN-GIF. Specifically, we first use graph-based intuitionistic fuzzy method to determine the membership and non-membership functions through the relative neighborhood density of imbalanced samples; moreover, an incremental self-ensemble deep stochastic configuration framework is presented to learn multi-level discriminative features, in which graph intuitionistic fuzzy cost-sensitive least squares loss function and weighted supervision mechanisms are applied to determine the parameters of DeSCN-GIF. Experimental results on Chinese syllable-based imbalanced depression voice datasets show that DeSCN-GIF has better binary classification performance compared to other learning models such as IFSCN, DSCN, SCN, GE-IFRVFL-CIL, IFRVFL, EDRVFL, DRVFL, RVFL, and DIFL-TSVM.
Chenglong Zhang 0001, Dawei Cheng, Jiankai Xue, Feng Wu 0005, David Zhang 0001
IEEE Trans. Fuzzy Syst.5
2025 Disentangled Representation Learning for Robust Brainprint Recognition
abstract
Electroencephalography (EEG) biometrics draws increasing attention in high-security requirements due to its advantages of anti-spoofing, live traits, and non-duplicated. However, existing EEG datasets, which rely on external stimuli or task-specific instructions for data collection, often intertwine identity-related information with biases such as emotional states, cognitive tasks, and disease markers. Besides, EEG signals are time-varying, while identity information within EEG signals is relatively fixed, which poses challenges for extracting identity features from EEG to perform accurate person identification. This high correlation hampers the promotion of brainprint recognition in real-life applications. In this paper, we propose a disentangled representation learning based identity recognition framework, which disentangles the EEG signal into intrinsic identity-related information and biased identity-invariant information, thus enhancing the performance of EEG biometrics. First, two parallel encoders are used to extract intrinsic identity-relevant and bias identity-irrelevant factors, respectively, and each encoder consists of a temporal filter module and a novel spatial-temporal attention module. Then, we further refine the disentanglement process through a correlation-driven loss that minimizes factor similarity across spatial-temporal and global representational domains. Adversarial training and reconstruction regularization are introduced to facilitate the identity and biased representations to be independent and complementary to each other. Additionally, we extend supervised contrastive learning to the component level, minimizing cross-component similarity and encouraging each component to independently reflect its unique information, thereby improving the disentanglement efficacy. Our proposed framework achieves state-of-the-art performance on diverse datasets encompassing emotional, motor imagery, and pathological conditions, demonstrating the robustness and effectiveness of our proposed brainprint identity recognition model.
Chuhang Zheng, Qi Zhu 0001, Lunke Fei, Shengrong Li, Xiangping Bryce Zhai, David Zhang 0001, Daoqiang Zhang
IEEE Trans. Inf. Forensics Secur.6
2025 A Perception CNN for Facial Expression Recognition
abstract
Convolutional neural networks (CNNs) can automatically learn data patterns to express face images for facial expression recognition (FER). However, they may ignore effect of facial segmentation of FER. In this paper, we propose a perception CNN for FER as well as PCNN. Firstly, PCNN can use five parallel networks to simultaneously learn local facial features based on eyes, cheeks and mouth to realize the sensitive capture of the subtle changes in FER. Secondly, we utilize a multi-domain interaction mechanism to register and fuse between local sense organ features and global facial structural features to better express face images for FER. Finally, we design a two-phase loss function to restrict accuracy of obtained sense information and reconstructed face images to guarantee performance of obtained PCNN in FER. Experimental results show that our PCNN achieves superior results on several lab and real-world FER benchmarks: CK+, JAFFE, FER2013, FERPlus, RAF-DB and Occlusion and Pose Variant Dataset. Its code is available at https://github.com/hellloxiaotian/PCNN.
Chunwei Tian, Jingyuan Xie, Lingjun Li, Wangmeng Zuo, Yanning Zhang 0001, David Zhang 0001
IEEE Trans. Image Process.6
2025 Caption Assisted Multimodal Large Language Model for Video Moment Retrieval
abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant potential across various multimodal tasks, including retrieval, summarization, and reasoning. However, it remains a substantial challenge for MLLMs to understand and precisely retrieve specific moments from a video, which require fine-grained spatial and temporal understanding of a video. To overcome this, we propose the Caption Assisted MLLM from Coarse to finE (CALCE), a novel two-stage framework designed for enhanced moment retrieval. Our pipeline begins with a first stage where captions extracted from the audio are utilized to assist the MLLM to provide a robust foundation for precise moment retrieval. To efficiently manage memory consumption from this additional data, a clustering algorithm is applied to the sparsely sampled video frames, categorizing them into key frames and non-key frames. The second stage focuses on recalling missed moments and achieving more fine-grained moment boundaries by adopting a higher sampling rate. In this process, predictions from the first stage cast votes for their correlated densely sampled frames, thereby filtering out less relevant frames. By repeating the process of the first stage with these selected frames, CALCE progressively retrieves video moments from coarse to precise. Experiments on QVHighlights and Charades-STA demonstrate the effectiveness of CALCE, which outperforms existing state-of-the-art methods. The code is available at https://github.com/tjhd1475/CALCE.
Peiyu Xie, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.5
2025 3DFACENet: 3D Facial Attractiveness Computation and Enhancement Network
abstract
The development of facial editing, virtual makeup, AR/VR technologies and 3D games applications underscore the need for advanced 3D facial attractiveness research. However, due to the lack of 3D beauty face data and the complexity of handling 3D face data, 3D facial aesthetics research remains largely unexplored. To fill this gap, we propose 3DFACENet, an innovative system designed for the computation and enhancement of 3D facial attractiveness. Our approach employs a 3D facial reconstruction encoder to generate encoded vectors from images and a render module to obtain 3D face models. To minimize computational load, we innovatively propose an attractiveness computation module which leverages 3D shape and texture coefficients rather than 3D mesh models to access facial attractiveness, achieving state-of-the-art results. To balance aesthetic enhancement and identity preservation, we design a controllable beautification decoder. For the first time, we introduce the concept of "attractive centers", demonstrating that an individual's distance to these centers is significantly negatively correlated with their beauty scores. Our beautification decoder edits 3D facial coefficients towards these centers, achieving a significant and controllable enhancement in facial attractiveness. Extensive experiments on the SCUT-FBP5500 and MEBeauty dataset validate the effectiveness and feasibility of 3DFACENet.
Tianhao Peng 0003, Mu Li 0005, Baoyuan Wu, David Zhang 0001
IEEE Trans. Image Process.5
2025 Focus Affinity Perception and Super-Resolution Embedding for Multifocus Image Fusion
abstract
Despite the fact that there is a remarkable achievement on multifocus image fusion, most of the existing methods only generate a low-resolution image if the given source images suffer from low resolution. Obviously, a naive strategy is to independently conduct image fusion and image super-resolution. However, this two-step approach would inevitably introduce and enlarge artifacts in the final result if the result from the first step meets artifacts. To address this problem, in this article, we propose a novel method to simultaneously achieve image fusion and super-resolution in one framework, avoiding step-by-step processing of fusion and super-resolution. Since a small receptive field can discriminate the focusing characteristics of pixels in detailed regions, while a large receptive field is more robust to pixels in smooth regions, a subnetwork is first proposed to compute the affinity of features under different types of receptive fields, efficiently increasing the discriminability of focused pixels. Simultaneously, in order to prevent from distortion, a gradient embedding-based super-resolution subnetwork is also proposed, in which the features from the shallow layer, the deep layer, and the gradient map are jointly taken into account, allowing us to get an upsampled image with high resolution. Compared with the existing methods, which implemented fusion and super-resolution independently, our proposed method directly achieves these two tasks in a parallel way, avoiding artifacts caused by the inferior output of image fusion or super-resolution. Experiments conducted on the real-world dataset substantiate the superiority of our proposed method compared with state of the arts.
Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023, Guangming Lu 0002, Yong Xu 0001, Zhengtao Yu 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.8
2025 To Combat Multiclass Imbalanced Problems by Aggregating Evolutionary Hierarchical Classifiers
abstract
Real-world datasets are often imbalanced, posing frequent challenges to canonical machine learning algorithms that assume a balanced class distribution. Moreover, the imbalance problem becomes more complicated when the dataset is multiclass. Although many approaches have been presented for imbalanced learning (IL), research on the multiclass imbalanced problem is relatively limited and deficient. To alleviate these issues, we propose a forest of evolutionary hierarchical classifiers (FEHC) method for multiclass IL (MCIL). FEHC can be seen as a classifier fusion framework with a forest structure, and it aggregates several evolutionary hierarchical multiclassifiers (EHMCs) to reduce generalization error. Specifically, a multichromosome genetic algorithm (MCGA) is designed to simultaneously select (sub)optimal features, classifiers, and hierarchical structures when generating these EHMCs. The MCGA adopts a dynamic weighting module to learn difficult classes and promote the diversity of FEHC. We also present the "stratified underbagging" (SUB) strategy to address class imbalance and the "soft tree traversal" (STT) strategy to make FEHC converge faster and better. We thoroughly evaluate the proposed algorithm using 14 multiclass imbalanced datasets with various properties. Compared with popular and state-of-the-art approaches, FEHC obtains better performance under different evaluation metrics. Codes have been made publicly available on GitHub.https://github.com/CUHKSZ-NING/FEHCClassifier.
Zhihan Ning, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Exploiting Meta-Learned Confidences for Imbalanced Multilabel Learning
abstract
Multilabel learning deals with datasets where each sample is associated with multiple labels. It is commonly assumed that label correlations should be well exploited to build an effective multilabel classifier. Moreover, the class imbalance problem occurs in many multilabel datasets and should be tackled to reduce the classification bias. While many multilabel learning methods have been proposed, research on imbalanced multilabel learning (IMLL) is relatively deficient. To address these issues, we exploit the value of meta-learned confidences, i.e., the prediction confidences iteratively updated over the out-of-bag samples, for IMLL. First, such meta-confidences can be fused to the original feature space to learn high-order label correlations. Second, meta-confidences can be used to calibrate the prediction results to alleviate class imbalance. Motivated by these, we propose an ensemble learning method named meta-confidence ensemble (MCE) for IMLL. Specifically, MCE iteratively makes bootstrap replicates of the multilabel training set, leverages the out-of-bag samples to generate meta-confidences, and fuses them to the original feature space to learn label correlations. A sparse projection method is presented to avoid overfitting and improve the ensemble diversity. Finally, the prediction result of an unseen sample is determined by the calibrated plurality vote of MCE's base classifiers. Extensive experiments demonstrated the effectiveness and superiority of MCE for IMLL. Codes have been made publicly available at https://github.com/ CUHKSZ-NING/MCEClassifier.
Zhihan Ning, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Federated Cross-Incremental Self-Supervised Learning for Medical Image Segmentation
abstract
Federated cross learning has shown impressive performance in medical image segmentation. However, it encounters the catastrophic forgetting issue caused by data heterogeneity across different clients and is particularly pronounced when simultaneously facing pixelwise label deficiency problem. In this article, we propose a novel federated cross-incremental self-supervised learning method, coined FedCSL, which not only can enable any client in the federation incrementally yet effectively learn from others without inducing knowledge forgetting or requiring massive labeled samples, but also preserve maximum data privacy. Specifically, to overcome the catastrophic forgetting issue, a novel cross-incremental collaborative distillation (CCD) mechanism is proposed, which distills explicit knowledge learned from previous clients to subsequent clients based on secure multiparty computation (MPC). Besides, an effective retrospect mechanism is designed to rearrange the training sequence of clients per round, further releasing the power of CCD by enforcing interclient knowledge propagation. In addition, to alleviate the need of large-scale densely annotated pretraining medical datasets, we also propose a two-stage training framework, in which federated cross-incremental self-supervised pretraining paradigm first extracts robust yet general image-level patterns across multi-institutional data silos via a novel round-robin distributed masked image modeling (MIM) pipeline; then, the resulting visual concepts, e.g., semantics, are transferred to the federated cross-incremental supervised fine-tuning paradigm, favoring various cross-silo medical image segmentation tasks. The experimental results on public datasets demonstrate the effectiveness of the proposed method as well as the consistently superior performance of our method over most state-of-the-art methods quantitatively and qualitatively.
Fan Zhang 0070, Chun-Mei Feng 0001, Binglu Wang, Shanshan Wang 0002, Junyu Dong, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.8
2025 Robust Decorrelated Stochastic Configuration Networks Ensemble via Weighted Negative Correlation Learning
abstract
Stochastic configuration network (SCN) is a kind of incremental random neural network that assigns input weights and biases through data-dependent supervisory mechanism. However, the robustness of SCN is significantly reduced when processing the data disturbed by outliers. Aiming at improve the noisy data regression performance of SCN, this article presents a novel robust decorrelated SCNs ensemble model (RDSCNE). Such a robust decorrelated ensemble framework adopts weighted negative correlation learning (WNCL) and a robust regularization technique, which can guarantee the generalization performance for noisy data processing. Specifically, we first present a WNCL framework based on kernel density estimation (KDE) to build SCNs ensemble model, so that the negative effects of noise can be suppressed through KDE to calculate penalty weights of each training sample for the computation of ensemble weights. Meanwhile,l1norm loss function combined withl2regularization technique is employed as the objective function of base components. This approach is designed to process outliers with sparse characteristics and alleviate the over-fitting phenomenon. Then, augmented Lagrange multiplier (ALM) method is used to calculate the objective function. Experimental results over some regression datasets with Gaussian outliers demonstrate that the proposed RDSCNE model has better robustness than the various SCN variants.
Chenglong Zhang 0001, Chaoxun Guo, Shifei Ding, Feng Wu 0005, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2024 ISFB-GAN: Interpretable semantic face beautification with generative adversarial network
Tianhao Peng 0003, Mu Li 0005, Fangmei Chen, Yong Xu 0001, Yahan Sun, David Zhang 0001
Expert Syst. Appl.7
2024 DSLSM: Dual-kernel-induced statistic level set model for image segmentation
Fan Zhang 0070, Xiaojun Duan, Binglu Wang, Huafeng Li 0001, Junyu Dong, David Zhang 0001
Expert Syst. Appl.8
2024 Context-aware graph embedding with gate and attention for session-based recommendation
Junlong Chi, Peilin Hong, Guangming Lu 0002, David Zhang 0001, Bingzhi Chen
Neurocomputing5
2024 Greedy deep stochastic configuration networks ensemble with boosting negative correlation learning
Chenglong Zhang 0001, Yang Wang 0028, David Zhang 0001
Inf. Sci.3
2024 Dual low-rank structure embedding for robust visual information processing
Jianhang Zhou, Hengmin Zhang, Shuyi Li 0003, Bob Zhang 0001, Leyuan Fang, David Zhang 0001
Knowl. Based Syst.6
2024 Contrastive feature decomposition for single image layer separation
Xin Feng 0005, Haobo Ji, Wenjie Pei, Guangming Lu 0002, David Zhang 0001
Neural Comput. Appl.6
2024 Exploiting sublimated deep features for image retrieval
Guanghai Liu 0001, Jing-Yu Yang 0001, David Zhang 0001
Pattern Recognit.4
2024 Cross co-teaching for semi-supervised medical image segmentation
Fan Zhang 0070, Jinjiang Wang, Huafeng Li 0001, Junyu Dong, David Zhang 0001
Pattern Recognit.8
2024 AMGNet: Aligned Multilevel Gabor Convolution Network for Palmprint Recognition
abstract
Palmprint recognition has seen significant advancements and garnered considerable attention recently. However, deep learning methods have yet to effectively incorporate insights from traditional approaches to extract palmprint-specific features. Moreover, intra-class spatial variation problems, which degrade the recognition performance, have not been adequately addressed. To tackle these limitations, this study proposes an Aligned Multilevel Gabor Convolution Network (AMGNet) to identify the informative and salient aspects of the palmprints. The network unifies a multilevel Gabor feature fusion branch with a spatial alignment branch, enabling the joint mining of aligned multilevel features specific to palmprints. Within the feature fusion branch, we incorporate two specialized Gabor convolution modules: one targets the principal lines of the palm, while the other focuses on the wrinkles, augmenting the discriminative power of the acquired features. To enhance the model’s robustness against within-class variations, we design a spatial alignment branch that specifically enables the rectification of palmprints’ spatial positions. In conjunction with this, we introduce a novel direction-based CosAngle loss function to facilitate geometric alignment among samples from same palms while spatially distancing those from different palms. Furthermore, we construct a palmprint database consisting of 3, 000 palms from 1, 500 individuals to explore large-scale population potential. Extensive experimental results on six benchmark datasets demonstrate that our proposed method outperforms other popular approaches in palmprint recognition tasks.
Wei Jia 0001, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 U²-Former: Nested U-Shaped Transformer for Image Restoration via Multi-View Contrastive Learning
abstract
While Transformer has achieved remarkable performance in various high-level vision tasks, it is still challenging to exploit the full potential of Transformer in image restoration. The crux lies in the limited depth of applying Transformer in the typical encoder-decoder framework for image restoration, resulting from heavy self-attention computation load and inefficient communications across different depth (scales) of layers. In this paper, we present a deep and effective Transformer-based network for image restoration, termed as U2-Former, which is able to employ self-attention of Transformer as the core operation for feature learning to perform image restoration in a deep encoding and decoding space. Specifically, it leverages the nested U-shaped structure to facilitate the interactions across different layers with different scales of feature maps. Furthermore, we optimize the computational efficiency for the basic Transformer block by introducing a simple yet effective feature-filtering mechanism to compress the token representation. Apart from the typical supervision ways for image restoration, our U2-Former also performs multi-view contrastive learning, which constructs positive pairs in various aspects, to learn noise-sensitive but content-irrelevant features and further decouple the noise component from the background image. Extensive experiments on various image restoration tasks, including reflection removal, rain streak removal and dehazing respectively, demonstrate the effectiveness of the proposed U2-Former.
Xin Feng 0005, Haobo Ji, Wenjie Pei, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Perceptive Self-Supervised Learning Network for Noisy Image Watermark Removal
abstract
Popular methods usually use a degradation model in a supervised way to learn a watermark removal model. However, it is true that reference images are difficult to obtain in the real world, as well as collected images by cameras suffer from noise. To overcome these drawbacks, we propose a perceptive self-supervised learning network for noisy image watermark removal (PSLNet) in this paper. PSLNet depends on a parallel network to remove noise and watermarks. The upper network uses task decomposition ideas to remove noise and watermarks in sequence. The lower network utilizes the degradation model idea to simultaneously remove noise and watermarks. Specifically, mentioned paired watermark images are obtained in a self-supervised way, and paired noisy images (i.e., noisy and reference images) are obtained in a supervised way. To enhance the clarity of obtained images, interacting two sub-networks and fusing obtained clean images are used to improve the effects of image watermark removal in terms of structural information and pixel enhancement. Taking into texture information account, a mixed loss uses obtained images and features to achieve a robust model of noisy image watermark removal. Comprehensive experiments show that our proposed method is very effective in comparison with popular convolutional neural networks (CNNs) for noisy image watermark removal. Codes can be obtained at https://github.com/hellloxiaotian/PSLNet.
Chunwei Tian, Menghua Zheng, Bo Li 0004, Yanning Zhang 0001, Shichao Zhang 0001, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Self-Supervised Multi-Scale Cropping and Simple Masked Attentive Predicting for Lung CT-Scan Anomaly Detection
abstract
Anomaly detection has been widely explored by training an out-of-distribution detector with only normal data for medical images. However, detecting local and subtle irregularities without prior knowledge of anomaly types brings challenges for lung CT-scan image anomaly detection. In this paper, we propose a self-supervised framework for learning representations of lung CT-scan images via both multi-scale cropping and simple masked attentive predicting, which is capable of constructing a powerful out-of-distribution detector. Firstly, we propose CropMixPaste, a self-supervised augmentation task for generating density shadow-like anomalies that encourage the model to detect local irregularities of lung CT-scan images. Then, we propose a self-supervised reconstruction block, named simple masked attentive predicting block (SMAPB), to better refine local features by predicting masked context information. Finally, the learned representations by self-supervised tasks are used to build an out-of-distribution detector. The results on real lung CT-scan datasets demonstrate the effectiveness and superiority of our proposed method compared with state-of-the-art methods.
Wei Li 0227, Guanghai Liu 0001, Haoyi Fan, David Zhang 0001
IEEE Trans. Medical Imaging5
2024 Multicontrast MRI Super-Resolution via Transformer-Empowered Multiscale Contextual Matching and Aggregation
abstract
Magnetic resonance imaging (MRI) possesses the unique versatility to acquire images under a diverse array of distinct tissue contrasts, which makes multicontrast super-resolution (SR) techniques possible and needful. Compared with single-contrast MRI SR, multicontrast SR is expected to produce higher quality images by exploiting a variety of complementary information embedded in different imaging contrasts. However, existing approaches still have two shortcomings: 1) most of them are convolution-based methods and, hence, weak in capturing long-range dependencies, which are essential for MR images with complicated anatomical patterns and 2) they ignore to make full use of the multicontrast features at different scales and lack effective modules to match and aggregate these features for faithful SR. To address these issues, we develop a novel multicontrast MRI SR network via transformer-empowered multiscale feature matching and aggregation, dubbed McMRSR$^{++}$. First, we tame transformers to model long-range dependencies in both reference and target images at different scales. Then, a novel multiscale feature matching and aggregation method is proposed to transfer corresponding contexts from reference features at different scales to the target features and interactively aggregate them Furthermore, a texture-preserving branch and a contrastive constraint are incorporated into our framework for enhancing the textural details in the SR images. Experimental results on both public and clinical in vivo datasets show that McMRSR$^{++}$outperforms state-of-the-art methods under peak signal to noise ratio (PSNR), structure similarity index measure (SSIM), and root mean square error (RMSE) metrics significantly. Visual results demonstrate the superiority of our method in restoring structures, demonstrating its great potential to improve scan efficiency in clinical practice.
Chengyan Wang, Qi Dou 0001, David Zhang 0001, Harry Qin
IEEE Trans. Neural Networks Learn. Syst.6
2024 A Heterogeneous Group CNN for Image Super-Resolution
abstract
Convolutional neural networks (CNNs) have obtained remarkable performance via deep architectures. However, these CNNs often achieve poor robustness for image super-resolution (SR) under complex scenes. In this article, we present a heterogeneous group SR CNN (HGSRCNN) via leveraging structure information of different types to obtain a high-quality image. Specifically, each heterogeneous group block (HGB) of HGSRCNN uses a heterogeneous architecture containing a symmetric group convolutional block and a complementary convolutional block in a parallel way to enhance the internal and external relations of different channels for facilitating richer low-frequency structure information of different types. To prevent the appearance of obtained redundant features, a refinement block (RB) with signal enhancements in a serial way is designed to filter useless information. To prevent the loss of original information, a multilevel enhancement mechanism guides a CNN to achieve a symmetric architecture for promoting expressive ability of HGSRCNN. Besides, a parallel upsampling mechanism is developed to train a blind SR model. Extensive experiments illustrate that the proposed HGSRCNN has obtained excellent SR performance in terms of both quantitative and qualitative analysis. Codes can be accessed at https://github.com/hellloxiaotian/HGSRCNN.
Chunwei Tian, Yanning Zhang 0001, Wangmeng Zuo, Chia-Wen Lin, David Zhang 0001, Yixuan Yuan
IEEE Trans. Neural Networks Learn. Syst.5
2024 Enhanced Spatial Feature Learning for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) has become an effective paradigm, which requires only class labels to train object detectors. However, WSOD detectors are prone to learn highly discriminative features corresponding to local objects rather than complete objects, resulting in imprecise object localization. To address the issue, designing backbones specifically for WSOD is a feasible solution. However, the redesigned backbone generally needs to be pretrained on large-scale ImageNet or trained from scratch, both of which require much more time and computational costs than fine-tuning. In this article, we explore to optimize the backbone without losing the availability of the original pretrained model. Since the pooling layer summarizes neighborhood features, it is crucial to spatial feature learning. In addition, it has no learnable parameters, so its modification will not change the pretrained model. Based on the above analysis, we further propose enhanced spatial feature learning (ESFL) for WSOD, which first takes full advantage of multiple kernels in a single pooling layer to handle multiscale objects and then enhances above-average activations within the rectangular neighborhood to alleviate the problem of ignoring unsalient object parts. The experimental results on the PASCAL VOC and the MS COCO benchmarks demonstrate that ESFL can bring significant performance improvement for the WSOD method and achieve state-of-the-art results.
Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 A Novel Hybrid Fusion Combining Palmprint and Palm Vein for Large-Scale Palm-Based Recognition
abstract
Palmprint and palm vein are emerging as unique biometric traits for identity authentication, each with its own advantages and limitations. Using these two traits jointly promises to enhance the discriminative and anti-spoofing capabilities. While existing research often combines these traits in parallel, such approaches lead to unnecessary increase in response time. Moreover, large-scale palm-based recognition, as a great potential task, raises higher requirements in accuracy and time efficiency. Nevertheless, few efforts have been dedicated to either data establishment or method investigation for this task. To this end, a large-scale palm-based multimodal dataset covering$ 20\,000 $palms is proposed, far larger than any of its kind. We also propose a hybrid fusion method to leverage these two diverse features. Our method employs a two-stage recognition process. First, a dual likelihood ratio test for coarse recognition is designed to assign palms into imposter certainty, genuine certainty or uncertainty classes. The coarse recognition narrows down the number of possible identities accurately using one trait, palm vein, consuming less recognition time. Then, in fine recognition, an adaptive weighted fusion of palmprint and palm vein is proposed to delicately rerecognize the uncertainty subsets that are in doubt in the coarse recognition, resulting in a more discriminative capacities. Experimental results confirm the effectiveness of our method, showing improved recognition performance with high-time efficiency.
Wei Jia 0001, Junan Chen 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2024 Toward Large-Scale Palmprint Image Analysis by a Rich Orientation Code
abstract
Palmprint recognition has gained considerable attention in recent years, accompanied by significant progress. However, large-scale palmprint recognition, which holds great potential for extensive civilian applications like university access control, remains underexplored. In this work, we propose the CUHK-T dataset, the largest palmprint dataset to date, containing over10k individuals’ palmprints. As the data volume expands, the heightened complexity, such as similar principal lines from different palms, necessitates a recognition method capable of extracting more discriminative features. Motivated by the rich palm lines distributed in palmprint, including not only nonintersecting, but also intersecting line segments, we model intersecting lines, investigate their properties, and propose a novel and explainable palmprint recognition method. The model treats the nonintersecting line segment as a special case, allowing for the extraction of orientation information from both types of line segments. In addition to the rich orientation information of the intersecting lines, the extracted feature accounts for the relative width of these lines. These advancements enable the extracted rich orientation code to be more discriminative and representative for palmprint. We then present a bitwise similarity measurement for efficiently and effectively comparing two rich orientation codes. Our extensive experiments and evaluations with popular palmprint recognition algorithms demonstrate the effectiveness and superior performance of our method on the large-scale dataset. These results also serve as a foundational baseline, facilitating the advancement of further research in the domain of large-scale palmprint recognition.
Wei Jia 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Heterogeneous Window Transformer for Image Denoising
abstract
Deep networks can usually depend on extracting more structural information to improve denoising results. However, they may ignore correlation between pixels from an image to pursue better-denoising performance. Window Transformer can use long- and short-distance modeling to interact pixels to address mentioned problem. To make a tradeoff between distance modeling and denoising time, we propose a heterogeneous window Transformer (HWformer) for image denoising. HWformer first designs heterogeneous global windows to capture global context information for improving denoising effects. To build a bridge between long and short-distance modeling, global windows are horizontally and vertically shifted to facilitate diversified information without increasing denoising time. To prevent the information loss phenomenon of independent patches, sparse idea is guided a feed-forward network to extract local information of neighboring patches. The proposed HWformer only takes 30% of popular restoration Transformer in terms of denoising time. Its codes can be obtained athttps://github.com/hellloxiaotian/HWformer.
Chunwei Tian, Menghua Zheng, Chia-Wen Lin, Zhiwu Li 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2023 M2SH: A Hybrid Approach to Table Structure Recognition using Two-Stage Multi-Modality Feature Fusion
abstract
Automatically recovering the original structure of tables from unstructured images is a challenging task, combining techniques from computer vision (CV) and natural language processing (NLP). Unfortunately, common feature extraction methods, naive fusion strategies, and rigid inductive biases have become roadblocks to the effective improvement of previous approaches. Distinguished from other modes of data representation, tables consist of many dispersed cells that are interdependent. Therefore, in this paper, we aim to propose a novel approach for recognizing table structures by mining the special properties of tables. The method begins by utilizing the adaptive fusion method to fuse visual and textual features acquired through a two-stream network. In the second stage, the layout features will be seamlessly integrated using a Kronecker-based strategy. The table elements with multi-modality features are then modeled based on spatial relationships. Interactions among them are established by a hybrid contextual aggregator that allows message passing at both local and global levels. Finally, table structure recognition is achieved by predicting the relationship between elements. We meticulously evaluate the proposed approach on various public datasets, including ICDAR2013, UNLV, WTW, SciTSR, and SciTSR-COMP, as well as a more complicated private dataset. The proposed method performs excellently on these datasets.
Zhihan Ning, Guopeng Wang, Yingjie Bai, David Zhang 0001
SMC7
2023 MDFN: Mask deep fusion network for visible and infrared image fusion without reference ground-truth
Chaoxun Guo, David Zhang 0001
Expert Syst. Appl.4
2023 Multi-adversarial Faster-RCNN with Paradigm Teacher for Unrestricted Object Detection
Zhenwei He, Lei Zhang 0038, Xinbo Gao 0001, David Zhang 0001
Int. J. Comput. Vis.4
2023 Sparse projection infinite selection ensemble for imbalanced classification
Zhihan Ning, David Zhang 0001
Knowl. Based Syst.3
2023 Deep adaptive hiding network for image hiding using attentive frequency extraction and gradual depth extraction
Le Zhang 0016, Yao Lu 0008, Jinxing Li 0003, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
Neural Comput. Appl.6
2023 Human Collective Intelligence Inspired Multi-View Representation Learning - Enabling View Communication by Simulating Human Communication Mechanism
abstract
In real-world applications, we often encounter multi-view learning tasks where we need to learn from multiple sources of data or use multiple sources of data to make decisions. Multi-view representation learning, which can learn a unified representation from multiple data sources, is a key pre-task of multi-view learning and plays a significant role in real-world applications. Accordingly, how to improve the performance of multi-view representation learning is an important issue. In this work, inspired by human collective intelligence shown in group decision making, we introduce the concept of view communication into multi-view representation learning. Furthermore, by simulating human communication mechanism, we propose a novel multi-view representation learning approach that can fulfill multi-round view communication. Thus, each view of our approach can exploit the complementary information from other views to help with modeling its own representation, and mutual help between views is achieved. Extensive experiment results on six datasets from three significant fields indicate that our approach substantially improves the average classification accuracy by 4.536% in medicine and bioinformatics fields as well as 4.115% in machine learning field.
Xiaodong Jia 0005, Xiaoyuan Jing, Qixing Sun, Songcan Chen, Bo Du 0001, David Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Learning efficient facial landmark model for human attractiveness analysis
Tianhao Peng 0003, Mu Li 0005, Fangmei Chen, Yong Xu 0001, David Zhang 0001
Pattern Recognit.5
2023 Multi-stage image denoising with the wavelet transform
Chunwei Tian, Menghua Zheng, Wangmeng Zuo, Bob Zhang 0001, Yanning Zhang 0001, David Zhang 0001
Pattern Recognit.6
2023 Facial Expression Recognition in the Wild Using Multi-Level Features and Attention Mechanisms
abstract
Learning discriminative features is of vital importance for automatic facial expression recognition (FER) in the wild. In this article, we propose a novel Slide-Patch and Whole-Face Attention model with SE blocks (SPWFA-SE), which jointly perceives the discriminative locality characteristics and informative global features of the face for effective FER. Specifically, the well-designed slide patches are proposed to extract local features. Different from the existing methods, our slide patches not only can maintain the information at the edge area of patches, but also do not need to detect facial landmarks. Moreover, to make the model adaptively focus on the distinguishable regions, an attention module is proposed in the patch level to learn the weight of each patch. Furthermore, squeeze-and-excitation blocks are explored in the channel level to learn the weight of each channel. As such, the proposed multi-level feature extraction and attention mechanisms can enhance the representative ability of the learned features. Extensive experiments on five challenging datasets demonstrate that our method can achieve state-of-the-art performance. Cross database experiments on another three databases show the superior generalization performance of our model. Furthermore, complexity analysis results show that our model contains fewer parameters with fast training advantages than other competing models.
Yingjian Li 0001, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, David Zhang 0001
IEEE Trans. Affect. Comput.5
2023 Neural Image Parts Group Search for Person Re-Identification
abstract
Employing partition strategy to explore fine-grained features has been verified to be beneficial for person re-identification in recent literature. However, existing methods primarily rely on expert experience to manually design various partition strategies, which may lead to a sub-optimal solution for fine-grained features exploration. In this paper, we propose a Neural Parts Group Search (NPGS) strategy that auto-searches the optimal parts group via evolutionary algorithm (EA) to facilitate the network to exploit the local details. And during search process, designing a high-quality search space is especially crucial for an efficient optimization. Considering the human top-down structure and the semantic coherence of parts, we design a coarse-to-fine parts search space (C2F-PSP) in NPGS, which effectively reduce the search complexity without the loss of parts expressivity. Additionally, since only employing the high-level semantic features is insufficient for the NPGS to search effective parts, we further develop an efficient feature aggregation strategy named hierarchical low-rank bilinear pooling that progressively integrates the high-level semantic property and the low-level fine-grained details to facilitate the NPGS to explore the fine-grained features. Furthermore, to relieve the interference of background during parts search process, we propose a novel Relational Attention Module (RAM) by exploiting the channel and spatial structural interdependence of pixels to strengthen the discriminative regions. Extensive experiments on the mainstream evaluation datasets demonstrate that our method outperforms the recent state-of-the-art Re-ID models.
Zhipu Liu, Lei Zhang 0038, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Contactless Palmprint Image Recognition Across Smartphones With Self-Paced CycleGAN
abstract
Contactless palmprint recognition, an emerging biometric technology, has attracted increasing attention due to its noninvasive and high practicability characteristics. Although it is naturally suitable for mobile application scenarios, the following two challenges severely limit its recognition performance: 1) the inconsistency in acquisition devices used in training and testing, and 2) many subjects are unable to be imaged on each device, resulting in incomplete data problems. To address these issues, we propose a self-paced CycleGAN with self-attention modules, which simultaneously synthesizes missing data and alleviates the influence of different imaging devices. Specifically, we develop CycleGAN with self-attention modules to generate missing training data by effectively mining the structural correlation among samples while capturing the cross-domain features. Furthermore, a self-paced learning strategy, which is a human cognitive-driven learning mechanism, is used to guide learning the robust cross-domain feature representation and recognition model, by which the relatively easy learning samples are gradually involved in the training process. To verify the effectiveness of the proposed method, we conduct experiments on contactless palmprint datasets collected using different smartphones. The results show that our approach outperforms state-of-the-art methods in classifying contactless palmprint images.
Qi Zhu 0001, Guangnan Xin, Lunke Fei, Dong Liang 0008, Zheng Zhang 0006, Daoqiang Zhang, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.7
2023 Multi-Spectral Palmprints Joint Attack and Defense With Adversarial Examples Learning
abstract
As an emerging biometric technology, multi-spectral palmprint recognition has attracted increasing attention in security due to its high accuracy and ease of use. Compared to single spectral case, multi-spectral palmprint model is more susceptible to the attack of adversarial examples. However, the previous adversarial example attack approaches cannot generate the most aggressive adversarial examples for multi-spectral palmprint recognition. In addition, most of them are dependent on the explicit architecture or need time-consuming queries about the network to be attacked, which significantly limits their application in the field of security. To solve the above problems, in this paper, we proposed the multi-spectral palmprints joint attack and defense framework based on multi-view adversarial examples learning. First, we respectively capture the multi-view deep common feature space for the different spectra and the discriminative feature space across the different subjects. Second, we introduce perturbation in the deep common space to achieve adversarial multi-spectral palmprints with gradient propagation. In addition, we pursue the manifold of the difference space and use it to suppress the discriminability of the recognition model with adversarial region theory. Finally, the generated adversarial examples are fed into the training model to enhance the robustness of the recognition algorithm. The experimental results on multi-spectral palmprint dataset demonstrate that the proposed multi-view joint attack approach is superior to the state-of-the-art adversarial example attack methods in attack accuracy and transferability. Moreover, the defense strategy with the adversarial examples by our method can significantly promote the robustness of multi-spectral palmprint recognition methods.
Qi Zhu 0001, Yuze Zhou, Lunke Fei, Daoqiang Zhang, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2023 HIPA: Hierarchical Patch Transformer for Single Image Super Resolution
abstract
Transformer-based architectures start to emerge in single image super resolution (SISR) and have achieved promising performance. However, most existing vision Transformer-based SISR methods still have two shortcomings: (1) they divide images into the same number of patches with a fixed size, which may not be optimal for restoring patches with different levels of texture richness; and (2) their position encodings treat all input tokens equally and hence, neglect the dependencies among them. This paper presents a HIPA, which stands for a novel Transformer architecture that progressively recovers the high resolution image using a hierarchical patch partition. Specifically, we build a cascaded model that processes an input image in multiple stages, where we start with tokens with small patch sizes and gradually merge them to form the full resolution. Such a hierarchical patch mechanism not only explicitly enables feature aggregation at multiple resolutions but also adaptively learns patch-aware features for different image regions, e.g., using a smaller patch for areas with fine details and a larger patch for textureless regions. Meanwhile, a new attention-based position encoding scheme for Transformer is proposed to let the network focus on which tokens should be paid more attention by assigning different weights to different tokens, which is the first time to our best knowledge. Furthermore, we also propose a multi-receptive field attention module to enlarge the convolution receptive field from different branches. The experimental results on several public datasets demonstrate the superior performance of the proposed HIPA over previous methods quantitatively and qualitatively. We will share our code and models when the paper is accepted.
Yiming Qian, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.7
2023 Pedestrian Detection by Exemplar-Guided Contrastive Learning
abstract
Typical methods for pedestrian detection focus on either tackling mutual occlusions between crowded pedestrians, or dealing with the various scales of pedestrians. Detecting pedestrians with substantial appearance diversities such as different pedestrian silhouettes, different viewpoints or different dressing, remains a crucial challenge. Instead of learning each of these diverse pedestrian appearance features individually as most existing methods do, we propose to perform contrastive learning to guide the feature learning in such a way that the semantic distance between pedestrians with different appearances in the learned feature space is minimized to eliminate the appearance diversities, whilst the distance between pedestrians and background is maximized. To facilitate the efficiency and effectiveness of contrastive learning, we construct an exemplar dictionary with representative pedestrian appearances as prior knowledge to construct effective contrastive training pairs and thus guide contrastive learning. Besides, the constructed exemplar dictionary is further leveraged to evaluate the quality of pedestrian proposals during inference by measuring the semantic distance between the proposal and the exemplar dictionary. Extensive experiments on both daytime and nighttime pedestrian detection validate the effectiveness of the proposed method.
Zebin Lin, Wenjie Pei, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Image Process.4
2023 Style Uncertainty Based Self-Paced Meta Learning for Generalizable Person Re-Identification
abstract
Domain generalizable person re-identification (DG ReID) is a challenging problem, because the trained model is often not generalizable to unseen target domains with different distribution from the source training domains. Data augmentation has been verified to be beneficial for better exploiting the source data to improve the model generalization. However, existing approaches primarily rely on pixel-level image generation that requires designing and training an extra generation network, which is extremely complex and provides limited diversity of augmented data. In this paper, we propose a simple yet effective feature based augmentation technique, named Style-uncertainty Augmentation (SuA). The main idea of SuA is to randomize the style of training data by perturbing the instance style with Gaussian noise during training process to increase the training domain diversity. And to better generalize knowledge across these augmented domains, we propose a progressive learning to learn strategy named Self-paced Meta Learning (SpML) that extends the conventional one-stage meta learning to multi-stage training process. The rationality is to gradually improve the model generalization ability to unseen target domains by simulating the mechanism of human learning. Furthermore, conventional person Re-ID loss functions are unable to leverage the valuable domain information to improve the model generalization. So we further propose a distance-graph alignment loss that aligns the feature relationship distribution among domains to facilitate the network to explore domain-invariant representations of images. Extensive experiments on four large-scale benchmarks demonstrate that our SuA-SpML achieves state-of-the-art generalization to unseen domains for person ReID.
Lei Zhang 0038, Zhipu Liu, Wensheng Zhang 0002, David Zhang 0001
IEEE Trans. Image Process.4
2023 Deep Margin-Sensitive Representation Learning for Cross-Domain Facial Expression Recognition
abstract
Cross-domain Facial Expression Recognition (FER) aims to safely transfer the learned knowledge from labeled source data to unlabeled target data, which is challenging due to the subtle difference between various expressions and the large discrepancy between domains. Existing methods mainly focus on reducing the domain shift for transferable features but fail to learn discriminative representations for recognizing facial expression, which may result in negative transfer under cross-domain settings. To this end, we propose a novel Deep Margin-Sensitive Representation Learning (DMSRL) framework, which can extract multi-level discriminative features during sematic-aware domain adaptation. Specifically, we design a semantic metric learning module based on the category prior of source data and generated pseudo labels of target data, which can facilitate discriminative intra-domain representation learning and transferable inter-domain knowledge discovery by enlarging the category margin. Moreover, we develop a mutual information minimization module by simultaneously distilling the domain-invariant components and eliminating the domain-sensitive ones, which benefits discriminative transferable feature learning by generating accurate pseudo target labels. Furthermore, instead of only utilizing the global features, we formulate a multi-level feature extracting module to concurrently get the local ones, which contain detailed information to distinguish the small changes among different expressions. These modules are jointly utilized in our DMSRL in an end-to-end manner to ensure the positive transfer of source knowledge. Extensive experimental results on seven databases demonstrate that our DMSRL can achieve superior performance against state-of-the-art baselines.
Yingjian Li 0001, Zheng Zhang 0006, Bingzhi Chen, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Multim.5
2023 Multiple Instance Detection Networks With Adaptive Instance Refinement
abstract
Weakly supervised object detection (WSOD) aims to train object detectors by using only image-level annotations. Many recent works on WSOD adopt multiple instance detection networks (MIDN), which usually generate a certain number of proposals and regard proposal classification as a latent model learning within image classification. However, these methods tend to detect salient object, salient object parts and clustered objects due to lack of instance-level annotations during training. Thus a core issue is how to guarantee that the network learn as many objects with precise bounding boxes as possible. In this paper, we address this issue by exploiting the potential of proposal scores during training. We propose an adaptive instance refinement (AIR) framework with three novel designs, which can be integrated with MIDN into a single network. Specifically, adaptive instance mining attempts to discover all positive instances according to the score distribution of proposals and their spatial similarity. Adaptive score modulation dynamically adjusts proposal scores to make the network focus more on instances with different difficulties in different training iterations. Adaptive knowledge refinement distills important information from all previous stages by the weighted average of proposal scores. The experimental results on the PASCAL VOC 2007 and 2012 benchmarks and the MS COCO benchmark demonstrate that AIR significantly improves the performance of the original MIDN and achieves the state-of-the-art results.
Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
IEEE Trans. Multim.5
2023 Self-Supervised Attentive Generative Adversarial Networks for Video Anomaly Detection
abstract
Video anomaly detection (VAD) refers to the discrimination of unexpected events in videos. The deep generative model (DGM)-based method learns the regular patterns on normal videos and expects the learned model to yield larger generative errors for abnormal frames. However, DGM cannot always do so, since it usually captures the shared patterns between normal and abnormal events, which results in similar generative errors for them. In this article, we propose a novel self-supervised framework for unsupervised VAD to tackle the above-mentioned problem. To this end, we design a novel self-supervised attentive generative adversarial network (SSAGAN), which is composed of the self-attentive predictor, the vanilla discriminator, and the self-supervised discriminator. On the one hand, the self-attentive predictor can capture the long-term dependences for improving the prediction qualities of normal frames. On the other hand, the predicted frames are fed to the vanilla discriminator and self-supervised discriminator for performing true-false discrimination and self-supervised rotation detection, respectively. Essentially, the role of the self-supervised task is to enable the predictor to encode semantic information into the predicted normal frames via adversarial training, in order for the angles of rotated normal frames can be detected. As a result, our self-supervised framework lessens the generalization ability of the model to abnormal frames, resulting in larger detection errors for abnormal frames. Extensive experimental results indicate that SSAGAN outperforms other state-of-the-art methods, which demonstrates the validity and advancement of SSAGAN.
Chao Huang 0008, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Jian Yang 0003, Yaowei Wang 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 Learning Context-Based Nonlocal Entropy Modeling for Image Compression
abstract
The entropy of the codes usually serves as the rate loss in the recent learned lossy image compression methods. Precise estimation of the probabilistic distribution of the codes plays a vital role in reducing the entropy and boosting the joint rate-distortion performance. However, existing deep learning based entropy models generally assume the latent codes are statistically independent or depend on some side information or local context, which fails to take the global similarity within the context into account and thus hinders the accurate entropy estimation. To address this issue, we propose a special nonlocal operation for context modeling by employing the global similarity within the context. Specifically, due to the constraint of context, nonlocal operation is incalculable in context modeling. We exploit the relationship between the code maps produced by deep neural networks and introduce the proxy similarity functions as a workaround. Then, we combine the local and the global context via a nonlocal attention block and employ it in masked convolutional networks for entropy modeling. Taking the consideration that the width of the transforms is essential in training low distortion models, we finally produce a U-net block in the transforms to increase the width with manageable memory consumption and time complexity. Experiments on Kodak and Tecnick datasets demonstrate the priority of the proposed context-based nonlocal attention block in entropy modeling and the U-net block in low distortion situations. On the whole, our model performs favorably against the existing image compression standards and recent deep image compression models.
Mu Li 0005, Kai Zhang 0008, Jinxing Li 0003, Wangmeng Zuo, Radu Timofte, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2022 Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-Identification
abstract
Thanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the probe-to- gallery, almost all existing RGB-IR based cross-modal person Re-ID methods focus on image-to-image matching, while the video-to-video matching which contains much richer spatial- and temporal-information remains under-explored. In this paper, we primarily study the video-based cross-modal per-son Re-ID method. To achieve this task, a video-based RGB-IR dataset is constructed, in which 927 valid identities with 463,259 frames and 21,863 tracklets captured by 12 RGB/IR cameras are collected. Based on our constructed dataset, we prove that with the increase of frames in a tracklet, the performance does meet more enhancement, demonstrating the significance of video-to-video matching in RGB-IR person Re-ID. Additionally, a novel method is further proposed, which not only projects two modalities to a modal-invariant subspace, but also extracts the temporal-memory for motion-invariant. Thanks to these two strategies, much better results are achieved on our video-based cross-modal person Re-ID. The code and dataset are released at: https://github.com/VCM-project233/MITML.
Jinxing Li 0003, Zeyu Ma 0001, Huafeng Li 0001, Kaixiong Xu, Guangming Lu 0002, David Zhang 0001
CVPR8
2022 BESS: Balanced evolutionary semi-stacking for disease detection using partially labeled imbalanced data
Zhihan Ning, Ziqing Ye, David Zhang 0001
Inf. Sci.4
2022 RVLSM: Robust variational level set method for image segmentation with intensity inhomogeneity and high noise
Fan Zhang 0070, Chuanshuo Cao, David Zhang 0001
Inf. Sci.5
2022 Real noise image adjustment networks for saliency-aware stylistic color retouch
Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, David Zhang 0001
Knowl. Based Syst.4
2022 Touchless palmprint recognition based on 3D Gabor template and block feature refinement
Zhaoqun Li, Jinxing Li 0003, Wei Jia 0001, David Zhang 0001
Knowl. Based Syst.6
2022 Image super-resolution with an enhanced group convolutional neural network
Chunwei Tian, Yixuan Yuan, Shichao Zhang 0001, Chia-Wen Lin, Wangmeng Zuo, David Zhang 0001
Neural Networks6
2022 MODENN: A Shallow Broad Neural Network Model Based on Multi-Order Descartes Expansion
abstract
Deep neural networks have achieved great success in almost every field of artificial intelligence. However, several weaknesses keep bothering researchers due to its hierarchical structure, particularly when large-scale parallelism, faster learning, better performance, and high reliability are required. Inspired by the parallel and large-scale information processing structures in the human brain, a shallow broad neural network model is proposed on a specially designed multi-order Descartes expansion operation. Such Descartes expansion acts as an efficient feature extraction method for the network, improve the separability of the original pattern by transforming the raw data pattern into a high-dimensional feature space, the multi-order Descartes expansion space. As a result, a single-layer perceptron network will be able to accomplish the classification task. The multi-order Descartes expansion neural network (MODENN) is thus created by combining the multi-order Descartes expansion operation and the single-layer perceptron together, and its capacity is proved equivalent to the traditional multi-layer perceptron and the deep neural networks. Three kinds of experiments were implemented, the results showed that the proposed MODENN model retains great potentiality in many aspects, including implementability, parallelizability, performance, robustness, and interpretability, indicating MODENN would be an excellent alternative to mainstream neural networks.
Haifeng Li 0001, Cong Xu 0004, Lin Ma 0003, Hongjian Bo, David Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Recursive Feature Diversity Network for audio super-resolution
Bo Jiang 0017, Mi-Xiao Hou, Yao Lu 0008, David Zhang 0001, Guangming Lu 0002
Speech Commun.5
2022 Multi-View Speech Emotion Recognition Via Collective Relation Construction
abstract
Automatic emotion recognition from speech plays a fundamental role towards advanced emotional intelligence in human-machine interaction systems. The discriminative knowledge from speech for effective emotion recognition may come from multiple physical properties such as energy spectrum, frequency, prosody, which could be collected as multi-view representations. However, the current works fail to fully explore the underlying interactive relations among multiple speech representations for emotion recognition. In this paper, we propose a novel Collective Multi-view Relation Network (CMRN) to exploit the intrinsic characteristics of multi-view speech representations for discriminative speech emotion recognition. Generally, the proposed CMRN consists of three sub-networks,i.e.,view-specific attention network, multi-view shared attention network and collective relation network. Specifically, the view-specific attention network is designed to excavate the distinguishable view-specific features deduced from the original speech. By contrast, the multi-view shared attention network is conceived to capture the collaborative knowledge from multiple views. Moreover, a well-designed collective relation network is explicitly constructed to characterize the shared-specific correlations, which could reflect the underlying physical interaction capabilities. As such, the decision phase can comprehensively leverage the shared and view-specific information of multiple representations, such that the final privileged deciding principle can aggregate the heterogeneous information of multi-view features to make accurate emotion recognition. Extensive experiments on two benchmark datasets demonstrate the superb performance of the proposed method in comparison with some state-of-the-art methods.
Mi-Xiao Hou, Zheng Zhang 0006, David Zhang 0001, Guangming Lu 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-Order Latent Domain
abstract
The crux of single-channel speech separation is how to encode the mixture of signals into such a latent embedding space that the signals from different speakers can be precisely separated. Existing methods for speech separation either transform the speech signals into frequency domain to perform separation or seek to learn a separable embedding space by constructing a latent domain based on convolutional filters. While the latter type of methods learning an embedding space achieves substantial improvement for speech separation, we argue that the embedding space defined by only one latent domain does not suffice to provide a thoroughly separable encoding space for speech separation. In this paper, we propose the Stepwise-Refining Speech Separation Network (SRSSN), which follows a coarse-to-fine separation framework. It first learns a 1-order latent domain to define an encoding space and thereby performs a rough separation in the coarse phase. Then the proposedSRSSNlearns a new latent domain along each basis function of the existing latent domain to obtain a high-order latent domain in the refining phase, which enables our model to perform a refining separation to achieve a more precise speech separation. We demonstrate the effectiveness of ourSRSSNby conducting extensive experiments, including speech separation in a clean (noise-free) setting on WSJ0-2/3mix datasets as well as in noisy/reverberant settings on WHAM!/WHAMR! datasets. Furthermore, we also perform experiments of speech recognition on separated speech signals by our model to evaluate the performance of speech separation indirectly.
Zengwei Yao, Wenjie Pei, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Multi-Label Chest X-Ray Image Classification via Semantic Similarity Graph Embedding
abstract
Automated multi-label chest X-ray (CXR) image classification has recently made significant progress in clinical diagnosis based on the advanced deep learning techniques. However, most existing methods mainly focus on analyzing locality visual cues from a single image but fail to leverage the underlying explicit correlations among different images for precise disease diagnosis. By contrast, an experienced radiologist expertizes in transferring knowledge from previous tasks to diagnose the present radiograph. To enable the machine like a radiologist, this paper proposes a novel Semantic Similarity Graph Embedding (SSGE) framework, which explicitly explores the semantic similarities among images to optimize the visual feature embedding for improving the performance of multi-label CXR images classification. Specifically, the proposed SSGE framework contains three main components: the image feature embedding (IFE) module, similarity graph construction (SGC) module, and semantic similarity learning (SSL) module. To realize interactive teaching and learning between visual and semantic information, the proposed SSGE framework is built on the “Teacher-Student” (semantic-visual) learning mechanism. With the guidance and supervision of the cross-image similarity graph generated by the SGC module, the SSL module leverages Graph Convolutional Network (GCN) to adaptively recalibrate the multi-image feature representations extracted from the IFE module, which guarantees their semantic consistency. Furthermore, we propose a novel re-weighting strategy to learn a more optimal semantic-similarity graph for the information propagation of the GCN layers. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in comparison with some state-of-the-art baselines.
Bingzhi Chen, Zheng Zhang 0006, Yingjian Li 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Generative Memory-Guided Semantic Reasoning Model for Image Inpainting
abstract
The critical challenge of single image inpainting stems from accurate semantic inference via limited information while maintaining image quality. Typical methods for semantic image inpainting train an encoder-decoder network by learning a one-to-one mapping from the corrupted image to the inpainted version. While such methods perform well on images with small corrupted regions, it is challenging for these methods to deal with images with large corrupted area due to two potential limitations. 1) Such one-to-one mapping paradigm tends to overfit each single training pair of images; 2) The inter-image prior knowledge about the general distribution patterns of visual semantics, which can be transferred across images sharing similar semantics, is not explicitly exploited. In this paper, we propose the Generative Memory-guided Semantic Reasoning Model (GM-SRM), which infers the content of corrupted regions based on not only the known regions of the corrupted image, but also the learned inter-image reasoning priors characterizing the generalizable semantic distribution patterns between similar images. In particular, the proposed GM-SRM first pre-learns a generative memory from the whole training data to explicitly learn the distribution of different semantic patterns. Then the learned memory are leveraged to retrieve the matching semantics for the current corrupted image to perform semantic reasoning during image inpainting. While the encoder-decoder network is used for guaranteeing the pixel-level content consistency, our generative priors are favorable for performing high-level semantic reasoning, which is particularly effective for inferring semantic content for large corrupted area. Extensive experiments on Paris Street View, CelebA-HQ, and Places2 benchmarks demonstrate that our GM-SRM outperforms the state-of-the-art methods for image inpainting in terms of both visual quality and quantitative metrics.
Xin Feng 0005, Wenjie Pei, Fengjun Li, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.5
2022 Deep Image Denoising With Adaptive Priors
abstract
Image denoising methods using deep neural networks have achieved a great progress in the image restoration. However, the recovered images restored by these deep denoising methods usually suffer from severe over-smoothness, artifacts, and detail loss. To improve the quality of restored images, we first propose Supplemental Priors (SP) method to adaptively predict depth-directed and sample-directed prior information for the reconstruction (decoder) networks. Furthermore, the over-parameterized deep neural networks and too precise supplemental prior information may cause an over-fitting, restricting the performance promotion. To improve the generalization of denoising networks, we further propose Regularization Priors (RP) method to flexibly learn depth-directed and dataset-directed regularization noise for the retrieving (encoder) networks. By respectively integrating the encoder and decoder with these plug-and-play RP block and SP block, we propose the final Adaptive Prior Denoising Networks, called APD-Nets. APD-Nets is the first attempt to simultaneously regularize and supplement denoising networks from the adaptive priors’ view with drawing learning-based mechanism into producing adaptive regularization noise and supplemental information. Extensive experiment results demonstrate our method significantly improves the generalization of denoising networks and the quality of restored images with greatly outperforming the traditional deep denoising methods both quantitatively and visually.The code will be released athttps://github.com/JiangBoCS/APD-Nets.
Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Self-Supervised Exclusive-Inclusive Interactive Learning for Multi-Label Facial Expression Recognition in the Wild
abstract
Facial Expression Recognition (FER) is a long-standing but challenging research problem in computer vision. Existing approaches mainly focus on single-label emotional prediction, which cannot handle the complex multi-label FER task because of the coupling behavior of multiple emotions on a single facial image. To this end, in this paper, we propose a novel Self-supervised Exclusive-Inclusive Interactive Learning (SEIIL) method to facilitate discriminative multi-label FER in the wild, which can effectively handle the coupled multiple sentiments with limited unconstrained training data. Specifically, we construct an emotion disentangling module to capture the inclusive and exclusive characteristics of facial expressions, which can decouple the compound numerous emotions on an image. Moreover, an adaptively-weighted ensemble technique is conceived to aggregate category-level latent exclusive embeddings, and then a conditional adversarial interactive learning module is designed to fully leverage the complementary between the inclusive and formulated latent representations. Furthermore, to tackle the insufficient data for training, we introduce a self-supervised learning strategy to augment the amount and diversity of facial images, which can endow the model with advanced generalization ability. Under this strategy, the proposed two modules can be concurrently utilized in our SEIIL to jointly handle the coupled emotions and alleviate the overfitting problem. Extensive experimental results on six databases illustrate the superb performance of our method against state-of-the-art baselines.
Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2022 Learning Informative and Discriminative Features for Facial Expression Recognition in the Wild
abstract
The informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models.
Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2022 Multiscale Conditional Regularization for Convolutional Neural Networks
abstract
With the increased model size of convolutional neural networks (CNNs), overfitting has become the main bottleneck to further improve the performance of networks. Currently, the weighting regularization methods have been proposed to address the overfitting problem and they perform satisfactorily. Since these regularization methods cannot be used in all the networks and they are usually not flexible enough in different phases of the training and test processes, this article proposes a multiscale conditional (MSC) regularization method. MSC divides the intermediate features into different scales and then generates new data for each scale features, respectively. In addition, the new data are generated by employing the information from two conditions: 1) each sample feature and 2) each layer pattern. Finally, a self-identity structure is proposed to supplement the features with the generated data. Therefore, MSC can adaptively and efficiently generate much finer and individualized data to make the entire regularization more flexible. Furthermore, MSC is more general and can be applied to all kinds of networks through the proposed self-identity structure. The experimental results on all the benchmark datasets showed that the proposed MSC regularization method achieves the best performances in all the networks.
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, Zheng Zhang 0006, David Zhang 0001
IEEE Trans. Cybern.6
2022 Addi-Reg: A Better Generalization-Optimization Tradeoff Regularization Method for Convolutional Neural Networks
abstract
In convolutional neural networks (CNNs), generating noise for the intermediate feature is a hot research topic in improving generalization. The existing methods usually regularize the CNNs by producing multiplicative noise (regularization weights), called multiplicative regularization (Multi-Reg). However, Multi-Reg methods usually focus on improving generalization but fail to jointly consider optimization, leading to unstable learning with slow convergence. Moreover, Multi-Reg methods are not flexible enough since the regularization weights are generated from a definite manual-design distribution. Besides, most popular methods are not universal enough, because these methods are only designed for the residual networks. In this article, we, for the first time, experimentally and theoretically explore the nature of generating noise in the intermediate features for popular CNNs. We demonstrate that injecting noise in the feature space can be transformed to generating noise in the input space, and these methods regularize the networks in a Mini-batch in Mini-batch (MiM) sampling manner. Based on these observations, this article further discovers that generating multiplicative noise can easily degenerate the optimization due to its high dependence on the intermediate feature. Based on these studies, we propose a novel additional regularization (Addi-Reg) method, which can adaptively produce additional noise with low dependence on intermediate feature in CNNs by employing a series of mechanisms. Particularly, these well-designed mechanisms can stabilize the learning process in training, and our Addi-Reg method can pertinently learn the noise distributions for every layer in CNNs. Extensive experiments demonstrate that the proposed Addi-Reg method is more flexible and universal, and meanwhile achieves better generalization performance with faster convergence against the state-of-the-art Multi-Reg methods.
Yao Lu 0008, Zheng Zhang 0006, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, David Zhang 0001
IEEE Trans. Cybern.6
2022 High Resolution Fingerprint Retrieval Based on Pore Indexing and Graph Comparison
abstract
Fingerprint retrieval aims to identify a query fingerprint image in a large database using indexing algorithms. Because of the abundant level 3 pore features within high-resolution fingerprint images, pore-based fingerprint retrieval algorithms have been rapidly developed. These retrieval algorithms, however, suffer from severe calculation-consuming problems with the pores increasing. This paper proposes a pore-based fingerprint retrieval method for high-resolution fingerprint images. The proposed method consists of two main steps. 1) In the pore indexing step, an indexing space is constructed using the binary codes of pores in enrolled images. Then, a designed graph-based searching algorithm searches the nearest neighbors of pores from the query image to construct one-to-many correspondences. 2) In the refinement step, the one-to-many correspondences are refined by a random walker-based graph comparison algorithm to remove the false correspondences. The remained nearest neighbors are used to calculate the similarities between the query image and the enrolled images. The proposed method is evaluated on two databases, showing that our method achieves better retrieval accuracies with a higher speed than the existing pore-based retrieval algorithms.
Yuanrong Xu, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2022 TDPN: Texture and Detail-Preserving Network for Single Image Super-Resolution
abstract
Single image super-resolution (SISR) using deep convolutional neural networks (CNNs) achieves the state-of-the-art performance. Most existing SISR models mainly focus on pursuing high peak signal-to-noise ratio (PSNR) and neglect textures and details. As a result, the recovered images are often perceptually unpleasant. To address this issue, in this paper, we propose a texture and detail-preserving network (TDPN), which focuses not only on local region feature recovery but also on preserving textures and details. Specifically, the high-resolution image is recovered from its corresponding low-resolution input in two branches. First, a multi-reception field based branch is designed to let the network fully learn local region features by adaptively selecting local region features in different reception fields. Then, a texture and detail-learning branch supervised by the textures and details decomposed from the ground-truth high resolution image is proposed to provide additional textures and details for the super-resolution process to improve the perceptual quality. Finally, we introduce a gradient loss into the SISR field and define a novel hybrid loss to strengthen boundary information recovery and to avoid overly smooth boundary in the final recovered high-resolution image caused by using only the MAE loss. More importantly, the proposed method is model-agnostic, which can be applied to most off-the-shelf SISR networks. The experimental results on public datasets demonstrate the superiority of our TDPN on most state-of-the-art SISR methods in PSNR, SSIM and perceptual quality. We will share our code on https://github.com/tocaiqing/TDPN.
Jinxing Li 0003, Huafeng Li 0001, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.6
2022 AVLSM: Adaptive Variational Level Set Model for Image Segmentation in the Presence of Severe Intensity Inhomogeneity and High Noise
abstract
Intensity inhomogeneity and noise are two common issues in images but inevitably lead to significant challenges for image segmentation and is particularly pronounced when the two issues simultaneously appear in one image. As a result, most existing level set models yield poor performance when applied to this images. To this end, this paper proposes a novel hybrid level set model, named adaptive variational level set model (AVLSM) by integrating an adaptive scale bias field correction term and a denoising term into one level set framework, which can simultaneously correct the severe inhomogeneous intensity and denoise in segmentation. Specifically, an adaptive scale bias field correction term is first defined to correct the severe inhomogeneous intensity by adaptively adjusting the scale according to the degree of intensity inhomogeneity while segmentation. More importantly, the proposed adaptive scale truncation function in the term is model-agnostic, which can be applied to most off-the-shelf models and improves their performance for image segmentation with severe intensity inhomogeneity. Then, a denoising energy term is constructed based on the variational model, which can remove not only common additive noise but also multiplicative noise often occurred in medical image during segmentation. Finally, by integrating the two proposed energy terms into a variational level set framework, the AVLSM is proposed. The experimental results on synthetic and real images demonstrate the superiority of AVLSM over most state-of-the-art level set models in terms of accuracy, robustness and running time.
Yiming Qian, Sanping Zhou, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.7
2022 End-to-End Optimized 360° Image Compression
abstract
The 360° image that offers a 360-degree scenario of the world is widely used in virtual reality and has drawn increasing attention. In 360° image compression, the spherical image is first transformed into a planar image with a projection such as equirectangular projection (ERP) and then saved with the existing codecs. The ERP images that represent different circles of latitude with the same number of pixels suffer from the unbalance sampling problem, resulting in inefficiency using planar compression methods, especially for the deep neural network (DNN) based codecs. To tackle this problem, we introduce a latitude adaptive coding scheme for DNNs by allocating variant numbers of codes for different regions according to the latitude on the sphere. Specifically, taking both the number of allocated codes for each region and their entropy into consideration, we introduce a flexible regional adaptive rate loss for region-wise rate controlling. Latitude adaptive constraints are then introduced to prevent spending too many codes on the over-sampling regions. Furthermore, we introduce viewport-based distortion loss by calculating the average distortion on a set of viewports. We optimize and test our model on a large 360° dataset containing 19,790 images collected from the Internet. The experiment results demonstrate the superiority of the proposed latitude adaptive coding scheme. On the whole, our model outperforms the existing image compression standards, including JPEG, JPEG2000, HEVC Intra Coding, and VVC Intra Coding, and helps to save around 15% bits compared to the baseline learned image compression model for planar images.
Mu Li 0005, Jinxing Li 0003, Shuhang Gu, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.5
2022 Joint Specifics and Consistency Hash Learning for Large-Scale Cross-Modal Retrieval
abstract
With the dramatic increase in the amount of multimedia data, cross-modal similarity retrieval has become one of the most popular yet challenging problems. Hashing offers a promising solution for large-scale cross-modal data searching by embedding the high-dimensional data into the low-dimensional similarity preserving Hamming space. However, most existing cross-modal hashing usually seeks a semantic representation shared by multiple modalities, which cannot fully preserve and fuse the discriminative modal-specific features and heterogeneous similarity for cross-modal similarity searching. In this paper, we propose a joint specifics and consistency hash learning method for cross-modal retrieval. Specifically, we introduce an asymmetric learning framework to fully exploit the label information for discriminative hash code learning, where 1) each individual modality can be better converted into a meaningful subspace with specific information, 2) multiple subspaces are semantically connected to capture consistent information, and 3) the integration complexity of different subspaces is overcome so that the learned collaborative binary codes can merge the specifics with consistency. Then, we introduce an alternatively iterative optimization to tackle the specifics and consistency hashing learning problem, making it scalable for large-scale cross-modal retrieval. Extensive experiments on five widely used benchmark databases clearly demonstrate the effectiveness and efficiency of our proposed method on both one-cross-one and one-cross-two retrieval tasks.
Jianyang Qin, Lunke Fei, Zheng Zhang 0006, Jie Wen 0001, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.6
2022 Multi-Feature Complementary Learning for Diabetes Mellitus Detection Using Pulse Signals
abstract
Computational pulse diagnosis is a convenient, non-invasive, and effective Diabetes Mellitus (DM) detection technique. Generally, diverse pulse features are extracted from different views to represent pulse signals and then used for achieving the pulse diagnosis. However, current pulse-based DM detection methods only used one pulse feature for detection, ignoring the fact that diverse pulse features can be combined together to boost the diagnosis performance. To this end, we propose a novel Multi-Feature Complementary Learning (MFCL) model for DM detection. By designing feature-specific projections, multiple features are separately projected into a shared observation space and effectively fused into one vector. Besides, a mapping function is built to correlate the fused vectors to category labels to make the fused vectors suitable for classification. Inspired by the graph Laplacian matrix, which effectively preserves the correlations among samples from different categories, we integrate it in MFCL and design a discriminative prior to make the fused vectors sufficiently discriminative. Finally, an optimization algorithm is proposed to alternatively optimize the projection variables and then generate fused feature vectors. The proposed method reaches an accuracy of 92.85% in DM detection, outperforming state-of-the-art methods.
Chaoxun Guo, David Zhang 0001
IEEE J. Biomed. Health Informatics3
2022 Deformable Template Network (DTN) for Object Detection
abstract
Objects often have different appearances because of viewpoint changes or part deformation. How to reasonably model these variations is still a big challenge for object detection. In this paper, we propose a novel Deformable Template Network (DTN), which exploits the pictorial structure to model possible variations of an object. DTN represents an object by virtue of a generated template in a deformable way. It has two key modules: the template generating module and the part matching module. The template generating module produces a template for a given object which defines the anchor positions of the$k{\times }k$parts. Based on such a template, the part matching module aims to perform part alignment around the anchor positions. In terms of each part, the matching process makes a trade-off between maximizing the detection score and minimizing the deformation cost relative to the anchor position. Moreover, DTN is a fully convolutional network which means it is competitive in terms of detection efficiency. We evaluate DTN on both the PASCAL VOC and MSCOCO datasets, achieving the state-of-the-art results, an accuracy of 82.7% for PASCAL VOC and of 44.9% for MSCOCO.
Shuai Wu 0001, Yong Xu 0001, Bob Zhang 0001, Jian Yang 0003, David Zhang 0001
IEEE Trans. Multim.5
2022 Jointly Heterogeneous Palmprint Discriminant Feature Learning
abstract
Heterogeneous palmprint recognition has attracted considerable research attention in recent years because it has the potential to greatly improve the recognition performance for personal authentication. In this article, we propose a simultaneous heterogeneous palmprint feature learning and encoding method for heterogeneous palmprint recognition. Unlike existing hand-crafted palmprint descriptors that usually extract features from raw pixels and require strong prior knowledge to design them, the proposed method automatically learns the discriminant binary codes from the informative direction convolution difference vectors of palmprint images. Differing from most heterogeneous palmprint descriptors that individually extract palmprint features from each modality, our method jointly learns the discriminant features from heterogeneous palmprint images so that the specific discriminant properties of different modalities can be better exploited. Furthermore, we present a general heterogeneous palmprint discriminative feature learning model to make the proposed method suitable for multiple heterogeneous palmprint recognition. Experimental results on the widely used PolyU multispectral palmprint database clearly demonstrate the effectiveness of the proposed method.
Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Chunwei Tian, Imad Rida, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2022 Semantic-Interactive Graph Convolutional Network for Multilabel Image Recognition
abstract
Multilabel image recognition, a critically practical task in computer vision, aims to predict multiple objects present in each image. The existing studies mainly focus on conceptual visual cues but fail to reconcile the visual information with their semantic guidance. Intuitively, humans can not only associate extra topological concepts but also imagine other approximate scenes based on a semantic description. Inspired by such semantic-interactive capability, two different types of semantic priors, i.e., the concept correlations of the same scene and semantic similarities among different scenes, should be further explored for the recognition decisions. To efficiently interact with these semantic relationships, in this article, we propose a novel semantic-interactive graph convolutional network (SI-GCN), which can leverage the topological information learned from knowledge graphs to boost the performance of multilabel recognition. Specifically, the proposed SI-GCN framework consists of two different GCN-based branches in parallel, i.e., concept correlations learning (CCL) branch and semantic similarity learning (SSL) branch. Inputting the semantic-embedding vectors of all the concepts, the CCL branch maps the label co-occurrence graph into a set of interdependent concept classifiers. Recalibrating the image feature embedding with the standardized supervision of the semantic similarity graph, the SSL branch learns the semantically consistent in-batch visual representations. Finally, a well-established interactive learning scheme is formulated to concurrently optimize the obtained concept classifiers and the visual representation learning in an end-to-end manner. Extensive experiments on the MS-COCO and Pascal VOC 2007 & 2012 benchmarks demonstrate the superiorities of the proposed SI-GCN method compared to the state-of-the-art baselines.
Bingzhi Chen, Zheng Zhang 0006, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2022 Innovative Contactless Palmprint Recognition System Based on Dual-Camera Alignment
abstract
Recently, contactless bimodal palmprint recognition technology has attracted increased attention due to the COVID-19 pandemic. Many dual-camera-based sensors have been proposed to capture palm vein and palmprint images synchronously. However, translations between captured palmprint and palm vein images differ depending on the distance between the hand and the sensors. To address this issue, we designed a low-cost method to align the bimodal palm regions for current dual-camera systems. In this study, we first implemented a contactless palm image acquisition device with a dual-camera module and a single-point time of flight (TOF) ranging sensor. Using this device, we collected a dataset named DCPD under different distances and light source intensities from 271 different palms. Then, a bimodal palm image alignment method is proposed based on the imaging and ranging models. After the system model is calibrated, the translation between the visible light and infrared light palm regions can be estimated quickly based on the palm distance. Finally, we designed a convolutional neural network (CNN) to effectively extract the fine- and coarse-grained palm features. Compared to widely used existing methods, the proposed networks achieved the lowest equal error rate (EER) on the Tongji, IITD, and DCPD datasets, and the average time cost of the system to perform one-time identification is approximately 0.15 s. The experimental results indicate that the proposed methods achieved high efficiency and comparable accuracy. In addition, the system’s EER and rank-1 on the DCPD dataset were 0.304% and 98.66%, respectively.
Zhaoqun Li, Bob Zhang 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2022 Asymmetric CNN for Image Superresolution
abstract
Deep convolutional neural networks (CNNs) have been widely applied for low-level vision over the past five years. According to the nature of different applications, designing appropriate CNN architectures is developed. However, customized architectures gather different features via treating all pixel points as equal to improve the performance of given application, which ignores the effects of local power pixel points and results in low training efficiency. In this article, we propose an asymmetric CNN (ACNet) comprising an asymmetric block (AB), a memory enhancement block (MEB), and a high-frequency feature enhancement block (HFFEB) for image superresolution (SR). The AB utilizes one-dimensional (1-D) asymmetric convolutions to intensify the square convolution kernels in horizontal and vertical directions for promoting the influences of local salient features for single image SR (SISR). The MEB fuses all hierarchical low-frequency features from AB via a residual learning technique to resolve the long-term dependency problem and transforms obtained low-frequency features into high-frequency features. The HFFEB exploits low- and high-frequency features to obtain more robust SR features and address the excessive feature enhancement problem. Additionally, it also takes charge of reconstructing a high-resolution image. Extensive experiments show that our ACNet can effectively address SISR, blind SISR, and blind SISR of blind noise problems. The code of the ACNet is shown athttps://github.com/hellloxiaotian/ACNet.
Chunwei Tian, Yong Xu 0001, Wangmeng Zuo, Chia-Wen Lin, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2021 BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion
Zhaoqun Li, Jinxing Li 0003, David Zhang 0001
ICONIP (6)5
2021 Highly shared Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001
Expert Syst. Appl.6
2021 Designing and training of a dual CNN for image denoising
Chunwei Tian, Yong Xu 0001, Wangmeng Zuo, Bo Du 0001, Chia-Wen Lin, David Zhang 0001
Knowl. Based Syst.6
2021 Learning Content-Weighted Deep Image Compression
abstract
Learning-based lossy image compression usually involves the joint optimization of rate-distortion performance, and requires to cope with the spatial variation of image content and contextual dependence among learned codes. Traditional entropy models can spatially adapt the local bit rate based on the image content, but usually are limited in exploiting context in code space. On the other hand, most deep context models are computationally very expensive and cannot efficiently perform decoding over the symbols in parallel. In this paper, we present a content-weighted encoder-decoder model, where the channel-wise multi-valued quantization is deployed for the discretization of the encoder features, and an importance map subnet is introduced to generate the importance masks for spatially varying code pruning. Consequently, the summation of importance masks can serve as an upper bound of the length of bitstream. Furthermore, the quantized representations of the learned code and importance map are still spatially dependent, which can be losslessly compressed using arithmetic coding. To compress the codes effectively and efficiently, we propose an upper-triangular masked convolutional network (triuMCN) for large context modeling. Experiments show that the proposed method can produce visually much better results, and performs favorably against deep and traditional lossy image compression approaches.
Mu Li 0005, Wangmeng Zuo, Shuhang Gu, Jane You, David Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Simultaneous Fidelity and Regularization Learning for Image Restoration
abstract
Most existing non-blind restoration methods are based on the assumption that a precise degradation model is known. As the degradation process can only be partially known or inaccurately modeled, images may not be well restored. Rain streak removal and image deconvolution with inaccurate blur kernels are two representative examples of such tasks. For rain streak removal, although an input image can be decomposed into a scene layer and a rain streak layer, there exists no explicit formulation for modeling rain streaks and the composition with scene layer. For blind deconvolution, as estimation error of blur kernel is usually introduced, the subsequent non-blind deconvolution process does not restore the latent image well. In this paper, we propose a principled algorithm within the maximum a posterior framework to tackle image restoration with a partially known or inaccurate degradation model. Specifically, the residual caused by a partially known or inaccurate degradation model is spatially dependent and complexly distributed. With a training set of degraded and ground-truth image pairs, we parameterize and learn the fidelity term for a degradation model in a task-driven manner. Furthermore, the regularization term can also be learned along with the fidelity term, thereby forming a simultaneous fidelity and regularization learning model. Extensive experimental results demonstrate the effectiveness of the proposed model for image deconvolution with inaccurate blur kernels, deconvolution with multiple degradations and rain streak removal.
Dongwei Ren, Wangmeng Zuo, David Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Jointly learning compact multi-view hash codes for few-shot FKP recognition
Lunke Fei, Bob Zhang 0001, Jie Wen 0001, Shaohua Teng, Shuyi Li 0003, David Zhang 0001
Pattern Recognit.6
2021 CompNet: Competitive Neural Network for Palmprint Recognition Using Learnable Gabor Kernels
abstract
Contactless palmprint recognition has recently made significant progress in palm-scanning payment and social security. However, most existing methods are based on handcrafted kernels and are sensitive to illumination and scale variations. To address this problem, a competitive convolutional neural network (CompNet) with constrained learnable Gabor filters is proposed for contactless palmprint recognition. The proposed CompNet is built on multisize competitive blocks, which are applied to effectively exploit the rich direction ordering information of the palmprint patterns by means of the ad-hoc softmax and channel-wise convolution operations. Compared to the current deep neural networks, the backbone of the proposed network contains only very few parameters, making it quite easy to train, especially on small-scale datasets. Experimental results obtained on four popular contactless palmprint datasets demonstrate that the proposed CompNet achieves the lowest equal error rate compared to the most commonly used methods.
Jinyang Yang, Guangming Lu 0002, David Zhang 0001
IEEE Signal Process. Lett.4
2021 Multimodal Emotion Recognition With Temporal and Semantic Consistency
abstract
Automated multimodal emotion recognition has become an emerging but challenging research topic in the fields of affective learning and sentiment analysis. The existing works mainly focus on developing multimodal fusion strategies to incorporate different emotion-related features. However, they fail to explore the inherent contextual consistency to reconcile the emotional information across modalities. In this paper, we propose a novel Time and Semantic Interaction Network (TSIN), which concurrently incorporates the advantages of temporal and semantic consistency into the multimodal emotion recognition task. Specifically, a well-designed Speech and Text Embedding (STE) module is devoted to formulating the initial embedding spaces by respectively building the modality-specific representations of speech and text. Instead of separately learning or directly fusing the acoustic and textual features, we propose a well-defined Time and Semantic Interaction (TSI) module to conduct the emotional parsing and sentiment refining by performing the fine-grained temporal alignment and cross-modal semantic interaction. Benefitting from temporal and semantic consistency constraints, both speech-text embeddings can be interactively optimized and fine-tuned in the learning process. In this way, the learnt acoustics and textual features can jointly and efficiently predict the final emotional state. Extensive experiments on the IEMOCAP dataset demonstrate the superiorities of our TSIN framework in comparison with state-of-the-art baselines.
Bingzhi Chen, Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Illuminance Compensation and Texture Enhancement via the Hodge Decomposition
abstract
Image brightness in color representation has caught active research interests, while its influence to texture features is relatively rarely studied. In this paper, we address the issue of illuminance, or brightness, interference to texture descriptors, especially to the Gabor filter based approaches. Firstly, we reveal the fact that the Gabor filter response linearly drifts with the illuminance. Secondly, the Hodge decomposition is introduced, and a linear regression model is presented to interpret the linear drift phenomenon. An interesting connection with the intrinsic image decomposition is established. Finally, a partial differential equation based method is proposed to compensate the drift issue. Consequently, the texture of the compensated image is enhanced. Extensive experiments are conducted to verify the proposed statements. The proposed linear regression model is verified empirically, and the compensation performances are demonstrated. Classification results on various databases indicate that the compensation can effectively improve the capacity of many existing texture description methods. Especially, the results show that the boosting effect on some deep learning based methods is also valid.
Bob Zhang 0001, Yong Xu 0001, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Adversarial View Confusion Feature Learning for Person Re-Identification
abstract
The performances of person re-identification tasks can be seriously degraded because of variations caused by view changes. In recent years, there are many methods focusing on how to solve cross view challenges which can be roughly divided into two categories: 1) learning view-invariant features without the help of view information. 2) combining view-wise features with the guide of view information. However, these methods are neither perfect enough. Methods of the first category are not roust enough for different kinds of view-invariants while methods of the other category can not generalize well in real-world applications. In this paper, we aim to learn view-invariant features with the help of view information. We proposed an end-to-end trainable framework, called View Confusion Feature Learning (VCFL), to learn view-invariant features by getting rid of view specific information. To the best of our knowledge, VCFL is originally proposed to learn view-invariant identity-wise features, and it is a kind of combination of view-generic and view-specific methods. The whole view confusion learning mechanism consists of three parts: 1) adversarial learning between feature extractor and the view classifier; 2) drawing the features with the same ID close to centers; 3) the guidance of SIFT, for seamlessly integration of hand-crafted features and deep features. In order to make the whole confusion mechanism work better, we further propose a VCFL+ model, which improves the fusion process in the feature map level through the thoughts of attention mechanism. Experiments on three benchmark datasets including Market1501, CUHK03, and DukeMTMC prove the superiority of our method over state-of-the-art approaches.
Lei Zhang 0038, Fangyi Liu, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Shared Linear Encoder-Based Multikernel Gaussian Process Latent Variable Model for Visual Classification
abstract
Multiview learning has been widely studied in various fields and achieved outstanding performances in comparison to many single-view-based approaches. In this paper, a novel multiview learning method based on the Gaussian process latent variable model (GPLVM) is proposed. In contrast to existing GPLVM methods which only assume that there are transformations from the latent variable to the multiple observed inputs, our proposed method simultaneously takes a back constraint into account, encoding multiple observations to the latent variable by enjoying the Gaussian process (GP) prior. Particularly, to overcome the difficulty of the covariance matrix calculation in the encoder, a linear projection is designed to map different observations to a consistent subspace first. The obtained variable in this subspace is then projected to the latent variable in the manifold space with the GP prior. Furthermore, different from most GPLVM methods which strongly assume that the covariance matrices follow a certain kernel function, for example, radial basis function (RBF), we introduce a multikernel strategy to design the covariance matrix, being more reasonable and adaptive for the data representation. In order to apply the presented approach to the classification, a discriminative prior is also embedded to the learned latent variables to encourage samples belonging to the same category to be close and those belonging to different categories to be far. Experimental results on three real-world databases substantiate the effectiveness and superiority of the proposed method compared with state-of-the-art approaches.
Jinxing Li 0003, Guangming Lu 0002, Bob Zhang 0001, Jane You, David Zhang 0001
IEEE Trans. Cybern.5
2021 Scaled Simplex Representation for Subspace Clustering
abstract
The self-expressive property of data points, that is, each data point can be linearly represented by the other data points in the same subspace, has proven effective in leading subspace clustering (SC) methods. Most self-expressive methods usually construct a feasible affinity matrix from a coefficient matrix, obtained by solving an optimization problem. However, the negative entries in the coefficient matrix are forced to be positive when constructing the affinity matrix via exponentiation, absolute symmetrization, or squaring operations. This consequently damages the inherent correlations among the data. Besides, the affine constraint used in these methods is not flexible enough for practical applications. To overcome these problems, in this article, we introduce a scaled simplex representation (SSR) for the SC problem. Specifically, the non-negative constraint is used to make the coefficient matrix physically meaningful, and the coefficient vector is constrained to be summed up to a scalar to make it more discriminative. The proposed SSR-based SC (SSRSC) model is reformulated as a linear equality-constrained problem, which is solved efficiently under the alternating direction method of multipliers framework. Experiments on benchmark datasets demonstrate that the proposed SSRSC algorithm is very efficient and outperforms the state-of-the-art SC methods on accuracy. The code can be found at https://github.com/csjunxu/SSRSC.
Jun Xu 0019, Mengyang Yu, Ling Shao 0001, Wangmeng Zuo, Deyu Meng, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Cybern.7
2021 AdvKin: Adversarial Convolutional Network for Kinship Verification
abstract
Kinship verification in the wild is an interesting and challenging problem. The goal of kinship verification is to determine whether a pair of faces are blood relatives or not. Most previous methods for kinship verification can be divided as handcrafted features-based shallow learning methods and convolutional neural network (CNN)-based deep-learning methods. Nevertheless, these methods are still facing the challenging task of recognizing kinship cues from facial images. The reason is that the family ID information and the distribution difference of pairwise kin-faces are rarely considered in kinship verification tasks. To this end, a family ID-based adversarial convolutional network (AdvKin) method focused on discriminative Kin features is proposed for both small-scale and large-scale kinship verification in this article. The merits of this article are four-fold: 1) for kin-relation discovery, a simple yet effective self-adversarial mechanism based on a negative maximum mean discrepancy (NMMD) loss is formulated as attacks in the first fully connected layer; 2) a pairwise contrastive loss and family ID-based softmax loss are jointly formulated in the second and third fully connected layer, respectively, for supervised training; 3) a two-stream network architecture with residual connections is proposed in AdvKin; and 4) for more fine-grained deep kin-feature augmentation, an ensemble of patch-wise AdvKin networks is proposed (E-AdvKin). Extensive experiments on 4 small-scale benchmark KinFace datasets and 1 large-scale families in the wild (FIW) dataset from the first Large-Scale Kinship Recognition Data Challenge, show the superiority of our proposed AdvKin model over other state-of-the-art approaches.
Lei Zhang 0038, Qingyan Duan, David Zhang 0001, Wei Jia 0001, Xizhao Wang
IEEE Trans. Cybern.3
2021 Deep-Masking Generative Network: A Unified Framework for Background Restoration From Superimposed Images
abstract
Restoring the clean background from the superimposed images containing a noisy layer is the common crux of a classical category of tasks on image restoration such as image reflection removal, image deraining and image dehazing. These tasks are typically formulated and tackled individually due to diverse and complicated appearance patterns of noise layers within the image. In this work we present the Deep-Masking Generative Network (DMGN), which is a unified framework for background restoration from the superimposed images and is able to cope with different types of noise. Our proposed DMGN follows a coarse-to-fine generative process: a coarse background image and a noise image are first generated in parallel, then the noise image is further leveraged to refine the background image to achieve a higher-quality background image. In particular, we design the novel Residual Deep-Masking Cell as the core operating unit for our DMGN to enhance the effective information and suppress the negative information during image generation via learning a gating mask to control the information flow. By iteratively employing this Residual Deep-Masking Cell, our proposed DMGN is able to generate both high-quality background image and noisy image progressively. Furthermore, we propose a two-pronged strategy to effectively leverage the generated noise image as contrasting cues to facilitate the refinement of the background image. Extensive experiments across three typical tasks for image background restoration, including image reflection removal, image rain steak removal and image dehazing, show that our DMGN consistently outperforms state-of-the-art methods specifically designed for each single task.
Xin Feng 0005, Wenjie Pei, Zihui Jia, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Image Process.5
2021 Layer-Output Guided Complementary Attention Learning for Image Defocus Blur Detection
abstract
Defocus blur detection (DBD), which has been widely applied to various fields, aims to detect the out-of-focus or in-focus pixels from a single image. Despite the fact that the deep learning based methods applied to DBD have outperformed the hand-crafted feature based methods, the performance cannot still meet our requirement. In this paper, a novel network is established for DBD. Unlike existing methods which only learn the projection from the in-focus part to the ground-truth, both in-focus and out-of-focus pixels, which are completely and symmetrically complementary, are taken into account. Specifically, two symmetric branches are designed to jointly estimate the probability of focus and defocus pixels, respectively. Due to their complementary constraint, each layer in a branch is affected by an attention obtained from another branch, effectively learning the detailed information which may be ignored in one branch. The feature maps from these two branches are then passed through a unique fusion block to simultaneously get the two-channel output measured by a complementary loss. Additionally, instead of estimating only one binary map from a specific layer, each layer is encouraged to estimate the ground truth to guide the binary map estimation in its linked shallower layer followed by a top-to-bottom combination strategy, gradually exploiting the global and local information. Experimental results on released datasets demonstrate that our proposed method remarkably outperforms state-of-the-art algorithms.
Jinxing Li 0003, Lingxiao Yang, Shuhang Gu, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.7
2021 Harmonization Shared Autoencoder Gaussian Process Latent Variable Model With Relaxed Hamming Distance
abstract
Multiview learning has shown its superiority in visual classification compared with the single-view-based methods. Especially, due to the powerful representation capacity, the Gaussian process latent variable model (GPLVM)-based multiview approaches have achieved outstanding performances. However, most of them only follow the assumption that the shared latent variables can be generated from or projected to the multiple observations but fail to exploit the harmonization in the back constraint and adaptively learn a classifier according to these learned variables, which would result in performance degradation. To tackle these two issues, in this article, we propose a novel harmonization shared autoencoder GPLVM with a relaxed Hamming distance (HSAGP-RHD). Particularly, an autoencoder structure with the Gaussian process (GP) prior is first constructed to learn the shared latent variable for multiple views. To enforce the agreement among various views in the encoder, a harmonization constraint is embedded into the model by making consistency for the view-specific similarity. Furthermore, we also propose a novel discriminative prior, which is directly imposed on the latent variable to simultaneously learn the fused features and adaptive classifier in a unit model. In detail, the centroid matrix corresponding to the centroids of different categories is first obtained. A relaxed Hamming distance (RHD)-based measurement is subsequently presented to measure the similarity and dissimilarity between the latent variable and centroids, not only allowing us to get the closed-form solutions but also encouraging the points belonging to the same class to be close, while those belonging to different classes to be far. Due to this novel prior, the category of the out-of-sample is also allowed to be simply assigned in the testing phase. Experimental results conducted on three real-world data sets demonstrate the effectiveness of the proposed method compared with state-of-the-art approaches.
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Yong Xu 0001, Feng Wu 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2021 A Novel Multicamera System for High-Speed Touchless Palm Recognition
abstract
Palm-related biometrics have been widely studied for a long time, as the palm contains many distinctive patterns. However, most of the existing systems are designed to work within an ideal environment, such as in front of a unicolor background or in a large enclosure. Those preconditions can avoid influences of ambient light and hand distance change, but at the same time, they also limit the applications of palm recognition. In the work reported in this paper, we designed a novel red-green-blue and depth-based four-camera system that can capture the palm-related images separately in real time. The techniques of region-of-interest (ROI) location, ROI alignment, and light-source intensity optimization were studied. The ROI location method is modified to increase the robustness of hand gesture variation. Based on the depth information, we proposed the coordinate mapping and inclination rectification methods to obtain aligned ROI pairs. Using this device, we collected a video-based multimodal palm image database. After the parameter optimization and information fusion, the equal-error-rate of our approach on this database is lower than 0.47%. The recognition rate obtained from the support-vector-machine-based fusion is higher than 99.8%. The experimental results prove that the proposed system achieves advantages of anti-spoofing, high speed, high accuracy, and small size.
David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001, Nan Luo
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Fast Pore Comparison for High Resolution Fingerprint Images Based on Multiple Co-Occurrence Descriptors and Local Topology Similarities
abstract
Pore-based fingerprint recognition has been researched for decades. Many algorithms have been proposed to improve the recognition accuracy of the system. However, the accuracies are always improved at the cost of speed. This article proposes a novel method to compare the pores in high-resolution fingerprint images using the popular coarse-to-fine strategy. A multiple spatial pairwise local co-occurrence descriptor is proposed to improve the calculation of the similarities between pores. It calculates multiple local co-occurrence statistics for each pore using its neighbors. The proposed method can establish correspondences between pores more accurately. The refinement of the correspondences is then achieved by using a local topology-preserving matching algorithm. The algorithm uses rotational invariant local structures and pore pair local topology similarities to calculate the cost of each correspondence. It can remove the mismatches more accurately and efficiently. The experimental results on two high-resolution fingerprint image databases show that the proposed algorithm perform well in both accuracy and speed comparing to the existing algorithms.
Yuanrong Xu, Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2020 Similarity and diversity induced paired projection for cross-modal retrieval
Jinxing Li 0003, Mu Li 0005, Guangming Lu 0002, Bob Zhang 0001, Hongpeng Yin, David Zhang 0001
Inf. Sci.6
2020 3D palmprint identification using blocked histogram and improved sparse representation-based classifier
Zhaozong Meng, Nan Gao 0002, Zonghua Zhang, David Zhang 0001
Neural Comput. Appl.5
2020 High-parameter-efficiency convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001
Neural Comput. Appl.5
2020 Optimal Projection Guided Transfer Hashing for Image Retrieval
abstract
Recently, learning to hash has been widely studied for image retrieval thanks to the computation and storage efficiency of binary codes. Most existing learning to hash methods have yielded significant performance. However, for most existing learning to hash methods, sufficient training images are required and used to learn precise hashing codes. In some real-world applications, there are not always sufficient training images in the domain of interest. In addition, some existing supervised approaches need a amount of labeled data, which is an expensive process in terms of time, labor and human expertise. To handle such problems, inspired by transfer learning, we propose a simple yet effective unsupervised hashing method named Optimal Projection Guided Transfer Hashing (GTH) where we borrow the images of other different but related domain i.e., source domain to help learn precise hashing codes for the domain of interest i.e., target domain. In GTH, we aim to learn domain-invariant hashing functions. To achieve that, we propose to minimize the error matrix between two hashing projections of target and source domains. We seek for the maximum likelihood estimation (MLE) solution of the error matrix between the two hashing projections due to the domain gap. Furthermore, an alternating optimization method is adopted to obtain the two projections of target and source domains. By doing so, two projections can be progressively aligned. Extensive experiments on various benchmark databases for cross-domain visual recognition verify that our method outperforms many state-of-the-art learning to hash methods. The source code is available at https://github.com/liuji93/GTH.
Lei Zhang 0038, Ji Liu 0002, Yang Yang 0002, Fuxiang Huang, Feiping Nie 0001, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2020 DRPL: Deep Regression Pair Learning for Multi-Focus Image Fusion
abstract
In this paper, a novel deep network is proposed for multi-focus image fusion, named Deep Regression Pair Learning (DRPL). In contrast to existing deep fusion methods which divide the input image into small patches and apply a classifier to judge whether the patch is in focus or not, DRPL directly converts the whole image into a binary mask without any patch operation, subsequently tackling the difficulty of the blur level estimation around the focused/defocused boundary. Simultaneously, a pair learning strategy, which takes a pair of complementary source images as inputs and generates two corresponding binary masks, is introduced into the model, greatly imposing the complementary constraint on each pair and making a large contribution to the performance improvement. Furthermore, as the edge or gradient does exist in the focus part while there is no similar property for the defocus part, we also embed a gradient loss to ensure the generated image to be all-in-focus. Then the structural similarity index (SSIM) is utilized to make a trade-off between the reference and fused images. Experimental results conducted on the synthetic and real-world datasets substantiate the effectiveness and superiority of DRPL compared with other state-of-the-art approaches. The testing code can be found in https://github.com/sasky1/DPRL.
Jinxing Li 0003, Xiaobao Guo, Guangming Lu 0002, Bob Zhang 0001, Yong Xu 0001, Feng Wu 0001, David Zhang 0001
IEEE Trans. Image Process.7
2020 Efficient and Effective Context-Based Convolutional Entropy Modeling for Image Compression
abstract
Precise estimation of the probabilistic structure of natural images plays an essential role in image compression. Despite the recent remarkable success of end-to-end optimized image compression, the latent codes are usually assumed to be fully statistically factorized in order to simplify entropy modeling. However, this assumption generally does not hold true and may hinder compression performance. Here we present contextbased convolutional networks (CCNs) for efficient and effective entropy modeling. In particular, a 3D zigzag scanning order and a 3D code dividing technique are introduced to define proper coding contexts for parallel entropy decoding, both of which boil down to place translation-invariant binary masks on convolution filters of CCNs. We demonstrate the promise of CCNs for entropy modeling in both lossless and lossy image compression. For the former, we directly apply a CCN to the binarized representation of an image to compute the Bernoulli distribution of each code for entropy estimation. For the latter, the categorical distribution of each code is represented by a discretized mixture of Gaussian distributions, whose parameters are estimated by three CCNs. We then jointly optimize the CCNbased entropy model along with analysis and synthesis transforms for rate-distortion performance. Experiments on the Kodak and Tecnick datasets show that our methods powered by the proposed CCNs generally achieve comparable compression performance to the state-of-the-art while being much faster.
Mu Li 0005, Kede Ma, Jane You, David Zhang 0001, Wangmeng Zuo
IEEE Trans. Image Process.4
2020 Remove Cosine Window From Correlation Filter-Based Visual Trackers: When and How
abstract
Correlation filters (CFs) have been continuously advancing the state-of-the-art tracking performance and have been extensively studied in the recent few years. Nonetheless, the existing CF trackers adopt a cosine window to spatially reweight base image to alleviate boundary discontinuity. However, cosine window emphasizes more on the central regions of base image and has the risk of contaminating negative training samples during model learning. On the other hand, spatial regularization deployed in many recent CF trackers plays a similar role as cosine window by enforcing spatial penalty on CF coefficients. Therefore, we in this paper investigate the feasibility to remove cosine window from CF trackers with spatial regularization. When simply removing cosine window, CF with spatial regularization still suffers from small degree of boundary discontinuity. To tackle this issue, binary and Gaussian shaped mask functions are further introduced for eliminating boundary discontinuity while reweighting the estimation error of each training sample, and can be incorporated with multiple CF trackers with spatial regularization. In comparison to the baseline methods with cosine window, our methods are effective in handling boundary discontinuity and sample contamination, thereby benefiting tracking performance. Extensive experiments on four benchmarks show that our methods perform favorably against the state-of-the-art trackers using either handcrafted or deep CNN features.
Feng Li 0031, Xiaohe Wu, Wangmeng Zuo, David Zhang 0001, Lei Zhang 0006
IEEE Trans. Image Process.4
2020 Deep-Like Hashing-in-Hash for Visual Retrieval: An Embarrassingly Simple Method
abstract
Existing hashing methods have yielded significant performance in image and multimedia retrieval, which can be categorized into two groups: shallow hashing and deep hashing. However, there still exist some intrinsic limitations among them. The former generally adopts a one-step strategy to learn the hashing codes for discovering the discriminative binary feature, but the latent discriminative information in the learned hashing codes is not well exploited. The latter, as deep neural network based hashing models, can learn highly discriminative and compact features, but relies on large-scale data and computation resources for numerous network parameters tuning with back-propagation optimization. Straightforward training of deep hashing models from scratch on small-scale data is almost impossible. Therefore, in order to develop efficient but effective learning to hash algorithm that depends only on small-scale data, we propose a novel non-neural network based deep-like learning framework, i.e. multi-level cascaded hashing (MCH) approach with hierarchical learning strategy, for image retrieval. The contributions are threefold. First, a hashing-in-hash architecture is designed in MCH, which inherits the excellent traits of traditional neural networks based deep learning, such that discriminative binary features that are beneficial to image retrieval can be effectively captured. Second, in each level the binary features of all preceding levels and the visual appearance feature are simultaneously cascaded as inputs of all subsequent levels to retrain, which fully exploits the implicated discriminative information. Third, a basic learning to hash (BLH) model with label constraint is proposed for hierarchical learning. Without loss of generality, the existing hashing models can be easily integrated into our MCH framework. We show experimentally on small- and large-scale visual retrieval tasks that our method outperforms several state-of-the-arts.
Lei Zhang 0038, Ji Liu 0002, Fuxiang Huang, Yang Yang 0002, David Zhang 0001
IEEE Trans. Image Process.5
2020 Deep Cascade Model-Based Face Recognition: When Deep-Layered Learning Meets Small Data
abstract
Sparse representation based classification (SRC), nuclear-norm matrix regression (NMR), and deep learning (DL) have achieved a great success in face recognition (FR). However, there still exist some intrinsic limitations among them. SRC and NMR based coding methods belong to one-step model, such that the latent discriminative information of the coding error vector cannot be fully exploited. DL, as a multi-step model, can learn powerful representation, but relies on large-scale data and computation resources for numerous parameters training with complicated back-propagation. Straightforward training of deep neural networks from scratch on small-scale data is almost infeasible. Therefore, in order to develop efficient algorithms that are specifically adapted for small-scale data, we propose to derive the deep models of SRC and NMR. Specifically, in this paper, we propose an end-to-end deep cascade model (DCM) based on SRC and NMR with hierarchical learning, nonlinear transformation and multi-layer structure for corrupted face recognition. The contributions include four aspects. First, an end-to-end deep cascade model for small-scale data without back-propagation is proposed. Second, a multi-level pyramid structure is integrated for local feature representation. Third, for introducing nonlinear transformation in layer-wise learning, softmax vector coding of the errors with class discrimination is proposed. Fourth, the existing representation methods can be easily integrated into our DCM framework. Experiments on a number of small-scale benchmark FR datasets demonstrate the superiority of the proposed model over state-of-the-art counterparts. Additionally, a perspective that deep-layered learning does not have to be convolutional neural network with back-propagation optimization is consolidated. The demo code is available in https://github.com/liuji93/DCM.
Lei Zhang 0038, Ji Liu 0002, Bob Zhang 0001, David Zhang 0001, Ce Zhu
IEEE Trans. Image Process.4
2020 Lesion Location Attention Guided Network for Multi-Label Thoracic Disease Classification in Chest X-Rays
abstract
Traditional clinical experiences have shown the benefit of lesion location attention for improving clinical diagnosis tasks. Inspired by this point of interest, in this paper we propose a novel lesion location attention guided network named LLAGnet to focus on the discriminative features from lesion locations for multi-label thoracic disease classification in chest X-rays (CXRs). By revealing the equivalence of the region-level attention (RLA) and channel-level attention (CLA), we find that the RLA is available as priors for object localization while the CLA implicitly provides high weights to the attractive channels, which both enable lesion location attention excitation. To integrate the advantages from both mechanisms, the proposed LLAGnet is structured with two corresponding attention modules, i.e., the RLA and CLA modules. Specifically, the RLA module consists of the global and local branches. And the weakly supervised attention mechanism embedded in the global branch can obtain visual regions of lesion locations by back-propagating gradients. Then the optimal attention region is amplified and applied to the local branch to provide more fine-grained features for the image classification. Finally, the CLA module adaptively enhances the weights of channel-wise features from the lesion locations by modeling interdependencies among channels. Extensive experiments on the ChestX-ray14 dataset clearly substantiate the effectiveness of LLAGnet as compared with the state-of-the-art baselines.
Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001
IEEE J. Biomed. Health Informatics4
2020 Label Co-Occurrence Learning With Graph Convolutional Networks for Multi-Label Chest X-Ray Image Classification
abstract
Existing multi-label medical image learning tasks generally contain rich relationship information among pathologies such as label co-occurrence and interdependency, which is of great importance for assisting in clinical diagnosis and can be represented as the graph-structured data. However, most state-of-the-art works only focus on regression from the input to the binary labels, failing to make full use of such valuable graph-structured information due to the complexity of graph data. In this paper, we propose a novel label co-occurrence learning framework based on Graph Convolution Networks (GCNs) to explicitly explore the dependencies between pathologies for the multi-label chest X-ray (CXR) image classification task, which we term the "CheXGCN". Specifically, the proposed CheXGCN consists of two modules, i.e., the image feature embedding (IFE) module and label co-occurrence learning (LCL) module. Thanks to the LCL model, the relationship between pathologies is generalized into a set of classifier scores by introducing the word embedding of pathologies and multi-layer graph information propagation. During end-to-end training, it can be flexibly integrated into the IFE module and then adaptively recalibrate multi-label outputs with these scores. Extensive experiments on the ChestX-Ray14 and CheXpert datasets have demonstrated the effectiveness of CheXGCN as compared with the state-of-the-art baselines.
Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, Hongbing Yu, David Zhang 0001
IEEE J. Biomed. Health Informatics5
2020 Relaxed Asymmetric Deep Hashing Learning: Point-to-Angle Matching
abstract
Due to the powerful capability of the data representation, deep learning has achieved a remarkable performance in supervised hash function learning. However, most of the existing hashing methods focus on point-to-point matching that is too strict and unnecessary. In this article, we propose a novel deep supervised hashing method by relaxing the matching between each pair of instances to a point-to-angle way. Specifically, an inner product is introduced to asymmetrically measure the similarity and dissimilarity between the real-valued output and the binary code. Different from existing methods that strictly enforce each element in the real-valued output to be either +1 or -1, we only encourage the output to be close to its corresponding semantic-related binary code under the cross-angle. This asymmetric product not only projects both the real-valued output and the binary code into the same Hamming space but also relaxes the output with wider choices. To further exploit the semantic affinity, we propose a novel Hamming-distance-based triplet loss, efficiently making a ranking for the positive and negative pairs. An algorithm is then designed to alternatively achieve optimal deep features and binary codes. Experiments on four real-world data sets demonstrate the effectiveness and superiority of our approach to the state of the art.
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Jane You, Yong Xu 0001, Feng Wu 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2020 SRGC-Nets: Sparse Repeated Group Convolutional Neural Networks
abstract
Group convolution is widely used in many mobile networks to remove the filter's redundancy from the channel extent. In order to further reduce the redundancy of group convolution, this article proposes a novel repeated group convolutional (RGC) kernel, which has M primary groups, and each primary group includes N tiny groups. In every primary group, the same convolutional kernel is repeated in all the tiny groups. The RGC filter is the first kernel to remove the redundancy from group extent. Based on RGC, a sparse RGC (SRGC) kernel is also introduced in this article, and its corresponding network is called SRGC neural networks (SRGC-Net). The SRGC kernel is the summation of RGC kernel and pointwise group convolutional (PGC) kernel. The number of PGC's groups is M . Accordingly, in each primary group, besides the center locations in all channels, the values of parameters located in other N-1 tiny groups are all zero. Therefore, SRGC can significantly reduce the parameters. Moreover, it can also effectively retrieve spatial and channel-difference features by utilizing RGC and PGC to preserve the richness of produced features. Comparative experiments were performed on the benchmark classification data sets. Compared with the traditional popular networks, SRGC-Nets can perform better with timely reducing the model size and computational complexity. Furthermore, it can also achieve better performances than other latest state-of-the-art mobile networks on most of the databases and effectively decrease the test and training runtime.
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2020 Guide Subspace Learning for Unsupervised Domain Adaptation
abstract
A prevailing problem in many machine learning tasks is that the training (i.e., source domain) and test data (i.e., target domain) have different distribution [i.e., non-independent identical distribution (i.i.d.)]. Unsupervised domain adaptation (UDA) was proposed to learn the unlabeled target data by leveraging the labeled source data. In this article, we propose a guide subspace learning (GSL) method for UDA, in which an invariant, discriminative, and domain-agnostic subspace is learned by three guidance terms through a two-stage progressive training strategy. First, the subspace-guided term reduces the discrepancy between the domains by moving the source closer to the target subspace. Second, the data-guided term uses the coupled projections to map both domains to a unified subspace, where each target sample can be represented by the source samples with a low-rank coefficient matrix that can preserve the global structure of data. In this way, the data from both domains can be well interlaced and the domain-invariant features can be obtained. Third, for improving the discrimination of the subspaces, the label-guided term is constructed for prediction based on source labels and pseudo-target labels. To further improve the model tolerance to label noise, a label relaxation matrix is introduced. For the solver, a two-stage learning strategy with teacher teaches and student feedbacks mode is proposed to obtain the discriminative domain-agnostic subspace. In addition, for handling nonlinear domain shift, a nonlinear GSL (NGSL) framework is formulated with kernel embedding, such that the unified subspace is imposed with nonlinearity. Experiments on various cross-domain visual benchmark databases show that our methods outperform many state-of-the-art UDA methods. The source code is available at https://github.com/Fjr9516/GSL.
Lei Zhang 0038, Jingru Fu, Shanshan Wang 0008, David Zhang 0001, Zhao Yang Dong, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.4
2019 Learning a Visual Tracker from a Single Movie without Annotation
abstract
The recent success of deep network in visual trackers learning largely relies on human labeled data, which are however expensive to annotate. Recently, some unsupervised methods have been proposed to explore the learning of visual trackers without labeled data, while their performance lags far behind the supervised methods. We identify the main bottleneck of these methods as inconsistent objectives between off-line training and online tracking stages. To address this problem, we propose a novel unsupervised learning pipeline which is based on the discriminative correlation filter network. Our method iteratively updates the tracker by alternating between target localization and network optimization. In particular, we propose to learn the network from a single movie, which could be easily obtained other than collecting thousands of video clips or millions of images. Extensive experiments demonstrate that our approach is insensitive to the employed movies, and the trained visual tracker achieves leading performance among existing unsupervised learning approaches. Even compared with the same network trained with human labeled bounding boxes, our tracker achieves similar results on many tracking benchmarks. Code is available at: https://github.com/ZjjConan/UL-Tracker-AAAI2019.
Lingxiao Yang, David Zhang 0001, Lei Zhang 0006
AAAI2
2019 Robust Deep Softmax Regression Against Label Noise for Unsupervised Domain Adaptation
abstract
Domain adaptation aims to generalize the classification model from a source domain to a different but related target domain. Recent studies have revealed the benefit of deep convolutional features trained on a large dataset (e.g. ImageNet) in alleviating domain discrepancy. However, literatures show that the transferability of features decreases as (i) the difference between the source and target domains increases, or (ii) the layers are toward the top layers. Therefore, even with deep features, domain adaptation remains necessary. In this paper, we propose a novel unsupervised domain adaptation (UDA) model for deep neural networks, which is learned with the labeled source samples and the unlabeled target ones simultaneously. For target samples without labels, pseudo labels are assigned to them according to their maximum classification scores during training of the UDA model. However, due to the domain discrepancy, label noise generally is inevitable, which degrades the performance of the domain adaptation model. Thus, to effectively utilize the target samples, three specific robust deep softmax regression (RDSR) functions are performed for them with high, medium and low classification confidence respectively. Extensive experiments show that our method yields the state-of-the-art results, demonstrating the effectiveness of the robust deep softmax regression classifier in UDA.
Guangbin Wu, David Zhang 0001, Weishan Chen, Wangmeng Zuo, Zhuang Xia
Int. J. Pattern Recognit. Artif. Intell.2
2019 Body surface feature-based multi-modal Learning for Diabetes Mellitus detection
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Jane You, David Zhang 0001
Inf. Sci.5
2019 Joint learning for voice based disease detection
Kebin Wu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001
Pattern Recognit.2
2019 Sparse, collaborative, or nonnegative representation: Which helps pattern classification?
Jun Xu 0019, Wangpeng An, Lei Zhang 0006, David Zhang 0001
Pattern Recognit.4
2019 High resolution fingerprint recognition using pore and edge descriptors
Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, David Zhang 0001
Pattern Recognit. Lett.4
2019 Structurally Incoherent Low-Rank 2DLPP for Image Classification
abstract
Preserving projection-based methods are good for finding the manifold structure embedded in data. As they use the Euclidean distance as a metric, which is sensitive to noise and outliers in data, nuclear norm-based 2D locality preserving projection (NN-2DLPP) is thus proposed to improve the robustness of 2DLPP. However, NN-2DLPP does not consider the discriminant ability of data. In order to improve the discriminant ability of preserving projection methods, in this paper, we use preserving projection learning with structurally incoherence of data and propose structurally incoherent low-rank 2DLPP (SILR-2DLPP) for image classification. This approach provides a discriminative representation of preserving projection learning by recovering the distinct different classes of the data. SILR-2DLPP searches the optimal subspace and low-rank representation simultaneously. We further extend SILR-2DLPP to a kernel case and propose kernel SILR-2DLPP (KSILR-2DLPP) to obtain a nonlinear representation. The theoretical analysis including the convergence and computational complexity of SILR-2DLPP are presented. To verify the performance of SILR-2DLPP and KSILR-2DLPP, six well-known image databases were used in the experiments. The experimental results show that the proposed methods are superior to the previous preserving projection methods for image classification.
Yuwu Lu, Chun Yuan 0003, Xuelong Li 0001, Zhihui Lai 0001, David Zhang 0001, LinLin Shen
IEEE Trans. Circuits Syst. Video Technol.5
2019 Horizontal and Vertical Nuclear Norm-Based 2DLDA for Image Representation
abstract
2-D linear discriminant analysis (2DLDA) has been widely used in pattern recognition and image classification. 2DLDA selects discriminative features from the up and left corner of images. However, 2DLDA uses the Frobenius norm (F-norm), which is sensitive to noise or outliers in data, as a metric. In this paper, we propose a novel framework, called horizontal and vertical nuclear norm-based 2DLDA (HVNN-2DLDA) for image representation. In the proposed framework, HVNN-2DLDA methods (i.e., HNN-2DLDA and VNN-2DLDA) are proposed, and both use the nuclear norm as a criterion. The nuclear norm can provide more structure and global information for the reconstruction of noisy images. HNN-2DLDA and VNN-2DLDA represent images in the row and column directions, respectively. In addition, by combining the row and column directions, we propose a bilateral nuclear norm-based 2DLDA method called BNN-2DLDA. The advantage of BNN-2DLDA over HNN-2DLDA and VNN-2DLDA is that an image sample can be represented by both the row and the column directions instead of only the row or column direction. HVNN-2DLDA learns a set of local optimal projection vectors by maximizing the ratio of the nuclear norm of the between-class scatter matrix and the nuclear norm of the within-class scatter matrix. To verify the robustness and recognition performance in image classification of HVNN-2DLDA, six public image databases are used for experiments. The experimental results demonstrate the effectiveness and the feasibility of the proposed framework.
Yuwu Lu, Chun Yuan 0003, Zhihui Lai 0001, Xuelong Li 0001, David Zhang 0001, Wai Keung Wong
IEEE Trans. Circuits Syst. Video Technol.5
2019 Fingerprint Pore Comparison Using Local Features and Spatial Relations
abstract
High-resolution fingerprint recognition has been a hot topic for many years. Compared with a traditional fingerprint image, a high-resolution fingerprint image can provide more features, such as pores and ridge contours. Introducing these features into fingerprint comparison and recognition can improve the recognition accuracy and reduce the risk of identification errors. This paper proposes a novel method for comparing pores on high-resolution fingerprint images. The method can be divided into two steps. In the first step, fingerprints are aligned using the pixel-category-distance-based data-driven descending algorithm. Traditionally, fingerprints are aligned based on feature points, such as minutiae and singular points. Such alignment methods are not suitable when dealing with partial fingerprints because small overlapping areas often do not contain enough features to guarantee a correct alignment. In this research, the ridges and valleys on fingerprints are used in combination with the orientation field for alignment. The proposed algorithm performs well when aligning both partial and full fingerprints. The common areas between the two images can be estimated based on the alignment result. In the second step, pores lying in the common areas are selected for comparison. To improve the comparison accuracy, pores are compared using local features and spatial relations. A graph comparison algorithm is designed in this step. The experimental results show that the proposed method is more accurate than other state-of-the-art pore comparison algorithms.
Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, Feng Liu 0013, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 Visual Classification With Multikernel Shared Gaussian Process Latent Variable Model
abstract
Multiview learning methods often achieve improvement compared with single-view-based approaches in many applications. Due to the powerful nonlinear ability and probabilistic perspective of Gaussian process (GP), some GP-based multiview efforts were presented. However, most of these methods make a strong assumption on the kernel function (e.g., radial basis function), which limits the capacity of the real data modeling. In order to address this issue, in this paper, we propose a novel multiview approach by combining a multikernel and GP latent variable model. Instead of designing a deterministic kernel function, multiple kernel functions are established to automatically adapt various types of data. Considering a simple way of obtaining latent variables at the testing stage, a projection from the observed space to the latent space as a back constraint has also been simultaneously introduced into the proposed method. Additionally, different from some existing methods which apply the classifiers off-line, a hinge loss is embedded into the model to jointly learn the classification hyperplane, encouraging the latent variables belonging to the different classes to be separated. An efficient algorithm based on the gradient decent technique is constructed to optimize our method. Finally, we apply the proposed approach to three real-world datasets and the associated results demonstrate the effectiveness and superiority of our model compared with other state-of-the-art methods.
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Hu Ren, David Zhang 0001
IEEE Trans. Cybern.5
2019 Low-Rank 2-D Neighborhood Preserving Projection for Enhanced Robust Image Representation
abstract
2-D neighborhood preserving projection (2DNPP) uses 2-D images as feature input instead of 1-D vectors used by neighborhood preserving projection (NPP). 2DNPP requires less computation time than NPP. However, both NPP and 2DNPP use the L2norm as a metric, which is sensitive to noise in data. In this paper, we proposed a novel NPP method called low-rank 2DNPP (LR-2DNPP). This method divided the input data into a component part that encoded low-rank features, and an error part that ensured the noise was sparse. Then, a nearest neighbor graph was learned from the clean data using the same procedure as 2DNPP. To ensure that the features learned by LR-2DNPP were optimal for classification, we combined the structurally incoherent learning and low-rank learning with NPP to form a unified model called discriminative LR-2DNPP (DLR2DNPP). By encoding the structural incoherence of the learned clean data, DLR-2DNPP could enhance the discriminative ability for feature extraction. Theoretical analyses on the convergence and computational complexity of LR-2DNPP and DLR-2DNPP were presented in details. We used seven public image databases to verify the performance of the proposed methods. The experimental results showed the effectiveness of our methods for robust image representation.
Yuwu Lu, Zhihui Lai 0001, Xuelong Li 0001, Wai Keung Wong, Chun Yuan 0003, David Zhang 0001
IEEE Trans. Cybern.6
2019 Manifold Criterion Guided Transfer Learning via Intermediate Domain Generation
abstract
In many practical transfer learning scenarios, the feature distribution is different across the source and target domains (i.e., nonindependent identical distribution). Maximum mean discrepancy (MMD), as a domain discrepancy metric, has achieved promising performance in unsupervised domain adaptation (DA). We argue that the MMD-based DA methods ignore the data locality structure, which, up to some extent, would cause the negative transfer effect. The locality plays an important role in minimizing the nonlinear local domain discrepancy underlying the marginal distributions. For better exploiting the domain locality, a novel local generative discrepancy metric-based intermediate domain generation learning called Manifold Criterion guided Transfer Learning (MCTL) is proposed in this paper. The merits of the proposed MCTL are fourfold: 1) the concept of manifold criterion (MC) is first proposed as a measure validating the distribution matching across domains, and DA is achieved if the MC is satisfied; 2) the proposed MC can well guide the generation of the intermediate domain sharing similar distribution with the target domain, by minimizing the local domain discrepancy; 3) a global generative discrepancy metric is presented, such that both the global and local discrepancies can be effectively and positively reduced; and 4) a simplified version of MCTL called MCTL-S is presented under a perfect domain generation assumption for more generic learning scenario. Experiments on a number of benchmark visual transfer tasks demonstrate the superiority of the proposed MC guided generative transfer method, by comparing with the other state-of-the-art methods. The source code is available in https://github.com/wangshanshanCQU/MCTL.
Lei Zhang 0038, Shanshan Wang 0008, Guang-Bin Huang, Wangmeng Zuo, Jian Yang 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2019 Feature Extraction Methods for Palmprint Recognition: A Survey and Evaluation
abstract
Palmprint processes a number of unique features for reliable personal recognition. However, different types of palmprint images contain different dominant features. Instead, only some features of the palmprint are visible in a palmprint image, whereas the other features may not be notable. For example, the low-resolution palmprint image has visible principal lines and wrinkles. By contrast, the high-resolution palmprint image contains clear ridge patterns and minutiae points. In addition, the three dimensional (3-D) palmprint image possesses curvatures of the palmprint surface. So far, there is no work to summarize the feature extraction of different types of palmprint images. In this paper, we have an aim to completely study the feature extraction and recognition of palmprint. We propose to use a unified framework to classify palmprint images into four categories: (1) the contact-based; (2) contactless; (3) high-resolution; and (4) 3-D palmprint images. Then, we analyze the motivations and theories of the representative extraction and matching methods for different types of palmprint images. Finally, we compare and test the state-of-the-art methods via the widely used palmprint databases, and point out some potential directions for future research.
Lunke Fei, Guangming Lu 0002, Wei Jia 0001, Shaohua Teng, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2018 A Probabilistic Hierarchical Model for Multi-View and Multi-Feature Classification
abstract
Some recent works in classification show that the data obtained from various views with different sensors for an object contributes to achieving a remarkable performance. Actually, in many real-world applications, each view often contains multiple features, which means that this type of data has a hierarchical structure, while most of existing works do not take these features with multi-layer structure into consideration simultaneously. In this paper, a probabilistic hierarchical model is proposed to address this issue and applied for classification. In our model, a latent variable is first learned to fuse the multiple features obtained from a same view, sensor or modality. Particularly, mapping matrices corresponding to a certain view are estimated to project the latent variable from a shared space to the multiple observations. Since this method is designed for the supervised purpose, we assume that the latent variables associated with different views are influenced by their ground-truth label. In order to effectively solve the proposed method, the Expectation-Maximization (EM) algorithm is applied to estimate the parameters and latent variables. Experimental results on the extensive synthetic and two real-world datasets substantiate the effectiveness and superiority of our approach as compared with state-of-the-art.
Jinxing Li 0003, Hongwei Yong, Bob Zhang 0001, Mu Li 0005, Lei Zhang 0006, David Zhang 0001
AAAI6
2018 Learning Convolutional Networks for Content-Weighted Image Compression
abstract
Lossy image compression is generally formulated as a joint rate-distortion optimization problem to learn encoder, quantizer, and decoder. Due to the non-differentiable quantizer and discrete entropy estimation, it is very challenging to develop a convolutional network (CNN)-based image compression system. In this paper, motivated by that the local information content is spatially variant in an image, we suggest that: (i) the bit rate of the different parts of the image is adapted to local content, and (ii) the content-aware bit rate is allocated under the guidance of a content-weighted importance map. The sum of the importance map can thus serve as a continuous alternative of discrete entropy estimation to control compression rate. The binarizer is adopted to quantize the output of encoder and a proxy function is introduced for approximating binary operation in backward propagation to make it differentiable. The encoder, decoder, binarizer and importance map can be jointly optimized in an end-to-end manner. And a convolutional entropy encoder is further presented for lossless compression of importance map and binary codes. In low bit rate image compression, experiments show that our system significantly outperforms JPEG and JPEG 2000 by structural similarity (SSIM) index, and can produce the much better visual result with sharp edges, rich textures, and fewer artifacts.
Mu Li 0005, Wangmeng Zuo, Shuhang Gu, Debin Zhao, David Zhang 0001
CVPR5
2018 A Hybrid l1-l0 Layer Decomposition Model for Tone Mapping
abstract
Tone mapping aims to reproduce a standard dynamic range image from a high dynamic range image with visual information preserved. State-of-the-art tone mapping algorithms mostly decompose an image into a base layer and a detail layer, and process them accordingly. These methods may have problems of halo artifacts and over-enhancement, due to the lack of proper priors imposed on the two layers. In this paper, we propose a hybrid ℓ1-ℓ0decomposition model to address these problems. Specifically, an ℓ1sparsity term is imposed on the base layer to model its piecewise smoothness property. An ℓ0sparsity term is imposed on the detail layer as a structural prior, which leads to piecewise constant effect. We further propose a multiscale tone mapping scheme based on our layer decomposition model. Experiments show that our tone mapping algorithm achieves visually compelling results with little halo artifacts, outperforming the state-of-the-art tone mapping algorithms in both subjective and objective evaluations.
Zhetong Liang, Jun Xu 0019, David Zhang 0001, Zisheng Cao, Lei Zhang 0006
CVPR3
2018 A Trilateral Weighted Sparse Coding Scheme for Real-World Image Denoising
Jun Xu 0019, Lei Zhang 0006, David Zhang 0001
ECCV (8)3
2018 Shared Linear Encoder-based Gaussian Process Latent Variable Model for Visual Classification
abstract
Multi-view learning has shown its powerful potential in many applications and achieved outstanding performances compared with the single-view based methods. In this paper, we propose a novel multi-view learning model based on the Gaussian Process Latent Variable Model (GPLVM) to learn a shared latent variable in the manifold space with a linear and gaussian process prior based back projection. Different from existing GPLVM methods which only consider a mapping from the latent space to the observed space, the proposed method simultaneously takes a back projection from the observation to the latent variable into account. Concretely, due to the various dimensions of different views, a projection for each view is first learned to linearly map its observation to a subspace. The gaussian process prior is then imposed on another transformation to non-linearly and efficiently map the learned subspace to a shared manifold space. In order to apply the proposed approach to the classification, a discriminative regularization is also embedded to exploit the label information. Experimental results on three real-world databases substantiate the effectiveness and superiority of the proposed approach as compared with several state-of-the-art approaches.
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, David Zhang 0001
ACM Multimedia4
2018 Learning acoustic features to detect Parkinson's disease
Kebin Wu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001
Neurocomputing2
2018 Two-phase linear reconstruction measure-based classification for face recognition
Jianping Gou, Yong Xu 0001, David Zhang 0001, Qirong Mao, Lan Du 0002, Yongzhao Zhan 0001
Inf. Sci.3
2018 Maximal granularity structure and generalized multi-view discriminant analysis for person re-identification
Cairong Zhao, Xuekuan Wang, Duoqian Miao 0001, Hanli Wang, Wei-Shi Zheng 0001, Yong Xu 0001, David Zhang 0001
Pattern Recognit.7
2018 Data-Driven Facial Beauty Analysis: Prediction, Retrieval and Manipulation
abstract
Facial beauty analysis becomes an emerging research area due to many potential applications, such as aesthetic surgery plan, cosmetic industry, photo retouching, and entertainment. In this paper, we propose a data-driven facial beauty analysis framework that contains three application modules: prediction, retrieval, and manipulation. A beauty model is the core of the framework. With carefully designed features, the model can be built for different purposes. For prediction, we combine several low-level face representations and high-level features to form a feature vector and perform feature selection to optimize the feature set. The model built with the optimized feature set outperforms state-of-the-art methods. Then, we discuss two scenarios of beauty-oriented face retrieval: for recommendation and for beautification. Finally, we propose two approaches for facial beauty manipulation. One is an exemplar-based approach that uses the retrieved results. The other is a model-based approach that modifies facial features along the gradient of the beauty model. In this case, the model is built with the shape or appearance feature. Experimental results show that the exemplar-based approach is better for shape beautification; the model-based approach is suitable for texture beautification; and the combination of them can increase the attractiveness of a query face robustly.
Fangmei Chen, Xihua Xiao, David Zhang 0001
IEEE Trans. Affect. Comput.3
2018 Learning Parts-Based and Global Representation for Image Classification
abstract
Nonnegative matrix factorization (NMF), known as a famous matrix factorization technique, has been widely used in pattern recognition and computer vision. NMF represents the input data matrix as a product of two nonnegative factors. As NMF is based on the Euclidean distance, which is sensitive to noise or errors in the data, some robust NMF methods are proposed. Mainly focusing on parts-based representation, these robust NMF methods often neglect global representation of data. In fact, the global geometry information of data is more robust than the local information about the noisy data in terms of image classification. In order to effectively improve the robustness of NMF and learn part-based and global representation of the data, a novel method low-rank nonnegative factorization (LRNF) is proposed in this paper. First, we assume that the data are grossly corrupted, and the$L_{1} $norm is used as a sparse constraint on the assumed noise matrix. Then, LRNF learns a low-rank matrix with the global representation ability. Finally, we make a nonnegative factorization of the learned low-rank matrix. We can obtain a base matrix, which preserves locality and globality properties of the data in the meantime. Extensive experiments have been conducted on nine real-world image databases to verify the performance of the proposed LRNF method by comparing with the state-of-the-art algorithms on robust dimensionality reduction.
Yuwu Lu, Zhihui Lai 0001, Xuelong Li 0001, David Zhang 0001, Wai Keung Wong, Chun Yuan 0003
IEEE Trans. Circuits Syst. Video Technol.4
2018 Robust Discriminant Regression for Feature Extraction
abstract
Ridge regression (RR) and its extended versions are widely used as an effective feature extraction method in pattern recognition. However, the RR-based methods are sensitive to the variations of data and can learn only limited number of projections for feature extraction and recognition. To address these problems, we propose a new method called robust discriminant regression (RDR) for feature extraction. In order to enhance the robustness, the L2,1-norm is used as the basic metric in the proposed RDR. The designed robust objective function in regression form can be solved by an iterative algorithm containing an eigenfunction, through which the optimal orthogonal projections of RDR can be obtained by eigen decomposition. The convergence analysis and computational complexity are presented. In addition, we also explore the intrinsic connections and differences between the RDR and some previous methods. Experiments on some well-known databases show that RDR is superior to the classical and very recent proposed methods reported in the literature, no matter the L2-norm or the L2,1-norm-based regression methods. The code of this paper can be downloaded from http://www.scholat.com/laizhihui.
Zhihui Lai 0001, Dongmei Mo, Wai Keung Wong, Yong Xu 0001, Duoqian Miao 0001, David Zhang 0001
IEEE Trans. Cybern.6
2018 Learning Domain-Invariant Subspace Using Domain Features and Independence Maximization
abstract
Domain adaptation algorithms are useful when the distributions of the training and the test data are different. In this paper, we focus on the problem of instrumental variation and time-varying drift in the field of sensors and measurement, which can be viewed as discrete and continuous distributional change in the feature space. We propose maximum independence domain adaptation (MIDA) and semi-supervised MIDA to address this problem. Domain features are first defined to describe the background information of a sample, such as the device label and acquisition time. Then, MIDA learns a subspace which has maximum independence with the domain features, so as to reduce the interdomain discrepancy in distributions. A feature augmentation strategy is also designed to project samples according to their backgrounds so as to improve the adaptation. The proposed algorithms are flexible and fast. Their effectiveness is verified by experiments on synthetic datasets and four real-world ones on sensors, measurement, and computer vision. They can greatly enhance the practicability of sensor systems, as well as extend the application scope of existing domain adaptation algorithms by uniformly handling different kinds of distributional change.
Ke Yan 0003, Lu Kou, David Zhang 0001
IEEE Trans. Cybern.3
2018 Partial Deconvolution With Inaccurate Blur Kernel
abstract
Most non-blind deconvolution methods are developed under the error-free kernel assumption, and are not robust to inaccurate blur kernel. Unfortunately, despite the great progress in blind deconvolution, estimation error remains inevitable during blur kernel estimation. Consequently, severe artifacts such as ringing effects and distortions are likely to be introduced in the non-blind deconvolution stage. In this paper, we tackle this issue by suggesting: 1) a partial map in the Fourier domain for modeling kernel estimation error, and 2) a partial deconvolution model for robust deblurring with inaccurate blur kernel. The partial map is constructed by detecting the reliable Fourier entries of estimated blur kernel. And partial deconvolution is applied to wavelet-based and learning-based models to suppress the adverse effect of kernel estimation error. Furthermore, an E-M algorithm is developed for estimating the partial map and recovering the latent sharp image alternatively. Experimental results show that our partial deconvolution model is effective in relieving artifacts caused by inaccurate blur kernel, and can achieve favorable deblurring quality on synthetic and real blurry images.
Dongwei Ren, Wangmeng Zuo, David Zhang 0001, Jun Xu 0019, Lei Zhang 0006
IEEE Trans. Image Process.3
2018 External Prior Guided Internal Prior Learning for Real-World Noisy Image Denoising
abstract
Most of existing image denoising methods learn image priors from either external data or the noisy image itself to remove noise. However, priors learned from external data may not be adaptive to the image to be denoised, while priors learned from the given noisy image may not be accurate due to the interference of corrupted noise. Meanwhile, the noise in real-world noisy images is very complex, which is hard to be described by simple distributions such as Gaussian distribution, making real-world noisy image denoising a very challenging problem. We propose to exploit the information in both external data and the given noisy image, and develop an external prior guided internal prior learning method for real-world noisy image denoising. We first learn external priors from an independent set of clean natural images. With the aid of learned external priors, we then learn internal priors from the given noisy image to refine the prior model. The external and internal priors are formulated as a set of orthogonal dictionaries to efficiently reconstruct the desired image. Extensive experiments are performed on several real-world noisy image datasets. The proposed method demonstrates highly competitive denoising performance, outperforming state-of-the-art denoising methods including those designed for real-world noisy images.
Jun Xu 0019, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Image Process.3
2018 Shared Autoencoder Gaussian Process Latent Variable Model for Visual Classification
abstract
Multiview learning reveals the latent correlation among different modalities and utilizes the complementary information to achieve a better performance in many applications. In this paper, we propose a novel multiview learning model based on the Gaussian process latent variable model (GPLVM) to learn a set of nonlinear and nonparametric mapping functions and obtain a shared latent variable in the manifold space. Different from the previous work on the GPLVM, the proposed shared autoencoder Gaussian process (SAGP) latent variable model assumes that there is an additional mapping from the observed data to the shared manifold space. Due to the introduction of the autoencoder framework, both nonlinear projections from and to the observation are considered simultaneously. Additionally, instead of fully connecting used in the conventional autoencoder, the SAGP achieves the mappings utilizing the GP, which remarkably reduces the number of estimated parameters and avoids the phenomenon of overfitting. To make the proposed method adaptive for classification, a discriminative regularization is embedded into the proposed method. In the optimization process, an efficient algorithm based on the alternating direction method and gradient decent techniques is designed to solve the encoder and decoder parts alternatively. Experimental results on three real-world data sets substantiate the effectiveness and superiority of the proposed approach as compared with the state of the art.
Jinxing Li 0003, Bob Zhang 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2018 F-SVM: Combination of Feature Transformation and SVM Learning via Convex Relaxation
abstract
The generalization error bound of the support vector machine (SVM) depends on the ratio of the radius and margin. However, conventional SVM only considers the maximization of the margin but ignores the minimization of the radius, which restricts its performance when applied to joint learning of feature transformation and the SVM classifier. Although several approaches have been proposed to integrate the radius and margin information, most of them either require the form of the transformation matrix to be diagonal, or are nonconvex and computationally expensive. In this paper, we suggest a novel approximation for the radius of the minimum enclosing ball in feature space, and then propose a convex radius-margin-based SVM model for joint learning of feature transformation and the SVM classifier, i.e., F-SVM. A generalized block coordinate descent method is adopted to solve the F-SVM model, where the feature transformation is updated via the gradient descent and the classifier is updated by employing the existing SVM solver. By incorporating with kernel principal component analysis, F-SVM is further extended for joint learning of nonlinear transformation and the classifier. F-SVM can also be incorporated with deep convolutional networks to improve image classification performance. Experiments on the UCI, LFW, MNIST, CIFAR-10, CIFAR-100, and Caltech101 data sets demonstrate the effectiveness of F-SVM.
Xiaohe Wu, Wangmeng Zuo, Liang Lin 0004, Wei Jia 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2018 Discriminative and Robust Competitive Code for Palmprint Recognition
abstract
Various palmprint recognition methods have been proposed based on orientation features of palmprints. Among them, the competitive code method using the dominant orientation of palmprint images achieves promising performance in palmprint recognition. In this paper, we propose a discriminative and robust competitive code based method, which uses a more accurate dominant orientation representation of palmprint images for palmprint authentication. Moreover, we propose to weight the orientation information of a neighbor area to improve the precision and stability of the discriminative and robust dominant orientation code. Experiments performed on three types of palmprint databases and a noisy dataset validate the effectiveness of the proposed method.
Yong Xu 0001, Lunke Fei, Jie Wen 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2018 Efficient Solutions for Discreteness, Drift, and Disturbance (3D) in Electronic Olfaction
abstract
In this paper, we aim at presenting the new challenges of electronic noses (E-noses) and proposing effective methods for handling the new challenging scientific issues to be solved, such as signal discreteness (reproducibility), systematical drift and nontarget disturbances. We first review the progress of E-noses in applications, systems, and algorithms during the past two decades. Recall a number of significant achievements and motivated by the current issues that hinder large-scale application pace of E-nose technology, we propose to address three key issues: 1) discreteness; 2) drift; and 3) disturbance (simplified as 3D issues), which are sensor induced and sensor specific. For each issue, a highly effective and efficient method is proposed. Specifically, for discreteness issue, a global affine transformation method is introduced for E-nose instruments batch calibration; for drift issue, an unsupervised feature adaptation model is proposed to achieve effective drift adaptation; additionally, for disturbance issue, we proposed a simple targets-to-targets self-representation classifier method for fast nontargets detection, without knowing any prior knowledge of thousands of nontarget disturbances in real world. For each method, a closed form solution can be analytically determined and the simplicity is guaranteed. Experiments demonstrate the effectiveness and efficiency of the proposed methods for addressing the proposed 3D issues in real applications of E-noses.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Multi-channel Weighted Nuclear Norm Minimization for Real Color Image Denoising
abstract
Most of the existing denoising algorithms are developed for grayscale images. It is not trivial to extend them for color image denoising since the noise statistics in R, G, and B channels can be very different for real noisy images. In this paper, we propose a multi-channel (MC) optimization model for real color image denoising under the weighted nuclear norm minimization (WNNM) framework. We concatenate the RGB patches to make use of the channel redundancy, and introduce a weight matrix to balance the data fidelity of the three channels in consideration of their different noise statistics. The proposed MC-WNNM model does not have an analytical solution. We reformulate it into a linear equality-constrained problem and solve it via alternating direction method of multipliers. Each alternative updating step has a closed-form solution and the convergence can be guaranteed. Experiments on both synthetic and real noisy image datasets demonstrate the superiority of the proposed MC-WNNM over state-of-the-art denoising methods.
Jun Xu 0019, Lei Zhang 0006, David Zhang 0001, Xiangchu Feng
ICCV3
2017 Part-based convolutional neural network for visual recognition
abstract
Mid-level element based representations have been proven to be very effective for visual recognition. We present a method to discover discriminative elements based on deep Convolutional Neural Networks (CNNs), namely Part-based CNN (P-CNN), which acts as the role of encoding module in part-based representation. The P-CNN can be attached at arbitrary layer of a pre-trained CNN and be trained using image-level labels. The training of P-CNN essentially corresponds to the optimization and selection of discriminative mid-level visual elements. For an input image, the output of P-CNN is naturally the part-based coding and can be directly used for image recognition. By applying P-CNN to multiple layers of a pretrained CNN, more diverse visual elements can be obtained for visual recognitions. Experiments are conducted on two recognition tasks and their results demonstrate the effectiveness of the proposed method.
Lingxiao Yang, Xiaohua Xie, Peihua Li, David Zhang 0001, Lei Zhang 0006
ICIP4
2017 Learning a real-time generic tracker using convolutional neural networks
abstract
This paper presents a novel frame-pair based method for visual object tracking. Instead of adopting two-stream Convolutional Neural Networks (CNNs) to represent each frame, we stack frame pairs as the input, resulting in a single-stream CNN tracker with much fewer parameters. The proposed tracker can learn generic motion patterns of objects with much less annotated videos than previous methods. Besides, it is found that trackers trained using two successive frames tend to predict the centers of searching windows as the locations of tracked targets. To alleviate this problem, we propose a novel sampling strategy for off-line training. Specifically, we construct a pair by sampling two frames with a random offset. The offset controls the moving smoothness of objects. Experiments on the challenging VOT14 and OTB datasets show that the proposed tracker performs on par with recently developed generic trackers, but with much less memory. In addition, our tracker can run in a speed of over 100 (30) fps with a GPU (CPU), much faster than most deep neural network based trackers.
Linnan Zhu, Lingxiao Yang, David Zhang 0001, Lei Zhang 0006
ICME3
2017 Deep Location-Specific Tracking
abstract
Convolutional Neural Network (CNN) based methods have shown significant performance gains in the problem of visual tracking in recent years. Due to many uncertain changes of objects online, such as abrupt motion, background clutter and large deformation, the visual tracking is still a challenging task. We propose a novel algorithm, namely Deep Location-Specific Tracking, which decomposes the tracking problem into a localization task and a classification task, and trains an individual network for each task. The localization network exploits the information in the current frame and provides a specific location to improve the probability of successful tracking, while the classification network finds the target among many examples generated around the target location in the previous frame, as well as the one estimated from the localization network in the current frame. CNN based trackers often have massive number of trainable parameters, and are prone to over-fitting to some particular object states, leading to less precision or tracking drift. We address this problem by learning a classification network based on 1 × 1 convolution and global average pooling. Extensive experimental results on popular benchmark datasets show that the proposed tracker achieves competitive results without using additional tracking videos for fine-tuning. The code is available at https://github.com/ZjjConan/DLST
Lingxiao Yang, Risheng Liu, David Zhang 0001, Lei Zhang 0006
ACM Multimedia3
2017 Unsupervised Domain Adaptation with Robust Deep Logistic Regression
Guangbin Wu, Weishan Chen, Wangmeng Zuo, David Zhang 0001
PSIVT4
2017 Joint discriminative and collaborative representation for fatty liver disease diagnosis
Jinxing Li 0003, Bob Zhang 0001, David Zhang 0001
Expert Syst. Appl.3
2017 Facial beauty analysis based on geometric feature: Toward attractiveness assessment application
Lei Zhang 0038, David Zhang 0001, Mingming Sun 0006, Fangmei Chen
Expert Syst. Appl.2
2017 Joint distance and similarity measure learning based on triplet-based constraints
Mu Li 0005, Qilong Wang 0001, David Zhang 0001, Peihua Li, Wangmeng Zuo
Inf. Sci.3
2017 Joint similar and specific learning for diabetes mellitus and impaired glucose regulation detection
Jinxing Li 0003, David Zhang 0001, Bob Zhang 0001
Inf. Sci.2
2017 Domain class consistency based transfer learning for image classification across domains
Lei Zhang 0038, Jian Yang 0003, David Zhang 0001
Inf. Sci.3
2017 Improving texture analysis performance in biometrics by adjusting image sharpness
Kunai Zhang, Bob Zhang 0001, David Zhang 0001
Pattern Recognit.4
2017 3D palmprint identification combining blocked ST and PCA
Nan Gao 0002, Zonghua Zhang, David Zhang 0001
Pattern Recognit. Lett.4
2017 An optimized palmprint recognition approach based on image sharpness
Kunai Zhang, David Zhang 0001
Pattern Recognit. Lett.3
2017 Nonnegative Discriminant Matrix Factorization
abstract
Nonnegative matrix factorization (NMF), which aims at obtaining the nonnegative low-dimensional representation of data, has received wide attention. To obtain more effective nonnegative discriminant bases from the original NMF, in this paper, a novel method called nonnegative discriminant matrix factorization (NDMF) is proposed for image classification. NDMF integrates the nonnegative constraint, orthogonality, and discriminant information in the objective function. NDMF considers the incoherent information of both factors in standard NMF and is proposed to enhance the discriminant ability of the learned base matrix. NDMF projects the low-dimensional representation of the subspace of the base matrix to regularize the NMF for discriminant subspace learning. Based on the Euclidean distance metric and the generalized Kullback-Leibler (KL) divergence, two kinds of iterative algorithms are presented to solve the optimization problem. The between- and within-class scatter matrices are divided into positive and negative parts for the update rules and the proofs of the convergence are also presented. Extensive experimental results demonstrate the effectiveness of the proposed method in comparison with the state-of-the-art discriminant NMF algorithms.
Yuwu Lu, Zhihui Lai 0001, Yong Xu 0001, Xuelong Li 0001, David Zhang 0001, Chun Yuan 0003
IEEE Trans. Circuits Syst. Video Technol.5
2017 Rotational Invariant Dimensionality Reduction Algorithms
abstract
A common intrinsic limitation of the traditional subspace learning methods is the sensitivity to the outliers and the image variations of the object since they use the norm as the metric. In this paper, a series of methods based on the -norm are proposed for linear dimensionality reduction. Since the -norm based objective function is robust to the image variations, the proposed algorithms can perform robust image feature extraction for classification. We use different ideas to design different algorithms and obtain a unified rotational invariant (RI) dimensionality reduction framework, which extends the well-known graph embedding algorithm framework to a more generalized form. We provide the comprehensive analyses to show the essential properties of the proposed algorithm framework. This paper indicates that the optimization problems have global optimal solutions when all the orthogonal projections of the data space are computed and used. Experimental results on popular image datasets indicate that the proposed RI dimensionality reduction algorithms can obtain competitive performance compared with the previous norm based subspace learning algorithms.
Zhihui Lai 0001, Yong Xu 0001, Jian Yang 0003, LinLin Shen, David Zhang 0001
IEEE Trans. Cybern.5
2017 Distance Metric Learning via Iterated Support Vector Machines
abstract
Distance metric learning aims to learn from the given training data a valid distance metric, with which the similarity between data samples can be more effectively evaluated for classification. Metric learning is often formulated as a convex or nonconvex optimization problem, while most existing methods are based on customized optimizers and become inefficient for large scale problems. In this paper, we formulate metric learning as a kernel classification problem with the positive semi-definite constraint, and solve it by iterated training of support vector machines (SVMs). The new formulation is easy to implement and efficient in training with the off-the-shelf SVM solvers. Two novel metric learning models, namely positive-semidefinite constrained metric learning (PCML) and nonnegative-coefficient constrained metric learning (NCML), are developed. Both PCML and NCML can guarantee the global optimality of their solutions. Experiments are conducted on general classification, face verification, and person re-identification to evaluate our methods. Compared with the state-of-the-art approaches, our methods can achieve comparable classification accuracy and are efficient in training.
Wangmeng Zuo, David Zhang 0001, Liang Lin 0004, Yuchi Huang, Deyu Meng, Lei Zhang 0006
IEEE Trans. Image Process.3
2017 Generalized Feature Extraction for Wrist Pulse Analysis: From 1-D Time Series to 2-D Matrix
abstract
Traditional Chinese pulse diagnosis, known as an empirical science, depends on the subjective experience. Inconsistent diagnostic results may be obtained among different practitioners. A scientific way of studying the pulse should be to analyze the objectified wrist pulse waveforms. In recent years, many pulse acquisition platforms have been developed with the advances in sensor and computer technology. And the pulse diagnosis using pattern recognition theories is also increasingly attracting attentions. Though many literatures on pulse feature extraction have been published, they just handle the pulse signals as simple 1-D time series and ignore the information within the class. This paper presents a generalized method of pulse feature extraction, extending the feature dimension from 1-D time series to 2-D matrix. The conventional wrist pulse features correspond to a particular case of the generalized models. The proposed method is validated through pattern classification on actual pulse records. Both quantitative and qualitative results relative to the 1-D pulse features are given through diabetes diagnosis. The experimental results show that the generalized 2-D matrix feature is effective in extracting both the periodic and nonperiodic information. And it is practical for wrist pulse analysis.
Dimin Wang, David Zhang 0001, Guangming Lu 0002
IEEE J. Biomed. Health Informatics2
2017 Nuclear Norm-Based 2DLPP for Image Classification
abstract
Two-dimensional locality preserving projections (2DLPP) that use 2D image representation in preserving projection learning can preserve the intrinsic manifold structure and local information of data. However, 2DLPP is based on the Euclidean distance, which is sensitive to noise and outliers in data. In this paper, we propose a novel locality preserving projection method called nuclear norm-based two-dimensional locality preserving projections (NN-2DLPP). First, NN-2DLPP recovers the noisy data matrix through low-rank learning. Second, noise in data is removed and the learned clean data points are projected on a new subspace. Without the disturbance of noise, data points belonging to the same class are kept as close to each other as possible in the new projective subspace. Experimental results on six public image databases with face recognition, object classification, and handwritten digit recognition tasks demonstrated the effectiveness of the proposed method.
Yuwu Lu, Chun Yuan 0003, Zhihui Lai 0001, Xuelong Li 0001, Wai Keung Wong, David Zhang 0001
IEEE Trans. Multim.6
2017 A Locality-Constrained and Label Embedding Dictionary Learning Algorithm for Image Classification
abstract
Locality and label information of training samples play an important role in image classification. However, previous dictionary learning algorithms do not take the locality and label information of atoms into account together in the learning process, and thus their performance is limited. In this paper, a discriminative dictionary learning algorithm, called the locality-constrained and label embedding dictionary learning (LCLE-DL) algorithm, was proposed for image classification. First, the locality information was preserved using the graph Laplacian matrix of the learned dictionary instead of the conventional one derived from the training samples. Then, the label embedding term was constructed using the label information of atoms instead of the classification error term, which contained discriminating information of the learned dictionary. The optimal coding coefficients derived by the locality-based and label-based reconstruction were effective for image classification. Experimental results demonstrated that the LCLE-DL algorithm can achieve better performance than some state-of-the-art algorithms.
Zhihui Lai 0001, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2017 A New Discriminative Sparse Representation Method for Robust Face Recognition via l2 Regularization
abstract
Sparse representation has shown an attractive performance in a number of applications. However, the available sparse representation methods still suffer from some problems, and it is necessary to design more efficient methods. Particularly, to design a computationally inexpensive, easily solvable, and robust sparse representation method is a significant task. In this paper, we explore the issue of designing the simple, robust, and powerfully efficient sparse representation methods for image classification. The contributions of this paper are as follows. First, a novel discriminative sparse representation method is proposed and its noticeable performance in image classification is demonstrated by the experimental results. More importantly, the proposed method outperforms the existing state-of-the-art sparse representation methods. Second, the proposed method is not only very computationally efficient but also has an intuitive and easily understandable idea. It exploits a simple algorithm to obtain a closed-form solution and discriminative representation of the test sample. Third, the feasibility, computational efficiency, and remarkable classification accuracy of the proposed l₂ regularization-based representation are comprehensively shown by extensive experiments and analysis. The code of the proposed method is available at http://www.yongxu.org/lunwen.html.
Yong Xu 0001, Zuofeng Zhong, Jian Yang 0003, Jane You, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2017 Evolutionary Cost-Sensitive Extreme Learning Machine
abstract
Conventional extreme learning machines (ELMs) solve a Moore-Penrose generalized inverse of hidden layer activated matrix and analytically determine the output weights to achieve generalized performance, by assuming the same loss from different types of misclassification. The assumption may not hold in cost-sensitive recognition tasks, such as face recognition-based access control system, where misclassifying a stranger as a family member may result in more serious disaster than misclassifying a family member as a stranger. Though recent cost-sensitive learning can reduce the total loss with a given cost matrix that quantifies how severe one type of mistake against another, in many realistic cases, the cost matrix is unknown to users. Motivated by these concerns, this paper proposes an evolutionary cost-sensitive ELM, with the following merits: 1) to the best of our knowledge, it is the first proposal of ELM in evolutionary cost-sensitive classification scenario; 2) it well addresses the open issue of how to define the cost matrix in cost-sensitive learning tasks; and 3) an evolutionary backtracking search algorithm is induced for adaptive cost matrix optimization. Experiments in a variety of cost-sensitive tasks well demonstrate the effectiveness of the proposed approaches, with about 5%-10% improvements.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2017 Door Knob Hand Recognition System
abstract
Biometric applications have been used globally in everyday life. However, conventional biometrics is created and optimized for high-security scenarios. Being used in daily life by ordinary untrained people is a new challenge. Facing this challenge, designing a biometric system with prior constraints of ergonomics, we propose ergonomic biometrics design model, which attains the physiological factors, the psychological factors, and the conventional security characteristics. With this model, a novel hand-based biometric system, door knob hand recognition system (DKHRS), is proposed. DKHRS has the identical appearance of a conventional door knob, which is an optimum solution in both physiological factors and psychological factors. In this system, a hand image is captured by door knob imaging scheme, which is a tailored omnivision imaging structure and is optimized for this predetermined door knob appearance. Then features are extracted by local Gabor binary pattern histogram sequence method and classified by projective dictionary pair learning. In the experiment on a large data set including 12 000 images from 200 people, the proposed system achieves competitive recognition performance comparing with conventional biometrics like face and fingerprint recognition systems, with an equal error rate of 0.091%. This paper shows that a biometric system could be built with a reliable recognition performance under the ergonomic constraints.
Xiaofeng Qu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2016 Joint Learning of Single-Image and Cross-Image Representations for Person Re-identification
abstract
Person re-identification has been usually solved as either the matching of single-image representation (SIR) or the classification of cross-image representation (CIR). In this work, we exploit the connection between these two categories of methods, and propose a joint learning frame-work to unify SIR and CIR using convolutional neural network (CNN). Specifically, our deep architecture contains one shared sub-network together with two sub-networks that extract the SIRs of given images and the CIRs of given image pairs, respectively. The SIR sub-network is required to be computed once for each image (in both the probe and gallery sets), and the depth of the CIR sub-network is required to be minimal to reduce computational burden. Therefore, the two types of representation can be jointly optimized for pursuing better matching accuracy with moderate computational cost. Furthermore, the representations learned with pairwise comparison and triplet comparison objectives can be combined to improve matching performance. Experiments on the CUHK03, CUHK01 and VIPeR datasets show that the proposed method can achieve favorable accuracy while compared with state-of-the-arts.
Wangmeng Zuo, Liang Lin 0004, David Zhang 0001, Lei Zhang 0006
CVPR4
2016 iPEEH: Improving pitch estimation by enhancing harmonics
Kebin Wu, David Zhang 0001, Guangming Lu 0002
Expert Syst. Appl.2
2016 Combining a causal effect criterion for evaluation of facial attractiveness models
Fangmei Chen, David Zhang 0001
Neurocomputing2
2016 Dorsal hand vein recognition via hierarchical combination of texture and shape clues
Di Huang 0001, Xiangrong Zhu 0002, Yunhong Wang 0001, David Zhang 0001
Neurocomputing4
2016 Individualized learning for improving kernel Fisher discriminant analysis
Zizhu Fan, Yong Xu 0001, Xiaozhao Fang, David Zhang 0001
Pattern Recognit.5
2016 Double-orientation code and nonlinear matching scheme for palmprint recognition
Lunke Fei, Yong Xu 0001, Wenliang Tang, David Zhang 0001
Pattern Recognit.4
2016 Robust single-object image segmentation based on salient transition region
Guanghai Liu 0001, David Zhang 0001, Yong Xu 0001
Pattern Recognit.3
2016 Compositional models and Structured learning for visual recognition
Liang Lin 0004, Jason J. Corso, Wangmeng Zuo, David Zhang 0001, Benjamin Z. Yao
Pattern Recognit.4
2016 Quadratic projection based feature extraction with its application to biometric recognition
Yan Yan 0001, Hanzi Wang, Si Chen 0002, Xiaochun Cao, David Zhang 0001
Pattern Recognit.5
2016 Half-orientation extraction of palmprint features
Lunke Fei, Yong Xu 0001, David Zhang 0001
Pattern Recognit. Lett.3
2016 Low-Rank Preserving Projections
abstract
As one of the most popular dimensionality reduction techniques, locality preserving projections (LPP) has been widely used in computer vision and pattern recognition. However, in practical applications, data is always corrupted by noises. For the corrupted data, samples from the same class may not be distributed in the nearest area, thus LPP may lose its effectiveness. In this paper, it is assumed that data is grossly corrupted and the noise matrix is sparse. Based on these assumptions, we propose a novel dimensionality reduction method, named low-rank preserving projections (LRPP) for image classification. LRPP learns a low-rank weight matrix by projecting the data on a low-dimensional subspace. We use the L21 norm as a sparse constraint on the noise matrix and the nuclear norm as a low-rank constraint on the weight matrix. LRPP keeps the global structure of the data during the dimensionality reduction procedure and the learned low rank weight matrix can reduce the disturbance of noises in the data. LRPP can learn a robust subspace from the corrupted data. To verify the performance of LRPP in image dimensionality reduction and classification, we compare LRPP with the state-of-the-art dimensionality reduction methods. The experimental results show the effectiveness and the feasibility of the proposed method with encouraging results.
Yuwu Lu, Zhihui Lai 0001, Yong Xu 0001, Xuelong Li 0001, David Zhang 0001, Chun Yuan 0003
IEEE Trans. Cybern.5
2016 A Level Set Approach to Image Segmentation With Intensity Inhomogeneity
abstract
It is often a difficult task to accurately segment images with intensity inhomogeneity, because most of representative algorithms are region-based that depend on intensity homogeneity of the interested object. In this paper, we present a novel level set method for image segmentation in the presence of intensity inhomogeneity. The inhomogeneous objects are modeled as Gaussian distributions of different means and variances in which a sliding window is used to map the original image into another domain, where the intensity distribution of each object is still Gaussian but better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying a bias field with the original signal within the window. A maximum likelihood energy functional is then defined on the whole image region, which combines the bias field, the level set function, and the piecewise constant function approximating the true image signal. The proposed level set method can be directly applied to simultaneous segmentation and bias correction for 3 and 7T magnetic resonance images. Extensive evaluation on synthetic and real-images demonstrate the superiority of the proposed method over other representative algorithms.
Kaihua Zhang 0001, Lei Zhang 0006, Kin-Man Lam 0001, David Zhang 0001
IEEE Trans. Cybern.4
2016 Multi-Label Dictionary Learning for Image Annotation
abstract
Image annotation has attracted a lot of research interest, and multi-label learning is an effective technique for image annotation. How to effectively exploit the underlying correlation among labels is a crucial task for multi-label learning. Most existing multi-label learning methods exploit the label correlation only in the output label space, leaving the connection between the label and the features of images untouched. Although, recently some methods attempt toward exploiting the label correlation in the input feature space by using the label information, they cannot effectively conduct the learning process in both the spaces simultaneously, and there still exists much room for improvement. In this paper, we propose a novel multi-label learning approach, named multi-label dictionary learning (MLDL) with label consistency regularization and partial-identical label embedding MLDL, which conducts MLDL and partial-identical label embedding simultaneously. In the input feature space, we incorporate the dictionary learning technique into multi-label learning and design the label consistency regularization term to learn the better representation of features. In the output label space, we design the partial-identical label embedding, in which the samples with exactly same label set can cluster together, and the samples with partial-identical label sets can collaboratively represent each other. Experimental results on the three widely used image datasets, including Corel 5K, IAPR TC12, and ESP Game, demonstrate the effectiveness of the proposed approach.
Xiaoyuan Jing, Fei Wu 0004, Zhiqiang Li 0003, Ruimin Hu, David Zhang 0001
IEEE Trans. Image Process.5
2016 Discriminative Transfer Subspace Learning via Low-Rank and Sparse Representation
abstract
In this paper, we address the problem of unsupervised domain transfer learning in which no labels are available in the target domain. We use a transformation matrix to transfer both the source and target data to a common subspace, where each target sample can be represented by a combination of source samples such that the samples from different domains can be well interlaced. In this way, the discrepancy of the source and target domains is reduced. By imposing joint low-rank and sparse constraints on the reconstruction coefficient matrix, the global and local structures of data can be preserved. To enlarge the margins between different classes as much as possible and provide more freedom to diminish the discrepancy, a flexible linear classifier (projection) is obtained by learning a non-negative label relaxation matrix that allows the strict binary label matrix to relax into a slack variable matrix. Our method can avoid a potentially negative transfer by using a sparse matrix to model the noise and, thus, is more robust to different types of noise. We formulate our problem as a constrained low-rankness and sparsity minimization problem and solve it by the inexact augmented Lagrange multiplier method. Extensive experiments on various visual domain adaptation tasks show the superiority of the proposed method over the state-of-the art methods. The MATLAB code of our method will be publicly available at http://www.yongxu.org/lunwen.html.
Yong Xu 0001, Xiaozhao Fang, Xuelong Li 0001, David Zhang 0001
IEEE Trans. Image Process.5
2016 Robust Visual Knowledge Transfer via Extreme Learning Machine-Based Domain Adaptation
abstract
We address the problem of visual knowledge adaptation by leveraging labeled patterns from source domain and a very limited number of labeled instances in target domain to learn a robust classifier for visual categorization. This paper proposes a new extreme learning machine based cross-domain network learning framework, that is called Extreme Learning Machine (ELM) based Domain Adaptation (EDA). It allows us to learn a category transformation and an ELM classifier with random projection by minimizing the -norm of the network output weights and the learning error simultaneously. The unlabeled target data, as useful knowledge, is also integrated as a fidelity term to guarantee the stability during cross domain learning. It minimizes the matching error between the learned classifier and a base classifier, such that many existing classifiers can be readily incorporated as base classifiers. The network output weights cannot only be analytically determined, but also transferrable. Additionally, a manifold regularization with Laplacian graph is incorporated, such that it is beneficial to semi-supervised learning. Extensively, we also propose a model of multiple views, referred as MvEDA. Experiments on benchmark visual datasets for video event recognition and object recognition, demonstrate that our EDA methods outperform existing cross-domain learning methods.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Image Process.2
2016 LSDT: Latent Sparse Domain Transfer Learning for Visual Adaptation
abstract
We propose a novel reconstruction-based transfer learning method called latent sparse domain transfer (LSDT) for domain adaptation and visual categorization of heterogeneous data. For handling cross-domain distribution mismatch, we advocate reconstructing the target domain data with the combined source and target domain data points based on ℓ1-norm sparse coding. Furthermore, we propose a joint learning model for simultaneous optimization of the sparse coding and the optimal subspace representation. In addition, we generalize the proposed LSDT model into a kernel-based linear/nonlinear basis transformation learning framework for tackling nonlinear subspace shifts in reproduced kernel Hilbert space. The proposed methods have three advantages: 1) the latent space and the reconstruction are jointly learned for pursuit of an optimal subspace transfer; 2) with the theory of sparse subspace clustering, a few valuable source and target data points are formulated to reconstruct the target data with noise (outliers) from source domain removed during domain adaptation, such that the robustness is guaranteed; and 3) a nonlinear projection of some latent space with kernel is easily generalized for dealing with highly nonlinear domain shift (e.g., face poses). Extensive experiments on several benchmark vision data sets demonstrate that the proposed approaches outperform other state-of-the-art representation-based domain adaptation methods.
Lei Zhang 0038, Wangmeng Zuo, David Zhang 0001
IEEE Trans. Image Process.3
2016 Learning Iteration-wise Generalized Shrinkage-Thresholding Operators for Blind Deconvolution
abstract
Salient edge selection and time-varying regularization are two crucial techniques to guarantee the success of maximum a posteriori (MAP)-based blind deconvolution. However, the existing approaches usually rely on carefully designed regularizers and handcrafted parameter tuning to obtain satisfactory estimation of the blur kernel. Many regularizers exhibit the structure-preserving smoothing capability, but fail to enhance salient edges. In this paper, under the MAP framework, we propose the iteration-wise ℓp-norm regularizers together with data-driven strategy to address these issues. First, we extend the generalized shrinkage-thresholding (GST) operator for ℓp-norm minimization with negative p value, which can sharpen salient edges while suppressing trivial details. Then, the iteration-wise GST parameters are specified to allow dynamical salient edge selection and time-varying regularization. Finally, instead of handcrafted tuning, a principled discriminative learning approach is proposed to learn the iterationwise GST operators from the training dataset. Furthermore, the multi-scale scheme is developed to improve the efficiency of the algorithm. Experimental results show that, negative p value is more effective in estimating the coarse shape of blur kernel at the early stage, and the learned GST operators can be well generalized to other dataset and real world blurry images. Compared with the state-of-the-art methods, our method achieves better deblurring results in terms of both quantitative metrics and visual quality, and it is much faster than the state-of-the-art patch-based blind deconvolution method.
Wangmeng Zuo, Dongwei Ren, David Zhang 0001, Shuhang Gu, Lei Zhang 0006
IEEE Trans. Image Process.3
2016 An Optimal Pulse System Design by Multichannel Sensors Fusion
abstract
Pulse diagnosis, recognized as an important branch of traditional Chinese medicine (TCM), has a long history for health diagnosis. Certain features in the pulse are known to be related with the physiological status, which have been identified as biomarkers. In recent years, an electronic equipment is designed to obtain the valuable information inside pulse. Single-point pulse acquisition platform has the benefit of low cost and flexibility, but is time consuming in operation and not standardized in pulse location. The pulse system with a single-type sensor is easy to implement, but is limited in extracting sufficient pulse information. This paper proposes a novel system with optimal design that is special for pulse diagnosis. We combine a pressure sensor with a photoelectric sensor array to make a multichannel sensor fusion structure. Then, the optimal pulse signal processing methods and sensor fusion strategy are introduced for the feature extraction. Finally, the developed optimal pulse system and methods are tested on pulse database acquired from the healthy subjects and the patients known to be afflicted with diabetes. The experimental results indicate that the classification accuracy is increased significantly under the optimal design and also demonstrate that the developed pulse system with multichannel sensors fusion is more effective than the previous pulse acquisition platforms.
Dimin Wang, David Zhang 0001, Guangming Lu 0002
IEEE J. Biomed. Health Informatics2
2016 Comparison of Three Different Types of Wrist Pulse Signals by Their Physical Meanings and Diagnosis Performance
abstract
Increasing interest has been focused on computational pulse diagnosis where sensors are developed to acquire pulse signals, and machine learning techniques are exploited to analyze health conditions based on the acquired pulse signals. By far, a number of sensors have been employed for pulse signal acquisition, which can be grouped into three major categories, i.e., pressure, photoelectric, and ultrasonic sensors. To guide the sensor selection for computational pulse diagnosis, in this paper, we analyze the physical meanings and sensitivities of signals acquired by these three types of sensors. The dependence and complementarity of the different sensors are discussed from both the perspective of cardiovascular fluid dynamics and comparative experiments by evaluating disease classification performance. Experimental results indicate that each sensor is more appropriate for the diagnosis of some specific disease that the changes of physiological factors can be effectively reflected by the sensor, e.g., ultrasonic sensor for diabetes and pressure sensor for arteriosclerosis, and improved diagnosis performance can be obtained by combining three types of signals.
Wangmeng Zuo, Peng Wang 0089, David Zhang 0001
IEEE J. Biomed. Health Informatics3
2016 Visual Understanding via Multi-Feature Shared Learning With Global Consistency
abstract
Image/video data is usually represented with multiple visual features. Fusion of multi-source information for establishing attributes has been widely recognized. Multi- feature visual recognition has recently received much attention in multimedia applications. This paper studies visual understanding via a newly proposed l2-norm-based multi-feature shared learning framework, which can simultaneously learn a global label matrix and multiple sub-classifiers with the labeled multi-feature data. Additionally, a group graph manifold regularizer composed of the Laplacian and Hessian graph is proposed. It can better preserve the manifold structure of each feature, such that the label prediction power is much improved through semi-supervised learning with global label consistency. For convenience, we call the proposed approach global-label- consistent classifier (GLCC). The merits of the proposed method include the following: 1) the manifold structure information of each feature is exploited in learning, resulting in a more faithful classification owing to the global label consistency; 2) a group graph manifold regularizer based on the Laplacian and Hessian regularization is constructed ; and 3) an efficient alternative optimization method is introduced as a fast solver owing its speed to convex sub-problems. Experiments on several benchmark visual datasets-the 17-category Oxford Flower dataset, the challenging 101- category Caltech dataset, the YouTube and Consumer Videos dataset, and the large-scale NUS-WIDE dataset-have been used for multimedia understanding . The results demonstrate that the proposed approach compares favorably with state-of-the-art algorithms. An extensive experiment using the deep convolutional activation features also shows the effectiveness of the proposed approach. The code will be available on http://www.escience.cn/people/lei/index.html.
Lei Zhang 0038, David Zhang 0001
IEEE Trans. Multim.2
2016 Approximate Orthogonal Sparse Embedding for Dimensionality Reduction
abstract
Locally linear embedding (LLE) is one of the most well-known manifold learning methods. As the representative linear extension of LLE, orthogonal neighborhood preserving projection (ONPP) has attracted widespread attention in the field of dimensionality reduction. In this paper, a unified sparse learning framework is proposed by introducing the sparsity or L1-norm learning, which further extends the LLE-based methods to sparse cases. Theoretical connections between the ONPP and the proposed sparse linear embedding are discovered. The optimal sparse embeddings derived from the proposed framework can be computed by iterating the modified elastic net and singular value decomposition. We also show that the proposed model can be viewed as a general model for sparse linear and nonlinear (kernel) subspace learning. Based on this general model, sparse kernel embedding is also proposed for nonlinear sparse feature extraction. Extensive experiments on five databases demonstrate that the proposed sparse learning framework performs better than the existing subspace learning algorithm, particularly in the cases of small sample sizes.
Zhihui Lai 0001, Wai Keung Wong, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2016 A Novel Line-Scan Palmprint Acquisition System
abstract
Biometric recognition systems have been widely used globally. However, one effective and highly accurate biometric authentication method, palmprint recognition, has not been popularly applied as it should have been, which could be due to the lack of small, flexible and user-friendly acquisition systems. To expand the use of palmprint biometrics, we propose a novel palmprint acquisition system based on the line-scan image sensor. The proposed system consists of a customized and highly integrated line-scan sensor, a self-adaptive synchronizing unit, and a field-programmable gate array controller with a cross-platform interface. The volume of the proposed system is over 94% smaller than the volume of existing palmprint systems, without compromising its verification performance. The verification performance of the proposed system was tested on a database of 8000 samples collected from 250 people, and the equal error rate is 0.048%, which is comparable to the best area camera-based systems.
Xiaofeng Qu, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2015 Patch Group Based Nonlocal Self-Similarity Prior Learning for Image Denoising
abstract
Patch based image modeling has achieved a great success in low level vision such as image denoising. In particular, the use of image nonlocal self-similarity (NSS) prior, which refers to the fact that a local patch often has many nonlocal similar patches to it across the image, has significantly enhanced the denoising performance. However, in most existing methods only the NSS of input degraded image is exploited, while how to utilize the NSS of clean natural images is still an open problem. In this paper, we propose a patch group (PG) based NSS prior learning scheme to learn explicit NSS models from natural images for high performance denoising. PGs are extracted from training images by putting nonlocal similar patches into groups, and a PG based Gaussian Mixture Model (PG-GMM) learning algorithm is developed to learn the NSS prior. We demonstrate that, owe to the learned PG-GMM, a simple weighted sparse coding model, which has a closed-form solution, can be used to perform image denoising effectively, resulting in high PSNR measure, fast speed, and particularly the best visual quality among all competing methods.
Jun Xu 0019, Lei Zhang 0006, Wangmeng Zuo, David Zhang 0001, Xiangchu Feng
ICCV4
2015 2D facial landmark model design by combining key points and inserted points
Fangmei Chen, Yong Xu 0001, David Zhang 0001, Kai Chen 0023
Expert Syst. Appl.3
2015 Robust tongue segmentation by fusing region-based and edge-based approaches
Kebin Wu, David Zhang 0001
Expert Syst. Appl.2
2015 A salt & pepper noise filter based on local and global image information
Yong Cheng 0001, Kezong Tang, Yong Xu 0001, David Zhang 0001
Neurocomputing5
2015 Study on novel Curvature Features for 3D fingerprint recognition
Feng Liu 0013, David Zhang 0001, LinLin Shen
Neurocomputing2
2015 Fast total-variation based image restoration based on derivative alternated direction optimization methods
Dongwei Ren, David Zhang 0001, Wangmeng Zuo
Neurocomputing3
2015 Ear-parotic face angle: A unique feature for 3D ear recognition
Bob Zhang 0001, David Zhang 0001
Pattern Recognit. Lett.3
2015 A Framework of Joint Graph Embedding and Sparse Regression for Dimensionality Reduction
abstract
Over the past few decades, a large number of algorithms have been developed for dimensionality reduction. Despite the different motivations of these algorithms, they can be interpreted by a common framework known as graph embedding. In order to explore the significant features of data, some sparse regression algorithms have been proposed based on graph embedding. However, the problem is that these algorithms include two separate steps: (1) embedding learning and (2) sparse regression. Thus their performance is largely determined by the effectiveness of the constructed graph. In this paper, we present a framework by combining the objective functions of graph embedding and sparse regression so that embedding learning and sparse regression can be jointly implemented and optimized, instead of simply using the graph spectral for sparse regression. By the proposed framework, supervised, semisupervised, and unsupervised learning algorithms could be unified. Furthermore, we analyze two situations of the optimization problem for the proposed framework. By adopting an ℓ2,1-norm regularization for the proposed framework, it can perform feature selection and subspace learning simultaneously. Experiments on seven standard databases demonstrate that joint graph embedding and sparse regression method can significantly improve the recognition performance and consistently outperform the sparse regression method.
Xiaoshuang Shi, Zhenhua Guo 0001, Zhihui Lai 0001, Yujiu Yang 0001, Zhifeng Bao, David Zhang 0001
IEEE Trans. Image Process.6
2015 Combining Left and Right Palmprint Images for More Accurate Personal Identification
abstract
Multibiometrics can provide higher identification accuracy than single biometrics, so it is more suitable for some real-world personal identification applications that need high-standard security. Among various biometrics technologies, palmprint identification has received much attention because of its good performance. Combining the left and right palmprint images to perform multibiometrics is easy to implement and can obtain better results. However, previous studies did not explore this issue in depth. In this paper, we proposed a novel framework to perform multibiometrics by comprehensively combining the left and right palmprint images. This framework integrated three kinds of scores generated from the left and right palmprint images to perform matching score-level fusion. The first two kinds of scores were, respectively, generated from the left and right palmprint images and can be obtained by any palmprint identification method, whereas the third kind of score was obtained using a specialized algorithm proposed in this paper. As the proposed algorithm carefully takes the nature of the left and right palmprint images into account, it can properly exploit the similarity of the left and right palmprints of the same subject. Moreover, the proposed weighted fusion scheme allowed perfect identification performance to be obtained in comparison with previous palmprint identification methods.
Yong Xu 0001, Lunke Fei, David Zhang 0001
IEEE Trans. Image Process.3
2015 A Kernel Classification Framework for Metric Learning
abstract
Learning a distance metric from the given training samples plays a crucial role in many machine learning tasks, and various models and optimization algorithms have been proposed in the past decade. In this paper, we generalize several state-of-the-art metric learning methods, such as large margin nearest neighbor (LMNN) and information theoretic metric learning (ITML), into a kernel classification framework. First, doublets and triplets are constructed from the training samples, and a family of degree-2 polynomial kernel functions is proposed for pairs of doublets or triplets. Then, a kernel classification framework is established to generalize many popular metric learning methods such as LMNN and ITML. The proposed framework can also suggest new metric learning methods, which can be efficiently implemented, interestingly, using the standard support vector machine (SVM) solvers. Two novel metric learning methods, namely, doublet-SVM and triplet-SVM, are then developed under the proposed framework. Experimental results show that doublet-SVM and triplet-SVM achieve competitive classification accuracies with state-of-the-art metric learning methods but with significantly less training time.
Wangmeng Zuo, Lei Zhang 0006, Deyu Meng, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2014 Fast Visual Tracking via Dense Spatio-temporal Context Learning
Kaihua Zhang 0001, Lei Zhang 0006, Qingshan Liu 0001, David Zhang 0001, Ming-Hsuan Yang 0001
ECCV (5)4
2014 Sparse Representation Based Fisher Discrimination Dictionary Learning for Image Classification
Meng Yang 0001, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001
Int. J. Comput. Vis.4
2014 Integrate the original face image and its mirror image for face recognition
Yong Xu 0001, Xuelong Li 0001, Jian Yang 0003, David Zhang 0001
Neurocomputing4
2014 Special issue on "New sensing and processing technologies for hand-based biometrics authentication"
David Zhang 0001, Lei Zhang 0006
Inf. Sci.1
2014 Combination of linear regression classification and collaborative representation classification
David Zhang 0001, Kuanquan Wang, Jingdong Liu
Neural Comput. Appl.4
2014 3D fingerprint reconstruction system using feature correspondences and prior estimated finger model
Feng Liu 0013, David Zhang 0001
Pattern Recognit.2
2014 A New Hypothesis on Facial Beauty Perception
abstract
In this article, a new hypothesis on facial beauty perception is proposed: the weighted average of two facial geometric features is more attractive than the inferior one between them. Extensive evidences support the new hypothesis. We collected 390 well-known beautiful face images (e.g., Miss Universe, movie stars, and super models) as well as 409 common face images from multiple sources. Dozens of volunteers rated the face images according to their attractiveness. Statistical regression models are trained on this database. Under the empirical risk principle, the hypothesis is tested on 318,801 pairs of images and receives consistently supportive results. A corollary of the hypothesis is attractive facial geometric features construct a convex set. This corollary derives a convex hull based face beautification method, which guarantees attractiveness and minimizes the before--after difference. Experimental results show its superiority to state-of-the-art geometric based face beautification methods. Moreover, the mainstream hypotheses on facial beauty perception (e.g., the averageness, symmetry, and golden ratio hypotheses) are proved to be compatible with the proposed hypothesis.
Fangmei Chen, Yong Xu 0001, David Zhang 0001
ACM Trans. Appl. Percept.3
2014 Human Gait Recognition via Sparse Discriminant Projection Learning
abstract
As an important biometric feature, human gait has great potential in video-surveillance-based applications. In this paper, we focus on the matrix representation-based human gait recognition and propose a novel discriminant subspace learning method called sparse bilinear discriminant analysis (SBDA). SBDA extends the recently proposed matrix-representation-based discriminant analysis methods to sparse cases. By introducing the L1and L2norms into the objective function of SBDA, two interrelated sparse discriminant subspaces can be obtained for gait feature extraction. Since the optimization problem has no closed-form solutions, an iterative method is designed to compute the optimal sparse subspace using the L1and L2norms sparse regression. Theoretical analyses reveal the close relationship between SBDA and previous matrix-representation-based discriminant analysis methods. Since each nonzero element in each subspace is selected from the most important variables/factors, SBDA is potential to perform equivalent to or even better than the state-of-the-art subspace learning methods in gait recognition. Moreover, using the strategy of SBDA plus linear discriminant analysis (LDA), we can further improve the performance. A set of experiments on the standard USF HumanID and CASIA gait databases demonstrate that the proposed SBDA and SBDA + LDA can obtain competitive performance.
Zhihui Lai 0001, Yong Xu 0001, Zhong Jin, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2014 Integrating Conventional and Inverse Representation for Face Recognition
abstract
Representation-based classification methods are all constructed on the basis of the conventional representation, which first expresses the test sample as a linear combination of the training samples and then exploits the deviation between the test sample and the expression result of every class to perform classification. However, this deviation does not always well reflect the difference between the test sample and each class. With this paper, we propose a novel representation-based classification method for face recognition. This method integrates conventional and the inverse representation-based classification for better recognizing the face. It first produces conventional representation of the test sample, i.e., uses a linear combination of the training samples to represent the test sample. Then it obtains the inverse representation, i.e., provides an approximation representation of each training sample of a subject by exploiting the test sample and training samples of the other subjects. Finally, the proposed method exploits the conventional and inverse representation to generate two kinds of scores of the test sample with respect to each class and combines them to recognize the face. The paper shows the theoretical foundation and rationale of the proposed method. Moreover, this paper for the first time shows that a basic nature of the human face, i.e., the symmetry of the face can be exploited to generate new training and test samples. As these new samples really reflect some possible appearance of the face, the use of them will enable us to obtain higher accuracy. The experiments show that the proposed conventional and inverse representation-based linear regression classification (CIRLRC), an improvement to linear regression classification (LRC), can obtain very high accuracy and greatly outperforms the naive LRC and other state-of-the-art conventional representation based face recognition methods. The accuracy of CIRLRC can be 10% greater than that of LRC.
Yong Xu 0001, Xuelong Li 0001, Jian Yang 0003, Zhihui Lai 0001, David Zhang 0001
IEEE Trans. Cybern.5
2014 Image Set-Based Collaborative Representation for Face Recognition
abstract
With the rapid development of digital imaging and communication technologies, image set-based face recognition (ISFR) is becoming increasingly important. One key issue of ISFR is how to effectively and efficiently represent the query face image set using the gallery face image sets. The set-to-set distance-based methods ignore the relationship between gallery sets, whereas representing the query set images individually over the gallery sets ignores the correlation between query set images. In this paper, we propose a novel image set-based collaborative representation and classification method for ISFR. By modeling the query set as a convex or regularized hull, we represent this hull collaboratively over all the gallery sets. With the resolved representation coefficients, the distance between the query set and each gallery set can then be calculated for classification. The proposed model naturally and effectively extends the image-based collaborative representation to an image set based one, and our extensive experiments on benchmark ISFR databases show the superiority of the proposed method to state-of-the-art ISFR methods under different set sizes in terms of both recognition rate and efficiency.
Pengfei Zhu 0001, Wangmeng Zuo, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2014 Gradient Histogram Estimation and Preservation for Texture Enhanced Image Denoising
abstract
Natural image statistics plays an important role in image denoising, and various natural image priors, including gradient-based, sparse representation-based, and nonlocal self-similarity-based ones, have been widely studied and exploited for noise removal. In spite of the great success of many denoising algorithms, they tend to smooth the fine scale image textures when removing noise, degrading the image visual quality. To address this problem, in this paper, we propose a texture enhanced image denoising method by enforcing the gradient histogram of the denoised image to be close to a reference gradient histogram of the original image. Given the reference gradient histogram, a novel gradient histogram preservation (GHP) algorithm is developed to enhance the texture structures while removing noise. Two region-based variants of GHP are proposed for the denoising of images consisting of regions with different textures. An algorithm is also developed to effectively estimate the reference gradient histogram from the noisy observation of the unknown image. Our experimental results demonstrate that the proposed GHP algorithm can well preserve the texture appearance in the denoised images, making them look more natural.
Wangmeng Zuo, Lei Zhang 0006, Chunwei Song, David Zhang 0001, Huijun Gao
IEEE Trans. Image Process.4
2014 Modified Principal Component Analysis: An Integration of Multiple Similarity Subspace Models
abstract
We modify the conventional principal component analysis (PCA) and propose a novel subspace learning framework, modified PCA (MPCA), using multiple similarity measurements. MPCA computes three similarity matrices exploiting the similarity measurements: 1) mutual information; 2) angle information; and 3) Gaussian kernel similarity. We employ the eigenvectors of similarity matrices to produce new subspaces, referred to as similarity subspaces. A new integrated similarity subspace is then generated using a novel feature selection approach. This approach needs to construct a kind of vector set, termed weak machine cell (WMC), which contains an appropriate number of the eigenvectors spanning the similarity subspaces. Combining the wrapper method and the forward selection scheme, MPCA selects a WMC at a time that has a powerful discriminative capability to classify samples. MPCA is very suitable for the application scenarios in which the number of the training samples is less than the data dimensionality. MPCA outperforms the other state-of-the-art PCA-based methods in terms of both classification accuracy and clustering result. In addition, MPCA can be applied to face image reconstruction. MPCA can use other types of similarity measurements. Extensive experiments on many popular real-world data sets, such as face databases, show that MPCA achieves desirable classification results, as well as has a powerful capability to represent data.
Zizhu Fan, Yong Xu 0001, Wangmeng Zuo, Jian Yang 0003, Jinhui Tang 0001, Zhihui Lai 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2014 Multilinear Sparse Principal Component Analysis
abstract
In this brief, multilinear sparse principal component analysis (MSPCA) is proposed for feature extraction from the tensor data. MSPCA can be viewed as a further extension of the classical principal component analysis (PCA), sparse PCA (SPCA) and the recently proposed multilinear PCA (MPCA). The key operation of MSPCA is to rewrite the MPCA into multilinear regression forms and relax it for sparse regression. Differing from the recently proposed MPCA, MSPCA inherits the sparsity from the SPCA and iteratively learns a series of sparse projections that capture most of the variation of the tensor data. Each nonzero element in the sparse projections is selected from the most important variables/factors using the elastic net. Extensive experiments on Yale, Face Recognition Technology face databases, and COIL-20 object database encoded the object images as second-order tensors, and Weizmann action database as third-order tensors demonstrate that the proposed MSPCA algorithm has the potential to outperform the existing PCA-based subspace learning algorithms.
Zhihui Lai 0001, Yong Xu 0001, Qingcai Chen, Jian Yang 0003, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2014 Introduction to the Special Section on Biometric Systems and Applications
abstract
Nowadays, biometrics is an important technological area receiving continuously growing interest from academia, industry, government, and the general public, due to the criticality and the social impact of its applications. Biometric systems are in fact rapidly being adopted in a wide variety of applications such as security, ambient intelligence, electronic and physical access control, digital rights management, background checking and defense, medical diagnosis as well as for adaptive environments.
Michele Nappi, Vincenzo Piuri, Tieniu Tan, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2013 Texture Enhanced Image Denoising via Gradient Histogram Preservation
abstract
Image denoising is a classical yet fundamental problem in low level vision, as well as an ideal test bed to evaluate various statistical image modeling methods. One of the most challenging problems in image denoising is how to preserve the fine scale texture structures while removing noise. Various natural image priors, such as gradient based prior, nonlocal self-similarity prior, and sparsity prior, have been extensively exploited for noise removal. The denoising algorithms based on these priors, however, tend to smooth the detailed image textures, degrading the image visual quality. To address this problem, in this paper we propose a texture enhanced image denoising (TEID) method by enforcing the gradient distribution of the denoised image to be close to the estimated gradient distribution of the original image. A novel gradient histogram preservation (GHP) algorithm is developed to enhance the texture structures while removing noise. Our experimental results demonstrate that the proposed GHP based TEID can well preserve the texture features of the denoised images, making them look more natural.
Wangmeng Zuo, Lei Zhang 0006, Chunwei Song, David Zhang 0001
CVPR4
2013 From Point to Set: Extend the Learning of Distance Metrics
abstract
Most of the current metric learning methods are proposed for point-to-point distance (PPD) based classification. In many computer vision tasks, however, we need to measure the point-to-set distance (PSD) and even set-to-set distance (SSD) for classification. In this paper, we extend the PPD based Mahalanobis distance metric learning to PSD and SSD based ones, namely point-to-set distance metric learning (PSDML) and set-to-set distance metric learning (SSDML), and solve them under a unified optimization framework. First, we generate positive and negative sample pairs by computing the PSD and SSD between training samples. Then, we characterize each sample pair by its covariance matrix, and propose a covariance kernel based discriminative function. Finally, we tackle the PSDML and SSDML problems by using standard support vector machine solvers, making the metric learning very efficient for multiclass visual classification tasks. Experiments on gender classification, digit recognition, object categorization and face recognition show that the proposed metric learning methods can effectively enhance the performance of PSD and SSD based classification.
Pengfei Zhu 0001, Lei Zhang 0006, Wangmeng Zuo, David Zhang 0001
ICCV4
2013 A Generalized Iterated Shrinkage Algorithm for Non-convex Sparse Coding
abstract
In many sparse coding based image restoration and image classification problems, using non-convex Ip-norm minimization (0 ≤ p1-norm minimization. A number of algorithms, e.g., iteratively reweighted least squares (IRLS), iteratively thresholding method (ITM-Ip), and look-up table (LUT), have been proposed for non-convex Ip-norm sparse coding, while some analytic solutions have been suggested for some specific values of p. In this paper, by extending the popular soft-thresholding operator, we propose a generalized iterated shrinkage algorithm (GISA) for Ip-norm non-convex sparse coding. Unlike the analytic solutions, the proposed GISA algorithm is easy to implement, and can be adopted for solving non-convex sparse coding problems with arbitrary p values. Compared with LUT, GISA is more general and does not need to compute and store the look-up tables. Compared with IRLS and ITM-Ip, GISA is theoretically more solid and can achieve more accurate solutions. Experiments on image restoration and sparse coding based face recognition are conducted to validate the performance of GISA.
Wangmeng Zuo, Deyu Meng, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001
ICCV5
2013 On brewing fresh espresso: LinkedIn's distributed data serving platform
abstract
Espresso is a document-oriented distributed data serving platform that has been built to address LinkedIn's requirements for a scalable, performant, source-of-truth primary store. It provides a hierarchical document model, transactional support for modifications to related documents, real-time secondary indexing, on-the-fly schema evolution and provides a timeline consistent change capture stream. This paper describes the motivation and design principles involved in building Espresso, the data model and capabilities exposed to clients, details of the replication and secondary indexing implementation and presents a set of experimental results that characterize the performance of the system along various dimensions.
Lin Qiao, Kapil Surlaker, Shirshanka Das, Tom Quiggle, Bob Schulman, Bhaskar Ghosh, Antony Curtis, Oliver Seeliger, Aditya Auradkar, Chris Beaver, Gregory Brandt, Mihir Gandhi, Kishore Gopalakrishna, Wai Ip, Swaroop Jagadish, Shi Lu, Alexander Pachev, Aditya Ramesh, Abraham Sebastian, Rupa Shanbhag, Subbu Subramaniam, Sajid Topiwala, Cuong Tran 0003, Jemiah Westerman, David Zhang 0001
SIGMOD Conference27
2013 A high quality color imaging system for computerized tongue image analysis
Xingzheng Wang, David Zhang 0001
Expert Syst. Appl.2
2013 Facial image medical analysis system using quantitative chromatic feature
Xingzheng Wang, Bob Zhang 0001, Zhenhua Guo 0001, David Zhang 0001
Expert Syst. Appl.4
2013 Is local dominant orientation necessary for the classification of rotation invariant texture?
Zhenhua Guo 0001, Qin Li 0001, Lin Zhang 0014, Jane You, David Zhang 0001, Wenhuang Liu
Neurocomputing5
2013 A sparse representation method of bimodal biometrics and palmprint recognition experiments
Yong Xu 0001, Zizhu Fan, Minna Qiu, David Zhang 0001, Jing-Yu Yang 0001
Neurocomputing4
2013 Using the idea of the sparse representation to perform coarse-to-fine face recognition
Yong Xu 0001, Qi Zhu 0001, Zizhu Fan, David Zhang 0001, Jian-Xun Mi, Zhihui Lai 0001
Inf. Sci.4
2013 Computerized facial diagnosis using both color and texture features
Bob Zhang 0001, Xingzheng Wang, Fakhri Karray, Zhimin Yang, David Zhang 0001
Inf. Sci.5
2013 Complete large margin linear discriminant analysis using mathematical programming approach
Xiaobo Chen 0001, Jian Yang 0003, David Zhang 0001, Jun Liang 0004
Pattern Recognit.3
2013 Joint discriminative dimensionality reduction and dictionary learning for face recognition
Zhizhao Feng, Meng Yang 0001, Lei Zhang 0006, Yan Liu 0004, David Zhang 0001
Pattern Recognit.5
2013 A survey of graph theoretical approaches to image segmentation
Bo Peng 0006, Lei Zhang 0006, David Zhang 0001
Pattern Recognit.3
2013 Gabor feature based robust representation and classification for face recognition with Gabor occlusion dictionary
Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001
Pattern Recognit.4
2013 Fast gradient vector flow computation based on augmented Lagrangian method
Dongwei Ren, Wangmeng Zuo, Zhouchen Lin, David Zhang 0001
Pattern Recognit. Lett.5
2013 Joint Registration and Active Contour Segmentation for Object Tracking
abstract
This paper presents a novel object tracking framework by joint registration and active contour segmentation (JRACS), which can robustly deal with the non-rigid shape changes of the target. The target region, which includes both foreground and background pixels, is implicitly represented by a level set. A Bhattacharyya similarity based metric is proposed to locate the region whose foreground and background distributions best match those of the tracked target. Based on this metric, a tracking framework that consists of a registration stage and a segmentation stage is then established. The registration step roughly locates the target object by modeling its motion as an affine transformation, and the segmentation step refines the registration result and computes the true contour of the target. The robust tracking performance of the proposed JRACS method is demonstrated by real video sequences where the objects have clear non-rigid shape changes.
Jifeng Ning, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2013 An Optimized Wavelength Band Selection for Heavily Pigmented Iris Recognition
abstract
Commercial iris recognition systems usually acquire images of the eye in 850-nm band of the electromagnetic spectrum. In this work, the heavily pigmented iris images are captured at 12 wavelengths, from 420 to 940 nm. The purpose is to find the most suitable wavelength band for the heavily pigmented iris recognition. A multispectral acquisition system is first designed for imaging the iris at narrow spectral bands in the range of 420-940 nm. Next, a set of 200 human black irises which correspond to the right and left eyes of 100 different subjects are acquired for an analysis. Finally, the most suitable wavelength for heavily pigmented iris recognition is found based on two approaches: 1) the quality assurance of texture; 2) matching performance-equal error rate (EER) and false rejection rate (FRR). This result is supported by visual observations of magnified detailed local iris texture information. The experimental results suggest that there exists a most suitable wavelength band for heavily pigmented iris recognition when using a single band of wavelength as illumination.
Yazhuo Gong, David Zhang 0001, Jingqi Yan
IEEE Trans. Inf. Forensics Secur.2
2013 Distal-Interphalangeal-Crease-Based User Authentication System
abstract
Touchless-based fingerprint recognition technology is thought to be an alternative to touch-based systems to solve problems of hygienic, latent fingerprints, and maintenance. However, there are few studies about touchless fingerprint recognition systems due to the lack of a large database and the intrinsic drawback of low ridge-valley contrast of touchless fingerprint images. This paper proposes an end-to-end solution for user authentication systems based on touchless fingerprint images in which a multiview strategy is adopted to collect images and the robust fingerprint feature of touchless image is extracted for matching with high recognition accuracy. More specifically, a touchless multiview fingerprint capture device is designed to generate three views of raw images followed by preprocessing steps including region of interest (ROI) extraction and image correction. The distal interphalangeal crease (DIP)-based feature is then extracted and matched to recognize the human's identity in which part selection is introduced to improve matching efficiency. Experiments are conducted on two sessions of touchless multiview fingerprint image database with 541 fingers acquired about two weeks apart. An EER of ~ 1.7% can be achieved by using the proposed DIP-based feature, which is much better than touchless fingerprint recognition by using scale invariant feature transformation (SIFT) and minutiae features. The given fusion results show that it is effective to combine the DIP-based feature, minutiae, and SIFT feature for touchless fingerprint recognition systems. The EER is as low as ~ 0.5%.
Feng Liu 0013, David Zhang 0001, Zhenhua Guo 0001
IEEE Trans. Inf. Forensics Secur.2
2013 Reconstruction Based Finger-Knuckle-Print Verification With Score Level Adaptive Binary Fusion
abstract
Recently, a new biometrics identifier, namely finger knuckle print (FKP), has been proposed for personal authentication with very interesting results. One of the advantages of FKP verification lies in its user friendliness in data collection. However, the user flexibility in positioning fingers also leads to a certain degree of pose variations in the collected query FKP images. The widely used Gabor filtering based competitive coding scheme is sensitive to such variations, resulting in many false rejections. We propose to alleviate this problem by reconstructing the query sample with a dictionary learned from the template samples in the gallery set. The reconstructed FKP image can reduce much the enlarged matching distance caused by finger pose variations; however, both the intra-class and inter-class distances will be reduced. We then propose a score level adaptive binary fusion rule to adaptively fuse the matching distances before and after reconstruction, aiming to reduce the false rejections without increasing much the false acceptances. Experimental results on the benchmark PolyU FKP database show that the proposed method significantly improves the FKP verification accuracy.
Guangwei Gao, Lei Zhang 0006, Jian Yang 0003, Lin Zhang 0014, David Zhang 0001
IEEE Trans. Image Process.5
2013 Sparse Tensor Discriminant Analysis
abstract
The classical linear discriminant analysis has undergone great development and has recently been extended to different cases. In this paper, a novel discriminant subspace learning method called sparse tensor discriminant analysis (STDA) is proposed, which further extends the recently presented multilinear discriminant analysis to a sparse case. Through introducing the L1 and L2 norms into the objective function of STDA, we can obtain multiple interrelated sparse discriminant subspaces for feature extraction. As there are no closed-form solutions, k-mode optimization technique and the L1 norm sparse regression are combined to iteratively learn the optimal sparse discriminant subspace along different modes of the tensors. Moreover, each non-zero element in each subspace is selected from the most important variables/factors, and thus STDA has the potential to perform better than other discriminant subspace methods. Extensive experiments on face databases (Yale, FERET, and CMU PIE face databases) and the Weizmann action database show that the proposed STDA algorithm demonstrates the most competitive performance against the compared tensor-based methods, particularly in small sample sizes.
Zhihui Lai 0001, Yong Xu 0001, Jian Yang 0003, Jinhui Tang 0001, David Zhang 0001
IEEE Trans. Image Process.5
2013 Statistical Analysis of Tongue Images for Feature Extraction and Diagnostics
abstract
In this paper, an in-depth analysis on the statistical distribution characteristics of human tongue color that aims to propose a mathematically described tongue color space for diagnostic feature extraction is presented. Three characteristics of tongue color space, i.e., tongue color gamut that defines the range of colors, color centers of 12 tongue color categories, and color distribution of typical image features in the tongue color gamut, are elaborately investigated in this paper. Based on a large database, which contains over 9000 tongue images collected by a specially designed noncontact colorimetric imaging system using a digital camera, the tongue color gamut is established in the CIE chromaticity diagram by an innovatively proposed color gamut boundary descriptor using one-class SVM algorithm. Thereafter, centers of 12 tongue color categories are defined accordingly. Furthermore, color distributions of several typical tongue features, such as red points and petechial points, are obtained to build a relationship between the tongue color space and color distributions of various tongue features. With the obtained tongue color space, a new color feature extraction method is proposed for diagnostic classification purposes, with experimental results validating its effectiveness.
Xingzheng Wang, Bob Zhang 0001, Zhimin Yang, Haoqian Wang, David Zhang 0001
IEEE Trans. Image Process.5
2013 Regularized Robust Coding for Face Recognition
abstract
Recently the sparse representation based classification (SRC) has been proposed for robust face recognition (FR). In SRC, the testing image is coded as a sparse linear combination of the training samples, and the representation fidelity is measured by the l2-norm or l1 -norm of the coding residual. Such a sparse coding model assumes that the coding residual follows Gaussian or Laplacian distribution, which may not be effective enough to describe the coding residual in practical FR systems. Meanwhile, the sparsity constraint on the coding coefficients makes the computational cost of SRC very high. In this paper, we propose a new face coding model, namely regularized robust coding (RRC), which could robustly regress a given signal with regularized regression coefficients. By assuming that the coding residual and the coding coefficient are respectively independent and identically distributed, the RRC seeks for a maximum a posterior solution of the coding problem. An iteratively reweighted regularized robust coding (IR(3)C) algorithm is proposed to solve the RRC model efficiently. Extensive experiments on representative face databases demonstrate that the RRC is much more effective and efficient than state-of-the-art sparse representation based methods in dealing with face occlusion, corruption, lighting, and expression changes, etc.
Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001
IEEE Trans. Image Process.4
2013 Reinitialization-Free Level Set Evolution via Reaction Diffusion
abstract
This paper presents a novel reaction-diffusion (RD) method for implicit active contours that is completely free of the costly reinitialization procedure in level set evolution (LSE). A diffusion term is introduced into LSE, resulting in an RD-LSE equation, from which a piecewise constant solution can be derived. In order to obtain a stable numerical solution from the RD-based LSE, we propose a two-step splitting method to iteratively solve the RD-LSE equation, where we first iterate the LSE equation, then solve the diffusion equation. The second step regularizes the level set function obtained in the first step to ensure stability, and thus the complex and costly reinitialization procedure is completely eliminated from LSE. By successfully applying diffusion to LSE, the RD-LSE model is stable by means of the simple finite difference method, which is very easy to implement. The proposed RD method can be generalized to solve the LSE for both variational level set method and partial differential equation-based level set method. The RD-LSE method shows very good performance on boundary antileakage. The extensive and promising experimental results on synthetic and real images validate the effectiveness of the proposed RD-LSE approach.
Kaihua Zhang 0001, Lei Zhang 0006, Huihui Song 0003, David Zhang 0001
IEEE Trans. Image Process.4
2013 Iris-Based Medical Analysis by Geometric Deformation Features
abstract
Iris analysis studies the relationship between human health and changes in the anatomy of the iris. Apart from the fact that iris recognition focuses on modeling the overall structure of the iris, iris diagnosis emphasizes the detecting and analyzing of local variations in the characteristics of irises. This paper focuses on studying the geometrical structure changes in irises that are caused by gastrointestinal diseases, and on measuring the observable deformations in the geometrical structures of irises that are related to roundness, diameter and other geometric forms of the pupil and the collarette. Pupil and collarette based features are defined and extracted. A series of experiments are implemented on our experimental pathological iris database, including manual clustering of both normal and pathological iris images, manual classification by non-specialists, manual classification by individuals with a medical background, classification ability verification for the proposed features, and disease recognition by applying the proposed features. The results prove the effectiveness and clinical diagnostic significance of the proposed features and a reliable recognition performance for automatic disease diagnosis. Our research results offer a novel systematic perspective for iridology studies and promote the progress of both theoretical and practical work in iris diagnosis.
Lin Ma 0003, David Zhang 0001, Naimin Li, Yan Cai 0009, Wangmeng Zuo, Kuanquan Wang
IEEE J. Biomed. Health Informatics2
2013 A New Tongue Colorchecker Design by Space Representation for Precise Correction
abstract
In order to improve the correction accuracy on tongue colors by use of Munsell colorchecker, this research aims to design a new colorchecker by aid of tongue color space. Three essential issues leading to the development of this space-based colorchecker are elaborately investigated in this study. Firstly, based on a large and comprehensive tongue database, tongue color space is established by which all visible colors can be classified as tongue or non-tongue colors. Hence, colors of the designed tongue colorchecker are selected from tongue colors to achieve high correction performance. Secondly, the minimum sufficient number of colors involved in colorchecker is yielded by comparing the correction accuracy when different number (ranged from 10 to 200) of colors are contained. Thereby, 24 colors are included because the obtained minimum number of colors is 20. Lastly, criteria for optimal color selection and its corresponding objective function are presented. Two color selection methods, i.e., greedy and clustering-based selection method, are proposed to solve the objective function. Experimental results show that clustering-based one outperforms its counterpart to generate the new tongue colorchecker. Compared to Munsell colorchecker, this proposed space-based colorchecker can greatly improve the correction accuracy by 48%. Further experimental results on more correction task also validate its effectiveness and superiority.
Xingzheng Wang, David Zhang 0001
IEEE J. Biomed. Health Informatics2
2013 Robust Kernel Representation With Statistical Local Features for Face Recognition
abstract
Factors such as misalignment, pose variation, and occlusion make robust face recognition a difficult problem. It is known that statistical features such as local binary pattern are effective for local feature extraction, whereas the recently proposed sparse or collaborative representation-based classification has shown interesting results in robust face recognition. In this paper, we propose a novel robust kernel representation model with statistical local features (SLF) for robust face recognition. Initially, multipartition max pooling is used to enhance the invariance of SLF to image registration error. Then, a kernel-based representation model is proposed to fully exploit the discrimination information embedded in the SLF, and robust regression is adopted to effectively handle the occlusion in face images. Extensive experiments are conducted on benchmark face databases, including extended Yale B, AR (A. Martinez and R. Benavente), multiple pose, illumination, and expression (multi-PIE), facial recognition technology (FERET), face recognition grand challenge (FRGC), and labeled faces in the wild (LFW), which have different variations of lighting, expression, pose, and occlusions, demonstrating the promising performance of the proposed method.
Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2013 Handheld System Design for Dual-Eye Multispectral Iris Capture With One Camera
abstract
This paper describes the design and implementation of a dual-eye multispectral iris capture handheld device with one camera, which consists of the following four parts: 1) capture unit; 2) illumination unit; 3) interaction unit; and 4) control unit. A multispectral iris image database is created by the proposed capture device, and then, we use the multispectral score-level fusion to further investigate the effectiveness of the proposed capture device by the 1-D Log-Gabor wavelet filter approach. Experimental results are also presented in this paper.
Yazhuo Gong, David Zhang 0001, Jingqi Yan
IEEE Trans. Syst. Man Cybern. Syst.2
2013 Palm-Print Classification by Global Features
abstract
Three-dimensional (3-D) palm print has proved to be a significant biometrics for personal authentication. Three-dimensional palm prints are harder to counterfeit than 2-D palm prints and more robust to variations in illumination and serious scrabbling on the palm surface. Previous work on 3-D palm-print recognition has concentrated on local features such as texture and lines. In this paper, we propose three novel global features of 3-D palm prints which describe shape information and can be used for coarse matching and indexing to improve the efficiency of palm-print recognition, particularly in very large databases. The three proposed shape features are maximum depth of palm center, horizontal cross-sectional area of different levels, and radial line length from the centroid to the boundary of 3-D palm-print horizontal cross section of different levels. We treat these features as a column vector and use orthogonal linear discriminant analysis to reduce their dimensionality. We then adopt two schemes: 1) coarse-level matching and 2) ranking support vector machine to improve the efficiency of palm-print recognition. We conducted a series of 3-D palm-print recognition experiments using an established 3-D palm-print database, and the results demonstrate that the proposed method can greatly reduce penetration rates.
Bob Zhang 0001, Wei Li 0016, Pei Qing, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2012 All aboard the Databus!: Linkedin's scalable consistent change data capture platform
abstract
In Internet architectures, data systems are typically categorized into source-of-truth systems that serve as primary stores for the user-generated writes, and derived data stores or indexes which serve reads and other complex queries. The data in these secondary stores is often derived from the primary data through custom transformations, sometimes involving complex processing driven by business logic. Similarly data in caching tiers is derived from reads against the primary data store, but needs to get invalidated or refreshed when the primary data gets mutated. A fundamental requirement emerging from these kinds of data architectures is the need to reliably capture, flow and process primary data changes.
Shirshanka Das, Chavdar Botev, Kapil Surlaker, Bhaskar Ghosh, Balaji Varadarajan, Sunil Nagaraj, David Zhang 0001, Jemiah Westerman, Phanindra Ganti, Boris Shkolnik, Sajid Topiwala, Alexander Pachev, Naveen Somasundaram, Subbu Subramaniam
SoCC7
2012 Relaxed collaborative representation for pattern classification
abstract
Regularized linear representation learning has led to interesting results in image classification, while how the object should be represented is a critical issue to be investigated. Considering the fact that the different features in a sample should contribute differently to the pattern representation and classification, in this paper we present a novel relaxed collaborative representation (RCR) model to effectively exploit the similarity and distinctiveness of features. In RCR, each feature vector is coded on its associated dictionary to allow flexibility of feature coding, while the variance of coding vectors is minimized to address the similarity among features. In addition, the distinctiveness of different features is exploited by weighting its distance to other features in the coding domain. The proposed RCR is simple, while our extensive experimental results on benchmark image databases (e.g., various face and flower databases) show that it is very competitive with state-of-the-art image classification methods.
Meng Yang 0001, Lei Zhang 0006, David Zhang 0001, Shenlong Wang
CVPR3
2012 Efficient Misalignment-Robust Representation for Real-Time Face Recognition
Meng Yang 0001, Lei Zhang 0006, David Zhang 0001
ECCV (1)3
2012 Data Infrastructure at LinkedIn
abstract
Linked In is among the largest social networking sites in the world. As the company has grown, our core data sets and request processing requirements have grown as well. In this paper, we describe a few selected data infrastructure projects at Linked In that have helped us accommodate this increasing scale. Most of those projects build on existing open source projects and are themselves available as open source. The projects covered in this paper include: (1) Voldemort: a scalable and fault tolerant key-value store, (2) Data bus: a framework for delivering database changes to downstream applications, (3) Espresso: a distributed data store that supports flexible schemas and secondary indexing, (4) Kafka: a scalable and efficient messaging system for collecting various user activity events and log data.
Aditya Auradkar, Chavdar Botev, Shirshanka Das, Dave De Maagd, Alex Feinberg, Phanindra Ganti, Bhaskar Ghosh, Kishore Gopalakrishna, Brendan Harris, Joel Koshy, Kevin Krawez, Jay Kreps, Shi Lu, Sunil Nagaraj, Neha Narkhede, Sasha Pachev, Igor Perisic, Lin Qiao, Tom Quiggle, Jun Rao, Bob Schulman, Abraham Sebastian, Oliver Seeliger, Adam Silberstein, Boris Shkolnik, Chinmay Soman, Roshan Sumbaly, Kapil Surlaker, Sajid Topiwala, Cuong Tran 0003, Balaji Varadarajan, Jemiah Westerman, Zach White, David Zhang 0001
ICDE35
2012 A comprehensive evaluation of full reference image quality assessment algorithms
abstract
Recent years have witnessed a growing interest in developing objective image quality assessment (IQA) algorithms that can measure the image quality consistently with subjective evaluations. For the full reference (FR) IQA problem, great progress has been made in the past decade. On the other hand, several new large scale image datasets have been released for evaluating FR IQA methods in recent years. Meanwhile, no work has been reported to evaluate and compare the performance of state-of-the-art and representative FR IQA methods on all the available datasets. In this paper, we aim to fulfill this task by reporting the performance of eleven selected FR IQA algorithms on all the seven public IQA image datasets. Our evaluation results and the associated discussions will be very helpful for relevant researchers to have a clearer understanding about the status of modern FR IQA indices. Evaluation results presented in this paper are also online available at http://sse.tongji.edu.cn/linzhang/IQA/IQA.htm.
Lin Zhang 0014, Lei Zhang 0006, Xuanqin Mou, David Zhang 0001
ICIP4
2012 Vessel segmentation and width estimation in retinal images using multiscale production of matched filter responses
Qin Li 0001, Jane You, David Zhang 0001
Expert Syst. Appl.3
2012 Impact of Full Rank Principal Component Analysis on Classification Algorithms for Face Recognition
abstract
Full rank principal component analysis (FR-PCA) is a special form of principal component analysis (PCA) which retains all nonzero components of PCA. Generally speaking, it is hard to estimate how the accuracy of a classifier will change after data are compressed by PCA. However, this paper reveals an interesting fact that the transformation by FR-PCA does not change the accuracy of many well-known classification algorithms. It predicates that people can safely use FR-PCA as a preprocessing tool to compress high-dimensional data without deteriorating the accuracies of these classifiers. The main contribution of the paper is that it theoretically proves that the transformation by FR-PCA does not change accuracies of the k nearest neighbor, the minimum distance, support vector machine, large margin linear projection, and maximum scatter difference classifiers. In addition, through extensive experimental studies conducted on several benchmark face image databases, this paper demonstrates that FR-PCA can greatly promote the efficiencies of above-mentioned five classification algorithms in appearance-based face recognition.
Fengxi Song, Jane You, David Zhang 0001, Yong Xu 0001
Int. J. Pattern Recognit. Artif. Intell.3
2012 Local directional derivative pattern for rotation invariant texture classification
Zhenhua Guo 0001, Qin Li 0001, Jane You, David Zhang 0001, Wenhuang Liu
Neural Comput. Appl.4
2012 Automatic tongue image segmentation based on gradient vector flow and region merging
Jifeng Ning, David Zhang 0001, Chengke Wu 0001
Neural Comput. Appl.2
2012 Hand shape recognition based on coherent distance shape contexts
Rong-Xiang Hu, Wei Jia 0001, David Zhang 0001, Jie Gui, Liang-Tu Song
Pattern Recognit.3
2012 Optimal subset-division based discrimination and its kernelization for face and palmprint recognition
Xiaoyuan Jing, Sheng Li 0001, David Zhang 0001, Chao Lan, Jing-Yu Yang 0001
Pattern Recognit.3
2012 Phase congruency induced local features for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Zhenhua Guo 0001
Pattern Recognit.3
2012 Face feature extraction and recognition based on discriminant subclass-center manifold preserving projection
Xiaoyuan Jing, Chao Lan, David Zhang 0001, Jing-Yu Yang 0001, Sheng Li 0001, Songhao Zhu
Pattern Recognit. Lett.3
2012 Supervised and Unsupervised Parallel Subspace Learning for Large-Scale Image Recognition
abstract
Subspace learning is an effective and widely used image feature extraction and classification technique. However, for the large-scale image recognition issue in real-world applications, many subspace learning methods often suffer from large computational burden. In order to reduce the computational time and improve the recognition performance of subspace learning technique under this situation, we introduce the idea of parallel computing which can reduce the time complexity by splitting the original task into several subtasks. We develop a parallel subspace learning framework. In this framework, we first divide the sample set into several subsets by designing two random data division strategies that are equal data division and unequal data division. These two strategies correspond to equal and unequal computational abilities of nodes under parallel computing environment. Next, we calculate projection vectors from each subset in parallel. The graph embedding technique is employed to provide a general formulation for parallel feature extraction. After combining the extracted features from all nodes, we present a unified criterion to select most distinctive features for classification. Under the developed framework, we separately propose supervised and unsupervised parallel subspace learning approaches, which are called parallel linear discriminant analysis (PLDA) and parallel locality preserving projection (PLPP). PLDA selects features with the largest Fisher scores by estimating the weighted and unweighted sample scatter, while PLPP selects features with the smallest Laplacian scores by constructing a whole affinity matrix. Theoretically, we analyze the time complexities of proposed approaches and provide the fundamental supports for applying random division strategies. In the experiments, we establish two real parallel computing environments and employ four public image and video databases as the test data. Experimental results demonstrate that the proposed approaches outperform several related supervised and unsupervised subspace learning methods, and significantly reduce the computational time.
Xiaoyuan Jing, Sheng Li 0001, David Zhang 0001, Jian Yang 0003, Jing-Yu Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2012 Feature Selection for Monotonic Classification
abstract
Monotonic classification is a kind of special task in machine learning and pattern recognition. Monotonicity constraints between features and decision should be taken into account in these tasks. However, most existing techniques are not able to discover and represent the ordinal structures in monotonic datasets. Thus, they are inapplicable to monotonic classification. Feature selection has been proven effective in improving classification performance and avoiding overfitting. To the best of our knowledge, no technique has been specially designed to select features in monotonic classification until now. In this paper, we introduce a function, which is called rank mutual information, to evaluate monotonic consistency between features and decision in monotonic tasks. This function combines the advantages of dominance rough sets in reflecting ordinal structures and mutual information in terms of robustness. Then, rank mutual information is integrated with the search strategy of min-redundancy and max-relevance to compute optimal subsets of features. A collection of numerical experiments are given to show the effectiveness of the proposed technique.
Qinghua Hu, Lei Zhang 0006, David Zhang 0001, Yanping Song, Maozu Guo 0001, Daren Yu
IEEE Trans. Fuzzy Syst.4
2012 On Robust Fuzzy Rough Set Models
abstract
Rough sets, especially fuzzy rough sets, are supposedly a powerful mathematical tool to deal with uncertainty in data analysis. This theory has been applied to feature selection, dimensionality reduction, and rule learning. However, it is pointed out that the classical model of fuzzy rough sets is sensitive to noisy information, which is considered as a main source of uncertainty in applications. This disadvantage limits the applicability of fuzzy rough sets. In this paper, we reveal why the classical fuzzy rough set model is sensitive to noise and how noisy samples impose influence on fuzzy rough computation. Based on this discussion, we study the properties of some current fuzzy rough models in dealing with noisy data and introduce several new robust models. The properties of the proposed models are also discussed. Finally, a robust classification algorithm is designed based on fuzzy lower approximations. Some numerical experiments are given to illustrate the effectiveness of the models. The classifiers that are developed with the proposed models achieve good generalization performance.
Qinghua Hu, Lei Zhang 0006, Shuang An, David Zhang 0001, Daren Yu
IEEE Trans. Fuzzy Syst.4
2012 Feature Band Selection for Online Multispectral Palmprint Recognition
abstract
A palmprint is a unique and reliable biometric feature with high usability. In the past decades, many palmprint recognition systems have been successfully developed. However, most of the previous work used the white light as the illumination source, and the recognition accuracy and anti-spoof capability is limited. Recently, multispectral imaging has attracted considerable research attention as it can acquire more discriminative information in a short time. One crucial step in developing online multispectral palmprint systems is how to determine the optimal number of spectral bands and select the most representative bands to build the system. This paper presents a study on feature band selection by analyzing hyperspectral palmprint data (520-1050 nm). Our experimental results showed that three spectral bands could provide most of the discriminate information of a palmprint. This finding could be used as the guidance for designing new online multispectral palmprint systems.
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wenhuang Liu
IEEE Trans. Inf. Forensics Secur.2
2012 Monogenic Binary Coding: An Efficient Local Feature Extraction Approach to Face Recognition
abstract
Local-feature-based face recognition (FR) methods, such as Gabor features encoded by local binary pattern, could achieve state-of-the-art FR results in large-scale face databases such as FERET and FRGC. However, the time and space complexity of Gabor transformation are too high for many practical FR applications. In this paper, we propose a new and efficient local feature extraction scheme, namely monogenic binary coding (MBC), for face representation and recognition. Monogenic signal representation decomposes an original signal into three complementary components: amplitude, orientation, and phase. We encode the monogenic variation in each local region and monogenic feature in each pixel, and then calculate the statistical features (e.g., histogram) of the extracted local features. The local statistical features extracted from the complementary monogenic components (i.e., amplitude, orientation, and phase) are then fused for effective FR. It is shown that the proposed MBC scheme has significantly lower time and space complexity than the Gabor-transformation-based local feature methods. The extensive FR experiments on four large-scale databases demonstrated the effectiveness of MBC, whose performance is competitive with and even better than state-of-the-art local-feature-based FR methods.
Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.4
2012 Rotation-Invariant Nonrigid Point Set Matching in Cluttered Scenes
abstract
This paper addresses the problem of rotation-invariant nonrigid point set matching. The shape context (SC) feature descriptor is used because of its strong discriminative nature, whereas edges in the graphs constructed by point sets are used to determine the orientations of SCs. Similar to lengths or directions, oriented SCs constructed this way can be regarded as attributes of edges. By matching edges between two point sets, rotation invariance is achieved. Two novel ways of constructing graphs on a model point set are proposed, aiming at making the orientations of SCs as robust to disturbances as possible. The structures of these graphs facilitate the use of dynamic programming (DP) for optimization. The strong discriminative nature of SC, the special structure of the model graphs, and the global optimality of DP make our methods robust to various types of disturbances, particularly clutters. The extensive experiments on both synthetic and real data validated the robustness of the proposed methods to various types of disturbances. They can robustly detect the desired shapes in complex and highly cluttered scenes.
Wei Lian, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Image Process.3
2012 Combination of Heterogeneous Features for Wrist Pulse Blood Flow Signal Diagnosis via Multiple Kernel Learning
abstract
Wrist pulse signal is of great importance in the analysis of the health status and pathologic changes of a person. A number of feature extraction methods have been proposed to extract linear and nonlinear, and time and frequency features of wrist pulse signal. These features are heterogeneous in nature and are likely to contain complementary information, which highlights the need for the integration of heterogeneous features for pulse classification and diagnosis. In this paper, we propose a novel effective method to classify the wrist pulse blood flow signals by using the multiple kernel learning (MKL) algorithm to combine multiple types of features. In the proposed method, seven types of features are first extracted from the wrist pulse blood flow signals using the state-of-the-art pulse feature extraction methods, and are then fed to an efficient MKL method, SimpleMKL, to combine heterogeneous features for more effective classification. Experimental results show that the proposed method is promising in integrating multiple types of pulse features to further enhance the classification performance.
Lei Liu 0049, Wangmeng Zuo, David Zhang 0001, Naimin Li
IEEE Trans. Inf. Technol. Biomed.3
2012 Rank Entropy-Based Decision Trees for Monotonic Classification
abstract
In many decision making tasks, values of features and decision are ordinal. Moreover, there is a monotonic constraint that the objects with better feature values should not be assigned to a worse decision class. Such problems are called ordinal classification with monotonicity constraint. Some learning algorithms have been developed to handle this kind of tasks in recent years. However, experiments show that these algorithms are sensitive to noisy samples and do not work well in real-world applications. In this work, we introduce a new measure of feature quality, called rank mutual information (RMI), which combines the advantage of robustness of Shannon's entropy with the ability of dominance rough sets in extracting ordinal structures from monotonic data sets. Then, we design a decision tree algorithm (REMT) based on rank mutual information. The theoretic and experimental analysis shows that the proposed algorithm can get monotonically consistent decision trees, if training samples are monotonically consistent. Its performance is still good when data are contaminated with noise.
Qinghua Hu, Xunjian Che, Lei Zhang 0006, David Zhang 0001, Maozu Guo 0001, Daren Yu
IEEE Trans. Knowl. Data Eng.4
2012 A Novel 3-D Palmprint Acquisition System
abstract
Palmprints have been widely studied for personal authentication because they are highly accurate and incur low costs. Most of the previous work has focused on two-dimensional (2-D) palmprint identification. However, the inner surfaces of palms contain not only texture information but also shape information. Unfortunately, 2-D palmprint systems lose the shape information when capturing palmprint images. Hence, three-dimensional (3-D) information is important for palmprint systems. In this paper, we have designed and developed a novel 3-D palmprint acquisition system based on structured-light imaging technology. The acquisition system can obtain 3-D palmprint information and, at the same time, the corresponding 2-D texture, which are used for personal authentication. A 3-D palmprint database that contains 8000 samples has been established by using the developed acquisition system, and the test results illustrate the effectiveness of our system.
Wei Li 0016, David Zhang 0001, Guangming Lu 0002, Nan Luo
IEEE Trans. Syst. Man Cybern. Part A2
2011 Robust sparse coding for face recognition
abstract
Recently the sparse representation (or coding) based classification (SRC) has been successfully used in face recognition. In SRC, the testing image is represented as a sparse linear combination of the training samples, and the representation fidelity is measured by the l2-norm or l1-norm of coding residual. Such a sparse coding model actually assumes that the coding residual follows Gaussian or Laplacian distribution, which may not be accurate enough to describe the coding errors in practice. In this paper, we propose a new scheme, namely the robust sparse coding (RSC), by modeling the sparse coding as a sparsity-constrained robust regression problem. The RSC seeks for the MLE (maximum likelihood estimation) solution of the sparse coding problem, and it is much more robust to outliers (e.g., occlusions, corruptions, etc.) than SRC. An efficient iteratively reweighted sparse coding algorithm is proposed to solve the RSC model. Extensive experiments on representative face databases demonstrate that the RSC scheme is much more effective than state-of-the-art methods in dealing with face occlusion, corruption, lighting and expression changes, etc.
Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001
CVPR4
2011 Fisher Discrimination Dictionary Learning for sparse representation
abstract
Sparse representation based classification has led to interesting image recognition results, while the dictionary used for sparse coding plays a key role in it. This paper presents a novel dictionary learning (DL) method to improve the pattern classification performance. Based on the Fisher discrimination criterion, a structured dictionary, whose dictionary atoms have correspondence to the class labels, is learned so that the reconstruction error after sparse coding can be used for pattern classification. Meanwhile, the Fisher discrimination criterion is imposed on the coding coefficients so that they have small within-class scatter but big between-class scatter. A new classification scheme associated with the proposed Fisher discrimination DL (FDDL) method is then presented by using both the discriminative information in the reconstruction error and sparse coding coefficients. The proposed FDDL is extensively evaluated on benchmark image databases in comparison with existing sparse representation and DL based classification methods.
Meng Yang 0001, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001
ICCV4
2011 A linear subspace learning approach via sparse coding
abstract
Linear subspace learning (LSL) is a popular approach to image recognition and it aims to reveal the essential features of high dimensional data, e.g., facial images, in a lower dimensional space by linear projection. Most LSL methods compute directly the statistics of original training samples to learn the subspace. However, these methods do not effectively exploit the different contributions of different image components to image recognition. We propose a novel LSL approach by sparse coding and feature grouping. A dictionary is learned from the training dataset, and it is used to sparsely decompose the training samples. The decomposed image components are grouped into a more discriminative part (MDP) and a less discriminative part (LDP). An unsupervised criterion and a supervised criterion are then proposed to learn the desired subspace, where the MDP is preserved and the LDP is suppressed simultaneously. The experimental results on benchmark face image databases validated that the proposed methods outperform many state-of-the-art LSL schemes.
Lei Zhang 0006, Pengfei Zhu 0001, Qinghua Hu, David Zhang 0001
ICCV4
2011 Adaptive Weighted Fusion of Local Kernel Classifiers for Effective Pattern Classification
Shixin Yang, Wangmeng Zuo, Lei Liu 0049, Yanlai Li, David Zhang 0001
ICIC (1)5
2011 Face recognition based on local uncorrelated and weighted global uncorrelated discriminant transforms
abstract
Feature extraction is one of the most important problems in image recognition tasks. In many applications such as face recognition, it is desirable to eliminate the redundancy among the extracted discriminant features. In this paper, we propose two novel feature extraction approaches named local uncorrelated discriminant transform (LUDT) and weighted global uncorrelated discriminant transform (WGUDT) for face recognition, respectively. LUDT and WGUDT separately construct the local uncorrelated constraints and the weighted global uncorrelated constraints. Then they iteratively calculate the optimal discriminant vectors that maximize the Fisher criterion under the corresponding statistical uncorrelated constraints, respectively. The proposed LUDT and WGUDT approaches are evaluated on the public AR and FERET face databases. Experimental results demonstrate that the proposed approaches outperform several representative feature extraction methods.
Xiaoyuan Jing, Sheng Li 0001, David Zhang 0001, Jing-Yu Yang 0001
ICIP3
2011 Discriminant subclass-center manifold preserving projection for face feature extraction
abstract
Manifold learning is an effective feature extraction technique, which seeks a low-dimensional space where the manifold structure, in terms of local neighborhood, of the data set can be well preserved. A typical manifold learning method constructs a local neighborhood centered at individual samples. In this paper, we propose to construct local neighborhoods that centered at subclass centers, and seek an embedded space where such neighborhood is well preserved. We show from a probability perspective that, neighbors of a subclass center would contain more intra-class data than inter-class data, which may be desirable for discrimination. Meanwhile, we simultaneously enhance the discriminative power of extracted features by maximizing the Fisher ratio of embedded data based on subclass centers. Experimental results on CAS-PEAL and FERET face databases demonstrate that our proposed approach is more effective than most typical manifold learning methods and their supervised extensions in classification performance.
Chao Lan, Xiaoyuan Jing, David Zhang 0001, Shi-Qiang Gao, Jing-Yu Yang 0001
ICIP3
2011 A novel kernel discriminant feature extraction framework based on mapped virtual samples for face recognition
abstract
In this paper, we propose a novel kernel discriminant feature extraction framework based on the mapped virtual samples (MVS) for face recognition. We calculate a non-symmetric kernel matrix by constructing a few virtual samples (including eigen-samples and common vector samples) in the input space, and then express kernel projection vectors by using mapped virtual samples (MVS). Under this framework, we realize two MVS-based representative kernel methods including kernel principal component analysis (KPCA) and generalized discriminant analysis (GDA). Experimental results on the AR and CAS-PEAL face databases demonstrate that the proposed framework can effectively improve the classification performance of kernel discriminant methods. In addition, the MVS-based kernel approaches have a lower computational cost in contrast with the related kernel methods.
Sheng Li 0001, Xiaoyuan Jing, David Zhang 0001, Yong-Fang Yao, Lu-Sha Bian
ICIP3
2011 An augmented Lagrangian method for fast gradient vector flow computation
abstract
Gradient vector flow (GVF) and its generalization have been widely applied in many image processing applications. The high cost of GVF computation, however, has restricted their potential applications to images with large size. In this paper, motivated by progress in fast image restoration algorithms, we reformulate the GVF computation problem as a convex optimization model with an equality constraint, and solve it using a fast algorithm, inexact augmented Lagrangian method (ALM). With fast Fourier transform (FFT), we provide a novel simple and efficient algorithm for GVF computation. Experimental results show that the proposed method can improve the computational speed by an order of magnitude, and is even more efficient for images with large sizes.
Jianfeng Lu 0003, Wangmeng Zuo, David Zhang 0001
ICIP4
2011 Sparse cost-sensitive classifier with application to face recognition
abstract
Sparse representation technique has been successfully employed to solve face recognition task. Though current sparse representation based classifier proves to achieve high classification accuracy, it implicitly assumes that the losses of all misclassifications are the same. However, in many real-world applications, different misclassifications could lead to different losses. Driven by this concern, we propose in this paper a sparse cost-sensitive classifier for face recognition. Our approach uses probabilistic model of sparse representation to estimate the posterior probabilities of a testing sample, calculates all the misclassification losses via the posterior probabilities and then predicts the class label by minimizing the losses. Experimental results on the public AR and FRGC face databases validate the efficacy of the proposed approach.
Jiang-Yue Man, Xiaoyuan Jing, David Zhang 0001, Chao Lan
ICIP3
2011 Measuring relevance between discrete and continuous features based on neighborhood mutual information
Qinghua Hu, Lei Zhang 0006, David Zhang 0001, Shuang An, Witold Pedrycz
Expert Syst. Appl.3
2011 Online joint palmprint and palmvein verification
David Zhang 0001, Zhenhua Guo 0001, Guangming Lu 0002, Lei Zhang 0006, Wangmeng Zuo
Expert Syst. Appl.1
2011 Combine crossing matching scores with conventional matching scores for bimodal biometrics and face and palmprint recognition experiments
Yong Xu 0001, Qi Zhu 0001, David Zhang 0001
Neurocomputing3
2011 Accelerating the kernel-method-based feature extraction procedure from the viewpoint of numerical approximation
Yong Xu 0001, David Zhang 0001
Neural Comput. Appl.2
2011 A novel hierarchical fingerprint matching approach
Feng Liu 0013, Qijun Zhao, David Zhang 0001
Pattern Recognit.3
2011 Image segmentation by iterated region merging with localized graph cuts
Bo Peng 0006, Lei Zhang 0006, David Zhang 0001, Jian Yang 0003
Pattern Recognit.3
2011 Facial expression recognition on multiple manifolds
Qijun Zhao, David Zhang 0001
Pattern Recognit.3
2011 From classifiers to discriminators: A nearest neighbor rule induced discriminant analysis
Jian Yang 0003, Lei Zhang 0006, Jing-Yu Yang 0001, David Zhang 0001
Pattern Recognit.4
2011 Quantitative analysis of human facial beauty using geometric features
David Zhang 0001, Qijun Zhao, Fangmei Chen
Pattern Recognit.1
2011 Ensemble of local and global information for finger-knuckle-print recognition
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Hailong Zhu
Pattern Recognit.3
2011 On accurate orientation extraction and appropriate distance measure for low-resolution palmprint recognition
Wangmeng Zuo, David Zhang 0001
Pattern Recognit.3
2011 Empirical study of light source selection for palmprint recognition
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wangmeng Zuo, Guangming Lu 0002
Pattern Recognit. Lett.2
2011 Fast palmprint identification with multiple templates per subject
Wangmeng Zuo, David Zhang 0001, Bin Li 0053
Pattern Recognit. Lett.3
2011 Color image canonical correlation analysis for face feature extraction and recognition
Xiaoyuan Jing, Sheng Li 0001, Chao Lan, David Zhang 0001, Jing-Yu Yang 0001, Qian Liu 0010
Signal Process.4
2011 A Two-Phase Test Sample Sparse Representation Method for Use With Face Recognition
abstract
In this paper, we propose a two-phase test sample representation method for face recognition. The first phase of the proposed method seeks to represent the test sample as a linear combination of all the training samples and exploits the representation ability of each training sample to determine M “nearest neighbors” for the test sample. The second phase represents the test sample as a linear combination of the determined M nearest neighbors and uses the representation result to perform classification. We propose this method with the following assumption: the test sample and its some neighbors are probably from the same class. Thus, we use the first phase to detect the training samples that are far from the test sample and assume that these samples have no effects on the ultimate classification decision. This is helpful to accurately classify the test sample. We will also show the probability explanation of the proposed method. A number of face recognition experiments show that our method performs very well.
Yong Xu 0001, David Zhang 0001, Jian Yang 0003, Jing-Yu Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 A Unified Framework for Contactless Hand Verification
abstract
Two-dimensional (2-D) hand-geometry features carry limited discriminatory information and therefore yield moderate performance when utilized for personal identification. This paper investigates a new approach to achieve performance improvement by simultaneously acquiring and combining three-dimensional (3-D) and 2-D features from the human hand. The proposed approach utilizes a 3-D digitizer to simultaneously acquire intensity and range images of the presented hands of the users in a completely contact-free manner. Two new representations that effectively characterize the local finger surface features are extracted from the acquired range images and are matched using the proposed matching metrics. In addition, the characterization of 3-D palm surface using SurfaceCode is proposed for matching a pair of 3-D palms. The proposed approach is evaluated on a database of 177 users acquired in two sessions. The experimental results suggest that the proposed 3-D hand-geometry features have significant discriminatory information to reliably authenticate individuals. Our experimental results demonstrate that consolidating 3-D and 2-D hand-geometry features results in significantly improved performance that cannot be achieved with the traditional 2-D hand-geometry features alone. Furthermore, this paper also investigates the performance improvement that can be achieved by integrating five biometric features, i.e., 2-D palmprint, 3-D palmprint, finger texture, along with 3-D and 2-D hand-geometry features, that are simultaneously extracted from the user's hand presented for authentication.
Vivek Kanhangad, Ajay Kumar 0001, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.3
2011 Contactless and Pose Invariant Biometric Identification Using Hand Surface
abstract
This paper presents a novel approach for hand matching that achieves significantly improved performance even in the presence of large hand pose variations. The proposed method utilizes a 3-D digitizer to simultaneously acquire intensity and range images of the user's hand presented to the system in an arbitrary pose. The approach involves determination of the orientation of the hand in 3-D space followed by pose normalization of the acquired 3-D and 2-D hand images. Multimodal (2-D as well as 3-D) palmprint and hand geometry features, which are simultaneously extracted from the user's pose normalized textured 3-D hand, are used for matching. Individual matching scores are then combined using a new dynamic fusion strategy. Our experimental results on the database of 114 subjects with significant pose variations yielded encouraging results. Consistent (across various hand features considered) performance improvement achieved with the pose correction demonstrates the usefulness of the proposed approach for hand based biometric systems with unconstrained and contact-free imaging. The experimental results also suggest that the dynamic fusion approach employed in this work helps to achieve performance improvement of 60% (in terms of EER) over the case when matching scores are combined using the weighted sum rule.
Vivek Kanhangad, Ajay Kumar 0001, David Zhang 0001
IEEE Trans. Image Process.3
2011 Automatic Image Segmentation by Dynamic Region Merging
abstract
This paper addresses the automatic image segmentation problem in a region merging style. With an initially oversegmented image, in which many regions (or superpixels) with homogeneous color are detected, an image segmentation is performed by iteratively merging the regions according to a statistical test. There are two essential issues in a region-merging algorithm: order of merging and the stopping criterion. In the proposed algorithm, these two issues are solved by a novel predicate, which is defined by the sequential probability ratio test and the minimal cost criterion. Starting from an oversegmented image, neighboring regions are progressively merged if there is an evidence for merging according to this predicate. We show that the merging order follows the principle of dynamic programming. This formulates the image segmentation as an inference problem, where the final segmentation is established based on the observed image. We also prove that the produced segmentation satisfies certain global properties. In addition, a faster algorithm is developed to accelerate the region-merging process, which maintains a nearest neighbor graph in each iteration. Experiments on real natural images are conducted to demonstrate the performance of the proposed dynamic region-merging algorithm.
Bo Peng 0006, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Image Process.3
2011 FSIM: A Feature Similarity Index for Image Quality Assessment
abstract
Image quality assessment (IQA) aims to use computational models to measure the image quality consistently with subjective evaluations. The well-known structural similarity index brings IQA from pixel- to structure-based stage. In this paper, a novel feature similarity (FSIM) index for full reference IQA is proposed based on the fact that human visual system (HVS) understands an image mainly according to its low-level features. Specifically, the phase congruency (PC), which is a dimensionless measure of the significance of a local structure, is used as the primary feature in FSIM. Considering that PC is contrast invariant while the contrast information does affect HVS' perception of image quality, the image gradient magnitude (GM) is employed as the secondary feature in FSIM. PC and GM play complementary roles in characterizing the image local quality. After obtaining the local quality map, we use PC again as a weighting function to derive a single quality score. Extensive experiments performed on six benchmark IQA databases demonstrate that FSIM can achieve much higher consistency with the subjective evaluations than state-of-the-art IQA metrics.
Lin Zhang 0014, Lei Zhang 0006, Xuanqin Mou, David Zhang 0001
IEEE Trans. Image Process.4
2011 Local Linear Discriminant Analysis Framework Using Sample Neighbors
abstract
The linear discriminant analysis (LDA) is a very popular linear feature extraction approach. The algorithms of LDA usually perform well under the following two assumptions. The first assumption is that the global data structure is consistent with the local data structure. The second assumption is that the input data classes are Gaussian distributions. However, in real-world applications, these assumptions are not always satisfied. In this paper, we propose an improved LDA framework, the local LDA (LLDA), which can perform well without needing to satisfy the above two assumptions. Our LLDA framework can effectively capture the local structure of samples. According to different types of local data structure, our LLDA framework incorporates several different forms of linear feature extraction approaches, such as the classical LDA and principal component analysis. The proposed framework includes two LLDA algorithms: a vector-based LLDA algorithm and a matrix-based LLDA (MLLDA) algorithm. MLLDA is directly applicable to image recognition, such as face recognition. Our algorithms need to train only a small portion of the whole training set before testing a sample. They are suitable for learning large-scale databases especially when the input data dimensions are very high and can achieve high classification accuracy. Extensive experiments show that the proposed algorithms can obtain good classification results.
Zizhu Fan, Yong Xu 0001, David Zhang 0001
IEEE Trans. Neural Networks3
2011 3-D Palmprint Recognition With Joint Line and Orientation Features
abstract
2-D palmprint has been recognized as an effective biometric identifier in the past decade. Recently, 3-D palmprint recognition was proposed to further improve the performance of palmprint systems. This paper presents a simple yet efficient scheme for 3-D palmprint recognition. After calculating and enhancing the mean-curvature image of the 3-D palmprint data, we extract both line and orientation features from it. The two types of features are then fused at either score level or feature level for the final 3-D palmprint recognition. The experiments on The Hong Kong Polytechnic University 3-D palmprint database, which contains 8000 samples from 400 palms show that the proposed feature extraction and fusion methods lead to promising performance.
Wei Li 0016, David Zhang 0001, Lei Zhang 0006, Guangming Lu 0002, Jingqi Yan
IEEE Trans. Syst. Man Cybern. Part C2
2010 Efficient joint 2D and 3D palmprint matching with alignment refinement
abstract
Palmprint verification is a relatively new but promising personal authentication technique for its high accuracy and fast matching speed. Two dimensional (2D) palmprint recognition has been well studied in the past decade, and recently three dimensional (3D) palmprint recognition techniques were also proposed. The 2D and 3D palmprint data can be captured simultaneously and they provide different and complementary information. 3D palmprint contains the depth information of the palm surface, while 2D palmprint contains plenty of textures. How to efficiently extract and fuse the 2D and 3D palmprint features to improve the recognition performance is a critical issue for practical palmprint systems. In this paper, an efficient joint 2D and 3D palmprint matching scheme is proposed. The principal line features and palm shape features are extracted and used to accurately align the palmprint, and a couple of matching rules are defined to efficiently use the 2D and 3D features for recognition. The experiments on a 2D+3D palmprint database which contains 8000 samples show that the proposed scheme can greatly improve the performance of palmprint verification.
Wei Li 0016, Lei Zhang 0006, David Zhang 0001, Guangming Lu 0002, Jingqi Yan
CVPR3
2010 The multiscale competitive code via sparse representation for palmprint verification
abstract
Palm lines are the most important features for palmprint recognition. They are best considered as typical multiscale features, where the principal lines can be represented at a larger scale while the wrinkles at a smaller scale. Motivated by the success of coding-based palmprint recognition methods, this paper investigates a compact representation of multiscale palm line orientation features, and proposes a novel method called the sparse multiscale competitive code (SMCC). The SMCC method first defines a filter bank of second derivatives of Gaussians with different orientations and scales, and then uses the l1-norm sparse coding to obtain a robust estimation of the multiscale orientation field. Finally, a generalized competitive code is used to encode the dominant orientation. Experimental results show that the SMCC achieves higher verification accuracy than state-of-the-art palmprint recognition methods, yet uses a smaller template size than other multiscale methods.
Wangmeng Zuo, Zhouchen Lin, Zhenhua Guo 0001, David Zhang 0001
CVPR4
2010 Video action recognition with spatio-temporal graph embedding and spline modeling
abstract
In recent years, video analysis and event recognition are becoming a popular research topic with wide applications in surveillance and security. In this paper, we proposed a video action appearance modeling based on spatio-temporal graph embedding and video action recognition based on video luminance field trajectory spline modeling and aligned matching. Graphs are computed from spline re-sampling of training video data set. Matching is achieved from minimizing the average projection distance between query clips and training groups. Simulation with the Cambridge hand gesture data set demonstrates the effectiveness of the proposed solution.
Yin Yuan, Haomian Zheng, Zhu Li 0001, David Zhang 0001
ICASSP4
2010 Hierarchical multiscale LBP for face and palmprint recognition
abstract
Local binary pattern (LBP), fast and simple for implementation, has shown its superiority in face and palmprint recognition. To extract representative features, “uniform” LBP was proposed and its effectiveness has been validated. However, all “non-uniform” patterns are clustered into one pattern, so a lot of useful information is lost. In this study, the authors propose to build a hierarchical multiscale LBP histogram for an image. The useful information of “non-uniform” patterns at large scale is dug out from its counterpart of small scale. The main advantage of the proposed scheme is that it can fully utilize LBP information while it does not need any training step, which may be sensitive to training samples. Experiments on one public face database and one palmprint database show the effectiveness of the proposed method.
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001, Xuanqin Mou
ICIP3
2010 Rotation invariant texture classification using adaptive LBP with directional statistical features
abstract
Local Binary Pattern (LBP) has been widely used in texture classification because of its simplicity and computational efficiency. Traditional LBP codes the sign of the local difference and uses the histogram of the binary code to model the given image. However, the directional statistical information is ignored in LBP. In this paper, some directional statistical features, specifically the mean and standard deviation of the local absolute difference are extracted and used to improve the LBP classification efficiency. In addition, the least square estimation is used to adaptively minimize the local difference for more stable directional statistical features, and we call this scheme the adaptive LBP (ALBP). By coupling the directional statistical features with ALBP, a new rotation invariant texture classification method is presented. Experiments on a large texture database show that the proposed texture feature extraction and classification scheme could significantly improve the classification accuracy of LBP.
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001
ICIP3
2010 Holistic orthogonal analysis of discriminant transforms for color face recognition
abstract
The key of color face recognition technique is how to effectively utilize the complementary information between color components and remove their redundancy. Present color face recognition methods generally reduce the correlations between color components in the image pixel level, and then extract the discriminant features from the uncorrelated color face images. In this paper, we propose a novel color face recognition approach based on the holistic orthogonal analysis (HOA) of discriminant transforms of color images. HOA can reduce the correlation of color information in the feature level. It in turn achieves the discriminant transforms of red, green and blue color images by using the Fisher criterion, and simultaneously makes the achieved transforms mutually orthogonal. Experimental results on the AR and FRGC-2 public color face image databases demonstrate that the proposed approach acquires better recognition performance than several representative color face recognition methods.
Xiaoyuan Jing, Qian Liu 0010, Chao Lan, Jiang-Yue Man, Sheng Li 0001, David Zhang 0001
ICIP6
2010 Texture classification via patch-based sparse texton learning
abstract
Texture classification is a classical yet still active topic in computer vision and pattern recognition. Recently, several new texture classification approaches by modeling texture images as distributions over a set of textons have been proposed. These textons are learned as the cluster centers in the image patch feature space using the K-means clustering algorithm. However, the Euclidian distance based the K-means clustering process may not be able to well characterize the intrinsic feature space of texture textons, which if often embedded into a low dimensional manifold. Inspired by the great success of l1-norm minimization based sparse representation (SR), in this paper we propose a novel texture classification method via patch-based sparse texton learning. Specifically, the dictionary of textons is learned by applying SR to image patches in the training dataset. The SR coefficients of the test images over the dictionary are used to construct the histograms for texture classification. Experimental results on benchmark database validate the effectiveness of the proposed method.
Lei Zhang 0006, Jane You, David Zhang 0001
ICIP4
2010 Metaface learning for sparse representation based face recognition
abstract
Face recognition (FR) is an active yet challenging topic in computer vision applications. As a powerful tool to represent high dimensional data, recently sparse representation based classification (SRC) has been successfully used for FR. This paper discusses the metaface learning (MFL) of face images under the framework of SRC. Although directly using the training samples as dictionary bases can achieve good FR performance, a well learned dictionary matrix can lead to higher FR rate with less dictionary atoms. An SRC oriented unsupervised MFL algorithm is proposed in this paper and the experimental results on benchmark face databases demonstrated the improvements brought by the proposed MFL algorithm over original SRC.
Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001
ICIP4
2010 ICP registration using principal line and orientation features for palmprint alignment
abstract
Image alignment is a crucial step for palmprint recognition. Current key point based palmprint pre-processing methods, however, can only provide coarse alignment. The rotation of extracted region of interest (ROI) often causes the failure of genuine matching. To solve this problem, in this paper we propose to use iterative closest point (ICP) algorithm for palmprint alignment before matching of features. Using both the palm line and orientation features extracted by high order steerable filters, the proposed method has fast convergence speed and high registration accuracy, which result in greatly improved verification accuracy. Experimental results on Hong Kong PolyU palmprint database demonstrate the effectiveness of the proposed method.
Wangmeng Zuo, David Zhang 0001
ICIP3
2010 Monogenic-LBP: A new approach for rotation invariant texture classification
abstract
Analysis of two-dimensional textures has many potential applications in computer vision. In this paper, we investigate the problem of rotation invariant texture classification, and propose a novel texture feature extractor, namely Monogenic-LBP (M-LBP). M-LBP integrates the traditional Local Binary Pattern (LBP) operator with the other two rotation invariant measures: the local phase and the local surface type computed by the 1st-order and 2nd-order Riesz transforms, respectively. The classification is based on the image's histogram of M-LBP responses. Extensive experiments conducted on the CUReT database demonstrate the overall superiority of M-LBP over the other state-of-the-art methods evaluated.
Lin Zhang 0014, Lei Zhang 0006, Zhenhua Guo 0001, David Zhang 0001
ICIP4
2010 A comparative study on quality assessment of high resolution fingerprint images
abstract
High resolution fingerprint images have been increasingly used in fingerprint recognition. They can provide more fine features (e.g. pores) than standard fingerprint images to improve the recognition accuracy. It is however still an open issue whether or not existing quality assessment methods are suitable for high resolution fingerprint images. This paper compares some typical quality indexes by analyzing the correlation between them and their prediction ability on minutia-based and pore-based high resolution fingerprint recognition accuracy. Experimental results show that the indexes based on ridge orientation are more effective for high resolution fingerprint recognition systems.
Qijun Zhao, Feng Liu 0013, Lei Zhang 0006, David Zhang 0001
ICIP4
2010 Feature Band Selection for Multispectral Palmprint Recognition
abstract
Palm print is a unique and reliable biometric characteristic with high usability. Many palm print recognition algorithms and systems have been successfully developed in the past decades. Most of the previous works use the white light sources for illumination. Recently, it has been attracting much research attention on developing new biometric systems with both high accuracy and high anti-spoof capability. Multispectral palm print imaging and recognition can be a potential solution to such systems because it can acquire more discriminative information for personal identity recognition. One crucial step in developing such systems is how to determine the minimal number of spectral bands and select the most representative bands to build the multispectral imaging system. This paper presents preliminary studies on feature band selection by analyzing hyper spectral palm print data (420nm~1100nm). Our experiments showed that 2 spectral bands at 700nm and 960nm could provide most discriminate information of palm print. This finding could be used as the guidance for designing multispectral palm print systems in the future.
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001
ICPR3
2010 Fingerprint Pore Matching Based on Sparse Representation
abstract
This paper proposes an improved direct fingerprint pore matching method. It measures the differences between pores by using the sparse representation technique. The coarse pore correspondences are then established and weighted based on the obtained differences. The false correspondences among them are finally removed by using the weighted RANSAC algorithm. Experimental results have shown that the proposed method can greatly improve the accuracy of existing methods.
Feng Liu 0013, Qijun Zhao, Lei Zhang 0006, David Zhang 0001
ICPR4
2010 Data Classification on Multiple Manifolds
abstract
Unlike most previous manifold-based data classification algorithms assume that all the data points are on a single manifold, we expect that data from different classes may reside on different manifolds of possible different dimensions. Therefore, better classification accuracy would be achieved by modeling the data by multiple manifolds each corresponding to a class. To this end, a general framework for data classification on multiple manifolds is presented. The manifolds are firstly learned for each class separately, and a stochastic optimization algorithm is then employed to get the near optimal dimensionality of each manifold from the classification viewpoint. Then, classification is performed under a newly defined minimum reconstruction error based classifier. Our method could be easily extended by involving various manifold learning methods and searching strategies. Experiments on both synthetic data and databases of facial expression images show the effectiveness of the proposed multiple manifold based approach.
Qijun Zhao, David Zhang 0001
ICPR3
2010 Monogenic Binary Pattern (MBP): A Novel Feature Extraction and Representation Model for Face Recognition
abstract
A novel feature extraction method, namely monogenic binary pattern (MBP), is proposed in this paper based on the theory of monogenic signal analysis, and the histogram of MBP (HMBP) is subsequently presented for robust face representation and recognition. MBP consists of two parts: one is monogenic magnitude encoded via uniform LBP, and the other is monogenic orientation encoded as quadrant-bit codes. The HMBP is established by concatenating the histograms of MBP of all sub-regions. Compared with the well-known and powerful Gabor filtering based LBP schemes, one clear advantage of HMBP is its lower time and space complexity because monogenic signal analysis needs fewer convolutions and generates more compact feature vectors. The experimental results on the AR and FERET face databases validate that the proposed MBP algorithm has better performance than or comparable performance with state-of-the-art local feature based methods but with significantly lower time and space complexity.
Meng Yang 0001, Lei Zhang 0006, Lin Zhang 0014, David Zhang 0001
ICPR4
2010 On the Dimensionality Reduction for Sparse Representation Based Face Recognition
abstract
Face recognition (FR) is an active yet challenging topic in computer vision applications. As a powerful tool to represent high dimensional data, recently sparse representation based classification (SRC) has been successfully used for FR. This paper discusses the dimensionality reduction (DR) of face images under the framework of SRC. Although one important merit of SRC is that it is insensitive to DR or feature extraction, a well trained projection matrix can lead to higher FR rate at a lower dimensionality. An SRC oriented unsupervised DR algorithm is proposed in this paper and the experimental results on benchmark face databases demonstrated the improvements brought by the proposed DR algorithm over PCA or random projection based DR under the SRC framework.
Lei Zhang 0006, Meng Yang 0001, Zhizhao Feng, David Zhang 0001
ICPR4
2010 Gaussian ERP Kernel Classifier for Pulse Waveforms Classification
abstract
While advances in sensor and signal processing techniques have provided effective tools for quantitative research on traditional Chinese pulse diagnosis (TCPD), the automatic classification of pulse waveforms is remained a difficult problem. To address this issue, this paper proposed a novel edit distance with real penalty (ERP)-based k-nearest neighbors (KNN) classifier by referring to recent progresses in time series matching and KNN classifier. Taking advantage of the metric property of ERP, we first develop a Gaussian ERP kernel, and then embed it into kernel difference-weighted KNN classifier. The proposed Gaussian ERP kernel classifier is evaluated on a dataset which includes 2470 pulse waveforms. Experimental results show that the proposed classifier is much more accurate than several other pulse waveform classification approaches.
Dongyu Zhang 0002, Wangmeng Zuo, David Zhang 0001, Yanlai Li, Naimin Li
ICPR3
2010 Time Series Classification Using Support Vector Machine with Gaussian Elastic Metric Kernel
abstract
Motivated by the great success of dynamic time warping (DTW) in time series matching, Gaussian DTW kernel had been developed for support vector machine (SVM)-based time series classification. Counter-examples, however, had been subsequently reported that Gaussian DTW kernel usually cannot outperform Gaussian RBF kernel in the SVM framework. In this paper, by extending the Gaussian RBF kernel, we propose one novel class of Gaussian elastic metric kernel (GEMK), and present two examples of GEMK: Gaussian time warp edit distance (GTWED) kernel and Gaussian edit distance with real penalty (GERP) kernel. Experimental results on UCR time series data sets show that, in terms of classification accuracy, SVM with GEMK is much superior to SVM with Gaussian RBF kernel and Gaussian DTW kernel, and the state-of-the-art similarity measure methods.
Dongyu Zhang 0002, Wangmeng Zuo, David Zhang 0001
ICPR3
2010 Parallel versus Hierarchical Fusion of Extended Fingerprint Features
abstract
Extended fingerprint features such as pores, dots and incipient ridges have been increasingly attracting attention from researchers and engineers working on automatic fingerprint recognition systems. A variety of methods have been proposed to combine these features with the traditional minutiae features. This paper comparatively analyses the parallel and hierarchical fusion approaches on a high resolution fingerprint image dataset. Based on the results, a novel and more effective hierarchical approach is presented for combining minutiae, pores, dots and incipient ridges.
Qijun Zhao, Feng Liu 0013, Lei Zhang 0006, David Zhang 0001
ICPR4
2010 Mining Hot Clusters of Similar Anomalies for System Management
David Zhang 0001, Zhongzhi Shi
PRICAI1
2010 Resource Allocation in LTE OFDMA Systems Using Genetic Algorithm and Semi-Smart Antennas
abstract
Orthogonal frequency division multiplexing (OFDMA) offers great spectrum efficiency and flexible frequency allocation to users without intra-cell interferences in LTE system. However, the cell edge users will experience high interferences from neighbouring cells. Many frequency reuse schemes have been proposed for improving the Signal to Interference and Noise Ratio (SINR) performance for cell edge users, most of them dividing available frequencies into groups for cell centre and cell edge users. In this research, we combine the traditional frequency scheme with novel semi-smart antennas and learning algorithm - Genetic Algorithm (GA). The semi-smart antennas can produce flexible coverage patterns for base stations (BSs) and the learning algorithm can coordinate the coverage patterns between BSs to minimise interferences for mobile units (MUs). A system level simulation contains 25 BSs and 1000 MUs has been developed. Simulation results show that the proposed scheme improves the total traffic load and reduces antenna propagation power for all cells compares to system with fixed coverage patterns.
Xu Yang 0010, Yapeng Wang 0001, David Zhang 0001, Laurie G. Cuthbert
WCNC3
2010 Fast and convergence-guaranteed algorithm for linear separation
David Zhang 0001, Yugang Li
Sci. China Inf. Sci.2
2010 A unified distance measurement for orientation coding in palmprint verification
Zhenhua Guo 0001, Wangmeng Zuo, Lei Zhang 0006, David Zhang 0001
Neurocomputing4
2010 An efficient method for computing orthogonal discriminant vectors
Yong Xu 0001, David Zhang 0001, Jane You
Neurocomputing3
2010 Tongue shape classification by geometric features
Bo Huang 0003, David Zhang 0001, Naimin Li
Inf. Sci.3
2010 Rotation invariant texture classification using LBP variance (LBPV) with global matching
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001
Pattern Recognit.3
2010 Interactive image segmentation by maximal similarity based region merging
Jifeng Ning, Lei Zhang 0006, David Zhang 0001, Chengke Wu 0001
Pattern Recognit.3
2010 A feature extraction method for use with bimodal biometrics
Yong Xu 0001, David Zhang 0001, Jing-Yu Yang 0001
Pattern Recognit.2
2010 LPP solution schemes for use with face recognition
Yong Xu 0001, Aini Zhong, Jian Yang 0003, David Zhang 0001
Pattern Recognit.4
2010 Two-stage image denoising by principal component analysis with local pixel grouping
Lei Zhang 0006, Weisheng Dong, David Zhang 0001, Guangming Shi
Pattern Recognit.3
2010 Robust palmprint verification using 2D and 3D features
David Zhang 0001, Vivek Kanhangad, Nan Luo, Ajay Kumar 0001
Pattern Recognit.1
2010 Dynamic tongueprint: A novel biometric identifier
David Zhang 0001, Zhi Liu 0004, Jingqi Yan
Pattern Recognit.1
2010 Online finger-knuckle-print verification for personal authentication
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001, Hailong Zhu
Pattern Recognit.3
2010 High resolution partial fingerprint alignment using pore-valley descriptors
Qijun Zhao, David Zhang 0001, Lei Zhang 0006, Nan Luo
Pattern Recognit.2
2010 Adaptive fingerprint pore modeling and extraction
Qijun Zhao, David Zhang 0001, Lei Zhang 0006, Nan Luo
Pattern Recognit.2
2010 Independent components extraction from image matrix
Quanxue Gao, Lei Zhang 0006, David Zhang 0001
Pattern Recognit. Lett.3
2010 Directional binary code with application to PolyU near-infrared face database
Baochang Zhang 0001, Lei Zhang 0006, David Zhang 0001, LinLin Shen
Pattern Recognit. Lett.3
2010 Post-processed LDA for face and palmprint recognition: What is the rationale
Wangmeng Zuo, David Zhang 0001, Kuanquan Wang
Signal Process.3
2010 A new framework for adaptive multimodal biometrics management
abstract
This paper presents a new evolutionary approach for adaptive combination of multiple biometrics to ensure the optimal performance for the desired level of security. The adaptive combination of multiple biometrics is employed to determine the optimal fusion strategy and the corresponding fusion parameters. The score-level fusion rules are adapted to ensure the desired system performance using a hybrid particle swarm optimization model. The rigorous experimental results presented in this paper illustrate that the proposed score-level approach can achieve significantly better and stable performance over the decision-level approach. There has been very little effort in the literature to investigate the performance of an adaptive multimodal fusion algorithm on real biometric data. This paper also presents the performance of the proposed approach from the real biometric samples which further validate the contributions from this paper.
Ajay Kumar 0001, Vivek Kanhangad, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.3
2010 A Completed Modeling of Local Binary Pattern Operator for Texture Classification
abstract
In this correspondence, a completed modeling of the local binary pattern (LBP) operator is proposed and an associated completed LBP (CLBP) scheme is developed for texture classification. A local region is represented by its center pixel and a local difference sign-magnitude transform (LDSMT). The center pixels represent the image gray level and they are converted into a binary code, namely CLBP-Center (CLBP_C), by global thresholding. LDSMT decomposes the image local differences into two complementary components: the signs and the magnitudes, and two operators, namely CLBP-Sign (CLBP_S) and CLBP-Magnitude (CLBP_M), are proposed to code them. The traditional LBP is equivalent to the CLBP_S part of CLBP, and we show that CLBP_S preserves more information of the local structure than CLBP_M, which explains why the simple LBP operator can extract the texture features reasonably well. By combining CLBP_S, CLBP_M, and CLBP_C features into joint or hybrid distributions, significant improvement can be made for rotation invariant texture classification.
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001
IEEE Trans. Image Process.3
2010 An Analysis of IrisCode
abstract
IrisCode is an iris recognition algorithm developed in 1993 and continuously improved by Daugman. It has been extensively applied in commercial iris recognition systems. IrisCode representing an iris based on coarse phase has a number of properties including rapid matching, binomial impostor distribution and a predictable false acceptance rate. Because of its successful applications and these properties, many similar coding methods have been developed for iris and palmprint identification. However, we lack a detailed analysis of IrisCode. The aim of this paper is to provide such an analysis as a way of better understanding IrisCode, extending the coarse phase representation to a precise phase representation, and uncovering the relationship between IrisCode and other coding methods. Our analysis demonstrates that IrisCode is a clustering algorithm with four prototypes; the locus of a Gabor function is a 2-D ellipse with respect to a phase parameter and can be approximated by a circle in many cases; Gabor function can be considered as a phase-steerable filter and the bitwise hamming distance can be regarded as a bitwise phase distance. We also discuss the theoretical foundation of the impostor binomial distribution. We use this analysis to develop a precise phase representation which can enhance accuracy. Finally, we relate IrisCode and other coding methods.
Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel
IEEE Trans. Image Process.2
2010 An Optimized Tongue Image Color Correction Scheme
abstract
The color images produced by digital cameras are usually device-dependent, i.e., the generated color information (usually presented in RGB color space) is dependent on the imaging characteristics of specific cameras. This is a serious problem in computer-aided tongue image analysis because it relies on the accurate rendering of color information. In this paper, we propose an optimized correction scheme that corrects the tongue images captured in different device-dependent color spaces to the target device-independent color space. The correction algorithm in this scheme is generated by comparing several popular correction algorithms, i.e., polynomial-based regression, ridge regression, support vector regression, and neural network mapping algorithms. We test the performance of the proposed scheme by computing the CIE L(*)a(*)b(*) color difference (∆E(ab)(*)) between estimated values and the target reference values. The experimental results on the colorchecker show that the color difference is less than 5 (∆E(ab)(*) < 5), while the experimental results on real tongue images show that the distorted tongue images (captured in various device-dependent color spaces) become more consistent with each other. In fact, the average color difference among them is greatly reduced by more than 95%.
Xingzheng Wang, David Zhang 0001
IEEE Trans. Inf. Technol. Biomed.2
2010 Studies on Hyperspectral Face Recognition in Visible Spectrum With Feature Band Selection
abstract
This correspondence paper studies face recognition by using hyperspectral imagery in the visible light bands. The spectral measurements over the visible spectrum have different discriminatory information for the task of face identification, and it is found that the absorption bands related to hemoglobin are more discriminative than the other bands. Therefore, feature band selection based on the physical absorption characteristics of face skin is performed, and two feature band subsets are selected. Then, three methods are proposed for hyperspectral face recognition, including whole band (2D)2PCA, single band (2D)2PCA with decision level fusion, and band subset fusion-based (2D)2PCA. A simple yet efficient decision level fusion strategy is also proposed for the latter two methods. To testify the proposed techniques, a hyperspectral face database was established which contains 25 subjects and has 33 bands over the visible light spectrum (0.4-0.72 μm). The experimental results demonstrated that hyperspectral face recognition with the selected feature bands outperforms that by using a single band, using the whole bands, or, interestingly, using the conventional RGB color bands.
Wei Di, Lei Zhang 0006, David Zhang 0001, Quan Pan 0001
IEEE Trans. Syst. Man Cybern. Part A3
2009 A Multi-scale Bilateral Structure Tensor Based Corner Detector
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001
ACCV (2)3
2009 Is White Light the Best Illumination for Palmprint Recognition?
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006
CAIP2
2009 Rotation Invariant Texture Classification Using Binary Filter Response Pattern (BFRP)
Zhenhua Guo 0001, Lei Zhang 0006, David Zhang 0001
CAIP3
2009 Second-Level Partition for Estimating FAR Confidence Intervals in Biometric Systems
Rongfeng Li 0002, Darun Tang, Wenxin Li 0005, David Zhang 0001
CAIP4
2009 Differential Feature Analysis for Palmprint Authentication
Xiangqian Wu 0002, Kuanquan Wang, Yong Xu 0001, David Zhang 0001
CAIP4
2009 Finger-Knuckle-Print Verification Based on Band-Limited Phase-Only Correlation
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001
CAIP3
2009 Curvature and singularity driven diffusion for oriented pattern enhancement with singular points
abstract
Oriented patterns, e.g. fingerprints, consist of smoothly varying flow-like patterns, together with important singular points (i.e. cores and deltas) where the orientation changes abruptly. Gabor filters and anisotropic diffusion methods have been widely used to enhance oriented patterns. However, none of them can well cope with regions of varying curvatures or regions surrounding singular points. By incorporating the ridge curvatures and the singularities into the diffusion model, we propose a new diffusion method to better exploit the global characteristics of oriented patterns. Specifically, we first locate the singular points, and regularize the estimated orientation field by using a singularity driven nonlinear diffusion process. We then enhance the oriented patterns by applying an oriented diffusion process which is driven by the curvature and singularity. Experiments on synthetic data and real fingerprint images validated that the proposed method is capable of consistently enhancing oriented patterns while well preserving the ridge structures in singular regions.
Qijun Zhao, Lei Zhang 0006, David Zhang 0001, Wenyi Huang
CVPR3
2009 Multi-view Ear Recognition Based on Moving Least Square Pose Interpolation
Heng Liu 0003, David Zhang 0001, Zhiyuan Zhang 0004
ICIC (2)2
2009 Palmprint verification using consistent orientation coding
abstract
Developing accurate and robust palmprint verification algorithms is one of the key issues in automatic palmprint recognition systems. Recently, orientation based coding algorithms, such as Competitive Code (CompCode) and Orthogonal Line Ordinal Features (OLOF), have been proposed and have been attracting much research attention. Such algorithms could achieve high accuracy with high feature matching speed for real time implementation. By investigating the relationship between these two different coding schemes, we propose in this paper a feature-level fusion scheme for palmprint verification. Only the stable features which are consistent between the two codes are extracted for matching. The experimental results on the public palmprint database show that the proposed fusion code could achieve at least 14% EER (Equal Error Rate) reduction compared with either of the original codes.
Zhenhua Guo 0001, Wangmeng Zuo, Lei Zhang 0006, David Zhang 0001
ICIP4
2009 Principal line based ICP alignment for palmprint verification
abstract
Image alignment is a crucial step in palmprint verification. However, most of the existing palmprint alignment methods use only some key points between fingers or in palm boundary to extract the region of interest (ROI), which is consequently used for feature extraction and matching. Such alignment methods can only give a coarse alignment of the palmprint images. This paper presents a new effective refinement method for palmprint alignment by adapting the iterative closest point (ICP) method to the palmprint principal lines. The proposed method offers a more accurate alignment of palmprints by correcting efficiently the shifting, rotation and scaling variations introduced in data acquisition. The experimental results show that the proposed method can greatly improve the palmprint verification accuracy in real time.
Wei Li 0016, Lei Zhang 0006, David Zhang 0001, Jingqi Yan
ICIP3
2009 Finger-knuckle-print: A new biometric identifier
abstract
This paper presents a new biometric identifier, namely finger-knuckle-print (FKP), for personal identity authentication. First a specific data acquisition device is constructed to capture the FKP images, and then an efficient FKP recognition algorithm is presented to process the acquired data. The local convex direction map of the FKP image is extracted, based on which a coordinate system is defined to align the images and a region of interest (ROI) is cropped for feature extraction. A competitive coding scheme, which uses 2D Gabor filters to extract the image local orientation information, is employed to extract and represent the FKP features. When matching, the angular distance is used to measure the similarity between two competitive code maps. An FKP database was established to examine the performance of the proposed system, and the experimental results demonstrated the efficiency and effectiveness of this new biometric characteristic.
Lin Zhang 0014, Lei Zhang 0006, David Zhang 0001
ICIP3
2009 Spatially Smooth Subspace Face Recognition Using LOG and DOG Penalties
Wangmeng Zuo, Lei Liu 0049, Kuanquan Wang, David Zhang 0001
ISNN (3)4
2009 Three Dimensional Palmprint Recognition
abstract
Palmprint has been widely studied as its high accuracy and low cost. Most of the previous studies are based on two dimensional (2D) image of the palmprint. However, 2D image can be easily forged, which will threaten the security of palmprint authentication system. Furthermore, 2D image can be easily affected by noise, such as scrabbling and dirty in the palm. To overcome these shortcomings, we develop a three dimensional (3D) palmprint identification system. The structured-light imaging technology is adopted to collect the 3D palmprint data, from which the stable mean curvature image (MCI) is extracted. Then the competitive coding (CompCode) technique is used to code the 3D palmprint pattern according the MCI. By using score level fusion of MCI and its CompCode, promising recognition performance is achieved on our established 3D palmprint database.
Lei Zhang 0006, David Zhang 0001
SMC3
2009 Object separation by polarimetric and spectral imagery fusion
Lei Zhang 0006, David Zhang 0001, Quan Pan 0001
Comput. Vis. Image Underst.3
2009 Improving the interest operator for face recognition
Yong Xu 0001, David Zhang 0001, Jing-Yu Yang 0001
Expert Syst. Appl.3
2009 Sequential row-column independent component analysis for face recognition
Quanxue Gao, Lei Zhang 0006, David Zhang 0001
Neurocomputing3
2009 Robust Object Tracking Using Joint Color-Texture Histogram
abstract
A novel object tracking algorithm is presented in this paper by using the joint color-texture histogram to represent a target and then applying it to the mean shift framework. Apart from the conventional color histogram features, the texture features of the object are also extracted by using the local binary pattern (LBP) technique to represent the object. The major uniform LBP patterns are exploited to form a mask for joint color-texture feature selection. Compared with the traditional color histogram based algorithms that use the whole target region for tracking, the proposed algorithm extracts effectively the edge and corner features in the target region, which characterize better and represent more robustly the target. The experimental results validate that the proposed method improves greatly the tracking accuracy and efficiency with fewer mean shift iterations than standard mean shift tracking. It can robustly track the target under complex scenes, such as similar target and background appearance, on which the traditional color based schemes may fail to track.
Jifeng Ning, Lei Zhang 0006, David Zhang 0001, Chengke Wu 0001
Int. J. Pattern Recognit. Artif. Intell.3
2009 A survey of palmprint recognition
Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel
Pattern Recognit.2
2009 Orientation selection using modified FCM for competitive code-based palmprint recognition
Wangmeng Zuo, David Zhang 0001, Kuanquan Wang
Pattern Recognit.3
2009 Palmprint verification using binary orientation co-occurrence vector
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wangmeng Zuo
Pattern Recognit. Lett.2
2009 PCA-Based Spatially Adaptive Denoising of CFA Images for Single-Sensor Digital Cameras
abstract
Single-sensor digital color cameras use a process called color demosiacking to produce full color images from the data captured by a color filter array (CAF). The quality of demosiacked images is degraded due to the sensor noise introduced during the image acquisition process. The conventional solution to combating CFA sensor noise is demosiacking first, followed by a separate denoising processing. This strategy will generate many noise-caused color artifacts in the demosiacking process, which are hard to remove in the denoising process. Few denoising schemes that work directly on the CFA images have been presented because of the difficulties arisen from the red, green and blue interlaced mosaic pattern, yet a well-designed "denoising first and demosiacking later" scheme can have advantages such as less noise-caused color artifacts and cost-effective implementation. This paper presents a principle component analysis (PCA)-based spatially-adaptive denoising algorithm, which works directly on the CFA data using a supporting window to analyze the local image statistics. By exploiting the spatial and spectral correlations existing in the CFA image, the proposed method can effectively suppress noise while preserving color edges and details. Experiments using both simulated and real CFA images indicate that the proposed scheme outperforms many existing approaches, including those sophisticated demosiacking and denoising schemes, in terms of both objective measurement and visual evaluation.
Lei Zhang 0006, Rastislav Lukac, Xiaolin Wu 0001, David Zhang 0001
IEEE Trans. Image Process.4
2009 A Modified Matched Filter With Double-Sided Thresholding for Screening Proliferative Diabetic Retinopathy
abstract
The early diagnosis of proliferative diabetic retinopathy (PDR), a common complication of diabetes that damages the retina, is crucial to the protection of the vision of diabetes sufferers. The onset of PDR is signaled by the appearance of neovascular net. Such neovascular nets might be identified using retinal vessel extraction techniques. The commonly used matched filter methods often produce false positive detections of neovascular nets due to their proneness to detect nonline edges as well as lines. In this paper, we propose a modified matched filter for retinal vessel extraction that applies a local vessel cross-section analysis using double-sided thresholding to reduce false responses to nonline edges. Our proposed modified matched filters demonstrated higher true positive rate and lesser false detection than existing matched-filter-based schemes in vessel extraction.
Lei Zhang 0006, Qin Li 0001, Jane You, David Zhang 0001
IEEE Trans. Inf. Technol. Biomed.4
2009 Palmprint Recognition Using 3-D Information
abstract
Palmprint has proved to be one of the most unique and stable biometric characteristics. Almost all the current palmprint recognition techniques capture the 2-D image of the palm surface and use it for feature extraction and matching. Although 2-D palmprint recognition can achieve high accuracy, the 2-D palmprint images can be counterfeited easily and much 3-D depth information is lost in the imaging process. This paper explores a 3-D palmprint recognition approach by exploiting the 3-D structural information of the palm surface. The structured light imaging is used to acquire the 3-D palmprint data, from which several types of unique features, including mean curvature image, Gaussian curvature image, and surface type, are extracted. A fast feature matching and score-level fusion strategy are proposed for palmprint matching and classification. With the established 3-D palmprint database, a series of verification and identification experiments is conducted to evaluate the proposed method. The results demonstrate that 3-D palmprint technique has high recognition performance. Although its recognition rate is a little lower than 2-D palmprint recognition, 3-D palmprint recognition has higher anticounterfeiting capability and is more robust to illumination variations and serious scrabbling in the palm surface. Meanwhile, by fusing the 2-D and 3-D palmprint information, much higher recognition rate can be achieved.
David Zhang 0001, Guangming Lu 0002, Wei Li 0016, Lei Zhang 0006, Nan Luo
IEEE Trans. Syst. Man Cybern. Part C1
2008 Directional independent component analysis with tensor representation
abstract
Conventional independent component analysis (ICA) learns the statistical independencies of 2D variables from the training images that are unfolded to vectors. The unfolded vectors, however, make the ICA suffer from the small sample size (SSS) problem that leads to the dimensionality dilemma. This paper presents a novel directional multilinear ICA method to solve those problems by encoding the input image or high dimensional data array as a general tensor. In addition, the mode-k matrix of the tensor is re-sampled and re-arranged to form a mode-k directional image to better exploit the directional information in training. An algorithm called mode-k directional ICA is then presented for feature extraction. Compared with the conventional ICA and other subspace analysis algorithms, the proposed method can greatly alleviate the SSS problem, reduce the computational cost in the learning stage by representing the data in lower dimension, and simultaneously exploit the directional information in the high dimensional dataset. Experimental results on well-known face and palmprint databases show that the proposed method has higher recognition accuracy than many existing ICA, PCA and even supervised FLD schemes while using a low dimension of features.
Lei Zhang 0006, Quanxue Gao, David Zhang 0001
CVPR3
2008 Incorporating user quality for performance improvement in hand identification
abstract
Personal authentication using hand images, especially with those acquired from the peg-free imaging, has very high user-acceptance and has invited lot of attention in the literature. This paper investigates a new approach to achieve the performance improvement by incorporating user quality into the matching stage. The proposed method of extracting user quality is based on the confidence of generating reliable matching scores from the user templates. The palmprint and hand-shape images are simultaneously extracted from the hand images and are used to ascertain the performance improvement for the individual trait. The experimental results presented in this paper show significant improvement in the performance while incorporating the proposed method of user quality in the matching stages. This user quality based fusion of two biometric modalities is also investigated and achieves significant improvement in the authentication performance.
Ajay Kumar 0001, David Zhang 0001
ICARCV2
2008 Dark line detection with line width extraction
abstract
Automated line detection is a classical image processing topic with many applications such as road detection in remote images and vessel detection in medical images. Many traditional line detectors, such as Gabor filter, the second order derivative of Gaussian and Radon transform will response not only to lines but also to edges, e.g. they will give high responses to the edges of bright lines or blobs when only dark lines are required. To reduce false detections when extracting only dark (or bright) lines, in this paper we propose a line detector by using the first derivative of Gaussian. It can detect dark lines without much false detection on blobs or bright lines. Meanwhile, the proposed method can estimate line width simultaneously. Experiments on various images are performed to test the proposed algorithm.
Qin Li 0001, Lei Zhang 0006, Jane You, David Zhang 0001, Prabir Bhattacharya
ICIP4
2008 Multimodal biometrics management using adaptive score-level combination
abstract
This paper presents a new evolutionary approach for adaptive combination of multiple biometrics to dynamically ensure the performance for the desired level of security. The adaptive combination of multiple biometrics is achieved at the matching score level. The score level fusion rules are adapted to ensure the required/desired system performance using particle swarm optimization. The experimental results presented in this paper illustrates two main advantages of the proposed score-level approach over the decision level approach; better performance and stable performance that require smaller number of iterations. There has not been any effort in the literature to investigate the performance of adaptive multimodal fusion algorithm on real biometric data. This paper also presents the performance of the proposed algorithm on real biometric data which further validates contributions from this paper.
Ajay Kumar 0001, Vivek Kanhangad, David Zhang 0001
ICPR3
2008 Tongue line extraction
abstract
Tongue line refers to the surface of the tongue covered with fissures or lines in deep or shallow shape and is one type of important features in clinical practice of Traditional Chinese Tongue Diagnosis (TCTD). However, it is hard to extract tongue lines completely due to the large variation of the widths of tongue lines and the strong noise caused by the rough surface of tongue and uneven illumination. In this paper, an improved wide line detector (WLD) is presented for tongue line extraction. Based on the characteristics of tongue lines, the original WLD is improved to avoid the undesired separation of a wide line and the influence of uneven lighting conditions. The proposed method has been tested on a total of 286 tongue line images and our experimental results demonstrate that the improved WLD significantly outperforms the original WLD for tongue line extraction by improving the TPR 16.5%, FPR 44.6% and PM 33.4%, respectively.
Laura Li Liu, David Zhang 0001, Ajay Kumar 0001, Xiangqian Wu 0002
ICPR2
2008 A cryptosystem based on palmprint feature
abstract
Biometric cryptography is a technique using biometric features to encrypt data, which can improve the security of the encrypted data and overcome the shortcomings of the traditional cryptography. This paper proposes a novel biometric cryptosystem based on palmprint features. In this system, the palmprint features, called DoG code, are extracted using Gaussian derivative filters. Then the Reed-Solomon error correcting technique and the logical XOR operation are employed to encrypt and decrypt the data. Experimental results show that this system can obtain a high security with a low false rejection rate.
Xiangqian Wu 0002, Kuanquan Wang, David Zhang 0001
ICPR3
2008 A performance evaluation of filter design and coding schemes for palmprint recognition
abstract
Palmprint recognition, as one of the most promising biometrics, has received considerable recent biometric research interest. Among various palmprint recognition techniques, coding based methods have been very successful since of its simplicity, high precision, small size of feature and rapidness for both feature extraction and matching. Several filters, such as Gabor and Gaussian, and coding schemes, such as competitive and ordinal measure, have been proposed for palmprint verification and identification. In this paper, we evaluate three filters, Gabor, Gaussian, and the second derivative of Gaussian filter, and two coding schemes, competitive code and ordinal measure on PolyU palmprint database. Results of verification experiment show that Gabor filter and competitive coding scheme is superior to other methods.
Wangmeng Zuo, Kuanquan Wang, David Zhang 0001
ICPR4
2008 FCM-based orientation selection for competitive coding-based palmprint recognition
abstract
Coding based methods are among the most promising palmprint recognition methods. As one representative coding method, the competitive code first convolves the palmprint image with a bank of Gabor filters with different orientations and then encodes the dominant orientation into its bitwise representation. Despite its effectiveness, few investigations have been given to study the influence of the number of filters and the orientation of each filter. In this paper, based on the statistical orientation distribution and the orientation separation principle, we propose a modified fuzzy C-means cluster algorithm to determine the orientations of filters. Experimental results indicate that, the proposed method achieves higher verification accuracy while compared with that of the original competitive code and several state-of-the-art methods. Considering both the computational complexity and the verification accuracy, six filters would be the optimal choice for proposed method.
Wangmeng Zuo, Kuanquan Wang, David Zhang 0001
ICPR4
2008 Adaptive pore model for fingerprint pore extraction
abstract
Sweat pores have been recently employed for automated fingerprint recognition, in which the pores are usually extracted by using a computationally expensive skeletonization method or a unitary scale isotropic pore model. In this paper, however, we show that real pores are not always isotropic. To accurately and robustly extract pores, we propose an adaptive anisotropic pore model, whose parameters are adjusted adaptively according to the fingerprint ridge direction and period. The fingerprint image is partitioned into blocks and a local pore model is determined for each block. With the local pore model, a matched filter is used to extract the pores within each block. Experiments on a high resolution (1200dpi) fingerprint dataset are performed and the results demonstrate that the proposed pore model and pore extraction method can locate pores more accurately and robustly in comparison with other state-of-the-art pore extractors.
Qijun Zhao, Lei Zhang 0006, David Zhang 0001, Nan Luo, Jing Bao
ICPR3
2008 Multiscale competitive code for efficient palmprint recognition
abstract
Coding-based method, which encodes the responses of a bank of filters into bitwise features, has been very successful in palmprint representation and matching. Palmprints, however, are typically multiscale features, where the palm lines can be represented at a higher scale while the wrinkles at a lower scale. In this work, we present a mutliscale competitive code method for efficient palmprint representation and matching. In filterbank design, we adopt the log-Gabor wavelets since of its less overlapping in the frequency domain. In palmprint representation, competitive code is used to encoding the dominant orientation of the filter responses in each scale. In palmprint matching, a fusion rule is proposed to combine the distances obtained using different scales. Experimental results indicate that the proposed method achieves better recognition accuracy and faster matching speed while compared with several state-of-the-art methods.
Wangmeng Zuo, Kuanquan Wang, David Zhang 0001
ICPR4
2008 Palmprint identification based on directional representation
abstract
In this paper, we propose a novel approach for palmprint identification, which contains two interesting components. Firstly, we propose the directional representation for appearance based approaches. The new representation is robust to drastic illumination changes and preserves important discriminative information for classification. We then generate virtual samples to enlarge the training set to compensate for matching errors caused by large rotations and translations. Based on these two strategies, the recognition performance of representative appearance based approaches can be improved significantly. Secondly, in order to improve the robustness of palmprint identification, we propose a fusion method combining proposed method and orientation based approaches e.g. Competitive Code, which can obtain very low Equal Error Rates.
Wei Jia 0001, De-Shuang Huang, Dacheng Tao, David Zhang 0001
SMC4
2008 Median Fisher Discriminator: a robust feature extraction method with applications to biometrics
Jian Yang 0003, Jing-Yu Yang 0001, David Zhang 0001
Frontiers Comput. Sci. China3
2008 A novel face recognition approach based on kernel discriminative common vectors (KDCV) feature extraction and RBF neural network
Xiaoyuan Jing, Yong-Fang Yao, Jing-Yu Yang 0001, David Zhang 0001
Neurocomputing4
2008 A highly scalable incremental facial feature extraction method
Fengxi Song, David Zhang 0001, Jing-Yu Yang 0001
Neurocomputing3
2008 An approach for directly extracting features from matrix data and its application in face recognition
Yong Xu 0001, David Zhang 0001, Jian Yang 0003, Jing-Yu Yang 0001
Neurocomputing2
2008 On kernel difference-weighted k-nearest neighbor classification
Wangmeng Zuo, David Zhang 0001, Kuanquan Wang
Pattern Anal. Appl.2
2008 Palmprint verification based on principal lines
De-Shuang Huang, Wei Jia 0001, David Zhang 0001
Pattern Recognit.3
2008 Palmprint verification based on robust line orientation code
Wei Jia 0001, De-Shuang Huang, David Zhang 0001
Pattern Recognit.3
2008 Three measures for secure palmprint identification
Adams Wai-Kin Kong, David Zhang 0001, Mohamed S. Kamel
Pattern Recognit.2
2008 Interest filter vs. interest operator: Face recognition using Fisher linear discriminant based on interest filter representation
Tuo Zhao, Zhizheng Liang, David Zhang 0001, Quan Zou 0001
Pattern Recognit. Lett.3
2008 Improved upper bounds on the L(2, 1) -labeling of the skew and converse skew product graphs
Zhendong Shao, David Zhang 0001
Theor. Comput. Sci.2
2008 Comments on "An Adaptive Multimodal Biometric Management Algorithm"
abstract
We note that there are some discrepancies in the results reported in the previous titled paper. Our experiments indicate that the authors have considered only a subset of all possible fusion rules, contradicting the statement that all possible rules have been considered. Moreover, the authors state that only monotonic rules can be optimal, and therefore, all other rules can be ignored. However, our experimental results examining all possible rules demonstrate that a nonmonotonic rule can also be an optimum fusion rule.
Vivek Kanhangad, Ajay Kumar 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Part C3
2007 Fusion of Palmprint and Iris for Personal Authentication
Xiangqian Wu 0002, David Zhang 0001, Kuanquan Wang, Ning Qi
ADMA2
2007 Biometric Recognition using Entropy-Based Discretization
abstract
The biometrics based recognition systems proposed in the literature have not yet exploited user-specific dependencies in the feature level representation. This paper suggests and investigates the performance improvement of the existing biometric systems using the discretization of extracted features. The performance improvement due to the unsupervised and supervised discretization schemes is compared on verity of classifiers; KNN, naive Bayes, SVM and FFN. The experimental results on the hand-geometry database of 100 users achieve significant improvement in the recognition accuracy and confirm the usefulness of discretization in biometrics systems.
Ajay Kumar 0001, David Zhang 0001
ICASSP (2)2
2007 A Spherical Rectification for Dual-PTZ-Camera System
abstract
In this paper we proposed a novel stereo rectification method for dual-PTZ-camera system, which is called spherical rectification. This method can be divided into two parts, offline pre-settings and online rectification. The offline pre-settings include camera calibration and building the spherical coordinate system. The online rectification only requires the pan, tilt and zoom values and does not need any image information such as feature points. So, compared with traditional rectification approaches, our method is more convenient and time-saving.
Dingrui Wan, Jie Zhou 0001, David Zhang 0001
ICASSP (1)3
2007 Kernel Difference-Weighted k-Nearest Neighbors Classification
Wangmeng Zuo, Kuanquan Wang, David Zhang 0001
ICIC (2)4
2007 Palmprint Verification using Complex Wavelet Transform
abstract
Palmprint is a unique and reliable biometric characteristic with high usability. With the increasing demand of automatic palmprint authentication systems, the development of accurate and robust palmprint verification algorithms has been attracting a lot of interests. The relative translation, rotation and distortion between two palmprint images will introduce much error in palmprint matching. However, an accurate registration of palmprint images is too time-consuming. In this paper, we propose a modified complex wavelet structural similarity index (CW-SSIM) to compute the matching score and hence identify the input palmprint. Since CW-SSIM is robust to translation, small rotation and distortion, a fast rough alignment of palmprint images is sufficient. CW-SSIM is also insensitive to luminance and contrast changes. Our experimental results show that the proposed scheme outperforms the state-of-the-art methods by achieving a higher genuine acceptance rate and a lower false acceptance rate simultaneously.
Lei Zhang 0006, Zhenhua Guo 0001, Zhou Wang 0001, David Zhang 0001
ICIP (2)4
2007 Iteratively Reweighted Fitting for Reduced Multivariate Polynomial Model
Wangmeng Zuo, Kuanquan Wang, David Zhang 0001
ISNN (2)3
2007 Automated Personal Authentication Using Both Palmprints
Xiangqian Wu 0002, Kuanquan Wang, David Zhang 0001
ICEC3