VLDB 2026 Research / reviewers in the wild / expert
Zhenan Sun
dblp:13/5916
· DBLP profile ↗
243ranked-venue papers
6as first author
103since 2021 · last 2026
0000-0003-4029-9935ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 143 · 3 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 141 · 3 first-author · 54 since 2021Security and privacy · 56 · 34 since 2021Human-computer interaction and ubiquitous computing · 27 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionabstractGenerating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motion variations and emotional intensity, especially in long-sequence modeling. Moreover, the lack of long-term and large-scale paired speaker-listener corpora incorporating head dynamics and fine-grained multi-modality annotations limits the application of dialogue modeling. Therefore, we first newly collect a large-scale multi-turn dataset of 3D dyadic conversation containing more than 1.4M valid frames for multi-modal responsive interaction, dubbed ListenerX. Additionally, we propose VividListener, a novel framework enabling fine-grained, expressive, and controllable listener dynamics modeling. This framework leverages multi-modal conditions as guiding principles for fostering coherent interactions between speakers and listeners. Specifically, we design the Responsive Interaction Module (RIM) to adaptively represent the multi-modal interactive embeddings. RIM ensures the listener dynamics achieve fine-grained semantic coordination with textual descriptions and adjustments, while preserving expressive reaction with speaker behavior. Meanwhile, we propose the Emotional Intensity Tags (EIT) for emotion intensity editing with multi-modal information integration, applying to both text descriptions and listener motion amplitude. Extensive experiments conducted on our newly collected ListenerX dataset demonstrate that VividListener achieves state-of-the-art performance, realizing expressive and controllable listener dynamics. Xingqun Qi, Bingkun Yang, Weile Chen, Zezhao Tian, Muyi Sun, Man Zhang 0005, Zhenan Sun |
AAAI | 9 |
| 2026 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks (Abstract Reprint)abstractDeep learning-based face recognition models are vulnerable to adversarial attacks. In contrast to general noises, the presence of imperceptible adversarial noises can lead to catastrophic errors in deep face recognition models. The primary difference between adversarial noise and general noise lies in its specificity. Adversarial attack methods give rise to noises tailored to the characteristics of the individual image and recognition model at hand. Diverse samples and recognition models can engender specific adversarial noise patterns, which pose significant challenges for adversarial defense. Addressing this challenge in the realm of face recognition presents a more formidable endeavor due to the inherent nature of face recognition as an open set task. In order to tackle this challenge, it is imperative to employ customized processing for each individual input sample. Drawing inspiration from the biological immune system, which can identify and respond to various threats, this paper aims to create an artificial immune system to provide adversarial defense for face recognition. The proposed defense model incorporates the principles of antibody cloning, mutation, selection, and memory mechanisms to generate a distinct antibody for each input sample, wherein the term antibody refers to a specialized noise removal manner. Furthermore, we introduce a self-supervised adversarial training mechanism that serves as a simulated rehearsal of immune system invasions. Extensive experimental results demonstrate the efficacy of the proposed method, surpassing state-of-the-art adversarial defense methods. Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
AAAI | 5 |
| 2026 | UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and PerceptionabstractThe remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual and textual modalities, especially in scenarios involving complex semantic instructions. However, existing approaches often rely heavily on vision-language models (VLMs) or modular designs for semantic guidance, leading to fragmented architectures and computational inefficiency. To address these challenges, we propose UniAlignment, a unified multimodal generation framework within a single diffusion transformer. UniAlignment introduces a dual-stream diffusion training strategy that incorporates both intrinsic-modal semantic alignment and cross-modal semantic alignment, thereby enhancing the model's cross-modal consistency and instruction-following robustness. Additionally, we present SemGen-Bench, a new benchmark specifically designed to evaluate multimodal semantic consistency under complex textual instructions. Extensive experiments across multiple tasks and benchmarks demonstrate that UniAlignment outperforms existing baselines, underscoring the significant potential of diffusion models in unified multimodal generation. Xinyang Song, Weining Wang 0001, Shaozhen Liu, Jingdong Chen, Qi Li 0005, Zhenan Sun |
AAAI | 8 |
| 2026 | DapQ-DiT: Distribution-Aware Post-Training Quantization for Efficient Generative Tasks in Diffusion TransformersabstractDiffusion Transformers (DiTs) have demonstrated remarkable performance in image and video generation tasks. However, their high computational and memory overheads severely restrict their practical deployment on resource-constrained devices. Post-training quantization (PTQ), an efficient and practical model compression technique, serves as a solution to alleviate this issue. Nevertheless, existing PTQ methods tailored for DiTs suffer from significant performance degradation when conducting low-bit weight–activation quantization. In this work, we identify two key factors responsible for such performance degradation. First, the weights of DiTs exhibit Gaussian-like distributions, which makes uniform quantization poorly matched to the actual weight density and introduces large quantization errors. Second, activation outliers with extremely large magnitudes, especially in specific linear layers, significantly widen the value range and severely reduce the suppression effectiveness of fixed Hadamard rotation. To address the above degradation issues, we propose DapQ-DiT, a novel distribution-aware post-training quantization framework tailored for DiTs. First, we introduce an arctan quantizer that explicitly adapts to Gaussian-like weight distributions, concentrating more quantization intervals in the high-density central region while preserving representation accuracy for critical weights. Second, we enhance the fixed Hadamard rotation by leveraging principal components derived from activation covariance, which allows the transformation to better align with real activation distributions and more effectively suppress diverse and extreme activation outliers. Extensive experiments conducted on text-to-image generation with PixArt and text-to-video generation with OpenSORA demonstrate that DapQ-DiT consistently outperforms existing PTQ methods across various prompt sets and diverse bit-width configurations. Lianwei Yang, Haokun Lin, Zhenan Sun, Qingyi Gu |
ICMR | 4 |
| 2026 | Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Zhenan Sun |
Int. J. Comput. Vis. | 6 |
| 2026 | CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database
Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Zhenan Sun, Tieniu Tan, Caifeng Shan |
Int. J. Comput. Vis. | 4 |
| 2026 | Reshape and rotate: Adaptive weight reshaping and fine-grained rotation for ultra-low-bit diffusion transformers quantization
Lianwei Yang, Haokun Lin, Caifeng Shan, Zhenan Sun, Qingyi Gu |
Neurocomputing | 5 |
| 2026 | Learning Knowledge-Based Prompts for Robust 3D Mask Presentation Attack Detectionabstract3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimodal features or remote photoplethysmography (rPPG) signals to distinguish between real faces and 3D masks, they face significant challenges, such as the high costs associated with multimodal sensors and limited generalization ability. Detection-related text descriptions offer concise, universal information and are cost-effective to obtain. However, the potential of vision-language multimodal features for 3D mask presentation attack detection remains unexplored. In this paper, we propose a novel knowledge-based prompt learning framework to explore the strong generalization capability of vision-language models for 3D mask presentation attack detection. Specifically, our approach incorporates entities and triples from knowledge graphs into the prompt learning process, generating fine-grained, task-specific explicit prompts that effectively harness the knowledge embedded in pre-trained vision-language models. Furthermore, considering different input images may emphasize distinct knowledge graph elements, we introduce a visual-specific knowledge filter based on an attention mechanism to refine relevant elements according to the visual context. Additionally, we leverage causal graph theory insights into the prompt learning process to further enhance the generalization ability of our method. During training, a spurious correlation elimination paradigm is employed, which removes category-irrelevant local image patches using guidance from knowledge-based text features, fostering the learning of generalized causal prompts that align with category-relevant local patches. Experimental results demonstrate that the proposed method achieves state-of-the-art intra- and cross-scenario detection performance on benchmark datasets. Fangling Jiang, Qi Li 0005, Weining Wang 0001, Caifeng Shan, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Attention-assisted multilevel fusion framework for generalized iris presentation attack detection
Caiyong Wang, Fukang Guo, Zhaofeng He 0001, Zhenan Sun |
Pattern Recognit. | 6 |
| 2026 | Procedure-Aware Hierarchical Alignment for Open Surgery Video-Language PretrainingabstractRecent advances in surgical robotics and computer vision have greatly improved intelligent systems' autonomy and perception in the operating room (OR), especially in endoscopic and minimally invasive surgeries. However, for open surgery, which is still the predominant form of surgical intervention worldwide, there has been relatively limited exploration due to its inherent complexity and the lack of large-scale, diverse datasets. To close this gap, we present OpenSurgery, by far the largest video-text pretraining and evaluation dataset for open surgery understanding. OpenSurgery consists of two subsets: OpenSurgery-Pretrain and OpenSurgery-EVAL. OpenSurgery-Pretrain consists of 843 publicly available open surgery videos for pretraining, spanning 102 hours and encompassing over 20 distinct surgical types. OpenSurgery-EVAL is a benchmark dataset for evaluating model performance in open surgery understanding, comprising 280 training and 120 test videos, totaling 49 hours. Each video in OpenSurgery is meticulously annotated by expert surgeons at three hierarchical levels of video, operation, and frame to ensure both high quality and strong clinical applicability. Next, we propose the Hierarchical Surgical Knowledge Pretraining (HierSKP) framework to facilitate large-scale multimodal representation learning for open surgery understanding. HierSKP leverages a granularity-aware contrastive learning strategy and enhances procedural comprehension by constructing hard negative samples and incorporating a Dynamic Time Warping (DTW)-based loss to capture fine-grained temporal alignment of visual semantics. Extensive experiments show that HierSKP achieves state-of-the-art performance on OpenSurgegy-EVAL across multiple tasks, including operation recognition, temporal action localization, and zero-shot cross-modal retrieval. This demonstrates its strong generalizability for further advances in open surgery understanding. Boqiang Xu, Jinlin Wu, Jian Liang 0001, Zhenan Sun, Hongbin Liu 0001, Jiebo Luo 0001, Zhen Lei 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Multi-View Images Suffice 3D Reasoning Through Chain-of-Thought Selection and Question-Guided Fusionabstract3D reasoning is crucial in areas like robotics and autonomous driving. Due to the high cost of 3D data acquisition, some recent methods attempt to enable LLMs to perform 3D reasoning through multi-view images, thereby transferring the powerful 2D reasoning capabilities of LLMs to 3D environments. However, these methods face challenges: either they use redundant views that contain many perspectives irrelevant to the question, or they rely on globally aggregated multi-view representations, losing the fine-grained vision-language correlations. To tackle these challenges, we propose 3DMulti-LLM, which mainly consists of three components: a COT selector, a question-guided fusion block, and pre-trained LLMs. Specifically, first, the COT selector leverages the powerful chain-of-thought reasoning capabilities of LLMs to identify question-related multi-view images. In this way, 3DMulti-LLM can eliminate a substantial amount of interference from unnecessary viewpoints. Then, we propose a question-guided fusion block for integrating multi-view features via question-guided interaction among various viewpoints. Finally, the pre-trained LLMs are utilized to reason in 3D scenes directly through multi-view features. Notably, our approach understands the 3D scene solely through multi-view images, without requiring the input of point cloud information or additional 3D feature extraction. Through our experiments, 3DMulti-LLM achieves impressive performance and surpasses existing 3D-input-free methods by + 12.2% and + 7.1% on ScanQA and 3DMV-VQA datasets, respectively. Boqiang Xu, Jinlin Wu, Wei Zhang 0255, Chenyang Su, Jian Liang 0001, Zhenan Sun, Zhen Lei 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Psyche-Wave: Fusing Vector-Quantized Morphology and LLM-Inferred Semantics from Millimeter-Wave SCG for Psychological State DecodingabstractThis paper introduces Psyche-Wave, a novel paradigm for non-contact psychological state assessment, addressing the challenge that existing methods struggle to reconcile signal representation robustness with deep physiological semantic understanding. The proposed framework is built upon high-fidelity Seismocardiogram (SCG) and respiratory signals, captured by a proprietary high-sampling-rate millimeter-wave (mmWave) radar system. Psyche-Wave features a parallel dual-branch architecture for complementary feature extraction. The first, a Data-Driven Morphological Branch, employs Vector Quantization (VQ) to encode the Mel spectrogram of the SCG signal into a codebook-based representation, yielding a noise-resilient morphological embedding. The second, a Knowledge-Driven Semantic Branch, leverages a Large Language Model (LLM) to infer deep contextual relationships from medically significant physiological parameters—including heart rate variability, cardiac time intervals, and cardiopulmonary coupling—outputting a rich semantic embedding. These complementary embeddings are then integrated through a dedicated fusion module and passed to a downstream classifier for precise emotion and personality trait evaluation. Comprehensive evaluations on a newly collected high-fidelity dataset, referred to as mmHeart-Pro, and the public ReMAP dataset demonstrate state-of-the-art performance. This work pioneers a new path that fuses data-driven morphological analysis with knowledge-driven semantic reasoning, significantly advancing the accuracy and interpretability of non-contact psychological sensing. Yiwei Ru, Zhenbo Xu, Yanlin Xu, Huijia Wu, Zhaofeng He 0001, Zhenan Sun |
BIBM | 7 |
| 2025 | Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action RecognitionabstractIn skeleton-based action recognition, a key challenge is distinguishing between actions with similar trajectories of joints due to the lack of image-level details in skeletal representations. Recognizing that the differentiation of similar actions relies on subtle motion details in specific body parts, we direct our approach to focus on the fine-grained motion of local skeleton components. To this end, we introduce ProtoGCN, a Graph Convolutional Network (GCN)-based model that breaks down the dynamics of entire skeleton sequences into a combination of learnable prototypes representing core motion patterns of action units. By contrasting the reconstruction of prototypes, ProtoGCN can effectively identify and enhance the discriminative representation of similar actions. Without bells and whistles, ProtoGCN achieves state-of-the-art performance on multiple benchmark datasets, including NTU RGB+D, NTU RGB+D 120, Kinetics-Skeleton, and FineGYM, which demonstrates the effectiveness of the proposed method. The code is available at https://github.com/firework8/ProtoGCN. Hongda Liu 0002, Yunfan Liu 0001, Yunlong Wang 0003, Zhenan Sun |
CVPR | 6 |
| 2025 | Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose GuidanceabstractText-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently in a scenario. To tackle this, we propose a novel multi-character video generation framework in a tuning-free manner, which is based on the separated text and pose guidance. Specifically, we first extract character masks from the pose sequence to identify the spatial position for each character, and then single prompts for each character are obtained with LLMs for precise text guidance. Moreover, the spatial-aligned cross attention and multi-branch control module are proposed to generate fine-grained controllable multi-character video. The visualized results of generating video demonstrate the precise controllability of our method for multicharacter generation. We also verify the generality of our method by applying it to various personalized T2I models. Moreover, the quantitative results show that our approach achieves superior performance compared with previous works. Beiyuan Zhang, Chunlei Fu, Xinyang Song, Zhenan Sun |
ICASSP | 5 |
| 2025 | BiommWave: A Non-Visual Approach for Biometric Recognition Using Millimeter-Wave RadarabstractThis paper explores the application of millimeter-wave (mmWave) radar in biometric recognition. As a non-visual human sensing technology, mmWave radar captures reflection properties and micro-movements, providing a complementary modality to visual appearance. We implemented a complete system pipeline to utilize mmWave sensing for individual recognition. Based on physiological mechanisms, we design preprocessing methods to extract intuitive biometric feature maps, concerning body reflections, cardiopulmonary activity, and micro-motion frequencies. To address the data uncertainty, a dynamic pole-based learning strategy is proposed to construct compact and discriminated feature distributions. In real-world evaluations, the system achieves 95.12% accuracy and an Equal Error Rate (EER) of 1.96%. This work leverages the advantages of mmWave radar for flexible, unconstrained, and private biometric systems. From a non-visual sensing perspective, it explores novel modalities as unique biometric cues, demonstrating significant value of research and applications. Mupei Li, Yunlong Wang 0003, Yiwei Ru, Kunbo Zhang, Zhenan Sun |
IJCB | 5 |
| 2025 | LAMAR: LLM-Guided Adaptive Perceptual Modeling for Micro-Action RecognitionabstractWe present LAMAR, a novel framework for Micro-Action Recognition that addresses the challenges of identifying subtle, ephemeral human movements lasting less than one-third of a second. Our approach leverages large language models to estimate semantic complexity of micro-actions and dynamically configure a hierarchical Vision Transformer architecture accordingly. LAMAR introduces: (1) a principled complexity estimation module that quantifies recognition difficulty by analyzing subtlety, noise susceptibility, and intra-class ambiguity; and (2) an adaptive perception pipeline that dynamically adjusts spatiotemporal resolution and attention mechanisms based on estimated complexity. Experiments on the MA-52 benchmark demonstrate that LAMAR outperforms state-of-the-art methods by 7.20% in accuracy while maintaining computational efficiency, establishing a new paradigm for context-aware visual analysis that intelligently allocates resources based on task difficulty. Yiwei Ru, Leyuan Wang, Ma He, Zhaofeng He 0001, Zhenan Sun |
IJCB | 5 |
| 2025 | Hierarchical Emotion-Guided Masked Transformer for Long-Sequence Co-Speech Gestures with Partial SupervisionabstractThis work proposes a method for co-speech gesture generation that produces natural, emotionally expressive body movements synchronized with speech. Existing methods typically assume complete motion supervision and static emotional states, limiting training data leverage and impairing fine-grained emotion transitions, ultimately leading to incoherent long-sequence gesture generation. In contrast, we propose a Hierarchical Emotion-Guided Masked Transformer that tackles these limitations. We introduce a Mask-Based Motion Modeling Strategy in a discrete latent space, enabling learning from partially annotated data and ensuring physically plausible motions under incomplete supervision, alongside a hierarchical Emotion Guidance Adaptor that injects time-varying emotional cues at multiple transformer levels, capturing both global emotion and subtle local nuances. An Alternating Optimization Mechanism is employed to decouple semantic alignment from emotion modulation during training, stabilizing learning and improving expression fidelity. At inference, a Mask-Guided Inference strategy seamlessly extends gestures over long sequences, mitigating boundary discontinuities and drift. Evaluations on a partially annotated dataset featuring natural emotional transitions show our method surpasses existing approaches in long-sequence co-speech gesture generation, yielding improved gestures-peech synchrony, enhanced motion diversity, and more faithful emotion-conveying ability. Kunbo Zhang, Zhenan Sun |
IJCB | 5 |
| 2025 | Contextualizing Borderline ECG Analysis via Multi-Modal Feature Extraction and Large Language Model InferenceabstractBorderline electrocardiograms (ECGs) pose a significant diagnostic challenge, as their waveforms often exhibit subtle deviations that overlap with both normal and pathological patterns. Conventional deep learning models, while adept at detecting common arrhythmias, struggle with these ambiguous cases due to sparse annotations and complex signal morphologies. To address this gap, we propose a multi-modal framework that combines structured feature extraction and large language model (LLM) inference for robust ECG classification. First, we employ specialized libraries to derive morphological markers, interval measurements, and heart rate variability (HRV) parameters from raw ECG data. These features, along with demographic metadata, are then seamlessly integrated into prompts for LLM-based few-shot learning. By embedding quantitative signals into textual templates, the model acquires a contextual understanding that transcends static, threshold-based judgments. Our method not only excels at detecting arrhythmias but also demonstrates enhanced performance in classifying borderline ECGs, guided by expert-validated annotations. Experimental results show that the proposed pipeline effectively mitigates class imbalance and improves interpretability, offering a scalable solution that bridges the divide between conventional signal processing and advanced medical-language reasoning. This integration of numeric features with textual prompts paves the way for more accurate, transparent, and adaptable cardiac diagnostics. Yanlin Xu, Yiwei Ru, Dongsen Zhang, Yongji Liu, Zhenan Sun |
ICME | 5 |
| 2025 | HDTPose: A Hierarchical Decoding Transformer for End-to-End Single-Stage Multi-Person Pose EstimationabstractThe estimation performance of human keypoints varies significantly across different keypoint types. Compared to prominent joints such as the head and shoulders, smaller and more flexible limb joints present greater challenges in identification and localization. Current single-stage methods typically treat all body joints uniformly, overlooking the inherent differences and structural relationships among keypoints. To address this limitation, we propose HDTPose, a fully end-to-end multi-person pose estimation framework based on a hierarchical decoding transformer. HDTPose formulates multi-person pose estimation as a hierarchical set prediction problem and employs a hierarchical decoder to progressively decode joints in an end-to-end manner. The decoding process consists of two stages: an instance-aware stage and a joint-aware stage, which explicitly model instance-wise and joint-wise relationships, respectively. Additionally, we introduce a structure-guided joint attention mechanism that leverages kinematic relationships to refine pose predictions. Extensive experiments on the COCO and MPII benchmarks demonstrate that HDTPose outperforms existing state-of-the-art single-stage methods, underscoring the effectiveness and superiority of our approach. Wei Zhang 0255, Qi Li 0005, Zhenan Sun |
IJCNN | 4 |
| 2025 | ReMeREC: Relation-aware and Multi-entity Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) aims to localize specified entities or regions from the source image according to the given natural language descriptions. While existing methods enable single-entity localization, they overlook modeling the complex inter-entity relationship in more practical multi-entity scenes, which limits their ability to produce accurate and reliable results. Moreover, the lack of high-quality multi-entity datasets incorporating fine-grained and paired image-text-relation annotations also limits addressing this challenge. To achieve this task, we first manually construct a relation-aware multi-entity REC dataset with fine-grained relation and text annotations, namely ReMeX. Additionally, we propose ReMeREC, a novel framework that effectively integrates textual and visual cues to localize multiple entities while capturing their inter-relationship. Specifically, to mitigate the semantic ambiguity arising from the absence of explicit entity boundaries in the source natural language description, we introduce a novel Text-adaptive Multi-entity Perceptron (TMP). TMP dynamically infers both the quantity and span of entities from corresponding fine-grained text cues, thus deriving representations that preserve the unique characteristics of each entity. Meanwhile, we design the Entity Inter-relationship Reasoner (EIR) to enhance semantic distinctiveness relationship modeling, leading to a more profound perception of the global scene. Furthermore, to better capture the fine-grained linguistic prompts for delineating multiple entity boundaries and inter-relationship, we leverage LLMs to generate a small-scale textual dataset, dubbed EntityText, which serves as an effective auxiliary resource and further improves the textual understanding. Extensive experiments conducted on four benchmark datasets demonstrate the superior performance of our framework. Remarkably, ReMeREC achieves outstanding results in multi-entity grounding and complex relationship prediction, outperforming other counterparts by a large margin. Yizhi Hu, Zezhao Tian, Xingqun Qi, Bingkun Yang, Junhui Yin, Muyi Sun, Man Zhang 0005, Zhenan Sun |
ACM Multimedia | 9 |
| 2025 | FingerVeinSyn-5M: A Million-Scale Dataset and Benchmark for Finger Vein RecognitionabstractA major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we introduce FVeinSyn, a synthetic generator capable of producing diverse finger vein patterns with rich intra-class variations. Using FVeinSyn, we created FingerVeinSyn-5M -- the largest available finger vein dataset -- containing 5 million samples from 50,000 unique fingers, each with 100 variations including shift, rotation, scale, roll, varying exposure levels, skin scattering blur, optical blur, and motion blur. FingerVeinSyn-5M is also the first to offer fully annotated finger vein images, supporting deep learning applications in this field. Models pretrained on FingerVeinSyn-5M and fine-tuned with minimal real data achieve an average 53.91% performance gain across multiple benchmarks. The dataset is publicly available at: https://github.com/EvanWang98/FingerVeinSyn-5M. Yifan Wang 0036, Jie Gui, Baosheng Yu, Qi Li 0005, Zhenan Sun, Juho Kannala, Guoying Zhao 0001 |
ACM Multimedia | 5 |
| 2025 | CLANet: A Denoising-Driven Framework for Robust mmWave Radar Vital Sign Monitoring
Yiwei Ru, Yongji Liu, Mupei Li, Dongsen Zhang, Zhaofeng He 0001, Zhenan Sun |
PRCV (3) | 6 |
| 2025 | Correction: Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2025 | AnyFace++: A Unified Framework for Free-Style Text-to-Face Synthesis and ManipulationabstractHuman faces contain rich semantic information that could hardly be described without a large vocabulary and complex sentence patterns. However, most existing text-to-image synthesis methods could only generate meaningful results based on limited sentence templates with words contained in the training set, which heavily impairs the generalization ability of these models. In this paper, we define a novel 'free-style' text-to-face generation and manipulation problem, and propose an effective solution, named AnyFace++, which is applicable to a much wider range of open-world scenarios. The CLIP model is involved in AnyFace++ for learning an aligned language-vision feature space, which also expands the range of acceptable vocabulary as it is trained on a large-scale dataset. To further improve the granularity of semantic alignment between text and images, a memory module is incorporated to convert the description with arbitrary length, format, and modality into regularized latent embeddings representing discriminative attributes of the target face. Moreover, the diversity and semantic consistency of generation results are improved by a novel semi-supervised training scheme and a series of newly proposed objective functions. Compared to state-of-the-art methods, AnyFace++ is capable of synthesizing and manipulating face images based on more flexible descriptions and producing realistic images with higher diversity. Jianxin Sun 0003, Qiyao Deng, Qi Li 0005, Muyi Sun, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | IrisFormer: A Dedicated Transformer Framework for Iris RecognitionabstractWhile Vision Transformer (ViT)-based methods have significantly improved the performance of various vision tasks in natural scenes, progress in iris recognition remains limited. In addition, the human iris contains unique characters that are distinct from natural scenes. To remedy this, this paper investigates a dedicated Transformer framework, termed IrisFormer, for iris recognition and attempts to improve the accuracy by combining the contextual modeling ability of ViT and iris-specific optimization to learn robust, fine-grained, and discriminative features. Specifically, to achieve rotation invariance in iris recognition, we employ relative position encoding instead of regular absolute position encoding for each iris image token, and a horizontal pixel-shifting strategy is utilized during training for data augmentation. Then, to enhance the model's robustness against local distortions such as occlusions and reflections, we randomly mask some tokens during training to force the model to learn representative identity features from only part of the image. Finally, considering that fine-grained features are more discriminative in iris recognition, we retain the entire token sequence for patch-wise feature matching instead of using the standard single classification token. Experiments on three popular datasets demonstrate that the proposed framework achieves competitive performance under both intra- and inter-dataset testing protocols. Xianyun Sun, Caiyong Wang, Yunlong Wang 0003, Jianze Wei, Zhenan Sun |
IEEE Signal Process. Lett. | 5 |
| 2025 | Enhancing Adversarial Transferability With Alignment NetworkabstractDeep neural networks (DNNs) have been confirmed to exhibit vulnerability, as they are susceptible to deception by adversarial examples. Transfer-based attacks perturb a surrogate model and use the transferability of adversarial examples to attack other models. The effectiveness of these attacks relies heavily on the surrogate model, which often focuses on non-critical regions like backgrounds or object edges, leading to poor transferability. The intrinsic properties of the surrogate model fundamentally determine the performance of transfer-based attacks, yet this aspect has rarely been the focus of research. Therefore, we respectively design image masking operations for Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), forcing the model to reallocate attention to the critical regions. The attention of the surrogate model on the masked image and the original image is then aligned by inserting an alignment network inside the model. The modified surrogate model becomes more proficient in capturing the critical regions within the image, thereby generating more powerful adversarial examples. The proposed alignment network can be integrated into existing transfer-based attacks, significantly enhancing their performance. In addition, we also propose a novel feature-level attack based on the aligned attention, demonstrating superior performance compared to existing state-of-the-art feature-level attacks. Qi Li 0005, Yiwei Ru, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Exploring Near-Infrared Iris Image Sequences for High Throughput Iris RecognitionabstractHigh throughput is demanding in real-world iris recognition applications. The challenges mainly originate from the variability in image quality under high-throughput capture conditions. Most of the degraded images are typically filtered out by traditional iris systems through Image Quality Assessment (IQA) module, adversely affecting efficiency and leading to low throughput and poor user experience. Therefore, a better and practical solution is to make the utmost of degraded iris images. In order to investigate the key problems of high-throughput iris recognition, we collect a novel iris sequence dataset under Near-infrared (NIR) illumination. This dataset is specifically constructed for high-throughput evaluation, which faithfully simulates the process of iris sequence acquisition in real-world iris systems. Comprehensive evaluations were conducted to figure out the deficiencies of current iris recognition algorithms. To this end, a testing methodology along with specific evaluation metrics is proposed. It is capable of assessing the throughput performance, e.g., the newly proposed Frame Consumption per Match (FCM). Through performance analysis, several insights were gathered to guide potential directions for developing high-throughput iris recognition algorithms. Furthermore, we consider to leverage iris sequence features for better throughput performance. Continuity sequence criteria and cumulative sequence feature strategy are proposed to enhance the throughput performance of existing algorithms with minimal cost. In summary, this work provides valuable data and rational insights for high-throughput iris recognition studies. The datasets and evaluation toolkit are publicly available on our website1. Mupei Li, Yunlong Wang 0003, Kunbo Zhang, Zhaofeng He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Cross-Optical Property Image Translation for Face Anti-Spoofing: From Visible to PolarizationabstractDespite the development of spectral sensors and spectral data-driven learning methods which have led to significant advances in face anti-spoofing (FAS), the singular dimensionality of spectral information often results in poor robustness and weak generalization. Polarization, another fundamental property of light, can reveal intrinsic differences between genuine and fake faces with advantaged performance in precision, robustness, and generalizability. In this paper, we propose a facial image translation method from visible light (VIS) to polarization (VPT), capable of generating valuable polarimetric optical characteristics for facial presentation attack detection using VIS spectrum information input only. Specifically, the VPT method adopts a multi-stream network structure, comprising a main network and two branch networks, to translate VIS images into degree of polarization (DoP) images and Stokes polarization parameters${S}_{1}$and${S}_{2}$. To further improve image translation quality, we introduce a frequency-domain consistency loss as a complement to the existing spatial losses to narrow the gap in the frequency domain. The physical mapping relations for the DoP and Stokes parameters are employed, and the Stokes loss is designed to ensure that the generated polarization modalities conform to objective physical laws. Extensive experiments on the CASIA-Polar and CASIA-SURF datasets demonstrate the superiority of VPT over other baseline methods in terms of polarization image quality and its remarkable performance in the FAS task. This work leverages the inherent physical advantages of polarization information in material discrimination tasks while addressing hardware limitations in polarization image collection, proposing a novel solution for face recognition system security control. Yu Tian 0017, Kunbo Zhang, Yalin Huang, Leyuan Wang, Yue Liu 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Uncertainty-Aware Bilateral Transformer for Accurate and Reliable Iris SegmentationabstractIris segmentation is a deterministic and critical part of the iris recognition system. However, its performance is usually degraded by data uncertainty in acquisition and annotation, impeding more accurate recognition of the iris recognition system. In the paper, we propose a bilateral self-attention by exploring spatial and visual relationships to effectively distinguish between iris and non-iris regions, then design a bilateral Transformer by enhancing spatial perception and hierarchical feature fusion to mitigate the impact of acquisition uncertainty. Besides, iris segmentation uncertainty learning is developed to estimate the uncertainty map according to prediction discrepancy. With the estimated uncertainty, a weighting scheme and a regularization term are designed to minimize the effect of annotation uncertainty. To investigate data uncertainty, the paper presents a challenging near-infrared iris dataset named UTIris. It comprises 3,690 images with high acquisition uncertainty and provides rich segmentation masks to explore annotation uncertainty. Furthermore, we manually label a large-scale iris dataset, ND-0405 [1], with additional binary maps of iris masks to evaluate segmentation performance. Experimental results on UTIris and four other databases demonstrate the effectiveness of the proposed method in iris segmentation, and its segmentation improvement consequently promotes recognition accuracy. Jianze Wei, Xingyu Gao 0001, Yunlong Wang 0003, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Multi-Scale Semantic-Guidance Networks: Robust Blind Face Restoration Against Adversarial AttacksabstractImage processing networks are known to be vulnerable to adversarial examples, where adding carefully crafted adversarial perturbations to the inputs can mislead the model. This paper addresses the problem of robust blind face restoration (BFR) against adversarial attacks. BFR refers to recovering the HQ images from the LQ images, which suffer from diverse unknown degradation, such as noise, blur, artifact removal, low resolution, etc. Although existing BFR methods exhibit good performance, they experience significant degradation when subtle distortions and perturbations are introduced into the input images. This paper is the first to investigate, improve comprehensively, and evaluate BFR methods towards adversarial attacks. Project Gradient Descent (PGD) is employed to generate adversarial examples, and multiple types of attacks were used to thoroughly assess the robustness of various BFR methods across different objectives, regions, and levels. We evaluate the robustness of multiple BFR methods and analyze the advantages of their structures and modules towards adversarial attacks. Experimental results demonstrate that the method utilizing latent feature encoding and pre-trained discrete HQ codebook achieves better robustness than other methods, with the latter outperforming the former. Similarly, multi-scale semantic guidance information also exhibits superior performance in enhancing robustness. Therefore, we propose a powerful BFR method to mitigate this issue while maintaining better performance. Extensive experiments on three real-world datasets demonstrate our method’s state-of-the-art robustness in different scenarios. Zhenyuan Zhang 0001, Xingqun Qi, Zhenbo Song, Zhiqin Yang, Jianfeng Lu 0003, Muyi Sun, Man Zhang 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Toward Generalized Iris Presentation Attack Detection: A Mask-and-Distill Mixture of Experts ApproachabstractIris Presentation Attack Detection (PAD) is critical for securing recognition systems, yet its practical deployment is severely hindered by the poor generalization of models across different acquisition devices and diverse datasets. To address this persistent cross-domain challenge, we first introduce a comprehensive evaluation framework, the Iris Presentation Attack Detection Cross-Domain-Testing (IPAD-CDT) Protocol, designed to evaluate the model robustness in these scenarios. Our core contribution is a novel Masked Mixture-of-Experts (MMoE) method, which enhances the generalization of Transformer-based architectures. MMoE introduces a structured information asymmetry, where "student" Experts learn robust features from masked inputs by distilling knowledge from an unmasked "teacher" Expert via a cosine distance loss. This mask-and-distill mechanism effectively mitigates overfitting and guides the model to learn domain-invariant cues. By integrating MMoE into a CLIP-based model, we conduct extensive experiments on our IPAD-CDT protocol. The results demonstrate that our method sets a new state-of-the-art, significantly outperforming existing models, especially in the challenging cross-dataset and cross-device settings. Hang Zou 0002, Chenxi Du, Ajian Liu 0001, Yuan Zhang 0023, Jing Liu 0062, Jun Wan 0001, Hui Zhang 0061, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2025 | CASIA-PR-V1: A Multi-Ethnic, Multi-Device and Cross-Spectral Dataset and a Multiscale Disentangled Model for Periocular RecognitionabstractPeriocular recognition is regarded as an alternative trait for biometric recognition that can effectively solve the identification problem under large occlusions. However, few datasets are tailored for periocular recognition. For most compromises, iris datasets at near-infrared wavelengths, miss information about the eyebrows or eyelids. In this paper, a challenging dataset for real scenarios named CASIA-PR-V1 with evaluation protocols is released for periocular recognition. It is collected from multiple types of mobile devices with different resolutions or wavelengths. A rich set of attributes, e.g., ethnicities, is tagged to support fine-grained classification tasks. Moreover, we consider a wide range of noisy data in unconstrained environment, especially for glasses. Superior to its counterparts, this periocular dataset is highly valuable for studying cross-device and cross-spectral periocular recognition with occlusions, as well as fine-grained attribute classification. Additionally, a multiscale disentangled model is proposed to extract discriminating representations for periocular recognition with severe occlusions. Extensive experiments are conducted on CASIA-PR-V1, and the results indicate the superiority of our model for unconstraint periocular recognition. Yiwei Ru, Yushan Han, Longteng Kong, Zijian Wang 0009, Yong He 0009, Zhenan Sun |
IEEE Trans. Multim. | 7 |
| 2024 | Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesabstractReconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive progress, they ignore formulating contacts in their frameworks, which results in producing less realistic object meshes. In this work, we explore how to model contacts in an explicit way to benefit the implicit reconstruction of hand-held objects. Our method consists of two components: explicit contact prediction and implicit shape reconstruction. In the first part, we propose a new subtask of directly estimating 3D hand-object contacts from a single image. The part-level and vertex-level graph-based transformers are cascaded and jointly learned in a coarse-to-fine manner for more accurate contact probabilities. In the second part, we introduce a novel method to diffuse estimated contact states from the hand mesh surface to nearby 3D space and leverage diffused contact probabilities to construct the implicit neural representation for the manipulated object. Benefiting from estimating the interaction patterns between the hand and the object, our method can reconstruct more realistic object meshes, especially for object parts that are in contact with hands. Extensive experiments on challenging benchmarks show that the proposed method outperforms the current state of the arts by a great margin. Our code is publicly available at https://junxinghu.github.io/projects/hoi.html. Junxing Hu, Hongwen Zhang 0001, Zerui Chen, Mengcheng Li, Yunlong Wang 0003, Yebin Liu, Zhenan Sun |
AAAI | 7 |
| 2024 | MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-Wise Pruning Error MetricabstractVision-language pretrained models have achieved impressive performance on various downstream tasks. However, their large model sizes hinder their utilization on platforms with limited computational resources. We find that directly using smaller pretrained models and applying magnitude-based pruning on CLIP models leads to in-flexibility and inferior performance. Recent efforts for VLP compression either adopt uni-modal compression metrics resulting in limited performance or involve costly mask-search processes with learnable masks. In this paper, we first propose the Module-wise Pruning Error (MoPE) met-ric, accurately assessing CLIP module importance by performance decline on cross-modal tasks. Using the MoPE metric, we introduce a unified pruning framework applica-ble to both pretraining and task-specific fine-tuning compression stages. For pretraining, MoPE-CLIP effectively leverages knowledge from the teacher model, significantly reducing pretraining costs while maintaining strong zero-shot capabilities. For fine-tuning, consecutive pruning from width to depth yields highly competitive task-specific models. Extensive experiments in two stages demonstrate the effectiveness of the MoPE metric, and MoPE-CLIP outperforms previous state-of-the-art VLP compression methods. Haokun Lin, Haoli Bai, Zhili Liu, Lu Hou 0002, Muyi Sun, Linqi Song, Ying Wei 0001, Zhenan Sun |
CVPR | 8 |
| 2024 | DCAPose: Improve One-Stage Multi-Person Pose Estimation with Dynamic Center AssignmentabstractSingle-stage methods for multi-person pose estimation have gained significant attention for their ability to concurrently localize person positions and perceive body structure in a single processing step. However, existing single-stage methods often rely on hand-crafted centers to represent the position of human instances. Such a simplification tends to overlook the intricate structure of the human body, resulting in a misalignment between the designated centers and the actual instance centers, which consequently leads to suboptimal performance. In this paper, we introduce DCAPose, a straightforward yet powerful pipeline to address this issue by redefining the process of center selection as a set prediction problem. Rather than directly supervising the center positions of instances, our approach considers each position on the center map as a potential instance candidate. We utilize a skeleton-aware bipartite matching loss to facilitate one-to-one matching between the poses of the candidate set and the ground truths. Additionally, we introduce a novel bidirectional hierarchical body representation to capture human structural information more accurately. Our method eliminates the need for Non-Maximum Suppression, greatly simplifying the processing pipeline and enabling end-to-end optimization. Extensive testing on challenging benchmarks COCO and CrowdPose confirms that DCAPose surpasses other leading single-stage methods, demonstrating the effectiveness and superiority of our framework. Wei Zhang 0255, Huiru Xie, Qi Li 0005, Zhenan Sun |
FG | 4 |
| 2024 | EditHuman: Fine-Grained Text-Driven Human Video EditingabstractRecently, video editing has made significant advances. Human character, as one of the core elements in video editing, has attracted great research attention. However, when editing characters with strong structural information, previous methods generally encounter blurring and distortion in the limbs. In this paper, we present EditHuman, a model to realize fine-grained text-driven human video editing tasks, which achieves continuous pose movements and high-quality limb expression. Considering complex body structures and continuity of motion, more precise designs are needed to obtain practical performance. Specifically, we propose a Cascaded UNet (CAU) to realize a coarse-to-fine denoising process and refined noise estimation. Meanwhile, we introduce two Heatmap-Centric Attention Modules called Key-Element Attention (KEA) and Key-Temporal Attention (KTA) to enhance the quality of human limb expression and inter-frame continuity. Moreover, we utilize the estimated heatmap to guide the noise prediction, which further refines the video quality. Extensive experiments show that EditHuman has achieved the SOTA performance. Kaiduo Zhang, Muyi Sun, Junxing Hu, Kunbo Zhang, Zhenan Sun |
IJCB | 5 |
| 2024 | DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsabstractQuantization of large language models (LLMs) faces significant challenges, particularly due to the presence of outlier activations that impede efficient low-bit representation. Traditional approaches predominantly address Normal Outliers, which are activations across all tokens with relatively large magnitudes. However, these methods struggle with smoothing Massive Outliers that display significantly larger values, which leads to significant performance degradation in low-bit quantization. In this paper, we introduce DuQuant, a novel approach that utilizes rotation and permutation transformations to more effectively mitigate both massive and normal outliers. First, DuQuant starts by constructing the rotation matrix, using specific outlier dimensions as prior knowledge, to redistribute outliers to adjacent channels by block-wise rotation. Second, We further employ a zigzag permutation to balance the distribution of outliers across blocks, thereby reducing block-wise variance. A subsequent rotation further smooths the activation landscape, enhancing model performance. DuQuant simplifies the quantization process and excels in managing outliers, outperforming the state-of-the-art baselines across various sizes and types of LLMs on multiple tasks, even with 4-bit weight-activation quantization. Our code is available at https://github.com/Hsu1023/DuQuant. Haokun Lin, Jingzhi Cui, Yingtao Zhang, Linzhan Mou, Linqi Song, Zhenan Sun, Ying Wei 0001 |
NeurIPS | 8 |
| 2024 | Open-Set Single-Domain Generalization for Robust Face Anti-Spoofing
Fangling Jiang, Qi Li 0005, Weining Wang 0001, Zhenan Sun |
Int. J. Comput. Vis. | 7 |
| 2024 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2024 | Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2024 | A Survey on Self-Supervised Learning: Algorithms, Applications, and Future TrendsabstractDeep supervised learning algorithms typically require a large volume of labeled data to achieve satisfactory performance. However, the process of collecting and labeling such data can be expensive and time-consuming. Self-supervised learning (SSL), a subset of unsupervised learning, aims to learn discriminative features from unlabeled data without relying on human-annotated labels. SSL has garnered significant attention recently, leading to the development of numerous related algorithms. However, there is a dearth of comprehensive studies that elucidate the connections and evolution of different SSL variants. This paper presents a review of diverse SSL methods, encompassing algorithmic aspects, application domains, three key trends, and open research questions. First, we provide a detailed introduction to the motivations behind most SSL algorithms and compare their commonalities and differences. Second, we explore representative applications of SSL in domains such as image processing, computer vision, and natural language processing. Lastly, we discuss the three primary trends observed in SSL research and highlight the open questions that remain. Jie Gui, Tuo Chen, Jing Zhang 0037, Qiong Cao, Zhenan Sun, Hao Luo 0004, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Learning Disentangled Representation for One-Shot Progressive Face SwappingabstractAlthough face swapping has attracted much attention in recent years, it remains a challenging problem. Existing methods leverage a large number of data samples to explore the intrinsic properties of face swapping without considering the semantic information of face images. Moreover, the representation of the identity information tends to be fixed, leading to suboptimal face swapping. In this paper, we present a simple yet efficient method named FaceSwapper, for one-shot face swapping based on Generative Adversarial Networks. Our method consists of a disentangled representation module and a semantic-guided fusion module. The disentangled representation module comprises an attribute encoder and an identity encoder, which aims to achieve the disentanglement of the identity and attribute information. The identity encoder is more flexible, and the attribute encoder contains more attribute details than its competitors. Benefiting from the disentangled representation, FaceSwapper can swap face images progressively. In addition, semantic information is introduced into the semantic-guided fusion module to control the swapped region and model the pose and expression more accurately. Experimental results show that our method achieves state-of-the-art results on benchmark datasets with fewer training samples. Qi Li 0005, Weining Wang 0001, Cheng-Zhong Xu 0001, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | r-FACE: Reference guided face component editing
Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Qi Li 0005, Zhenan Sun |
Pattern Recognit. | 5 |
| 2024 | An Automated Framework for Histopathological Nucleus Segmentation With Deep Attention Integrated NetworksabstractClinical management and accurate disease diagnosis are evolving from qualitative stage to the quantitative stage, particularly at the cellular level. However, the manual process of histopathological analysis is lab-intensive and time-consuming. Meanwhile, the accuracy is limited by the experience of the pathologist. Therefore, deep learning-empowered computer-aided diagnosis (CAD) is emerging as an important topic in digital pathology to streamline the standard process of automatic tissue analysis. Automated accurate nucleus segmentation can not only help pathologists make more accurate diagnosis, save time and labor, but also achieve consistent and efficient diagnosis results. However, nucleus segmentation is susceptible to staining variation, uneven nucleus intensity, background noises, and nucleus tissue differences in biopsy specimens. To solve these problems, we propose Deep Attention Integrated Networks (DAINets), which mainly built on self-attention based spatial attention module and channel attention module. In addition, we also introduce a feature fusion branch to fuse high-level representations with low-level features for multi-scale perception, and employ the mark-based watershed algorithm to refine the predicted segmentation maps. Furthermore, in the testing phase, we design Individual Color Normalization (ICN) to settle the dyeing variation problem in specimens. Quantitative evaluations on the multi-organ nucleus dataset indicate the priority of our automated nucleus segmentation framework. Muyi Sun, Wenxuan Zou, Song Wang 0006, Zhenan Sun |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Personalized Graph Generation for Monocular 3D Human Pose and Shape Estimationabstract3D human pose and shape estimation from a single RGB image is an appealing yet challenging task. Due to the graph-like nature of human parametric models, a growing number of graph neural network-based approaches have been proposed and achieved promising results. However, existing methods build graphs for different instances based on the same template SMPL mesh, neglecting the geometric perception of individual properties. In this work, we propose an end-to-end method named Personalized Graph Generation (PGG) to construct the geometry-aware graph from an intermediate predicted human mesh. Specifically, a convolutional module initially regresses a coarse SMPL mesh tailored for each sample. Guided by the 3D structure of this personalized mesh, PGG extracts the local features from the 2D feature map. Then, these geometry-aware features are integrated with the specific coarse SMPL parameters as vertex features. Furthermore, a body-oriented adjacency matrix is adaptively generated according to the coarse mesh. It considers individual full-body relations between vertices, enhancing the perception of body geometry. Finally, a graph attentional module is utilized to predict the residuals to get the final results. Quantitative experiments across four benchmarks and qualitative comparisons on more datasets show that the proposed method outperforms state-of-the-art approaches for 3D human pose and shape estimation. Junxing Hu, Hongwen Zhang 0001, Yunlong Wang 0003, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Improving Transferability of Adversarial Samples via Critical Region-Oriented Feature-Level AttackabstractDeep neural networks (DNNs) have received a lot of attention because of their impressive progress in computer vision. However, it has been recently shown that DNNs are vulnerable to being spoofed by carefully crafted adversarial samples. These samples are generated by specific attack algorithms that can obfuscate the target model without being detected by humans. Recently, feature-level attacks have been the focus of research due to their high transferability. Existing state-of-the-art feature-level attacks all improve the transferability by greedily changing the attention of the model. However, for images that contain multiple target class objects, the attention of different models may differ significantly. Thus greedily changing attention may cause the adversarial samples corresponding to these images to fall into the local optimum of the surrogate model. Furthermore, due to the great structural differences between vision transformers (ViTs) and convolutional neural networks (CNNs), adversarial samples generated on CNNs with feature-level attacks are more difficult to successfully attack ViTs. To overcome these drawbacks, we perform the Critical Region-oriented Feature-level Attack (CRFA) in this paper. Specifically, we first propose the Perturbation Attention-aware Weighting (PAW), which destroys critical regions of the image by performing feature-level attention weighting on the adversarial perturbations without changing the model attention as much as possible. Then we propose the Region ViT-critical Retrieval (RVR), which enables the generator to accommodate the transferability of adversarial samples on ViTs by adding extra prior knowledge of ViTs to the decoder. Extensive experiments demonstrate significant performance improvements achieved by our approach, i.e., improving the fooling rate by 19.9% against CNNs and 25.0% against ViTs as compared to state-of-the-art feature-level attack method. Qi Li 0005, Fangling Jiang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Understanding Deep Face Representation via Attribute RecoveryabstractDeep neural networks have proven to be highly effective in the face recognition task, as they can map raw samples into a discriminative high-dimensional representation space. However, understanding this complex space proves to be challenging for human observers. In this paper, we propose a novel approach that interprets deep face recognition models via facial attributes. To achieve this, we introduce a two-stage framework that recovers attributes from the deep face representations. This framework allows us to quantitatively measure the significance of facial attributes in relation to the recognition model. Moreover, this framework enables us to generate sample-specific explanations through counterfactual methodology. These explanations are not only understandable but also quantitative. Through the proposed approach, we are able to acquire a deeper understanding of how the recognition model conceptualizes the notion of “identity” and understand the reasons behind the error decisions made by the deep models. By utilizing attributes as an interpretable interface, the proposed method marks a paradigm shift in our comprehension of deep face recognition models. It allows a complex model, obtained through gradient backpropagation, to effectively “communicate” with humans. The source code is available here, or you can visit this website:https://github.com/RenMin1991/Facial-Attribute-Recovery. Yuhao Zhu 0003, Yunlong Wang 0003, Yongzhen Huang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Multi-Faceted Knowledge-Driven Graph Neural Network for Iris SegmentationabstractAccurate iris segmentation, especially around the iris inner and outer boundaries, is still a formidable challenge. Pixels within these areas are difficult to semantically distinguish since they have similar visual characteristics and close spatial positions. To tackle this problem, the paper proposes an iris segmentation graph neural network (ISeGraph) for accurate segmentation. ISeGraph regards individual pixels as nodes within the graph and constructs self-adaptive edges according to multi-faceted knowledge, including visual similarity, positional correlation, and semantic consistency for feature aggregation. Specifically, visual similarity strengthens the connections between nodes sharing similar visual characteristics, while positional correlation assigns weights according to the spatial distance between nodes. In contrast to the above knowledge, semantic consistency maps nodes into a semantic space and learns pseudo-labels to define relationships based on label consistency. ISeGraph leverages multi-faceted knowledge to generate self-adaptive relationships for accurate iris segmentation. Furthermore, a pixel-wise adaptive normalization module is developed to increase the feature discriminability. It takes informative features in the shallow layer as a reference to improve the segmentation features from a statistical perspective. Experimental results on three iris datasets illustrate that the proposed method achieves superior performance in iris segmentation, increasing the segmentation accuracy in areas near the iris boundaries. Jianze Wei, Yunlong Wang 0003, Xingyu Gao 0001, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Exploring Generalizable Distillation for Efficient Medical Image SegmentationabstractEfficient medical image segmentation aims to provide accurate pixel-wise predictions with a lightweight implementation framework. However, existing lightweight networks generally overlook the generalizability of the cross-domain medical segmentation tasks. In this paper, we propose Generalizable Knowledge Distillation (GKD), a novel framework for enhancing the performance of lightweight networks on cross-domain medical segmentation by generalizable knowledge distillation from powerful teacher networks. Considering the domain gaps between different medical datasets, we propose the Model-Specific Alignment Networks (MSAN) to obtain the domain-invariant representations. Meanwhile, a customized Alignment Consistency Training (ACT) strategy is designed to promote the MSAN training. Based on the domain-invariant vectors in MSAN, we propose two generalizable distillation schemes, Dual Contrastive Graph Distillation (DCGD) and Domain-Invariant Cross Distillation (DICD). In DCGD, two implicit contrastive graphs are designed to model the intra-coupling and inter-coupling semantic correlations. Then, in DICD, the domain-invariant semantic vectors are reconstructed from two networks (i.e., teacher and student) with a crossover manner to achieve simultaneous generalization of lightweight networks, hierarchically. Moreover, a metric named Fréchet Semantic Distance (FSD) is tailored to verify the effectiveness of the regularized domain-invariant features. Extensive experiments conducted on the Liver, Retinal Vessel and Colonoscopy segmentation datasets demonstrate the superiority of our method, in terms of performance and generalization ability on lightweight networks. Xingqun Qi, Zhuojie Wu, Wenxuan Zou, Yifan Gao 0003, Muyi Sun, Shanghang Zhang, Caifeng Shan, Zhenan Sun |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | Contextualized Relation Predictive Model for Self-Supervised Group Activity Representation LearningabstractGroup activity analysis has attracted remarkable attention recently due to the widespread applications in security, entertainment and military. This article targets at learning group activity representations with self-supervision, which differs from the majorities relying heavily on manually annotated labels. Moreover, existing Self-Supervised Learning (SSL) methods for videos are sub-optimal to generate such representations because of the complex context dynamics in group activities. In this article, an end-to-end framework termed Contextualized Relation Predictive Model (Con-RPM) is proposed for self-supervised group activity representation learning with predictive coding. It involves the Serial-Parallel Transformer Encoder (SPTrans-Encoder) to model the context of spatial interactions and temporal variations, and the Hybrid Context Transformer Decoder (HConTrans-Decoder) to predict the future spatio-temporal relations guided by holistic scene context. Additionally, to improve the discriminability and consistency of prediction, we introduce a united loss integrating group-wise and person-wise contrastive losses in frame-level as well as the adversarial loss in global sequence-level. Consequently, our Con-RPM learns robust group representations via describing temporal evolutions of individual relationships and scene semantics explicitly. Extensive experimental results on downstream tasks indicate the effectiveness and generalization of our model in self-supervised learning, and present state-of-the-art performance on the Volleyball, Collective Activity, VolleyTactic, and Choi's New datasets. Longteng Kong, Yushan Han, Jie Qin 0004, Zhenan Sun |
IEEE Trans. Multim. | 5 |
| 2024 | Deep Learning Based Occluded Person Re-Identification: A SurveyabstractOccluded person re-identification (Re-ID) focuses on addressing the occlusion problem when retrieving the person of interest across non-overlapping cameras. With the increasing demand for intelligent video surveillance and the application of person Re-ID technology, the real-world occlusion problem draws considerable interest from researchers. Although a large number of occluded person Re-ID methods have been proposed, there are few surveys that focus on occlusion. To fill this gap and help boost future research, this article provides a systematic survey of occluded person Re-ID. In this work, we review recent deep learning based occluded person Re-ID research. First, we summarize the main issues caused by occlusion as four groups: position misalignment, scale misalignment, noisy information, and missing information. Second, we categorize existing methods into six solution groups: matching, image transformation, multi-scale features, attention mechanism, auxiliary information, and contextual recovery. We also discuss the characteristics of each approach, as well as the issues they address. Furthermore, we present the performance comparison of recent occluded person Re-ID methods on four public datasets: Partial-ReID, Partial-iLIDS, Occluded-ReID, and Occluded-DukeMTMC. We conclude the study with thoughts on promising future research directions. Yunjie Peng, Jinlin Wu, Boqiang Xu, Chunshui Cao, Xu Liu 0008, Zhenan Sun, Zhiqiang He 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Sclera-TransFuse: Fusing Swin Transformer and CNN for Accurate Sclera SegmentationabstractSclera segmentation is a crucial step in sclera recognition, which has been greatly advanced by Convolutional Neural Networks (CNNs). However, when dealing with non-ideal eye images, many existing CNN-based approaches are still prone to failure. One major reason is that due to the limited range of receptive fields, CNNs are difficult to effectively model global semantic relevance and thus robustly resist noise interference. To solve this problem, this paper proposes a novel two-stream hybrid model, named Sclera-TransFuse, to integrate classical ResNet-34 and recently emerging Swin Transformer encoders. Specially, the self-attentive Swin Transformer has shown a strong ability in capturing long-range spatial dependencies and has a hierarchical structure similar to CNNs. The dual encoders firstly extract coarse- and fine-grained feature representations at hierarchical stages, separately. Then a novel Cross-Domain Fusion (CDF) module based on information interaction and self-attention mechanism is introduced to efficiently fuse the multi-scale features extracted from dual encoders. Finally, the fused features are progressively upsampled and aggregated to predict the sclera masks in the decoder meanwhile deep supervision strategies are employed to learn intermediate feature representations better and faster. Experimental results show that Sclera-TransFuse achieves state-of-the-art performance on various sclera segmentation benchmarks. Additionally, a UBIRIS.v2 subset of 683 eye images with manually labeled sclera masks, and our codes are publicly available to the community through https://github.com/Ihqqq/Sclera-TransFuse. Caiyong Wang, Guangzhe Zhao, Zhaofeng He 0001, Yunlong Wang 0003, Zhenan Sun |
IJCB | 6 |
| 2023 | DFGC-VRA: DeepFake Game Competition on Visual Realism AssessmentabstractThis paper presents the summary report on the DeepFake Game Competition on Visual Realism Assessment (DFGC-VRA). Deep-learning based face-swap videos, also known as deepfakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. There is a ongoing deepfake game between its creators and detectors, with the human in the loop. The research community has been focusing on the automatic detection of these fake videos, but the assessment of their visual realism, as perceived by human eyes, is still an unexplored dimension. Visual realism assessment, or VRA, is essential for assessing the potential impact that may be brought by a specific face-swap video, and it is also useful as a quality metric to compare different face-swap methods. This is the third edition of DFGC competitions, which focuses on the new visual realism assessment topic, different from previous ones that compete creators versus detectors. With this competition, we conduct a comprehensive study of the SOTA performance on the new task. We also release our MindSpore codes to further facilitate research in this field (https://github.com/bomb2peng/DFGC-VRA-benckmark). Bo Peng 0002, Xianyun Sun, Caiyong Wang, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Rongyu Zhang, Heng Cong, Lingzhi Fu, Yusheng Zhang, Boyuan Liu, Luka Dragar, Borut Batagelj, Peter Peer, Vitomir Struc, Xinghui Zhou, Kunlin Liu, Wenxiu Diao |
IJCB | 6 |
| 2023 | Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological IntelligenceabstractAffective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP. Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun |
ACM Multimedia | 8 |
| 2023 | Adversarial Learning Domain-Invariant Conditional Features for Robust Face Anti-spoofing
Fangling Jiang, Qi Li 0005, Zhenan Sun |
Int. J. Comput. Vis. | 5 |
| 2023 | GAN-Based Facial Attribute ManipulationabstractFacial Attribute Manipulation (FAM) aims to aesthetically modify a given face image to render desired attributes, which has received significant attention due to its broad practical applications ranging from digital entertainment to biometric forensics. In the last decade, with the remarkable success of Generative Adversarial Networks (GANs) in synthesizing realistic images, numerous GAN-based models have been proposed to solve FAM with various problem formulation approaches and guiding information representations. This paper presents a comprehensive survey of GAN-based FAM methods with a focus on summarizing their principal motivations and technical details. The main contents of this survey include: (i) an introduction to the research background and basic concepts related to FAM, (ii) a systematic review of GAN-based FAM methods in three main categories, and (iii) an in-depth discussion of important properties of FAM methods, open issues, and future research directions. This survey not only builds a good starting point for researchers new to this field but also serves as a reference for the vision community. Yunfan Liu 0001, Qi Li 0005, Qiyao Deng, Zhenan Sun, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Multiscale Dynamic Graph Representation for Biometric Recognition With OcclusionsabstractOcclusion is a common problem with biometric recognition in the wild. The generalization ability of CNNs greatly decreases due to the adverse effects of various occlusions. To this end, we propose a novel unified framework integrating the merits of both CNNs and graph models to overcome occlusion problems in biometric recognition, called multiscale dynamic graph representation (MS-DGR). More specifically, a group of deep features reflected on certain subregions is recrafted into a feature graph (FG). Each node inside the FG is deemed to characterize a specific local region of the input sample, and the edges imply the co-occurrence of non-occluded regions. By analyzing the similarities of the node representations and measuring the topological structures stored in the adjacent matrix, the proposed framework leverages dynamic graph matching to judiciously discard the nodes corresponding to the occluded parts. The multiscale strategy is further incorporated to attain more diverse nodes representing regions of various sizes. Furthermore, the proposed framework exhibits a more illustrative and reasonable inference by showing the paired nodes. Extensive experiments demonstrate the superiority of the proposed framework, which boosts the accuracy in both natural and occlusion-simulated cases by a large margin compared with that of baseline methods. Yunlong Wang 0003, Yuhao Zhu 0003, Kunbo Zhang, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | PyMAF-X: Towards Well-Aligned Full-Body Model Regression From Monocular ImagesabstractWe present PyMAF-X, a regression-based approach to recovering a parametric full-body model from a single image. This task is very challenging since minor parametric deviation may lead to noticeable misalignment between the estimated mesh and the input image. Moreover, when integrating part-specific estimations into the full-body model, existing solutions tend to either degrade the alignment or produce unnatural wrist poses. To address these issues, we propose a Pyramidal Mesh Alignment Feedback (PyMAF) loop in our regression network for well-aligned human mesh recovery and extend it as PyMAF-X for the recovery of expressive full-body models. The core idea of PyMAF is to leverage a feature pyramid and rectify the predicted parameters explicitly based on the mesh-image alignment status. Specifically, given the currently predicted parameters, mesh-aligned evidence will be extracted from finer-resolution features accordingly and fed back for parameter rectification. To enhance the alignment perception, an auxiliary dense supervision is employed to provide mesh-image correspondence guidance while spatial alignment attention is introduced to enable the awareness of the global contexts for our network. When extending PyMAF for full-body mesh recovery, an adaptive integration strategy is proposed in PyMAF-X to produce natural wrist poses while maintaining the well-aligned performance of the part-specific estimations. The efficacy of our approach is validated on several benchmark datasets for body, hand, face, and full-body mesh recovery, where PyMAF and PyMAF-X effectively improve the mesh-image alignment and achieve new The project page with code and video results can be found at https://www.liuyebin.com/pymaf-x. Hongwen Zhang 0001, Yating Tian, Yuxiang Zhang 0006, Mengcheng Li, Liang An 0001, Zhenan Sun, Yebin Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Semantic-based conditional generative adversarial hashing with pairwise labels
Qi Li 0005, Weining Wang 0001, Yuan Yan Tang, Cheng-Zhong Xu 0001, Zhenan Sun |
Pattern Recognit. | 5 |
| 2023 | Towards Spatially Disentangled Manipulation of Face Images With Pre-Trained StyleGANsabstractGenerative Adversarial Networks with style-based generators could successfully synthesize realistic images from input latent code. Moreover, recent studies have revealed that interpretable translations of generated images could be obtained by linearly traversing in the latent space. However, in most existing latent spaces, linear interpolation often leads to ‘spatially entangled modification’ in the manipulation result, which is undesirable in many real-world applications where local editing is required. To solve this problem, we propose to manipulate the latent code in the ‘style space’ and analyze its advantage in achieving spatial disentanglement. Furthermore, we point out the weakness of simply interpolating in the style space and propose ‘Style Intervention’, a lightweight optimization-based algorithm, to further improve the visual fidelity of manipulation results. The performance of our method is verified with the task of attribute editing on high-resolution face images. Both qualitative and quantitative results demonstrate the advantage of image translation in the style space and the effectiveness of our method on both real and synthetic images. Yunfan Liu 0001, Qi Li 0005, Qiyao Deng, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | AIF-LFNet: All-in-Focus Light Field Super-Resolution Method Considering the Depth-Varying DefocusabstractAs an aperture-divided computational imaging system, microlens array (MLA) -based light field (LF) imaging is playing an increasingly important role in computer vision. As the trade-off between the spatial and angular resolutions, deep learning (DL) -based image super-resolution (SR) methods have been applied to enhance the spatial resolution. However, in existing DL-based methods, the depth-varying defocus is not considered both in dataset development and algorithm design, which restricts many applications such as depth estimation and object recognition. To overcome this shortcoming, a super-resolution task that reconstructs all-in-focus high-resolution (HR) LF images from low-resolution (LR) LF images is proposed by designing a large dataset and proposing a convolutional neural network (CNN) -based SR method. The dataset is constructed by using Blender software, consisting of 150 light field images used as training data, and 15 light field images used as validation and testing data. The proposed network is designed by proposing the dilated deformable convolutional network (DCN) -based feature extraction block and the LF subaperture image (SAI) Deblur-SR block. The experimental results demonstrate that the proposed method achieves more appealing results both quantitatively and qualitatively. Shubo Zhou, Yunlong Wang 0003, Zhenan Sun, Kunbo Zhang, Xueqin Jiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | IrisGuideNet: Guided Localization and Segmentation Network for Unconstrained Iris BiometricsabstractIn recent years, unconstraint iris biometric is becoming more prevalent due to its wide range of user applications. But since it allows less user co-operation, it presents numerous challenges to the iris preprocessing task of localisation and segmentation (ILS). To address these challenges, many ILS techniques have been proposed with the deep learning CNN based approaches been the most effective. Training the CNN is data intensive and most of the existing CNN based ILS adopt general purpose CNN without any iris specific guidance. However, the available iris dataset comprises of small subsets with labelled images. As such, the existing CNN models can be less effective as they are trained with these dataset. Hence, in this paper, we propose a guided CNN based ILS technique termed IrisGuideNet by incorporating known iris specific heuristics into the network pipeline. IrisGuideNet has an encoder-decoder structure designed to be invariant to translation and rotation and can capture iris at multiple scales. To address the iris limited data problem, unlike the existing CNN based ILS, during the training process, we adopt the deep supervision technique, employ hybrid losses and introduce a novel iris specific heuristics named Iris Regularization Term (IRT) in other to effectively train the network. At inference, we introduce a novel Iris Infusion Module (IIM) that utilise the geometrical relationships between the ILS outputs to refine the predicted outputs through logical operations. Our models were trained and evaluated with the recently published NIR-ISL Challange * datasets and has proven to be effective as it has outperformed most of the participating models across all the database categories in the competition. Jawad Muhammad, Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Polarized Image Translation From Nonpolarized Cameras for Multimodal Face Anti-SpoofingabstractIn face antispoofing, it is desirable to have multimodal images to demonstrate liveness cues from various perspectives. However, in most face recognition scenarios, only a single modality, namely visible lighting (VIS) facial images is available. This paper first investigates the possibility of generating polarized (Polar) images from VIS cameras without changing the existing recognition devices to improve the accuracy and robustness of Presentation Attack Detection (PAD) in face biometrics. A novel multimodal face antispoofing framework is proposed based on the machine-learning relationship between VIS and Polar images of genuine faces. Specifically, a dual-modal central differential convolutional network (CDCN) is developed to capture the inherent spoofing features between the VIS and the generated Polar modalities. Quantitative and qualitative experimental results show that our proposed framework not only generates realistic Polar face images but also improves the state-of-the-art face anti-spoofing results on the VIS modal database (i.e. CASIA-SURF). Moreover, a polar face database, CASIA-Polar, has been constructed and will be shared with the public at http://biometrics.idealtest.org to inspire future applications within the biometric anti-spoofing field. Yu Tian 0017, Yalin Huang, Kunbo Zhang, Yue Liu 0005, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Exploring Bias in Sclera Segmentation Models: A Group Evaluation ApproachabstractBias and fairness of biometric algorithms have been key topics of research in recent years, mainly due to the societal, legal and ethical implications of potentially unfair decisions made by automated decision-making models. A considerable amount of work has been done on this topic across different biometric modalities, aiming at better understanding the main sources of algorithmic bias or devising mitigation measures. In this work, we contribute to these efforts and present the first study investigating bias and fairness of sclera segmentation models. Although sclera segmentation techniques represent a key component of sclera-based biometric systems with a considerable impact on the overall recognition performance, the presence of different types of biases in sclera segmentation methods is still underexplored. To address this limitation, we describe the results of a group evaluation effort (involving seven research groups), organized to explore the performance of recent sclera segmentation models within a common experimental framework and study performance differences (and bias), originating from various demographic as well as environmental factors. Using five diverse datasets, we analyze seven independently developed sclera segmentation models in different experimental configurations. The results of our experiments suggest that there are significant differences in the overall segmentation performance across the seven models and that among the considered factors, ethnicity appears to be the biggest cause of bias. Additionally, we observe that training with representative and balanced data does not necessarily lead to less biased results. Finally, we find that in general there appears to be a negative correlation between the amount of bias observed (due to eye color, ethnicity and acquisition device) and the overall segmentation performance, suggesting that advances in the field of semantic segmentation may also help with mitigating bias. Matej Vitek, Abhijit Das 0001, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Jalil Nourmohammadi-Khiarak, Mohsen Akbari Shahpar, Meysam Asgari-Chenaghlu, Farhang Jaryani, Juan E. Tapia, Andres Valenzuela, Caiyong Wang, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Kiran B. Raja, Gourav Gupta, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, S. V. Aruna Kumar, B. S. Harish, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 15 |
| 2023 | Contextual Measures for Iris RecognitionabstractThe iris patterns of the human contain a large amount of randomly distributed and irregularly shaped microstructures. These microstructures make the human iris informative biometric traits. To learn identity representation from them, this paper regards each iris region as a potential microstructure and proposes contextual measures (CM) to model the correlations between them. CM adopts two parallel branches to learn global and local contexts in iris image. The first one is the globally contextual measure branch. It measures the global context involving the relationships between all regions for feature aggregation and is robust to local occlusions. Besides, we improve its spatial perception considering the positional randomness of the microstructures. The other one is the locally contextual measure branch. This branch considers the role of local details in the phenotypic distinctiveness of iris patterns and learns a series of relationship atoms to capture contextual information from a local perspective. In addition, we develop the perturbation bottleneck to make sure that the two branches learn divergent contexts. It introduces perturbation to limit the information flow from input images to identity features, forcing CM to learn discriminative contextual information for iris recognition. Experimental results suggest that global and local contexts are two different clues critical for accurate iris recognition. The superior performance on four benchmark iris datasets demonstrates the effectiveness of the proposed approach in within-database and cross-database scenarios. Jianze Wei, Yunlong Wang 0003, Huaibo Huang, Ran He 0001, Zhenan Sun, Xingyu Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Joint Holistic and Masked Face RecognitionabstractWith the widespread use of face masks due to the COVID-19 pandemic, accurate masked face recognition has become more crucial than ever. While several studies have investigated masked face recognition using convolutional neural networks (CNNs), there is a paucity of research exploring the use of plain Vision Transformers (ViTs) for this task. Unlike ViT models used in image classification, object detection, and semantic segmentation, the model trained by modern face recognition losses struggles to converge when trained from scratch. To this end, this paper initializes the model parameters via a proxy task of patch reconstruction and observes that the ViT backbone exhibits improved training stability with satisfactory performance for face recognition. Beyond the training stability, two strategies based on prompts are proposed to integrate holistic and masked face recognition in a single framework, namely FaceT. Along with popular holistic face recognition benchmarks, several open-sourced masked face recognition benchmarks are collected for evaluation. Our extensive experiments demonstrate that the proposed FaceT performs on par or better than state-of-the-art CNNs on both holistic and masked face recognition benchmarks. Codes will be made available at https://github.com/zyainfal/Joint-Holistic-and-Masked-Face-Recognition. Yuhao Zhu 0003, Hui Jing, Linlin Dai, Zhenan Sun, Ping Li 0038 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Pose-Appearance Relational Modeling for Video Action RecognitionabstractRecent studies of video action recognition can be classified into two categories: the appearance-based methods and the pose-based methods. The appearance-based methods generally cannot model temporal dynamics of large motion well by virtue of optical flow estimation, while the pose-based methods ignore the visual context information such as typical scenes and objects, which are also important cues for action understanding. In this paper, we tackle these problems by proposing a Pose-Appearance Relational Network (PARNet), which models the correlation between human pose and image appearance, and combines the benefits of these two modalities to improve the robustness towards unconstrained real-world videos. There are three network streams in our model, namely pose stream, appearance stream and relation stream. For the pose stream, a Temporal Multi-Pose RNN module is constructed to obtain the dynamic representations through temporal modeling of 2D poses. For the appearance stream, a Spatial Appearance CNN module is employed to extract the global appearance representation of the video sequence. For the relation stream, a Pose-Aware RNN module is built to connect pose and appearance streams by modeling action-sensitive visual context information. Through jointly optimizing the three modules, PARNet achieves superior performances compared with the state-of-the-arts on both the pose-complete datasets (KTH, Penn-Action, UCF11) and the challenging pose-incomplete datasets (UCF101, HMDB51, JHMDB), demonstrating its robustness towards complex environments and noisy skeletons. Its effectiveness on NTU-RGBD dataset is also validated even compared with 3D skeleton-based methods. Furthermore, an appearance-enhanced PARNet equipped with a RGB-based I3D stream is proposed, which outperforms the Kinetics pre-trained competitors on UCF101 and HMDB51. The better experimental results verify the potentials of our framework by integrating various modules. Mengmeng Cui, Wei Wang 0115, Kunbo Zhang, Zhenan Sun, Liang Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | A Review on Generative Adversarial Networks: Algorithms, Theory, and ApplicationsabstractGenerative adversarial networks (GANs) have recently become a hot research topic; however, they have been studied since 2014, and a large number of algorithms have been proposed. Nevertheless, few comprehensive studies explain the connections among different GAN variants and how they have evolved. In this paper, we attempt to provide a review of the various GAN methods from the perspectives of algorithms, theory, and applications. First, the motivations, mathematical representations, and structures of most GAN algorithms are introduced in detail, and we compare their commonalities and differences. Second, theoretical issues related to GANs are investigated. Finally, typical applications of GANs in image processing and computer vision, natural language processing, music, speech and audio, the medical field, and data science are discussed. Jie Gui, Zhenan Sun, Yonggang Wen 0001, Dacheng Tao, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Graph Flow: Cross-Layer Graph Flow Distillation for Dual Efficient Medical Image SegmentationabstractWith the development of deep convolutional neural networks, medical image segmentation has achieved a series of breakthroughs in recent years. However, high-performance convolutional neural networks always mean numerous parameters and high computation costs, which will hinder the applications in resource-limited medical scenarios. Meanwhile, the scarceness of large-scale annotated medical image datasets further impedes the application of high-performance networks. To tackle these problems, we propose Graph Flow, a comprehensive knowledge distillation framework, for both network-efficiency and annotation-efficiency medical image segmentation. Specifically, the Graph Flow Distillation transfers the essence of cross-layer variations from a well-trained cumbersome teacher network to a non-trained compact student network. In addition, an unsupervised Paraphraser Module is integrated to purify the knowledge of the teacher, which is also beneficial for the training stabilization. Furthermore, we build a unified distillation framework by integrating the adversarial distillation and the vanilla logits distillation, which can further refine the final predictions of the compact network. With different teacher networks (traditional convolutional architecture or prevalent transformer architecture) and student networks, we conduct extensive experiments on four medical image datasets with different modalities (Gastric Cancer, Synapse, BUSI, and CVC-ClinicDB). We demonstrate the prominent ability of our method on these datasets, which achieves competitive performances. Moreover, we demonstrate the effectiveness of our Graph Flow through a novel semi-supervised paradigm for dual efficient medical image segmentation. Our code will be available at Graph Flow. Wenxuan Zou, Xingqun Qi, Muyi Sun, Zhenan Sun, Caifeng Shan |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Semantic-Aware Noise Driven Portrait Synthesis and ManipulationabstractSemantic portrait synthesis has drawn consistent attention and has made significant progress, yet achieving style diversity and semantic controllability simultaneously is still a challenge. Existing methods either 1) directly take a semantic label map as input, ignoring various possibilities of semantic styles, or 2) sample global noise as input, ignoring controllability of local semantics. To fill this gap, we propose semantic-aware noise, a simple but effective input that tackles both issues and shows improved results over baselines. Semantic-aware noise introduces semantic information into noise, and each semantic is sampled from the noise separately, combining the semantic controllability and the noise sampling diversity. To further expand and manipulate real images, we propose a novel ternary network structure, allowing simultaneous diverse semantic image synthesis and real image manipulation in a unified framework. Extensive experiments demonstrate that the proposed method achieves quantitatively superior and perceptually pleasing results compared to state-of-the-art methods. We also analyze the performance of our method with respect to different noise structures and real-life applications in diverse synthesis, interactive manipulation, and extreme pose scenarios. Qiyao Deng, Qi Li 0005, Jie Cao 0002, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Multim. | 5 |
| 2023 | Dilated Convolution-based Feature Refinement Network for Crowd LocalizationabstractAs an emerging computer vision task, crowd localization has received increasing attention due to its ability to produce more accurate spatially predictions. However, continuous scale variations in complex crowd scenes lead to tiny individuals at the edges, so that existing methods cannot achieve precise crowd localization. Aiming at alleviating the above problems, we propose a novel Dilated Convolution-based Feature Refinement Network (DFRNet) to enhance the representation learning capability. Specifically, the DFRNet is built with three branches that can capture the information of each individual in crowd scenes more precisely. More specifically, we introduce a Feature Perception Module to model long-range contextual information at different scales by adopting multiple dilated convolutions, thus providing sufficient feature information to perceive tiny individuals at the edge of images. Afterwards, a Feature Refinement Module is deployed at multiple stages of the three branches to facilitate the mutual refinement of feature information at different scales, thus further improving the expression capability of multi-scale contextual information. By incorporating the above modules, DFRNet can locate individuals in complex scenes more precisely. Extensive experiments on multiple datasets demonstrate that the proposed method has more advanced performance compared to existing methods and can be more accurately adapted to complex crowd scenes. Xingyu Gao 0001, Jinyang Xie, Zhenyu Chen 0003, Anan Liu, Zhenan Sun, Lei Lyu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Color-Unrelated Head-Shoulder Networks for Fine-Grained Person Re-identificationabstractPerson re-identification (re-id) attempts to match pedestrian images with the same identity across non-overlapping cameras. Existing methods usually study person re-id by learning discriminative features based on the clothing attributes (e.g., color, texture). However, the clothing appearance is not sufficient to distinguish different persons especially when they are in similar clothes, which is known as the fine-grained (FG) person re-id problem. By contrast, this paper proposes to exploit the color-unrelated feature along with the head-shoulder feature for FG person re-id. Specifically, a color-unrelated head-shoulder network (CUHS) is developed, which is featured in three aspects: (1) It consists of a lightweight head-shoulder segmentation layer for localizing the head-shoulder region and learning the corresponding feature. (2) It exploits instance normalization (IN) for learning color-unrelated features. (3) As IN inevitably reduces inter-class differences, we propose to explore richer visual cues for IN by an attention exploration mechanism to ensure high discrimination. We evaluate our model on the FG-reID, Market1501, and DukeMTMC-reID datasets, and the results show that CUHS surpasses previous methods on both the FG and conventional person re-id problems. Boqiang Xu, Jian Liang 0001, Lingxiao He, Jinlin Wu, Zhenan Sun |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | ShowFace: Coordinated Face Inpainting with Memory-Disentangled Refinement Networks
Zhuojie Wu, Xingqun Qi, Zijian Wang 0009, Kun Yuan 0003, Muyi Sun, Zhenan Sun |
BMVC | 7 |
| 2022 | AnyFace: Free-style Text-to-Face Synthesis and ManipulationabstractExisting text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face method namely AnyFace enabling much wider open world applications such as metaverse, social media, cosmetics, forensics, etc. AnyFace has a novel two-stream framework for face image synthesis and manipulation given arbitrary descriptions of the human face. Specifically, one stream performs text-to-face generation and the other conducts face image reconstruction. Facial text and image features are extracted using the CLIP (Contrastive Language-Image Pre-training) encoders. And a collaborative Cross Modal Distillation (CMD) module is designed to align the linguistic and visual features across these two streams. Furthermore, a Diverse Triplet Loss (DT loss) is developed to model fine-grained features and improve facial diversity. Extensive experiments on Multi-modal CelebA-HQ and CelebAText-HQ demonstrate significant advantages of AnyFace over state-of-the-art methods. AnyFace can achieve high-quality, high-resolution, and high-diversity face synthesis and manipulation results without any constraints on the number and content of input captions. Jianxin Sun 0003, Qiyao Deng, Qi Li 0005, Muyi Sun, Zhenan Sun |
CVPR | 6 |
| 2022 | Mimic Embedding via Adaptive Aggregation: Learning Generalizable Person Re-identification
Boqiang Xu, Jian Liang 0001, Lingxiao He, Zhenan Sun |
ECCV (14) | 4 |
| 2022 | DFGC 2022: The Second DeepFake Game CompetitionabstractThis paper presents the summary report on our DFGC 2022 competition. The DeepFake is rapidly evolving, and realistic face-swaps are becoming more deceptive and difficult to detect. On the other hand, methods for detecting DeepFakes are also improving. There is a two-party game between DeepFake creators and defenders. This competition provides a common platform for benchmarking the game between the current state-of-the-arts in Deep-Fake creation and detection methods. The main research question to be answered by this competition is the current state of the two adversaries when competed with each other. This is the second edition after the last year's DFGC 2021, with a new, more diverse video dataset, a more realistic game setting, and more reasonable evaluation metrics. With this competition, we aim to stimulate research ideas for building better defenses against the DeepFake threats. We also release our DFGC 2022 dataset contributed by both our participants and ourselves to enrich the DeepFake data resources for the research community (https://github.com/NiCE-X/DFGC-2022). Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Zhen Lei 0001, Siwei Lyu |
IJCB | 6 |
| 2022 | PDVN: A Patch-based Dual-view Network for Face Liveness Detection using Light Field Focal StackabstractLight Field Focal Stack (LFFS) can be efficiently rendered from a light field (LF) image captured by plenoptic cameras. Differences in the 3D surface and texture of biometric samples are internally reflected in the defocus blur and local patterns between the rendered slices of LFFS. This unique property makes LFFS quite appropriate to differentiate presentation attack instruments (PAIs) from bona fide samples. A patch-based dual-view network (PDVN) is proposed in this paper to leverage the merits of LFFS for face presentation attack detection (PAD). First, original LFFS data are divided into various local patches along spatial dimensions, which distracts the model from learning the useless facial semantics and greatly relieve the problem of insufficient samples. The strategy of dual-view branches is innovatively proposed, wherein the original view and microscopic view can simultaneously contribute to liveness detection. Separable 3D convolution on the focal dimension is verified to be more effective than vanilla 3D convolution for extracting discriminative features from LFFS data. The voting mechanism on predictions of patch LFFS samples further strengthens the robustness of the proposed framework. PDVN is compared with other face PAD methods on IST LLFFSD dataset and achieves perfect performance, i.e., ACER drops to 0. Yunlong Wang 0003, Mupei Li, Zhengquan Luo, Zhenan Sun |
IJCB | 4 |
| 2022 | D-ESRGAN: A Dual-Encoder GAN with Residual CNN and Vision Transformer for Iris Image Super-ResolutionabstractIris images captured in less-constrained environments, especially at long distances often suffer from the interference of low resolution, resulting in the loss of much valid iris texture information for iris recognition. In this paper, we propose a dual-encoder super-resolution generative adversarial network (D-ESRGAN) for compensating texture lost of the raw image meanwhile maintaining the newly generated textures more natural. Specifically, the proposed D-ESRGAN not only integrates the residual CNN encoder to extract local features, but also employs an emerging vision transformer encoder to capture global associative information. The local and global features from two encoders are further fused for the subsequent reconstruction of high-resolution features. During the training, we develop a three-stage strategy to alleviate the problem that generative adversarial networks are prone to collapse. Moreover, to boost the iris recognition performance, we introduce a triplet loss to push away the distance of super-resolved iris images with different IDs, and pull the distance of super-resolved iris images with the same ID much closer. Experimental results on the public CASIA-Iris-distance and CASIA-Iris-M1 datasets show that D-ESRGAN archives better performance than state-of-the-art baselines in terms of both super-resolution image quality metrics and iris recognition metric. Caiyong Wang, Gaosheng Wu, Yunlong Wang 0003, Zhenan Sun |
IJCB | 5 |
| 2022 | Disentangled Federated Learning for Tackling Attributes Skew via Invariant Aggregation and Diversity TransferringabstractAttributes skew hinders the current federated learning (FL) frameworks from consistent optimization directions among the clients, which inevitably leads to performance reduction and unstable convergence. The core problems lie in that: 1) Domain-specific attributes, which are non-causal and only locally valid, are indeliberately mixed into global aggregation. 2) The one-stage optimizations of entangled attributes cannot simultaneously satisfy two conflicting objectives, i.e., generalization and personalization. To cope with these, we proposed disentangled federated learning (DFL) to disentangle the domain-specific and cross-invariant attributes into two complementary branches, which are trained by the proposed alternating local-global optimization independently. Importantly, convergence analysis proves that the FL system can be stably converged even if incomplete client models participate in the global aggregation, which greatly expands the application scope of FL. Extensive experiments verify that DFL facilitates FL with higher performance, better interpretability, and faster convergence rate, compared with SOTA FL methods on both manually synthesized and realistic attributes skew datasets. Zhengquan Luo, Yunlong Wang 0003, Zilei Wang, Zhenan Sun, Tieniu Tan |
ICML | 4 |
| 2022 | MOST-Net: A Memory Oriented Style Transfer Network for Face Sketch SynthesisabstractFace sketch synthesis has been widely used in multimedia entertainment and law enforcement. Despite the recent developments in deep neural networks, accurate and realistic face sketch synthesis is still a challenging task due to the diversity and complexity of human faces. Current image-to-image translation-based face sketch synthesis frequently encounters over-fitting problems when it comes to small-scale datasets. To tackle this problem, we present an end-to-end Memory Oriented Style Transfer Network (MOST-Net) for face sketch synthesis which can produce high-fidelity sketches with limited data. Specifically, an external self-supervised dynamic memory module is introduced to capture the domain alignment knowledge in the long term. In this way, our proposed model could obtain the domain-transfer ability by establishing the durable relationship between faces and corresponding sketches on the feature level. Furthermore, we design a novel Memory Refinement Loss (MR Loss) for feature alignment in the memory module, which enhances the accuracy of memory slots in an unsupervised manner. Extensive experiments on the CUFS and the CUFSF datasets show that our MOST-Net achieves state-of-the-art performance, especially in terms of the Structural Similarity Index(SSIM). Fan Ji, Muyi Sun, Xingqun Qi, Qi Li 0005, Zhenan Sun |
ICPR | 5 |
| 2022 | Gender and ethnicity recognition based on visual attention-driven deep architectures
Souad Khellat-Kihel, Jawad Muhammad, Zhenan Sun, Massimo Tistarelli |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Learning 3D Human Shape and Pose From Dense Body PartsabstractReconstructing 3D human shape and pose from monocular images is challenging despite the promising results achieved by the most recent learning-based methods. The commonly occurred misalignment comes from the facts that the mapping from images to the model space is highly non-linear and the rotation-based pose representation of the body model is prone to result in the drift of joint positions. In this work, we investigate learning 3D human shape and pose from dense correspondences of body parts and propose a Decompose-and-aggregate Network (DaNet) to address these issues. DaNet adopts the dense correspondence maps, which densely build a bridge between 2D pixels and 3D vertexes, as intermediate representations to facilitate the learning of 2D-to-3D mapping. The prediction modules of DaNet are decomposed into one global stream and multiple local streams to enable global and fine-grained perceptions for the shape and pose predictions, respectively. Messages from local streams are further aggregated to enhance the robust prediction of the rotation-based poses, where a position-aided rotation feature refinement strategy is proposed to exploit spatial relationships between body joints. Moreover, a Part-based Dropout (PartDrop) strategy is introduced to drop out dense information from intermediate representations during training, encouraging the network to focus on more complementary body parts as well as neighboring position features. The efficacy of the proposed method is validated on both indoor and real-world datasets including Human3.6M, UP3D, COCO, and 3DPW, showing that our method could significantly improve the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://hongwenzhang.github.io/dense2mesh. Hongwen Zhang 0001, Jie Cao 0002, Guo Lu, Wanli Ouyang, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Cross-Spectral Iris Recognition by Learning Device-Specific BandabstractCross-spectral recognition is still an open challenge in iris recognition. In cross-spectral iris recognition, there exist distinct device-specific bands between near-infrared (NIR) and visible (VIS) images, resulting in the distribution gap between samples from different spectra and thus severe degradation in recognition performance. To tackle this problem, we propose a new cross-spectral iris recognition method to learn spectral-invariant features by estimating device-specific bands. In the proposed method,GaborTridentNetwork (GTN) first utilizes the Gabor function’s priors to perceive iris textures under different spectra, and then codes the device-specific band as the residual component to assist the generation of spectral-invariant features. By investigating the device-specific band, GTN effectively reduces the impact of device-specific bands on identity features. Besides, we make three efforts to further reduce the distribution gap. First,SpectralAdversarialNetwork (SAN) adopts a class-level adversarial strategy to align feature distributions. Second,Sample-Anchor (SA) loss upgrades triplet loss by pulling samples to their class center and pushing away from other class centers. Third, we develop a higher-order alignment loss to measures the distribution gap according to space bases and distribution shapes. Extensive experiments on five iris datasets demonstrate the efficacy of our proposed method for cross-spectral iris recognition. Jianze Wei, Yunlong Wang 0003, Yi Li 0018, Ran He 0001, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Perturbation Inactivation Based Adversarial Defense for Face RecognitionabstractDeep learning-based face recognition models are vulnerable to adversarial attacks. To curb these attacks, most defense methods aim to improve the robustness of recognition models against adversarial perturbations. However, the generalization capacities of these methods are quite limited. In practice, they are still vulnerable to unseen adversarial attacks. Deep learning models are fairly robust to general perturbations, such as Gaussian noises. A straightforward approach is to inactivate the adversarial perturbations so that they can be easily handled as general perturbations. In this paper, a plug-and-play adversarial defense method, named perturbation inactivation (PIN), is proposed to inactivate adversarial perturbations for adversarial defense. We discover that the perturbations in different subspaces have different influences on the recognition model. There should be a subspace, called the immune space, in which the perturbations have fewer adverse impacts on the recognition model than in other subspaces. Hence, our method estimates the immune space and inactivates the adversarial perturbations by restricting them to this subspace. The proposed method can be generalized to unseen adversarial perturbations since it does not rely on a specific kind of adversarial attack method. This approach not only outperforms several state-of-the-art adversarial defense methods but also demonstrates a superior generalization capacity through exhaustive experiments. Moreover, the proposed method can be successfully applied to four commercial APIs without additional training, indicating that it can be easily generalized to existing face recognition systems. The source code is available at https://github.com/RenMin1991/Perturbation-Inactivate.. Yuhao Zhu 0003, Yunlong Wang 0003, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | A Unified Framework for Biphasic Facial Age Translation With Noisy-Semantic Guided Generative Adversarial NetworksabstractBiphasic facial age translation aims at predicting the appearance of the input face at any age. Facial age translation has received considerable research attention in the last decade due to its practical value in cross-age face recognition and various entertainment applications. However, most existing methods model age changes between holistic images, regardless of the human face structure and the age-changing patterns of individual facial components. Consequently, the lack of semantic supervision will cause infidelity of generated faces in detail. To this end, we propose a unified framework for biphasic facial age translation with noisy-semantic guided generative adversarial networks. Structurally, we project the class-aware noisy semantic layouts to “soft” latent maps for the following injection operation on the individual facial parts. In particular, we introduce two sub-networks, ProjectionNet and ConstraintNet. ProjectionNet introduces the low-level structural semantic information with noise map and produces “soft” latent maps. ConstraintNet disentangles the high-level spatial features to constrain the “soft” latent maps, which endows more age-related context into the “soft” latent maps. Specifically, attention mechanism is employed in ConstraintNet for feature disentanglement. Meanwhile, in order to mine the strongest mapping ability of the network, we embed two types of learning strategies in the training procedure, supervised self-driven generation and unsupervised condition-driven cycle-consistent generation. As a result, extensive experiments conducted on MORPH and CACD datasets demonstrate the prominent ability of our proposed method which achieves state-of-the-art performance. Muyi Sun, Jianshu Li, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | Towards More Discriminative and Robust Iris Recognition by Learning Uncertain FactorsabstractThe uncontrollable acquisition process limits the performance of iris recognition. In the acquisition process, various inevitable factors, including eyes, devices, and environment, hinder the iris recognition system from learning a discriminative identity representation. This leads to severe performance degradation. In this paper, we explore uncertain acquisition factors and propose uncertainty embedding (UE) and uncertainty-guided curriculum learning (UGCL) to mitigate the influence of acquisition factors. UE represents an iris image using a probabilistic distribution rather than a deterministic point (binary template or feature vector) that is widely adopted in iris recognition methods. Specifically, UE learns identity and uncertainty features from the input image, and encodes them as two independent components of the distribution, mean and variance. Based on this representation, an input image can be regarded as an instantiated feature sampled from the UE, and we can also generate various virtual features through sampling. UGCL is constructed by imitating the progressive learning process of newborns. Particularly, it selects virtual features to train the model in an easy-to-hard order at different training stages according to their uncertainty. In addition, an instance-level enhancement method is developed by utilizing local and global statistics to mitigate the data uncertainty from image noise and acquisition conditions in the pixel-level space. The experimental results on six benchmark iris datasets verify the effectiveness and generalization ability of the proposed method on same-sensor and cross-sensor recognition. Jianze Wei, Huaibo Huang, Yunlong Wang 0003, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Learning Feature Recovery Transformer for Occluded Person Re-IdentificationabstractOne major issue that challenges person re-identification (Re-ID) is the ubiquitous occlusion over the captured persons. There are two main challenges for the occluded person Re-ID problem, i.e. , the interference of noise during feature matching and the loss of pedestrian information brought by the occlusions. In this paper, we propose a new approach called Feature Recovery Transformer (FRT) to address the two challenges simultaneously, which mainly consists of visibility graph matching and feature recovery transformer. To reduce the interference of the noise during feature matching, we mainly focus on visible regions that appear in both images and develop a visibility graph to calculate the similarity. In terms of the second challenge, based on the developed graph similarity, for each query image, we propose a recovery transformer that exploits the feature sets of its k -nearest neighbors in the gallery to recover the complete features. Extensive experiments across different person Re-ID datasets, including occluded, partial and holistic datasets, demonstrate the effectiveness of FRT. Specifically, FRT significantly outperforms state-of-the-art results by at least 6.2% Rank- 1 accuracy and 7.2% mAP scores on the challenging Occluded-Duke dataset. Boqiang Xu, Lingxiao He, Jian Liang 0001, Zhenan Sun |
IEEE Trans. Image Process. | 4 |
| 2021 | ReMix: Towards Image-to-Image Translation With Limited DataabstractImage-to-image (I2I) translation methods based on generative adversarial networks (GANs) typically suffer from overfitting when limited training data is available. In this work, we propose a data augmentation method (ReMix) to tackle this issue. We interpolate training samples at the feature level and propose a novel content loss based on the perceptual relations among samples. The generator learns to translate the in-between samples rather than memorizing the training set, and thereby forces the discriminator to generalize. The proposed approach effectively reduces the ambiguity of generation and renders content-preserving results. The ReMix method can be easily incorporated into existing GAN models with minor modifications. Experimental results on numerous tasks demonstrate that GAN models equipped with the ReMix method achieve significant improvements. Jie Cao 0002, Luanxuan Hou, Ming-Hsuan Yang 0001, Ran He 0001, Zhenan Sun |
CVPR | 5 |
| 2021 | One Shot Face Swapping on MegapixelsabstractFace swapping has both positive applications such as entertainment, human-computer interaction, etc., and negative applications such as DeepFake threats to politics, economics, etc. Nevertheless, it is necessary to understand the scheme of advanced methods for high-quality face swapping and generate enough and representative face swapping images to train DeepFake detection algorithms. This paper proposes the first Megapixel level method for one shot Face Swapping (or MegaFS for short). Firstly, MegaFS organizes face representation hierarchically by the proposed Hierarchical Representation Face Encoder (HieRFE) in an extended latent space to maintain more facial details, rather than compressed representation in previous face swapping methods. Secondly, a carefully designed Face Transfer Module (FTM) is proposed to transfer the identity from a source image to the target by a non-linear trajectory without explicit feature disentanglement. Finally, the swapped faces can be synthesized by StyleGAN2 with the benefits of its training stability and powerful generative capability. Each part of MegaFS can be trained separately so the requirement of our model for GPU memory can be satisfied for megapixel face swapping. In summary, complete face representation, stable training, and limited memory usage are the three novel contributions to the success of our method. Extensive experiments demonstrate the superiority of MegaFS and the first megapixel level face swapping database is released for research on DeepFake detection and face image editing in the public domain. Yuhao Zhu 0003, Qi Li 0005, Cheng-Zhong Xu 0001, Zhenan Sun |
CVPR | 5 |
| 2021 | A Large-scale Database for Less Cooperative Iris RecognitionabstractSince the outbreak of the COVID-19 pandemic, iris recognition has been used increasingly as contactless and unaffected by face masks. Although less user cooperation is an urgent demand for existing systems, corresponding manually annotated databases could hardly be obtained. This paper presents a large-scale database of near-infrared iris images named CASIA-Iris-Degradation Version 1.0 (DV1), which consists of 15 subsets of various degraded images, simulating less cooperative situations such as illumination, off-angle, occlusion, and nonideal eye state. A lot of open-source segmentation and recognition methods are compared comprehensively on the DV1 using multiple evaluations, and the best among them are exploited to conduct ablation studies on each subset. Experimental results show that even the best deep learning frameworks are not robust enough on the database, and further improvements are recommended for challenging factors such as half-open eyes, off-angle, and pupil dilation. Therefore, we publish the DV1 with manual annotations online to promote iris recognition. (http://www.cripacsir.cn/dataset/) Junxing Hu, Leyuan Wang, Zhengquan Luo, Yunlong Wang 0003, Zhenan Sun |
IJCB | 5 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 8 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 9 |
| 2021 | An End-to-End Autofocus Camera for Iris on the MoveabstractFor distant iris recognition, a long focal length lens is generally used to ensure the resolution of iris images, which reduces the depth of field and leads to potential defocus blur. To accommodate users standing statically at different distances, it is necessary to control focus quickly and accurately. And for users in motion, it is also expected to acquire a sufficient amount of accurately focused iris images. In this paper, we introduced a novel rapid auto-focus camera for active refocusing of the iris area of the moving objects with a focus-tunable lens. Our end-to-end computational algorithm can predict the best focus position from one single blurred image and generate the proper lens diopter control signal automatically. This scene-based active manipulation method enables real-time focus tracking of the iris area of a moving object. We built a testing bench to collect real-world focal stacks for evaluation of the autofocus methods. Our camera has reached an autofocus speed of over 50 fps. The results demonstrate the advantages of our proposed camera for biometric perception in static and dynamic scenes. The code is available at https://github.com/Debatrix/AquulaCam. Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun |
IJCB | 4 |
| 2021 | Contrastive Uncertainty Learning for Iris Recognition with Insufficient Labeled SamplesabstractCross-database recognition is still an unavoidable challenge when deploying an iris recognition system to a new environment. In the paper, we present a compromise problem that resembles the real-world scenario, named iris recognition with insufficient labeled samples. This new problem aims to improve the recognition performance by utilizing partially-or un-labeled data. To address the problem, we propose Contrastive Uncertainty Learning (CUL) by integrating the merits of uncertainty learning and contrastive self-supervised learning. CUL makes two efforts to learn a discriminative and robust feature representation. On the one hand, CUL explores the uncertain acquisition factors and adopts a probabilistic embedding to represent the iris image. In the probabilistic representation, the identity information and acquisition factors are disentangled into the mean and variance, avoiding the impact of uncertain acquisition factors on the identity information. On the other hand, CUL utilizes probabilistic embeddings to generate virtual positive and negative pairs. Then CUL builds its contrastive loss to group the similar samples closely and push the dissimilar samples apart. The experimental results demonstrate the effectiveness of the proposed CUL for iris recognition with insufficient labeled samples. Jianze Wei, Ran He 0001, Zhenan Sun |
IJCB | 3 |
| 2021 | PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopabstractRegression-based methods have recently shown promising results in reconstructing human meshes from monocular images. By directly mapping raw pixels to model parameters, these methods can produce parametric models in a feed-forward manner via neural networks. However, minor deviation in parameters may lead to noticeable mis-alignment between the estimated meshes and image evidences. To address this issue, we propose a Pyramidal Mesh Alignment Feedback (PyMAF) loop to leverage a feature pyramid and rectify the predicted parameters explicitly based on the mesh-image alignment status in our deep regressor. In PyMAF, given the currently predicted parameters, mesh-aligned evidences will be extracted from finer-resolution features accordingly and fed back for parameter rectification. To reduce noise and enhance the reliability of these evidences, an auxiliary pixel-wise supervision is imposed on the feature encoder, which provides mesh-image correspondence guidance for our network to preserve the most related information in spatial features. The efficacy of our approach is validated on several benchmarks, including Human3.6M, 3DPW, LSP, and COCO, where experimental results show that our approach consistently improves the mesh-image alignment of the reconstruction. The project page with code and video results can be found at https://hongwenzhang.github.io/pymaf. Hongwen Zhang 0001, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang 0002, Zhenan Sun |
ICCV | 7 |
| 2021 | Multi-caption Text-to-Face Synthesis: Dataset and AlgorithmabstractText-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm. Jianxin Sun 0003, Qi Li 0005, Weining Wang 0001, Jian Zhao 0006, Zhenan Sun |
ACM Multimedia | 5 |
| 2021 | Boosting End-to-end Multi-Object Tracking and Person Search via Knowledge DistillationabstractMulti-Object Tracking (MOT) and Person Search both demand to localize and identify specific targets from raw image frames. Existing methods can be classified into two categories, namely two-step strategy and end-to-end strategy. Two-step approaches have high accuracy but suffer from costly computations, while end-to-end methods show greater efficiency with limited performance. In this paper, we dissect the gap between two-step and end-to-end strategy and propose a simple yet effective end-to-end framework with knowledge distillation. Our proposed framework is simple in concept and easy to benefit from external datasets. Experimental results demonstrate that our model performs competitively with other sophisticated two-step and end-to-end methods in multi-object tracking and person search. Wei Zhang 0255, Lingxiao He, Xingyu Liao, Wu Liu 0005, Qi Li 0005, Zhenan Sun |
ACM Multimedia | 7 |
| 2021 | Deep Semantic Reconstruction Hashing for Similarity RetrievalabstractHashing has shown enormous potentials in preserving semantic similarity for large-scale data retrieval. Existing methods widely retain the similarity within two binary codes towards their discrete semantic affinity, i.e., 1 or -1. However, such a discrete reconstruction approach has obvious drawbacks. First, two unrelated dissimilar samples would have similar binary codes when both of them are the most dissimilar with an anchor sample. Second, the fine-grained semantic similarity cannot be shown in the generated binary codes among data with multiple semantic concepts. Furthermore, existing approaches generally adopt a point-wise error-minimizing strategy to enforce the real-valued codes close to its associated discrete codes, resulting in the well-learned paired semantic similarity being unintentionally damaged when performing quantization. To address these issues, we propose a novel deep hashing method with pairwise similarity-preserving quantization constraint, termed Deep Semantic Reconstruction Hashing (DSRH), which defines a high-level semantic affinity within each data pair to learn compact binary codes. Specifically, DSRH is expected to learn the specific binary codes whose similarity can reconstruct their high-level semantic similarity. Besides, we adopt a pairwise similarity-preserving quantization constraint instead of the traditional point-wise quantization technique, which is conducive to maintain the well-learned paired semantic similarity when performing quantization. Extensive experiments are conducted on four representative image retrieval benchmarks, and the proposed DSRH outperforms the state-of-the-art deep-learning methods with respect to different evaluation metrics. Yunbo Wang, Xianfeng Ou, Jian Liang 0001, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Controllable Multi-Attribute Editing of High-Resolution Face ImagesabstractIn recent years, significant progress has been achieved in face image editing due to the success of Generative Adversarial Network (GAN). However, state-of-the-art face editing methods mainly suffer from the following two limitations: 1) they are only applicable to face images with relative low-resolutions and 2) multi-attribute face editing may generate uncontrollable changes in non-target face attribute categories. To solve these problems, we propose a novel High-Quality Generative Adversarial Network (HQ-GAN) for controllable editing of multiple face attributes in high-resolution images. HQ-GAN has two novel ideas to break the limitations of resolution and controllability correspondingly: 1) fine-grained textures and realistic details of high-resolution face images are better preserved with the aid of textural features extracted by the wavelet transform module and 2) desired multi-attribute targets of face editing are emphasized using a weighted binary cross-entropy (BCE) loss so that the influence on non-target attributes is greatly reduced. To the best of our knowledge, HQ-GAN is the first attempt to achieve continuous editing of multiple face attributes on high-resolution images of the CelebA-HQ using only 28 000 training samples. Extensive qualitative results demonstrate the superiority of the proposed method in rendering realistic high-resolution face images with accurate attribute modification, and comprehensive quantitative results show that the proposed method significantly outperforms state-of-the-art face editing methods. Qiyao Deng, Qi Li 0005, Jie Cao 0002, Yunfan Liu 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | A3GAN: An Attribute-Aware Attentive Generative Adversarial Network for Face AgingabstractFace aging has received significant research attention in recent years. Although great progress has been achieved with the success of Generative Adversarial Networks (GANs) in synthesizing realistic images, most existing GAN-based face aging methods have two main problems: 1) unnatural changes of high-level semantic information due to the insufficient consideration of prior knowledge of input faces, and 2) distortions of low-level image content (e.g. modifications in age-irrelevant regions). In this article, we introduce A3GAN, an Attribute-Aware Attentive face aging model to address the above issues. Facial attribute vectors are regarded as the conditional information and embedded into both the generator and discriminator, encouraging synthesized faces to be faithful to attributes of corresponding inputs. To improve the visual fidelity of generation results, we leverage the attention mechanism to restrict modifications to age-related areas and preserve image details. Unlike previous works with attention modules, we introduce face parsing maps to help the generator distinguish image regions of interest and suppress attention activation elsewhere. Moreover, the wavelet packet transform is employed to capture textural features at multiple scales in the frequency space. Extensive experimental results demonstrate the effectiveness of our model in synthesizing photo-realistic aged face images and achieving state-of-the-art performance on popular datasets. Yunfan Liu 0001, Qi Li 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | CASIA-Face-Africa: A Large-Scale African Face Image DatabaseabstractFace recognition is a popular and well-studied area with wide applications in our society. However, racial bias had been proven to be inherent in most State Of The Art (SOTA) face recognition systems. Many investigative studies on face recognition algorithms have reported higher false positive rates of African subjects cohorts than the other cohorts. Lack of large-scale African face image databases in public domain is one of the main restrictions in studying the racial bias problem of face recognition. To this end, we collect a face image database namely CASIA-Face-Africa which contains 38,546 images of 1,183 African subjects. Multi-spectral cameras are utilized to capture the face images under various illumination settings. Demographic attributes and facial expressions of the subjects are also carefully recorded. For landmark detection, each face image in the database is manually labeled with 68 facial keypoints. A group of evaluation protocols are constructed according to different applications, tasks, partitions and scenarios. The performances of SOTA face recognition algorithms without re-training are reported as baselines. The proposed database along with its face landmark annotations, evaluation protocols and preliminary results form a good benchmark to study the essential aspects of face biometrics for African subjects, especially face image preprocessing, face feature analysis and matching, facial expression recognition, sex/age estimation, ethnic classification, face image generation, etc. The database can be downloaded from our website. Jawad Muhammad, Yunlong Wang 0003, Caiyong Wang, Kunbo Zhang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | New Joint-Drift-Free Scheme Aided with Projected ZNN for Motion Generation of Redundant Robot Manipulators Perturbed by DisturbancesabstractJoint-drift problems could result in failures in executing task or even damage robots in actual applications and different schemes have been presented to deal with such a knotty problem. However, in these existing schemes, there exists the coupling in coefficients for eliminating the drift in the joint space and the equality constraint for completing the given task in the Cartesian space, thereby, theoretically, leading to a paradox in achieving zero joint drift in the joint space and zero position error in the Cartesian space simultaneously. A novel joint-drift-free (JDF) scheme synthesized by a projected zeroing neural network (PZNN) model for the motion generation and control of redundant robot manipulators perturbed by disturbances is proposed and analyzed in this article. Besides, the PZNN model could adopt saturated or even nonconvex projection functions. The proposed scheme completely decouples the interferences of joint errors in the joint space and position errors in the Cartesian space for the first time. Beyond that, theoretical analysis is conducted in order to validate that the PZNN model is of global convergence to the theoretical kinematics solution to the motion generation of robots, and that the joint-drift problems are thus remedied. Moreover, several simulations and physical experiments on the strength of different robot manipulators are carried out to confirm the superiority, efficiency, and accuracy of the proposed JDF scheme synthesized by the PZNN model for remedying joint-drift problems of redundant robot manipulators in noisy environments. Huiyan Lu, Long Jin 0001, Jiliang Zhang 0001, Zhenan Sun, Shuai Li 0002, Zhijun Zhang 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | A Local Consensus Index Scheme for Random-Valued Impulse Noise Detection SystemsabstractThe issue of impulse noise detection and reduction is a critical problem for image processing application systems. In order to detect impulse noises in corrupted images, a statistic named local consensus index (LCI) is proposed for quantitatively evaluating how noise free a pixel is, and then an impulse noise detection scheme based on LCI is introduced. First, the similarity between arbitrary two pixels in an image is quantified based on both their geometric distance and intensity difference, and the LCI of arbitrary pixel is calculated by summing all the similarity values of pixels in its neighborhood. As a new statistic, the value of LCI indicates the local consensus of the concerned pixel regarding its neighbors and could also tell whether a pixel is noise free or impulsive. Therefore, LCI can be directly used as an efficient indicator of impulse noise. Furthermore, to improve the performance of impulse noise detection, different strategies are applied to the pixels at flat regions and the ones with complex textures, since distributions of LCI value within those regions are totally different. As for impulse noise filtering, a hybrid graph Laplacian regularization (HGLR) method is introduced to restore the intensities of those pixels degraded by impulse noise. We conduct extensive experiments to verify the effectiveness of our impulsive noise detection and reduction method, and the results show that the proposed method outperforms the state-of-the-art techniques in terms of impulse detection and noise removal. Xiuchun Xiao, Naixue Xiong, Jian-Huang Lai, Chang-Dong Wang 0001, Zhenan Sun |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | Age Progression and Regression with Spatial Attention ModulesabstractAge progression and regression refers to aesthetically rendering a given face image to present effects of face aging and rejuvenation, respectively. Although numerous studies have been conducted in this topic, there are two major problems: 1) multiple models are usually trained to simulate different age mappings, and 2) the photo-realism of generated face images is heavily influenced by the variation of training images in terms of pose, illumination, and background. To address these issues, in this paper, we propose a framework based on conditional Generative Adversarial Networks (cGANs) to achieve age progression and regression simultaneously. Particularly, since face aging and rejuvenation are largely different in terms of image translation patterns, we model these two processes using two separate generators, each dedicated to one age changing process. In addition, we exploit spatial attention mechanisms to limit image modifications to regions closely related to age changes, so that images with high visual fidelity could be synthesized for in-the-wild cases. Experiments on multiple datasets demonstrate the ability of our model in synthesizing lifelike face images at desired ages with personalized features well preserved, and keeping age-irrelevant regions unchanged. Qi Li 0005, Yunfan Liu 0001, Zhenan Sun |
AAAI | 3 |
| 2020 | Dynamic Graph Representation for Occlusion Handling in BiometricsabstractThe generalization ability of Convolutional neural networks (CNNs) for biometrics drops greatly due to the adverse effects of various occlusions. To this end, we propose a novel unified framework integrated the merits of both CNNs and graphical models to learn dynamic graph representations for occlusion problems in biometrics, called Dynamic Graph Representation (DGR). Convolutional features onto certain regions are re-crafted by a graph generator to establish the connections among the spatial parts of biometrics and build Feature Graphs based on these node representations. Each node of Feature Graphs corresponds to a specific part of the input image and the edges express the spatial relationships between parts. By analyzing the similarities between the nodes, the framework is able to adaptively remove the nodes representing the occluded parts. During dynamic graph matching, we propose a novel strategy to measure the distances of both nodes and adjacent matrixes. In this way, the proposed method is more convincing than CNNs-based methods because the dynamic graph method implies a more illustrative and reasonable inference of the biometrics decision. Experiments conducted on iris and face demonstrate the superiority of the proposed framework, which boosts the accuracy of occluded biometrics recognition by a large margin comparing with baseline methods. Yunlong Wang 0003, Zhenan Sun, Tieniu Tan |
AAAI | 3 |
| 2020 | Informative Sample Mining Network for Multi-domain Image-to-Image Translation
Jie Cao 0002, Huaibo Huang, Yi Li 0018, Ran He 0001, Zhenan Sun |
ECCV (19) | 5 |
| 2020 | Hierarchical Face Aging Through Disentangled Latent Characteristics
Peipei Li 0002, Huaibo Huang, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun |
ECCV (3) | 6 |
| 2020 | A Lightweight Multi-Label Segmentation Network for Mobile Iris BiometricsabstractThis paper proposes a novel, lightweight deep convolutional neural network specifically designed for iris segmentation of noisy images acquired by mobile devices. Unlike previous studies, which only focused on improving the accuracy of segmentation mask using the popular CNN technology, our method is a complete end-to-end iris segmentation solution, i.e., segmentation mask and parameterized pupillary and limbic boundaries of the iris are obtained simultaneously, which further enables CNN-based iris segmentation to be applied in any regular iris recognition systems. By introducing an intermediate pictorial boundary representation, predictions of iris boundaries and segmentation mask have collectively formed a multi-label semantic segmentation problem, which could be well solved by a carefully adapted stacked hourglass network. Experimental results show that our method achieves competitive or state-of-the-art performance in both iris segmentation and localization on two challenging mobile iris databases. Caiyong Wang, Yunlong Wang 0003, Boqiang Xu, Yong He 0009, Zhiwei Dong, Zhenan Sun |
ICASSP | 6 |
| 2020 | SSBC 2020: Sclera Segmentation Benchmarking Competition in the Mobile EnvironmentabstractThe paper presents a summary of the 2020 Sclera Segmentation Benchmarking Competition (SSBC), the 7th in the series of group benchmarking efforts centred around the problem of sclera segmentation. Different from previous editions, the goal of SSBC 2020 was to evaluate the performance of sclera-segmentation models on images captured with mobile devices. The competition was used as a platform to assess the sensitivity of existing models to i) differences in mobile devices used for image capture and ii) changes in the ambient acquisition conditions. 26 research groups registered for SSBC 2020, out of which 13 took part in the final round and submitted a total of 16 segmentation models for scoring. These included a wide variety of deep-learning solutions as well as one approach based on standard image processing techniques. Experiments were conducted with three recent datasets. Most of the segmentation models achieved relatively consistent performance across images captured with different mobile devices (with slight differences across devices), but struggled most with low-quality images captured in challenging ambient conditions, i.e., in an indoor environment and with poor lighting. Matej Vitek, Abhijit Das 0001, Yann Pourcenoux, Alexandre Missler, C. Paumier, Sumanta Das, Ishita De Ghosh, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Junxing Hu, Yong He 0009, Caiyong Wang, Yunlong Wang 0003, Zhenan Sun, Dailé Osorio Roig, Christian Rathgeb, Christoph Busch 0001, Juan E. Tapia, Andres Valenzuela, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, Sabari Nathan, R. Suganya 0001, Vineet Mehta, Abhinav Dhall, Kiran B. Raja, Gourav Gupta, Jalil Nourmohammadi-Khiarak, Mohsen Akbari-Shahper, Farhang Jaryani, Meysam Asgari-Chenaghlu, Ritesh Vyas, Sristi Dakshit, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 20 |
| 2020 | Recognition Oriented Iris Image Quality Assessment in the Feature SpaceabstractA large portion of iris images captured in real world scenarios are poor quality due to the uncontrolled environment and the non-cooperative subject. To ensure that the recognition algorithm is not affected by low-quality images, traditional hand-crafted factors based methods discard most images, which will cause system timeout and disrupt user experience. In this paper, we propose a recognition-oriented quality metric and assessment method for iris image to deal with the problem. The method regards the iris image em-beddings Distance in Feature Space (DFS) as the quality metric and the prediction is based on deep neural networks with the attention mechanism. The quality metric proposed in this paper can significantly improve the performance of the recognition algorithm while reducing the number of images discarded for recognition, which is advantageous over hand-crafted factors based iris quality assessment methods. The relationship between Image Rejection Rate (IRR) and Equal Error Rate (EER) is proposed to evaluate the performance of the quality assessment algorithm under the same image quality distribution and the same recognition algorithm. Compared with hand-crafted factors based methods, the proposed method is a trial to bridge the gap between the image quality assessment and biometric recognition. Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun |
IJCB | 5 |
| 2020 | All-in-Focus Iris Camera With a Great Capture VolumeabstractImaging volume of an iris recognition system has been restricting the throughput and cooperation convenience in biometric applications. Numerous improvement trials are still impractical to supersede the dominant fixed-focus lens in stand-off iris recognition due to incremental performance increase and complicated optical design. In this study, we develop a novel all-in-focus iris imaging system using a focus-tunable lens and a 2D steering mirror to greatly extend capture volume by spatiotemporal multiplexing method. Our iris imaging depth of field extension system requires no mechanical motion and is capable to adjust the focal plane at extremely high speed. In addition, the motorized reflection mirror adaptively steers the light beam to extend the horizontal and vertical field of views in an active manner. The proposed all-in-focus iris camera increases the depth of field up to 3.9 m which is afactor of 37.5 compared with conventional long focal lens. We also experimentally demonstrate the capability of this 3D light beam steering imaging system in real-time multi-person iris refocusing using dynamic focal stacks and the potential of continuous iris recognition for moving participants. Kunbo Zhang, Zhenteng Shen, Yunlong Wang 0003, Zhenan Sun |
IJCB | 4 |
| 2020 | A Novel Deep-learning Pipeline for Light Field Image Based Material RecognitionabstractThe primitive basis of image based material recognition builds upon the fact that discrepancies in the reflectances of distinct materials lead to imaging differences under multiple viewpoints. LF cameras possess coherent abilities to capture multiple sub-aperture views (SAIs) within one exposure, which can provide appropriate multi-view sources for material recognition. In this paper, a unified “Factorize-Connect-Merge” (FCM) deep-learning pipeline is proposed to solve problems of light field image based material recognition. 4D light-field data as input is initially decomposed into consecutive 3D light-field slices. Shallow CNN is leveraged to extract low-level visual features of each view inside these slices. As to establish correspondences between these SAIs, Bidirectional Long-Short Term Memory (Bi-LSTM) network is built upon these low-level features to model the imaging differences. After feature selection including concatenation and dimension reduction, effective and robust feature representations for material recognition can be extracted from 4D light-field data. Experimental results indicate that the proposed pipeline can obtain remarkable performances on both tasks of single-pixel material classification and full-image material segmentation. In addition, the proposed pipeline can potentially benefit and inspire other researchers who may also take LF images as input and need to extract 4D light-field representations for computer vision tasks such as object classification, semantic segmentation and edge detection. Yunlong Wang 0003, Kunbo Zhang, Zhenan Sun |
ICPR | 3 |
| 2020 | Reference Guided Face Component EditingabstractFace portrait editing has achieved great progress in recent years. However, previous methods either 1) operate on pre-defined face attributes, lacking the flexibility of controlling shapes of high-level semantic facial components (e.g., eyes, nose, mouth), or 2) take manually edited mask or sketch as an intermediate representation for observable changes, but such additional input usually requires extra efforts to obtain. To break the limitations (e.g. shape, mask or sketch) of the existing methods, we propose a novel framework termed r FACE (Reference Guided FAce Component Editing) for diverse and controllable face component editing with geometric changes. Specifically, r-FACE takes an image inpainting model as the backbone, utilizing reference images as conditions for controlling the shape of face components. In order to encourage the framework to concentrate on the target face components, an example-guided attention module is designed to fuse attention features and the target face component features extracted from the reference image. Through extensive experimental validation and comparisons, we verify the effectiveness of the proposed framework. Qiyao Deng, Jie Cao 0002, Yunfan Liu 0001, Zhenhua Chai, Qi Li 0005, Zhenan Sun |
IJCAI | 6 |
| 2020 | Dual-Structure Disentangling Variational Generation for Data-Limited Face ParsingabstractDeep learning based face parsing methods have attained state-of-the-art performance in recent years. Their superior performance heavily depends on the large-scale annotated training data. However, it is expensive and time-consuming to construct a large-scale pixel-level manually annotated dataset for face parsing. To alleviate this issue, we propose a novel Dual-Structure Disentangling Variational Generation (D2VG) network. Benefiting from the interpretable factorized latent disentanglement in VAE, D2VG can learn a joint structural distribution of facial image and its corresponding parsing map. Owing to these, it can synthesize large-scale paired face images and parsing maps from a standard Gaussian distribution. Then, we adopt both manually annotated and synthesized data to train a face parsing model in a supervised way. Since there are inaccurate pixel-level labels in synthesized parsing maps, we introduce a coarseness-tolerant learning algorithm, to effectively handle these noisy or uncertain labels. In this way, we can significantly boost the performance of face parsing. Extensive quantitative and qualitative results on HELEN, CelebAMask-HQ and LaPa demonstrate the superiority of our methods. Peipei Li 0002, Yinglu Liu, Hailin Shi, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun |
ACM Multimedia | 7 |
| 2020 | Black Re-ID: A Head-shoulder Descriptor for the Challenging Problem of Person Re-IdentificationabstractPerson re-identification (Re-ID) aims at retrieving an input person image from a set of images captured by multiple cameras. Although recent Re-ID methods have made great success, most of them extract features in terms of the attributes of clothing (e.g., color, texture). However, it is common for people to wear black clothes or be captured by surveillance systems in low light illumination, in which cases the attributes of the clothing are severely missing. We call this problem the Black Re-ID problem. To solve this problem, rather than relying on the clothing information, we propose to exploit head-shoulder features to assist person Re-ID. The head-shoulder adaptive attention network (HAA) is proposed to learn the head-shoulder feature and an innovative ensemble method is designed to enhance the generalization of our model. Given the input person image, the ensemble method would focus on the head-shoulder feature by assigning a larger weight if the individual insides the image is in black clothing. Due to the lack of a suitable benchmark dataset for studying the Black Re-ID problem, we also contribute the first Black-reID dataset, which contains 1274 identities in training set. Extensive evaluations on the Black-reID, Market1501 and DukeMTMC-reID datasets show that our model achieves the best result compared with the state-of-the-art Re-ID methods on both Black and conventional Re-ID problems. Furthermore, our method is also proved to be effective in dealing with person Re-ID in similar clothing. Our code and dataset are avaliable on https://github.com/xbq1994/. Boqiang Xu, Lingxiao He, Xingyu Liao, Wu Liu 0005, Zhenan Sun, Tao Mei 0001 |
ACM Multimedia | 5 |
| 2020 | Towards High Fidelity Face Frontalization in the Wild
Jie Cao 0002, Yibo Hu 0001, Hongwen Zhang 0001, Ran He 0001, Zhenan Sun |
Int. J. Comput. Vis. | 5 |
| 2020 | A General Framework for Deep Supervised Discrete Hashing
Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 2 |
| 2020 | Learning an Evolutionary Embedding via Massive Knowledge Distillation
Xiang Wu 0001, Ran He 0001, Yibo Hu 0001, Zhenan Sun |
Int. J. Comput. Vis. | 4 |
| 2020 | Adversarial Cross-Spectral Face Completion for NIR-VIS Face RecognitionabstractNear infrared-visible (NIR-VIS) heterogeneous face recognition refers to the process of matching NIR to VIS face images. Current heterogeneous methods try to extend VIS face recognition methods to the NIR spectrum by synthesizing VIS images from NIR images. However, due to the self-occlusion and sensing gap, NIR face images lose some visible lighting contents so that they are always incomplete compared to VIS face images. This paper models high-resolution heterogeneous face synthesis as a complementary combination of two components: a texture inpainting component and a pose correction component. The inpainting component synthesizes and inpaints VIS image textures from NIR image textures. The correction component maps any pose in NIR images to a frontal pose in VIS images, resulting in paired NIR and VIS textures. A warping procedure is developed to integrate the two components into an end-to-end deep network. A fine-grained discriminator and a wavelet-based discriminator are designed to improve visual quality. A novel 3D-based pose correction loss, two adversarial losses, and a pixel loss are imposed to ensure synthesis results. We demonstrate that by attaching the correction component, we can simplify heterogeneous face synthesis from one-to-many unpaired image translation to one-to-one paired image translation, and minimize the spectral and pose discrepancy during heterogeneous recognition. Extensive experimental results show that our network not only generates high-resolution VIS face images but also facilitates the accuracy improvement of heterogeneous face recognition. Ran He 0001, Jie Cao 0002, Lingxiao Song, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Learning disentangling and fusing networks for face completion under structured occlusions
Zhihang Li, Yibo Hu 0001, Ran He 0001, Zhenan Sun |
Pattern Recognit. | 4 |
| 2020 | Deep label refinement for age estimation
Peipei Li 0002, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun |
Pattern Recognit. | 5 |
| 2020 | Facial Age Synthesis With Label Distribution-Guided Generative Adversarial NetworkabstractThe existing research work on facial age synthesis has been mostly focused on long-term aging (e.g., over an age span of 10 years or more). In this paper, we employ generative adversarial networks (GANs) as a tool to investigate age synthesis over different age spans. Compared with long-term aging, short-term age synthesis suffers from the reduced amount of available training data, which can severely hinder the model training. We conduct a series of experiments to validate this. To facilitate short-term age synthesis, we further propose label distribution-guided generative adversarial network (ldGAN), where each sample is associated with an age label distribution (ALD) rather than a single age group. Accordingly, each sample can contribute not only to the learning of its own age group but also to neighbouring groups' learning. This is useful when addressing short-term aging to cope with the reduced amount of training data. In addition, unlike one-hot encoding which treats age groups as independent from one another, ldGAN can well capture the correlation among different age groups, so that smooth aging sequences can be achieved. The ALD model is integrated into GAN with a two-step process. Firstly, instead of the traditional one-hot encoding, ALD is applied as the condition of the generator. Secondly, we add a sequence of label distribution learners on top of several multi-scale discriminators, with the aim of minimizing the label distribution learning loss when optimizing both the generator and discriminators. Both qualitative and quantitative evaluations are conducted to assess ldGAN's ability in dealing with two core issues of face aging, i.e., aging effect generation and identity preservation. The obtained experimental results demonstrate the effectiveness of ldGAN in both learning short-term aging patterns and coping with the lack of training data. Yunlian Sun, Jinhui Tang 0001, Xiangbo Shu, Zhenan Sun, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Facial Age and Expression Synthesis Using Ordinal Ranking Adversarial NetworksabstractFacial image synthesis has been extensively studied, for a long time, in both computer graphics and computer vision. Particularly, the synthesis of face images with varying ages, expressions and poses has received an increasing attention owing to several real-world applications. In this paper, facial age and expression synthesis are addressed. While previous and current research papers on facial age synthesis mostly adopt an age span of 10 years, this paper investigates face aging with a shorter time span. For expression synthesis, given a neutral face, we work on synthesizing faces with varying expression intensities (e.g., from zero to high). Note that both human ages and expression intensities are inherently ordinal. To fully exploit this ordinal nature, we devise ordinal ranking generative adversarial networks (ranking GAN). For each face, a one-hot label is assigned to define its age range/expression intensity. By exploiting the relative order information among age ranges/expression intensities, a binary ranking vector is further computed for each face. In ranking GAN, one-hot labels are used as the condition of the generator for synthesizing faces with target age groups/expression intensities. Moreover, we add a sequence of cost-sensitive ordinal rankers on top of several multi-scale discriminators, with the aim of minimizing age/intensity rank estimation loss when optimizing both the generator and discriminators. In order to evaluate the proposed ranking GAN, extensive experiments are carried out on several public face databases. As demonstrated by the experimental testing, this ranking scheme performs well even when the amount of available labeled training data is limited. The reported experimental results well demonstrate the effectiveness of ranking GAN on synthesizing face aging sequences and faces with varying expression intensities. Yunlian Sun, Jinhui Tang 0001, Zhenan Sun, Massimo Tistarelli |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Towards Complete and Accurate Iris Segmentation Using Deep Multi-Task Attention Network for Non-Cooperative Iris RecognitionabstractIris images captured in non-cooperative environments often suffer from adverse noise, which challenges many existing iris segmentation methods. To address this problem, this paper proposes a high-efficiency deep learning based iris segmentation approach, named IrisParseNet. Different from many previous CNN-based iris segmentation methods, which only focus on predicting accurate iris masks by following popular semantic segmentation frameworks, the proposed approach is a complete iris segmentation solution, i.e., iris mask and parameterized inner and outer iris boundaries are jointly achieved by actively modeling them into a unified multi-task network. Moreover, an elaborately designed attention module is incorporated into it to improve the segmentation performance. To train and evaluate the proposed approach, we manually label three representative and challenging iris databases, i.e., CASIA.v4-distance, UBIRIS.v2, and MICHE-I, which involve multiple illumination (NIR, VIS) and imaging sensors (long-range and mobile iris cameras), along with various types of noises. Additionally, several unified evaluation protocols are built for fair comparisons. Extensive experiments are conducted on these newly annotated databases, and results show that the proposed approach achieves state-of-the-art performance on various benchmarks. Further, as a general drop-in replacement, the proposed iris segmentation method can be used for any iris recognition methodology, and would significantly improve the performance of non-cooperative iris recognition. Caiyong Wang, Jawad Muhammad, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Binocular Light-Field: Imaging Theory and Occlusion-Robust Depth Perception ApplicationabstractBinocular stereo vision (SV) has been widely used to reconstruct the depth information, but it is quite vulnerable to scenes with strong occlusions. As an emerging computational photography technology, light-field (LF) imaging brings about a novel solution to passive depth perception by recording multiple angular views in a single exposure. In this paper, we explore binocular SV and LF imaging to form the binocular-LF imaging system. An imaging theory is derived by modeling the imaging process and analyzing disparity properties based on the geometrical optics theory. Then an accurate occlusion-robust depth estimation algorithm is proposed by exploiting multibaseline stereo matching cues and defocus cues. The occlusions caused by binocular SV and LF imaging are detected and handled to eliminate the matching ambiguities and outliers. Finally, we develop a binocular-LF database and capture realworld scenes by our binocular-LF system to test the accuracy and robustness. The experimental results demonstrate that the proposed algorithm definitely recovers high quality depth maps with smooth surfaces and precise geometric shapes, which tackles the drawbacks of binocular SV and LF imaging simultaneously. Fei Liu 0031, Shubo Zhou, Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 5 |
| 2019 | Disentangled Variational Representation for Heterogeneous Face RecognitionabstractVisible (VIS) to near infrared (NIR) face matching is a challenging problem due to the significant domain discrepancy between the domains and a lack of sufficient data for training cross-modal matching algorithms. Existing approaches attempt to tackle this problem by either synthesizing visible faces from NIR faces, extracting domain-invariant features from these modalities, or projecting heterogeneous data onto a common latent space for cross-modal matching. In this paper, we take a different approach in which we make use of the Disentangled Variational Representation (DVR) for crossmodal matching. First, we model a face representation with an intrinsic identity information and its within-person variations. By exploring the disentangled latent variable space, a variational lower bound is employed to optimize the approximate posterior for NIR and VIS representations. Second, aiming at obtaining more compact and discriminative disentangled latent space, we impose a minimization of the identity information for the same subject and a relaxed correlation alignment constraint between the NIR and VIS modality variations. An alternative optimization scheme is proposed for the disentangled variational representation part and the heterogeneous face recognition network part. The mutual promotion between these two parts effectively reduces the NIR and VIS domain discrepancy and alleviates over-fitting. Extensive experiments on three challenging NIR-VIS heterogeneous face recognition databases demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods. Xiang Wu 0001, Huaibo Huang, Vishal M. Patel, Ran He 0001, Zhenan Sun |
AAAI | 5 |
| 2019 | Distant Supervised Centroid Shift: A Simple and Efficient Approach to Visual Domain AdaptationabstractConventional domain adaptation methods usually resort to deep neural networks or subspace learning to find invariant representations across domains. However, most deep learning methods highly rely on large-size source domains and are computationally expensive to train, while subspace learning methods always have a quadratic time complexity that suffers from the large domain size. This paper provides a simple and efficient solution, which could be regarded as a well-performing baseline for domain adaptation tasks. Our method is built upon the nearest centroid classifier, seeking a subspace where the centroids in the target domain are moderately shifted from those in the source domain. Specifically, we design a unified objective without accessing the source domain data and adopt an alternating minimization scheme to iteratively discover the pseudo target labels, invariant subspace, and target centroids. Besides its privacy-preserving property (distant supervision), the algorithm is provably convergent and has a promising linear time complexity. In addition, the proposed method can be readily extended to multi-source setting and domain generalization, and it remarkably enhances popular deep adaptation methods by borrowing the learned transferable features. Extensive experiments on several benchmarks including object, digit, and face recognition datasets validate that our methods yield state-of-the-art results in various domain adaptation tasks. Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
CVPR | 3 |
| 2019 | Attribute-Aware Face Aging With Wavelet-Based Generative Adversarial NetworksabstractSince it is difficult to collect face images of the same subject over a long range of age span, most existing face aging methods resort to unpaired datasets to learn age mappings. However, the matching ambiguity between young and aged face images inherent to unpaired training data may lead to unnatural changes of facial attributes during the aging process, which could not be solved by only enforcing identity consistency like most existing studies do. In this paper, we propose an attribute-aware face aging model with wavelet based Generative Adversarial Networks (GANs) to address the above issues. To be specific, we embed facial attribute vectors into both the generator and discriminator of the model to encourage each synthesized elderly face image to be faithful to the attribute of its corresponding input. In addition, a wavelet packet transform (WPT) module is incorporated to improve the visual fidelity of generated images by capturing age-related texture details at multiple scales in the frequency space. Qualitative results demonstrate the ability of our model in synthesizing visually plausible face images, and extensive quantitative evaluation results show that the proposed method achieves state-of-the-art performance on existing datasets. Yunfan Liu 0001, Qi Li 0005, Zhenan Sun |
CVPR | 3 |
| 2019 | Foreground-Aware Pyramid Reconstruction for Alignment-Free Occluded Person Re-IdentificationabstractRe-identifying a person across multiple disjoint camera views is important for intelligent video surveillance, smart retailing and many other applications. However, existing person re-identification methods are challenged by the ubiquitous occlusion over persons and suffer performance degradation. This paper proposes a novel occlusion-robust and alignment-free model for occluded person ReID and extends its application to realistic and crowded scenarios. The proposed model first leverages the fully convolution network (FCN) and pyramid pooling to extract spatial pyramid features. Then an alignment-free matching approach namely Foreground-aware Pyramid Reconstruction (FPR) is developed to accurately compute matching scores between occluded persons, regardless of their different scales and sizes. FPR uses the error from robust reconstruction over spatial pyramid features to measure similarities between two persons. More importantly, we design a occlusion-sensitive foreground probability generator that focuses more on clean human body parts to robustify the similarity computation with less contamination from occlusion. The FPR is easily embedded into any end-to-end person ReID models. The effectiveness of the proposed method is clearly demonstrated by the experimental results (Rank-1 accuracy) on three occluded person datasets: Partial REID (78.30%), Partial iLIDS (68.08%), Occluded REID (81.00%), and three benchmark person datasets: Market1501 (95.42%), DukeMTMC (88.64%), CUHK03 (76.08%). Lingxiao He, Yinggang Wang, Wu Liu 0005, Zhenan Sun, Jiashi Feng |
ICCV | 5 |
| 2019 | M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose AnalysisabstractFacial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-scale Multi-yaw Multi-pitch high-quality database is proposed for Facial Pose Analysis (M2FPA), including face frontalization, face rotation, facial pose estimation and pose-invariant face recognition. It contains 397,544 images of 229 subjects with yaw, pitch, attribute, illumination and accessory. M2FPA is the most comprehensive multi-view face database for facial pose analysis. Further, we provide an effective benchmark for face frontalization and pose-invariant face recognition on M2FPA with several state-of-the-art methods, including DR-GAN, TP-GAN and CAPG-GAN. We believe that the new database and benchmark can significantly push forward the advance of facial pose analysis in real-world applications. Moreover, a simple yet effective parsing guided discriminator is introduced to capture the local consistency during GAN optimization. Extensive quantitative and qualitative results on M2FPA and Multi-PIE demonstrate the superiority of our face frontalization method. Baseline results for both face synthesis and face recognition from state-of-the-art methods demonstrate the challenge offered by this new database. Peipei Li 0002, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun |
ICCV | 5 |
| 2019 | Towards Joint Multiply Semantics Hashing for Visual Search
Yunbo Wang, Zhenan Sun |
ICIG (3) | 2 |
| 2019 | DaNet: Decompose-and-aggregate Network for 3D Human Shape and Pose EstimationabstractReconstructing 3D human shape and pose from a monocular image is challenging despite the promising results achieved by most recent learning based methods. The commonly occurred misalignment comes from the facts that the mapping from image to model space is highly non-linear and the rotation-based pose representation of the body model is prone to result in drift of joint positions. In this work, we present the Decompose-and-aggregate Network (DaNet) to address these issues. DaNet includes three new designs, namely UVI guided learning, decomposition for fine-grained perception, and aggregation for robust prediction. First, we adopt the UVI maps, which densely build a bridge between 2D pixels and 3D vertexes, as an intermediate representation to facilitate the learning of image-to-model mapping. Second, we decompose the prediction task into one global stream and multiple local streams so that the network not only provides global perception for the camera and shape prediction, but also has detailed perception for part pose prediction. Lastly, we aggregate the message from local streams to enhance the robustness of part pose prediction, where a position-aided rotation feature refinement strategy is proposed to exploit the spatial relationship between body parts. Such a refinement strategy is more efficient since the correlations between position features are stronger than that in the original rotation feature space. The effectiveness of our method is validated on the Human3.6M and UP-3D datasets. Experimental results show that the proposed method significantly improves the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://github.com/HongwenZhang/DaNet-3DHumanReconstrution . Hongwen Zhang 0001, Jie Cao 0002, Guo Lu, Wanli Ouyang, Zhenan Sun |
ACM Multimedia | 5 |
| 2019 | Wavelet Domain Generative Adversarial Network for Multi-scale Face Hallucination
Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 3 |
| 2019 | Toward practical remote iris recognition: A boosting based framework
Man Zhang 0005, Zhaofeng He 0001, Hui Zhang 0061, Tieniu Tan, Zhenan Sun |
Neurocomputing | 5 |
| 2019 | Wasserstein CNN: Learning Invariant Features for NIR-VIS Face RecognitionabstractHeterogeneous face recognition (HFR) aims at matching facial images acquired from different sensing modalities with mission-critical applications in forensics, security and commercial sectors. However, HFR presents more challenging issues than traditional face recognition because of the large intra-class variation among heterogeneous face images and the limited availability of training samples of cross-modality face image pairs. This paper proposes the novel Wasserstein convolutional neural network (WCNN) approach for learning invariant features between near-infrared (NIR) and visual (VIS) face images (i.e., NIR-VIS face recognition). The low-level layers of the WCNN are trained with widely available face images in the VIS spectrum, and the high-level layer is divided into three parts: the NIR layer, the VIS layer and the NIR-VIS shared layer. The first two layers aim at learning modality-specific features, and the NIR-VIS shared layer is designed to learn a modality-invariant feature subspace. The Wasserstein distance is introduced into the NIR-VIS shared layer to measure the dissimilarity between heterogeneous feature distributions. W-CNN learning is performed to minimize the Wasserstein distance between the NIR distribution and the VIS distribution for invariant deep feature representations of heterogeneous face images. To avoid the over-fitting problem on small-scale heterogeneous face data, a correlation prior is introduced on the fully-connected WCNN layers to reduce the size of the parameter space. This prior is implemented by a low-rank constraint in an end-to-end network. The joint formulation leads to an alternating minimization for deep feature representation at the training stage and an efficient computation for heterogeneous data at the testing stage. Extensive experiments using three challenging NIR-VIS face recognition databases demonstrate the superiority of the WCNN method over state-of-the-art methods. Ran He 0001, Xiang Wu 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Aggregating Randomized Clustering-Promoting Invariant Projections for Domain AdaptationabstractUnsupervised domain adaptation aims to leverage the labeled source data to learn with the unlabeled target data. Previous trandusctive methods tackle it by iteratively seeking a low-dimensional projection to extract the invariant features and obtaining the pseudo target labels via building a classifier on source data. However, they merely concentrate on minimizing the cross-domain distribution divergence, while ignoring the intra-domain structure especially for the target domain. Even after projection, possible risk factors like imbalanced data distribution may still hinder the performance of target label inference. In this paper, we propose a simple yet effective domain-invariant projection ensemble approach to tackle these two issues together. Specifically, we seek the optimal projection via a novel relaxed domain-irrelevant clustering-promoting term that jointly bridges the cross-domain semantic gap and increases the intra-class compactness in both domains. To further enhance the target label inference, we first develop a 'sampling-and-fusion' framework, under which multiple projections are independently learned based on various randomized coupled domain subsets. Subsequently, aggregating models such as majority voting are utilized to leverage multiple projections and classify unlabeled target data. Extensive experimental results on six visual benchmarks including object, face, and digit images, demonstrate that the proposed methods gain remarkable margins over state-of-the-art unsupervised domain adaptation methods. Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Exploring uncertainty in pseudo-label guided unsupervised domain adaptation
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
Pattern Recognit. | 3 |
| 2019 | 3D Aided Duet GANs for Multi-View Face Image SynthesisabstractMulti-view face synthesis from a single image is an ill-posed computer vision problem. It often suffers from appearance distortions if it is not well-defined. Producing photo-realistic and identity preserving multi-view results is still a not well-defined synthesis problem. This paper proposes 3D aided duet generative adversarial networks (AD-GAN) to precisely rotate the yaw angle of an input face image to any specified angle. AD-GAN decomposes the challenging synthesis problem into two well-constrained subtasks that correspond to a face normalizer and a face editor. The normalizer first frontalizes an input image, and then the editor rotates the frontalized image to a desired pose guided by a remote code. In the meantime, the face normalizer is designed to estimate a novel dense UV correspondence field, making our model aware of 3D face geometry information. In order to generate photo-realistic local details and accelerate convergence process, the normalizer and the editor are trained in a two-stage manner and regulated by a conditional self-cycle loss and a perceptual loss. Exhaustive experiments on both controlled and uncontrolled environments demonstrate that the proposed method not only improves the visual realism of multi-view synthetic images but also preserves identity information well. Jie Cao 0002, Yibo Hu 0001, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2019 | Global and Local Consistent Wavelet-Domain Age SynthesisabstractAge synthesis is a challenging task due to the complicated and non-linear transformation in the human aging process. Aging information is usually reflected in local facial parts, such as wrinkles at the eye corners. However, these local facial parts contribute less in previous GAN-based methods for age synthesis. To address this issue, we propose a wavelet-domain global and local consistent age generative adversarial network (WaveletGLCA-GAN), in which one global specific network and three local specific networks are integrated together to capture both global topology information and local texture details of human faces. Different from the most existing methods that modeling age synthesis in image domain, we adopt wavelet transform to depict the textual information in frequency domain. Moreover, five types of losses are adopted: 1) adversarial loss aims to generate realistic wavelets; 2) identity preserving loss aims to better preserve identity information; 3) age preserving loss aims to enhance the accuracy of age synthesis; 4) pixel-wise loss aims to preserve the background information of the input face; and 5) the total variation regularization aims to remove ghosting artifacts. Our method is evaluated on three face aging datasets, including CACD2000, Morph, and FG-NET. Qualitative and quantitative experiments show the superiority of the proposed method over other state-of-the-arts. Peipei Li 0002, Yibo Hu 0001, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Dynamic Feature Matching for Partial Face RecognitionabstractPartial face recognition (PFR) in an unconstrained environment is a very important task, especially in situations where partial face images are likely to be captured due to occlusions, out-of-view, and large viewing angle, e.g., video surveillance and mobile devices. However, little attention has been paid to PFR so far and thus, the problem of recognizing an arbitrary patch of a face image remains largely unsolved. This study proposes a novel partial face recognition approach, called Dynamic Feature Matching (DFM), which combines Fully Convolutional Networks (FCNs) and Sparse Representation Classification (SRC) to address partial face recognition problem regardless of various face sizes. DFM does not require prior position information of partial faces against a holistic face. By sharing computation, the feature maps are calculated from the entire input image once, which yields a significant speedup. Experimental results demonstrate the effectiveness and advantages of DFM in comparison with state-of-the-art PFR methods on several partial face databases, including CAISA-NIR-Distance, CASIA-NIR-Mobile, and LFW databases. The performance of DFM is also impressive in partial person re-identification on Partial RE-ID and iLIDS databases. The source code of DFM can be found at https://github.com/lingxiao-he/dfm new. Lingxiao He, Qi Zhang 0015, Zhenan Sun |
IEEE Trans. Image Process. | 4 |
| 2019 | Local Semantic-Aware Deep Hashing With Hamming-Isometric QuantizationabstractHashing has attracted increasing attention due to its tremendous potential for efficient image retrieval and data storage. Compared with conventional hashing methods with a handcrafted feature, emerging deep hashing approaches employ deep neural networks to learn feature representations as well as hash functions, which have already been proved to be more powerful and robust in real-world applications. Currently, most of the existing deep hashing methods construct pairwise or triplet-wise constraint to obtain similar binary codes between similar data pair or relative similar binary codes within a triplet. However, some critical local structures of the data are lack of exploiting, thus the effectiveness of hash learning is not fully shown. To address this limitation, we propose a novel deep hashing method named local semantic-aware deep hashing with Hamming-isometric quantization (LSDH), where local similarity of the data is intentionally integrated into hash learning. Specifically, in the Hamming space, we exploit the potential semantic relation of the data to robustly preserve their local similarity. In addition to reducing the error introduced by binary quantizing, we further develop a Hamming-isometric objective to maximize the consistency of similarity between the pairwise binary-like feature and its binary codes pair, which is shown to be able to enhance the quality of binary codes. Extensive experimental results on several benchmark datasets, including three singlelabel datasets (i.e., CIFAR-10, CIFAR-20, and SUN397) and one multi-label dataset (NUS-WIDE), demonstrate that the proposed LSDH achieves superior performance over the latest state-of-theart hashing methods. Yunbo Wang, Jian Liang 0001, Zhenan Sun |
IEEE Trans. Image Process. | 4 |
| 2019 | Adversarial Learning Semantic Volume for 2D/3D Face Shape Regression in the WildabstractRegression based methods have revolutionized 2D landmark localization with the exploitation of deep neural networks and massive annotated datasets in the wild. However, it remains challenging for 3D landmark localization due to the lack of annotated datasets and the ambiguous nature of landmarks under 3D perspective. This paper revisits regression based methods and proposes an adversarial voxel and coordinate regression framework for 2D and 3D facial landmark localization in real-world scenarios. First, a semantic volumetric representation is introduced to encode the per-voxel likelihood of positions being the 3D landmarks. Then, an end-to-end pipeline is designed to jointly regress the proposed volumetric representation and the coordinate vector. Such a pipeline not only enhances the robustness and accuracy of the predictions but also unifies the 2D and 3D landmark localization so that 2D and 3D datasets could be utilized simultaneously. Further, an adversarial learning strategy is exploited to distill 3D structure learned from synthetic datasets to real-world datasets under weakly supervised settings, where an auxiliary regression discriminator is proposed to encourage the network to produce plausible predictions for both synthetic and real-world images. The effectiveness of our method is validated on benchmark datasets 3DFAW and AFLW2000-3D for both 2D and 3D facial landmark localization tasks. Experimental results show that the proposed method achieves significant improvements over previous state-of-the-art methods. Hongwen Zhang 0001, Qi Li 0005, Zhenan Sun |
IEEE Trans. Image Process. | 3 |
| 2018 | Dynamic Feature Learning for Partial Face RecognitionabstractPartial face recognition (PFR) in unconstrained environment is a very important task, especially in video surveillance, mobile devices, etc. However, a few studies have tackled how to recognize an arbitrary patch of a face image. This study combines Fully Convolutional Network (FCN) with Sparse Representation Classification (SRC) to propose a novel partial face recognition approach, called Dynamic Feature Matching (DFM), to address partial face images regardless of size. Based on DFM, we propose a sliding loss to optimize FCN by reducing the intra-variation between a face patch and face images of a subject, which further improves the performance of DFM. The proposed DFM is evaluated on several partial face databases, including LFW, YTF and CASIA-NIR-Distance databases. Experimental results demonstrate the effectiveness and advantages of DFM in comparison with state-of-the-art PFR methods. Lingxiao He, Qi Zhang 0015, Zhenan Sun |
CVPR | 4 |
| 2018 | Deep Spatial Feature Reconstruction for Partial Person Re-Identification: Alignment-Free ApproachabstractPartial person re-identification (re-id) is a challenging problem, where only several partial observations (images) of people are available for matching. However, few studies have provided flexible solutions to identifying a person in an image containing arbitrary part of the body. In this paper, we propose a fast and accurate matching method to address this problem. The proposed method leverages Fully Convolutional Network (FCN) to generate fix-sized spatial feature maps such that pixel-level features are consistent. To match a pair of person images of different sizes, a novel method called Deep Spatial feature Reconstruction (DSR) is further developed to avoid explicit alignment. Specifically, DSR exploits the reconstructing error from popular dictionary learning models to calculate the similarity between different spatial feature maps. In that way, we expect that the proposed FCN can decrease the similarity of coupled images from different persons and increase that from the same person. Experimental results on two partial person datasets demonstrate the efficiency and effectiveness of the proposed method in comparison with several state-of-the-art partial person re-id approaches. Additionally, DSR achieves competitive results on a benchmark person dataset Market1501 with 83.58% Rank-1 accuracy. Lingxiao He, Jian Liang 0001, Zhenan Sun |
CVPR | 4 |
| 2018 | Pose-Guided Photorealistic Face RotationabstractFace rotation provides an effective and cheap way for data augmentation and representation learning of face recognition. It is a challenging generative learning problem due to the large pose discrepancy between two face images. This work focuses on flexible face rotation of arbitrary head poses, including extreme profile views. We propose a novel Couple-Agent Pose-Guided Generative Adversarial Network (CAPG-GAN) to generate both neutral and profile head pose face images. The head pose information is encoded by facial landmark heatmaps. It not only forms a mask image to guide the generator in learning process but also provides a flexible controllable condition during inference. A couple-agent discriminator is introduced to reinforce on the realism of synthetic arbitrary view faces. Besides the generator and conditional adversarial loss, CAPG-GAN further employs identity preserving loss and total variation regularization to preserve identity information and refine local textures respectively. Quantitative and qualitative experimental results on the Multi-PIE and LFW databases consistently show the superiority of our face rotation method over the state-of-the-art. Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun |
CVPR | 5 |
| 2018 | End-to-End View Synthesis for Light Field Imaging with Pseudo 4DCNN
Yunlong Wang 0003, Fei Liu 0031, Zilei Wang, Guangqi Hou, Zhenan Sun, Tieniu Tan |
ECCV (2) | 5 |
| 2018 | Global and Local Consistent Age Generative Adversarial NetworksabstractAge progression/regression is a challenging task due to the complicated and non-linear transformation in human aging process. Many researches have shown that both global and local facial features are essential for face representation [1], but previous GAN based methods mainly focused on the global feature in age synthesis. To utilize both global and local facial information, we propose a Global and Local Consistent Age Generative Adversarial Network (GLCA-GAN). In our generator, a global network learns the whole facial structure and simulates the aging trend of the whole face, while three crucial facial patches are progressed or regressed by three local networks aiming at imitating subtle changes of crucial facial subregions. To preserve most of the details in age-attribute-irrelevant areas, our generator learns the residual face. Moreover, we employ an identity preserving loss to better preserve the identity information, as well as age preserving loss to enhance the accuracy of age synthesis. A pixel loss is also adopted to preserve detailed facial information of the input face. Our proposed method is evaluated on three face aging datasets, i.e., CACD dataset, Morph dataset and FG-NET dataset. Experimental results show appealing performance of the proposed method by comparing with the state-of-the-art. Peipei Li 0002, Yibo Hu 0001, Qi Li 0005, Ran He 0001, Zhenan Sun |
ICPR | 5 |
| 2018 | Joint Voxel and Coordinate Regression for Accurate 3D Facial Landmark Localizationabstract3D face shape is more expressive and viewpoint-consistent than its 2D counterpart. However, 3D facial landmark localization in a single image is challenging due to the ambiguous nature of landmarks under 3D perspective. Existing approaches typically adopt a suboptimal two-step strategy, performing 2D landmark localization followed by depth estimation. In this paper, we propose the Joint Voxel and Coordinate Regression (JVCR) method for 3D facial landmark localization, addressing it more effectively in an end-to-end fashion. First, a compact volumetric representation is proposed to encode the per-voxel likelihood of positions being the 3D landmarks. The dimensionality of such a representation is fixed regardless of the number of target landmarks, so that the curse of dimensionality could be avoided. Then, a stacked hourglass network is adopted to estimate the volumetric representation from coarse to fine, followed by a 3D convolution network that takes the estimated volume as input and regresses 3D coordinates of the face shape. In this way, the 3D structural constraints between landmarks could be learned by the neural network in a more efficient manner. Moreover, the proposed pipeline enables end-to-end training and improves the robustness and accuracy of 3D facial landmark localization. The effectiveness of our approach is validated on the 3DFAW and AFLW2000-3D datasets. Experimental results show that the proposed method achieves state-of-the-art performance in comparison with existing methods. Hongwen Zhang 0001, Qi Li 0005, Zhenan Sun |
ICPR | 3 |
| 2018 | Geometry Guided Adversarial Facial Expression SynthesisabstractFacial expression synthesis has drawn much attention in the field of computer graphics and pattern recognition. It has been widely used in face animation and recognition. However, it is still challenging due to the high-level semantic presence of large and non-linear face geometry variations. This paper proposes a Geometry-Guided Generative Adversarial Network (G2-GAN) for continuously-adjusting and identity-preserving facial expression synthesis. We employ facial geometry (fiducial points) as a controllable condition to guide facial texture synthesis with specific expression. A pair of generative adversarial subnetworks is jointly trained towards opposite tasks: expression removal and expression synthesis. The paired networks form a mapping cycle between neutral expression and arbitrary expressions, with which the proposed approach can be conducted among unpaired data. The proposed paired networks also facilitate other applications such as face transfer, expression interpolation and expression-invariant face recognition. Experimental results on several facial expression databases show that our method can generate compelling perceptual results on different expression editing tasks. Lingxiao Song, Zhihe Lu, Ran He 0001, Zhenan Sun, Tieniu Tan |
ACM Multimedia | 4 |
| 2018 | Learning a High Fidelity Pose Invariant Model for High-resolution Face FrontalizationabstractFace frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a High Fidelity Pose Invariant Model (HF-PIM) to produce photographic and identity-preserving results. HF-PIM frontalizes the profiles through a novel texture warping procedure and leverages a dense correspondence field to bind the 2D and 3D surface spaces. We decompose the prerequisite of warping into dense correspondence field estimation and facial texture map recovering, which are both well addressed by deep networks. Different from those reconstruction methods relying on 3D data, we also propose Adversarial Residual Dictionary Learning (ARDL) to supervise facial texture map recovering with only monocular images. Exhaustive experiments on both controlled and uncontrolled environments demonstrate that the proposed method not only boosts the performance of pose-invariant face recognition but also dramatically improves high-resolution frontalization appearances. Jie Cao 0002, Yibo Hu 0001, Hongwen Zhang 0001, Ran He 0001, Zhenan Sun |
NeurIPS | 5 |
| 2018 | IntroVAE: Introspective Variational Autoencoders for Photographic Image SynthesisabstractWe present a novel introspective variational autoencoder (IntroVAE) model for synthesizing high-resolution photographic images. IntroVAE is capable of self-evaluating the quality of its generated samples and improving itself accordingly. Its inference and generator models are jointly trained in an introspective way. On one hand, the generator is required to reconstruct the input images from the noisy outputs of the inference model as normal VAEs. On the other hand, the inference model is encouraged to classify between the generated and real samples while the generator tries to fool it as GANs. These two famous generative frameworks are integrated in a simple yet efficient single-stream architecture that can be trained in a single stage. IntroVAE preserves the advantages of VAEs, such as stable training and nice latent manifold. Unlike most other hybrid models of VAEs and GANs, IntroVAE requires no extra discriminators, because the inference model itself serves as a discriminator to distinguish between the generated and real samples. Experiments demonstrate that our method produces high-resolution photo-realistic images (e.g., CELEBA images at (1024^{2})), which are comparable to or better than the state-of-the-art GANs. Huaibo Huang, Zhihang Li, Ran He 0001, Zhenan Sun, Tieniu Tan |
NeurIPS | 4 |
| 2018 | Fast Supervised Discrete HashingabstractLearning-based hashing algorithms are "hot topics" because they can greatly increase the scale at which existing methods operate. In this paper, we propose a new learning-based hashing method called "fast supervised discrete hashing" (FSDH) based on "supervised discrete hashing" (SDH). Regressing the training examples (or hash code) to the corresponding class labels is widely used in ordinary least squares regression. Rather than adopting this method, FSDH uses a very simple yet effective regression of the class labels of training examples to the corresponding hash code to accelerate the algorithm. To the best of our knowledge, this strategy has not previously been used for hashing. Traditional SDH decomposes the optimization into three sub-problems, with the most critical sub-problem - discrete optimization for binary hash codes - solved using iterative discrete cyclic coordinate descent (DCC), which is time-consuming. However, FSDH has a closed-form solution and only requires a single rather than iterative hash code-solving step, which is highly efficient. Furthermore, FSDH is usually faster than SDH for solving the projection matrix for least squares regression, making FSDH generally faster than SDH. For example, our results show that FSDH is about 12-times faster than SDH when the number of hashing bits is 128 on the CIFAR-10 data base, and FSDH is about 151-times faster than FastHash when the number of hashing bits is 64 on the MNIST data-base. Our experimental results show that FSDH is not only fast, but also outperforms other comparative methods. Jie Gui, Tongliang Liu, Zhenan Sun, Dacheng Tao, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Demographic Analysis from Biometric Data: Achievements, Challenges, and New FrontiersabstractBiometrics is the technique of automatically recognizing individuals based on their biological or behavioral characteristics. Various biometric traits have been introduced and widely investigated, including fingerprint, iris, face, voice, palmprint, gait and so forth. Apart from identity, biometric data may convey various other personal information, covering affect, age, gender, race, accent, handedness, height, weight, etc. Among these, analysis of demographics (age, gender, and race) has received tremendous attention owing to its wide real-world applications, with significant efforts devoted and great progress achieved. This survey first presents biometric demographic analysis from the standpoint of human perception, then provides a comprehensive overview of state-of-the-art advances in automated estimation from both academia and industry. Despite these advances, a number of challenging issues continue to inhibit its full potential. We second discuss these open problems, and finally provide an outlook into the future of this very active field of research by sharing some promising opportunities. Yunlian Sun, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Learning structured ordinal measures for video based face recognition
Ran He 0001, Tieniu Tan, Larry Davis 0001, Zhenan Sun |
Pattern Recognit. | 4 |
| 2018 | Efficient auto-refocusing for light field camera
Chi Zhang 0060, Guangqi Hou, Zhaoxiang Zhang 0001, Zhenan Sun, Tieniu Tan |
Pattern Recognit. | 4 |
| 2018 | A Light CNN for Deep Face Representation With Noisy LabelsabstractThe volume of convolutional neural network (CNN) models proposed for face recognition has been continuously growing larger to better fit the large amount of training data. When training data are obtained from the Internet, the labels are likely to be ambiguous and inaccurate. This paper presents a Light CNN framework to learn a compact embedding on the large-scale face data with massive noisy labels. First, we introduce a variation of maxout activation, called max-feature-map (MFM), into each convolutional layer of CNN. Different from maxout activation that uses many feature maps to linearly approximate an arbitrary convex activation function, MFM does so via a competitive relationship. MFM can not only separate noisy and informative signals but also play the role of feature selection between two feature maps. Second, three networks are carefully designed to obtain better performance, meanwhile, reducing the number of parameters and computational costs. Finally, a semantic bootstrapping method is proposed to make the prediction of the networks more consistent with noisy labels. Experimental results show that the proposed framework can utilize large-scale noisy data to learn a Light model that is efficient in computational costs and storage spaces. The learned single network with a 256-D representation achieves state-of-the-art results on various face benchmarks without fine-tuning. Xiang Wu 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | DeMeshNet: Blind Face Inpainting for Deep MeshFace VerificationabstractMeshFace photos have been widely used in many Chinese business organizations to protect ID face photos from being misused. The occlusions incurred by random meshes severely degenerate the performance of face verification systems, which raises the MeshFace verification problem between MeshFace and daily photos. Previous methods cast this problem as a typical low-level vision problem, i.e., blind inpainting. They recover perceptually pleasing clear ID photos from MeshFaces by enforcing pixel level similarity between the recovered ID images and the ground-truth clear ID images and then perform face verification on them. Essentially, face verification is conducted on a compact feature space rather than the image pixel space. Therefore, this paper argues that pixel level similarity and feature level similarity jointly offer the key to improve the verification performance. Based on this insight, we offer a novel feature oriented blind face inpainting framework. Specifically, we implement this by establishing a novel DeMeshNet, which consists of three parts. The first part addresses blind inpainting of the MeshFaces by implicitly exploiting extra supervision from the occlusion position to enforce pixel level similarity. The second part explicitly enforces a feature level similarity in the compact feature space, which can explore informative supervision from the feature space to produce better inpainting results for verification. The last part copes with face alignment within the net via a customized spatial transformer module when extracting deep facial features. All three parts are implemented within an end-to-end network that facilitates efficient optimization. Extensive experiments on two MeshFace data sets demonstrate the effectiveness of the proposed DeMeshNet as well as the insight of this paper. Shu Zhang 0015, Ran He 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Combining Data-Driven and Model-Driven Methods for Robust Facial Landmark DetectionabstractFacial landmark detection is an important yet challenging task for real-world computer vision applications. This paper proposes an effective and robust approach for facial landmark detection by combining data- and model-driven methods. First, a fully convolutional network (FCN) is trained to compute response maps of all facial landmark points. Such a data-driven method could make full use of holistic information in a facial image for global estimation of facial landmarks. After that, the maximum points in the response maps are fitted with a pre-trained point distribution model (PDM) to generate the initial facial shape. This model-driven method is able to correct the inaccurate locations of outliers by considering the shape prior information. Finally, a weighted version of regularized landmark mean-shift (RLMS) is employed to fine-tune the facial shape iteratively. This estimation-correction-tuning process perfectly combines the advantages of the global robustness of the data-driven method (FCN), outlier correction capability of the model-driven method (PDM), and non-parametric optimization of RLMS. Results of extensive experiments demonstrate that our approach achieves state-of-the-art performances on challenging data sets, including 300W, AFLW, AFW, and COFW. The proposed method is able to produce satisfying detection results on face images with exaggerated expressions, large head poses, and partial occlusions. Hongwen Zhang 0001, Qi Li 0005, Zhenan Sun, Yunfan Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Deep Feature Fusion for Iris and Periocular Biometrics on Mobile DevicesabstractThe quality of iris images on mobile devices is significantly degraded due to hardware limitations and less constrained environments. Traditional iris recognition methods cannot achieve high identification rate using these low-quality images. To enhance the performance of mobile identification, we develop a deep feature fusion network that exploits the complementary information presented in iris and periocular regions. The proposed method first applies maxout units into the convolutional neural networks (CNNs) to generate a compact representation for each modality and then fuses the discriminative features of two modalities through a weighted concatenation. The parameters of convolutional filters and fusion weights are simultaneously learned to optimize the joint representation of iris and periocular biometrics. To promote the iris recognition research on mobile devices under near-infrared (NIR) illumination, we publicly release the CASIA-Iris-Mobile-V1.0 database, which in total includes 11 000 NIR iris images of both eyes from 630 Asians. It is the largest NIR mobile iris database as far as we know. On the newly built CASIA-Iris-M1-S3 data set, the proposed method achieves 0.60% equal error rate and 2.32% false non-match rate at false match rate =10-5, which are obviously better than unimodal biometrics as well as traditional fusion methods. Moreover, the proposed model requires much fewer storage spaces and computational resources than general CNNs. Qi Zhang 0015, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | LFNet: A Novel Bidirectional Recurrent Convolutional Neural Network for Light-Field Image Super-ResolutionabstractThe low spatial resolution of light-field image poses significant difficulties in exploiting its advantage. To mitigate the dependency of accurate depth or disparity information as priors for light-field image super-resolution, we propose an implicitly multi-scale fusion scheme to accumulate contextual information from multiple scales for super-resolution reconstruction. The implicitly multi-scale fusion scheme is then incorporated into bidirectional recurrent convolutional neural network, which aims to iteratively model spatial relations between horizontally or vertically adjacent sub-aperture images of light-field data. Within the network, the recurrent convolutions are modified to be more effective and flexible in modeling the spatial correlations between neighboring views. A horizontal sub-network and a vertical sub-network of the same network structure are ensembled for final outputs via stacked generalization. Experimental results on synthetic and real-world data sets demonstrate that the proposed method outperforms other state-of-the-art methods by a large margin in peak signal-to-noise ratio and gray-scale structural similarity indexes, which also achieves superior quality for human visual systems. Furthermore, the proposed method can enhance the performance of light field applications such as depth estimation. Yunlong Wang 0003, Fei Liu 0031, Kunbo Zhang, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 5 |
| 2018 | Supervised Discrete Hashing With RelaxationabstractData-dependent hashing has recently attracted attention due to being able to support efficient retrieval and storage of high-dimensional data, such as documents, images, and videos. In this paper, we propose a novel learning-based hashing method called "supervised discrete hashing with relaxation" (SDHR) based on "supervised discrete hashing" (SDH). SDH uses ordinary least squares regression and traditional zero-one matrix encoding of class label information as the regression target (code words), thus fixing the regression target. In SDHR, the regression target is instead optimized. The optimized regression target matrix satisfies a large margin constraint for correct classification of each example. Compared with SDH, which uses the traditional zero-one matrix, SDHR utilizes the learned regression target matrix and, therefore, more accurately measures the classification error of the regression model and is more flexible. As expected, SDHR generally outperforms SDH. Experimental results on two large-scale image data sets (CIFAR-10 and MNIST) and a large-scale and challenging face data set (FRGC) demonstrate the effectiveness and efficiency of SDHR. Jie Gui, Tongliang Liu, Zhenan Sun, Dacheng Tao, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Learning Invariant Deep Representation for NIR-VIS Face RecognitionabstractVisual versus near infrared (VIS-NIR) face recognition is still a challenging heterogeneous task due to large appearance difference between VIS and NIR modalities. This paper presents a deep convolutional network approach that uses only one network to map both NIR and VIS images to a compact Euclidean space. The low-level layers of this network are trained only on large-scale VIS data. Each convolutional layer is implemented by the simplest case of maxout operator. The high-level layer is divided into two orthogonal subspaces that contain modality-invariant identity information and modality-variant spectrum information respectively. Our joint formulation leads to an alternating minimization approach for deep representation at the training time and an efficient computation for heterogeneous data at the testing time. Experimental evaluations show that our method achieves 94% verification rate at FAR=0.1% on the challenging CASIA NIR-VIS 2.0 face recognition dataset. Compared with state-of-the-art methods, it reduces the error rate by 58% only with a compact 64-D representation. Ran He 0001, Xiang Wu 0001, Zhenan Sun, Tieniu Tan |
AAAI | 3 |
| 2017 | Fast multi-view face alignment via multi-task auto-encodersabstractFace alignment is an important problem in computer vision. It is still an open problem due to the variations of facial attributes (e.g., head pose, facial expression, illumination variation). Many studies have shown that face alignment and facial attribute analysis are often correlated. This paper develops a two-stage multi-task Auto-encoders framework for fast face alignment by incorporating head pose information to handle large view variations. In the first and second stages, multi-task Auto-encoders are used to roughly locate and further refine facial landmark locations with related pose information, respectively. Besides, the shape constraint is naturally encoded into our two-stage face alignment framework to preserve facial structures. A coarse-to-fine strategy is adopted to refine the facial landmark results with the shape constraint. Furthermore, the computational cost of our method is much lower than its deep learning competitors. Experimental results on various challenging datasets show the effectiveness of the proposed method. Qi Li 0005, Zhenan Sun, Ran He 0001 |
IJCB | 2 |
| 2017 | LivDet iris 2017 - Iris liveness detection competition 2017abstractPresentation attacks such as using a contact lens with a printed pattern or printouts of an iris can be utilized to bypass a biometric security system. The first international iris liveness competition was launched in 2013 in order to assess the performance of presentation attack detection (PAD) algorithms, with a second competition in 2015. This paper presents results of the third competition, LivDet-Iris 2017. Three software-based approaches to Presentation Attack Detection were submitted. Four datasets of live and spoof images were tested with an additional cross-sensor test. New datasets and novel situations of data have resulted in this competition being of a higher difficulty than previous competitions. Anonymous received the best results with a rate of rejected live samples of 3.36% and rate of accepted spoof samples of 14.71%. The results show that even with advances, printed iris attacks as well as patterned contacts lenses are still difficult for software-based systems to detect. Printed iris images were easier to be differentiated from live images in comparison to patterned contact lenses as was also seen in previous competitions. David Yambay, Benedict Becker, Naman Kohli, Daksha Yadav, Adam Czajka, Kevin W. Bowyer, Stephanie Schuckers, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Diego Gragnaniello, Carlo Sansone, Luisa Verdoliva, Lingxiao He, Yiwei Ru, Nianfeng Liu, Zhenan Sun, Tieniu Tan |
IJCB | 18 |
| 2017 | Wavelet-SRNet: A Wavelet-Based CNN for Multi-scale Face Super ResolutionabstractMost modern face super-resolution methods resort to convolutional neural networks (CNN) to infer highresolution (HR) face images. When dealing with very low resolution (LR) images, the performance of these CNN based methods greatly degrades. Meanwhile, these methods tend to produce over-smoothed outputs and miss some textural details. To address these challenges, this paper presents a wavelet-based CNN approach that can ultra-resolve a very low resolution face image of 16 × 16 or smaller pixelsize to its larger version of multiple scaling factors (2×, 4×, 8× and even 16×) in a unified framework. Different from conventional CNN methods directly inferring HR images, our approach firstly learns to predict the LR's corresponding series of HR's wavelet coefficients before reconstructing HR images from them. To capture both global topology information and local texture details of human faces, we present a flexible and extensible convolutional neural network with three types of loss: wavelet prediction loss, texture loss and full-image loss. Extensive experiments demonstrate that the proposed approach achieves more appealing results both quantitatively and qualitatively than state-ofthe- art super-resolution methods. Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan |
ICCV | 3 |
| 2017 | Deep Supervised Discrete HashingabstractWith the rapid growth of image and video data on the web, hashing has been extensively studied for image or video search in recent years. Benefiting from recent advances in deep learning, deep hashing methods have achieved promising results for image retrieval. However, there are some limitations of previous deep hashing methods (e.g., the semantic information is not fully exploited). In this paper, we develop a deep supervised discrete hashing algorithm based on the assumption that the learned binary codes should be ideal for classification. Both the pairwise label information and the classification information are used to learn the hash codes within one stream framework. We constrain the outputs of the last layer to be binary codes directly, which is rarely investigated in deep hashing algorithm. Because of the discrete nature of hash codes, an alternating minimization method is used to optimize the objective function. Experimental results have shown that our method outperforms current state-of-the-art methods on benchmark datasets. Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan |
NIPS | 2 |
| 2017 | High quality depth map estimation of object surface from light-field images
Fei Liu 0031, Guangqi Hou, Zhenan Sun, Tieniu Tan |
Neurocomputing | 3 |
| 2017 | Bin-based classifier fusion of iris and face biometrics
Di Miao, Man Zhang 0005, Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
Neurocomputing | 3 |
| 2017 | Editorial: Special issue on ubiquitous biometrics
Ran He 0001, Brian C. Lovell, Rama Chellappa, Anil K. Jain 0001, Zhenan Sun |
Pattern Recognit. | 5 |
| 2017 | A Code-Level Approach to Heterogeneous Iris RecognitionabstractMatching heterogeneous iris images in less constrained applications of iris biometrics is becoming a challenging task. The existing solutions try to reduce the difference between heterogeneous iris images in pixel intensities or filtered features. In contrast, this paper proposes a code-level approach in heterogeneous iris recognition. The non-linear relationship between binary feature codes of heterogeneous iris images is modeled by an adapted Markov network. This model transforms the number of iris templates in the probe into a homogenous iris template corresponding to the gallery sample. In addition, a weight map on the reliability of binary codes in the iris template can be derived from the model. The learnt iris template and weight map are jointly used in building a robust iris matcher against the variations of imaging sensors, capturing distance, and subject conditions. Extensive experimental results of matching cross-sensor, high-resolution versus low-resolution and, clear versus blurred iris images demonstrate the code-level approach can achieve the highest accuracy in compared with the existing pixel-level, feature-level, and score-level solutions. Nianfeng Liu, Jing Liu 0062, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Feature Selection Based on Structured Sparsity: A Comprehensive StudyabstractFeature selection (FS) is an important component of many pattern recognition tasks. In these tasks, one is often confronted with very high-dimensional data. FS algorithms are designed to identify the relevant feature subset from the original features, which can facilitate subsequent analysis, such as clustering and classification. Structured sparsity-inducing feature selection (SSFS) methods have been widely studied in the last few years, and a number of algorithms have been proposed. However, there is no comprehensive study concerning the connections between different SSFS methods, and how they have evolved. In this paper, we attempt to provide a survey on various SSFS methods, including their motivations and mathematical representations. We then explore the relationship among different formulations and propose a taxonomy to elucidate their evolution. We group the existing SSFS methods into two categories, i.e., vector-based feature selection (feature selection based on lasso) and matrix-based feature selection (feature selection based on lr,p-norm). Furthermore, FS has been combined with other machine learning algorithms for specific applications, such as multitask learning, multilabel learning, multiview learning, classification, and clustering. This paper not only compares the differences and commonalities of these methods based on regression and regularization strategies, but also provides useful guidelines to practitioners working in related fields to guide them how to do feature selection. Jie Gui, Zhenan Sun, Shuiwang Ji, Dacheng Tao, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Simultaneous Feature and Sample Reduction for Image-Set ClassificationabstractImage-set classification is the assignment of a label to a given image set. In real-life scenarios such as surveillance videos, each image set often contains much redundancy in terms of features and samples. This paper introduces a joint learning method for image-set classification that simultaneously learns compact binary codes and removes redundant samples. The joint objective function of our model mainly includes two parts. The first part seeks a hashing function to generate binary codes that have larger inter-class and smaller intra-class distances. The second one reduces redundant samples with discrete constraints in a low-rank way. A kernel method based on anchor points is further used to reduce sample variations. The proposed discrete objective function is simplified to a series of sub-problems that admit an analytical solution, resulting in a high-quality discrete solution with a low computational cost. Experiments on three commonly used image-set datasets show that the proposed method for the tasks of face recognition from image sets is efficient and effective. Man Zhang 0005, Ran He 0001, Zhenan Sun, Tieniu Tan |
AAAI | 4 |
| 2016 | A simple and robust super resolution method for light field imagesabstractLight field cameras generate low-resolution images due to the tradeoff between spatial and angular resolution. Traditional light field super-resolution (LFSR) methods depend on prior knowledge of depth information. This paper presents a projection-based LFSR solution without prior information based on redefinition of the mapping function between disparity and shearing shift. Moreover, simplified variational regularization is imposed in global optimization formulation to the rendered high-resolution images. Both a synthetic dataset and a real-world dataset of light field images captured by a self-developed light field camera are used to demonstrate the state-of-the-art performance of the proposed method. Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Zilei Wang, Tieniu Tan |
ICIP | 3 |
| 2016 | Group-Invariant Cross-Modal Subspace Learning
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
IJCAI | 3 |
| 2016 | Transformation invariant subspace clustering
Qi Li 0005, Zhenan Sun, Zhouchen Lin, Ran He 0001, Tieniu Tan |
Pattern Recognit. | 2 |
| 2016 | DeepIris: Learning pairwise filter bank for heterogeneous iris verification
Nianfeng Liu, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
Pattern Recognit. Lett. | 4 |
| 2016 | Representative Vector Machines: A Unified Framework for Classical ClassifiersabstractClassifier design is a fundamental problem in pattern recognition. A variety of pattern classification methods such as the nearest neighbor (NN) classifier, support vector machine (SVM), and sparse representation-based classification (SRC) have been proposed in the literature. These typical and widely used classifiers were originally developed from different theory or application motivations and they are conventionally treated as independent and specific solutions for pattern classification. This paper proposes a novel pattern classification framework, namely, representative vector machines (or RVMs for short). The basic idea of RVMs is to assign the class label of a test example according to its nearest representative vector. The contributions of RVMs are twofold. On one hand, the proposed RVMs establish a unified framework of classical classifiers because NN, SVM, and SRC can be interpreted as the special cases of RVMs with different definitions of representative vectors. Thus, the underlying relationship among a number of classical classifiers is revealed for better understanding of pattern classification. On the other hand, novel and advanced classifiers are inspired in the framework of RVMs. For example, a robust pattern classification method called discriminant vector machine (DVM) is motivated from RVMs. Given a test example, DVM first finds its k -NNs and then performs classification based on the robust M-estimator and manifold regularization. Extensive experimental evaluations on a variety of visual recognition tasks such as face recognition (Yale and face recognition grand challenge databases), object categorization (Caltech-101 dataset), and action recognition (Action Similarity LAbeliNg) demonstrate the advantages of DVM over other classifiers. Jie Gui, Tongliang Liu, Dacheng Tao, Zhenan Sun, Tieniu Tan |
IEEE Trans. Cybern. | 4 |
| 2016 | Complementary Cohort Strategy for Multimodal Face Pair MatchingabstractFace pair matching is the task of determining whether two face images represent the same person. Due to the limited expressive information embedded in the two face images as well as various sources of facial variations, it becomes a quite difficult problem. Toward the issue of few available images provided to represent each face, we propose to exploit an extra cohort set (identities in the cohort set are different from those being compared) by a series of cohort list comparisons. Useful cohort coefficients are then extracted from both sorted cohort identities and sorted cohort images for complementary information. To augment its robustness to complicated facial variations, we further employ multiple face modalities owing to their complementary value to each other for the face pair matching task. The final decision is made by fusing the extracted cohort coefficients with the direct matching score for all the available face modalities. To investigate the capacity of each individual modality on matching faces, the cohort behavior, and the performance achieved using our complementary cohort strategy, we conduct a set of experiments on two recently collected multimodal face databases. It is shown that using different modalities leads to different face pair matching performance. For each modality, employing our cohort scheme significantly reduces the equal error rate. By applying the proposed multimodal complementary cohort strategy, we achieve the best performance on our face pair matching task. Yunlian Sun, Kamal Nasrollahi, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Information Theoretic Subspace ClusteringabstractThis paper addresses the problem of grouping the data points sampled from a union of multiple subspaces in the presence of outliers. Information theoretic objective functions are proposed to combine structured low-rank representations (LRRs) to capture the global structure of data and information theoretic measures to handle outliers. In theoretical part, we point out that group sparsity-induced measures (ℓ2,1-norm, ℓα-norm, and correntropy) can be justified from the viewpoint of halfquadratic (HQ) optimization, which facilitates both convergence study and algorithmic development. In particular, a general formulation is accordingly proposed to unify HQ-based group sparsity methods into a common framework. In algorithmic part, we develop information theoretic subspace clustering methods via correntropy. With the help of Parzen window estimation, correntropy is used to handle either outliers under any distributions or sample-specific errors in data. Pairwise link constraints are further treated as a prior structure of LRRs. Based on the HQ framework, iterative algorithms are developed to solve the nonconvex information theoretic loss functions. Experimental results on three benchmark databases show that our methods can further improve the robustness of LRR subspace clustering and outperform other state-of-the-art subspace clustering methods. Ran He 0001, Liang Wang 0001, Zhenan Sun, Yingya Zhang, Bo Li 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Iris Texture Description Using Ordinal Co-occurrence Matrix Features
Yasser Chacon-Cabrera, Man Zhang 0005, Eduardo Garea Llano, Zhenan Sun |
CIARP | 4 |
| 2015 | Albedo assisted high-quality shape recovery from 4D light fieldsabstractOver the past decade, shape reconstruction methods have been limited to Lambertian reflectance with uniform albedo and controlled lighting environment. In this paper, we present an approach for recovering high-quality shapes from 4D light fields, which can handle non-Lambertian and multi-albedo scenes with shadows and inter-reflections. 4D light fields represent all light rays that hit the sensor plane from different directions, and the depth map from light fields is robust to non-Lambertian objects. Specifically, we estimate the albedos by eliminating shadows and inter-reflections with the edge and chromaticity. Then the lighting environment is analyzed from albedos and shading. Finally, the high-quality surface geometry is exactly recovered through normal refinement. We evaluate the effectiveness and robustness on the public 4D light fields database with both synthetic and real-world scenes. Fei Liu 0031, Guangqi Hou, Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2015 | Code Consistent Hashing Based on Information-Theoretic CriterionabstractLearning based hashing techniques have attracted broad research interests in the Big Media research area. They aim to learn compact binary codes which can preserve semantic similarity in the Hamming embedding. However, the discrete constraints imposed on binary codes typically make hashing optimizations very challenging. In this paper, we present a code consistent hashing (CCH) algorithm to learn discrete binary hash codes. To form a simple yet efficient hashing objective function, we introduce a new code consistency constraint to leverage discriminative information and propose to utilize the Hadamard code which favors an information-theoretic criterion as the class prototype. By keeping the discrete constraint and introducing an orthogonal constraint, our objective function can be minimized efficiently. Experimental results on three benchmark datasets demonstrate that the proposed CCH outperforms state-of-the-art hashing methods in both image retrieval and classification tasks, especially with short binary codes. Shu Zhang 0015, Jian Liang 0001, Ran He 0001, Zhenan Sun |
IEEE Trans. Big Data | 4 |
| 2015 | Robust Subspace Clustering With Complex NoiseabstractSubspace clustering has important and wide applications in computer vision and pattern recognition. It is a challenging task to learn low-dimensional subspace structures due to complex noise existing in high-dimensional data. Complex noise has much more complex statistical structures, and is neither Gaussian nor Laplacian noise. Recent subspace clustering methods usually assume a sparse representation of the errors incurred by noise and correct these errors iteratively. However, large corruptions incurred by complex noise cannot be well addressed by these methods. A novel optimization model for robust subspace clustering is proposed in this paper. Its objective function mainly includes two parts. The first part aims to achieve a sparse representation of each high-dimensional data point with other data points. The second part aims to maximize the correntropy between a given data point and its low-dimensional representation with other points. Correntropy is a robust measure so that the influence of large corruptions on subspace clustering can be greatly suppressed. An extension of pairwise link constraints is also proposed as prior information to deal with complex noise. Half-quadratic minimization is provided as an efficient solution to the proposed robust subspace clustering formulations. Experimental results on three commonly used data sets show that our method outperforms state-of-the-art subspace clustering methods. Ran He 0001, Yingya Zhang, Zhenan Sun, Qiyue Yin |
IEEE Trans. Image Process. | 3 |
| 2014 | Impact of sensor ageing on iris recognitionabstractSimilar to the impact of ageing on human beings, digital image sensors develop ageing effects over time. Since these imager's ageing effects (commonly denoted as pixel defects) leave marks in the captured images, it is not clear whether this affects the accuracy of iris recognition systems. This paper proposes a method to investigate the influence of sensor ageing on iris recognition by simulative ageing of an iris test database. A pixel model is introduced and an ageing algorithm is discussed to create the test database. To establish practical relevance, the simulation parameters are estimated from the observed ageing effects of a real iris scanner over the timespan of 4 years. Thomas Bergmüller, Luca Debiasi, Andreas Uhl, Zhenan Sun |
IJCB | 4 |
| 2014 | An optimal set of code words and correntropy for rotated least squares regressionabstractThis paper presents a robust feature extraction method for face recognition based on least squares regression (LSR). Our focus is to enhance the robustness and discriminability of the LSR. First, an optimal set of code words is introduced in LSR. Compared to the traditional set of code words, this new set uses less number of code words. Furthermore, it can make the distance of the regression targets of different classes as large as possible. Then, correntropy is integrated into the LSR model for better robustness. Furthermore, considering the commonly used distance metrics such as Euclidean distance and Cosine distance in the subspace are invariant to rotation transformation, rotation is introduced as additional freedom to promote flexibility without sacrificing accuracy. Our objective function is optimized using half-quadratic (HQ) optimization, which facilitates algorithm development and convergence study. Experimental results show that our method outperforms several subspace methods for face recognition, which indicates the validity of the proposed method. Jie Gui, Zhenan Sun, Guangqi Hou, Tieniu Tan |
IJCB | 2 |
| 2014 | Efficient auto-refocusing of iris images for light-field camerasabstractLight field photography provides a revolutionary possibility to reconstruct well-focused iris region from a 4D light-field image. However, such a “shoot and refocus” scheme is time-consuming in practice because it commonly needs to render an image sequence for finding the optimally refocused frame. This paper presents an efficient auto-refocusing iris imaging solution for lenselet-based light-field cameras. Firstly, a refocusing point spread function (R-PSF) is derived by detailed analysis of the relationship between refocusing depth and defocus blurriness. Secondly, an initial image is rendered at arbitrary depth. Thirdly, a content independent blurriness assessment method based on SVR (support vector regression) modeling is performed on the rendered image to locate depth shift from optimal focusing plane based on R-PSF. Finally, the optimally focused iris image is selected from a frontal candidate and a back candidate. Because our method only involves three times of image rendering based on precise localization of the optimal focusing plane, it is much more efficient than conventional “rendering and selection” solutions which need to render a large number of refocused images. Chi Zhang 0060, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IJCB | 3 |
| 2014 | The first ICB* competition on iris recognitionabstractIris recognition becomes an important technology in our society. Visual patterns of human iris provide rich texture information for personal identification. However, it is greatly challenging to match intra-class iris images with large variations in unconstrained environments because of noises, illumination variation, heterogeneity and so on. To track current state-of-the-art algorithms in iris recognition, we organized the first ICB* Competition on Iris Recognition in 2013 (or ICIR2013 shortly). In this competition, 8 participants from 6 countries submitted 13 algorithms totally. All the algorithms were trained on a public database (e.g. CASIA-Iris-Thousand [3]) and evaluated on an unpublished database. The testing results in terms of False Non-match Rate (FNMR) when False Match Rate (FMR) is 0.0001 are taken to rank the submitted algorithms. Man Zhang 0005, Jing Liu 0062, Zhenan Sun, Tieniu Tan, Wu Su, Fernando Alonso-Fernandez, Valérian Némesin, Nadia Othman, Koichi Noda, Peihua Li, Edmundo Hoyle, Akanksha Joshi |
IJCB | 3 |
| 2014 | Transform-invariant dictionary learning for face recognitionabstractDictionary learning has important applications in face recognition. However, large transformation variations of face images pose a grand challenge to conventional dictionary learning methods. A large portion of misleading dictionary atoms are usually learned to represent transformation factors, which will cause ambiguity in face recognition. To address this problem, this paper proposes a general framework for transform-invariant basis matrix learning. Specifically, we present a transform-invariant dictionary learning method which explicitly incorporates an appearance consistent error term to the original objective function in dictionary learning. The unified objective function is effectively optimized in an alternating iterative way. An ensemble of aligned images and a discriminative transform-invariant dictionary for sparse coding can be obtained by solving the formulated objective function. Experimental results on two public face databases demonstrate our algorithm's superiority compared with two state-of-the-art dictionary learning methods and the recently proposed transform-invariant PCA method. Shu Zhang 0015, Man Zhang 0005, Ran He 0001, Zhenan Sun |
ICIP | 4 |
| 2014 | Fusion of Multibiometrics Based on a New Robust Linear ProgrammingabstractMultibiometrics provides a reliable method for identity authentication and has the potential to be widely applied. The success of a multibiometrics method depends critically on its ability to fuse complementary information supplied by different modalities, where the most challenging problem is to evaluate the importance of different modalities. In addition, identity authentication at a distance has become a development trend of multibiometrics. In this paper, we propose a new robust linear programming method to fuse multibiometrics by combining the modalities optimally. The proposed method can provide a reasonable trade off between conservatism and robustness. Experimental results on CASIA-Iris-Distance, a public and challenging multibiometric database, demonstrate the effectiveness and robustness of this method. Di Miao, Zhenan Sun, Yongzhen Huang |
ICPR | 2 |
| 2014 | Slice representation of range data for head pose estimation
Yunqi Tang, Zhenan Sun, Tieniu Tan |
Comput. Vis. Image Underst. | 2 |
| 2014 | Distance metric learning for recognizing low-resolution iris images
Jing Liu 0062, Zhenan Sun, Tieniu Tan |
Neurocomputing | 2 |
| 2014 | Half-Quadratic-Based Iterative Minimization for Robust Sparse RepresentationabstractRobust sparse representation has shown significant potential in solving challenging problems in computer vision such as biometrics and visual surveillance. Although several robust sparse models have been proposed and promising results have been obtained, they are either for error correction or for error detection, and learning a general framework that systematically unifies these two aspects and explores their relation is still an open problem. In this paper, we develop a half-quadratic (HQ) framework to solve the robust sparse representation problem. By defining different kinds of half-quadratic functions, the proposed HQ framework is applicable to performing both error correction and error detection. More specifically, by using the additive form of HQ, we propose an ℓ1-regularized error correction method by iteratively recovering corrupted data from errors incurred by noises and outliers; by using the multiplicative form of HQ, we propose an ℓ1-regularized error detection method by learning from uncorrupted data iteratively. We also show that the ℓ1-regularization solved by soft-thresholding function has a dual relationship to Huber M-estimator, which theoretically guarantees the performance of robust sparse representation in terms of M-estimation. Experiments on robust face recognition under severe occlusion and corruption validate our framework and findings. Ran He 0001, Wei-Shi Zheng 0001, Tieniu Tan, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Iris Image Classification Based on Hierarchical Visual CodebookabstractIris recognition as a reliable method for personal identification has been well-studied with the objective to assign the class label of each iris image to a unique subject. In contrast, iris image classification aims to classify an iris image to an application specific category, e.g., iris liveness detection (classification of genuine and fake iris images), race classification (e.g., classification of iris images of Asian and non-Asian subjects), coarse-to-fine iris identification (classification of all iris images in the central database into multiple categories). This paper proposes a general framework for iris image classification based on texture analysis. A novel texture pattern representation method called Hierarchical Visual Codebook (HVC) is proposed to encode the texture primitives of iris images. The proposed HVC method is an integration of two existing Bag-of-Words models, namely Vocabulary Tree (VT), and Locality-constrained Linear Coding (LLC). The HVC adopts a coarse-to-fine visual coding strategy and takes advantages of both VT and LLC for accurate and sparse representation of iris texture. Extensive experimental results demonstrate that the proposed iris image classification method achieves state-of-the-art performance for iris liveness detection, race classification, and coarse-to-fine iris identification. A comprehensive fake iris image database simulating four types of iris spoof attacks is developed as the benchmark for research of iris liveness detection. Zhenan Sun, Hui Zhang 0061, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | How to Estimate the Regularization Parameter for Spectral Regression Discriminant Analysis and its Kernel Version?abstractSpectral regression discriminant analysis (SRDA) has recently been proposed as an efficient solution to large-scale subspace learning problems. There is a tunable regularization parameter in SRDA, which is critical to algorithm performance. However, how to automatically set this parameter has not been well solved until now. So this regularization parameter was only set to be a constant in SRDA, which is obviously suboptimal. This paper proposes to automatically estimate the optimal regularization parameter of SRDA based on the perturbation linear discriminant analysis (PLDA). In addition, two parameter estimation methods for the kernel version of SRDA are also developed. One is derived from the method of optimal regularization parameter estimation for SRDA. The other is to utilize the kernel version of PLDA. Experiments on a number of publicly available databases demonstrate the effectiveness of the proposed methods for face recognition, spoken letter recognition, handwritten digit recognition, and text categorization. Jie Gui, Zhenan Sun, Shuiwang Ji, Xindong Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Gabor Ordinal Measures for Face RecognitionabstractGreat progress has been achieved in face recognition in the last three decades. However, it is still challenging to characterize the identity related features in face images. This paper proposes a novel facial feature extraction method named Gabor ordinal measures (GOM), which integrates the distinctiveness of Gabor features and the robustness of ordinal measures as a promising solution to jointly handle inter-person similarity and intra-person variations in face images. In the proposal, different kinds of ordinal measures are derived from magnitude, phase, real, and imaginary components of Gabor images, respectively, and then are jointly encoded as visual primitives in local regions. The statistical distributions of these visual primitives in face image blocks are concatenated into a feature vector and linear discriminant analysis is further used to obtain a compact and discriminative feature representation. Finally, a two-stage cascade learning method and a greedy block selection method are used to train a strong classifier for face recognition. Extensive experiments on publicly available face image databases, such as FERET, AR, and large scale FRGC v2.0, demonstrate state-of-the-art face recognition performance of GOM. Zhenhua Chai, Zhenan Sun, Heydi Mendez Vazquez, Ran He 0001, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Group Sparse Multiview Patch Alignment Framework With View Consistency for Image ClassificationabstractNo single feature can satisfactorily characterize the semantic concepts of an image. Multiview learning aims to unify different kinds of features to produce a consensual and efficient representation. This paper redefines part optimization in the patch alignment framework (PAF) and develops a group sparse multiview patch alignment framework (GSM-PAF). The new part optimization considers not only the complementary properties of different views, but also view consistency. In particular, view consistency models the correlations between all possible combinations of any two kinds of view. In contrast to conventional dimensionality reduction algorithms that perform feature extraction and feature selection independently, GSM-PAF enjoys joint feature extraction and feature selection by exploiting l(2,1)-norm on the projection matrix to achieve row sparsity, which leads to the simultaneous selection of relevant features and learning transformation, and thus makes the algorithm more discriminative. Experiments on two real-world image data sets demonstrate the effectiveness of GSM-PAF for image classification. Jie Gui, Dacheng Tao, Zhenan Sun, Yong Luo 0002, Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 3 |
| 2014 | Ordinal Feature Selection for Iris and Palmprint RecognitionabstractOrdinal measures have been demonstrated as an effective feature representation model for iris and palmprint recognition. However, ordinal measures are a general concept of image analysis and numerous variants with different parameter settings, such as location, scale, orientation, and so on, can be derived to construct a huge feature space. This paper proposes a novel optimization formulation for ordinal feature selection with successful applications to both iris and palmprint recognition. The objective function of the proposed feature selection method has two parts, i.e., misclassification error of intra and interclass matching samples and weighted sparsity of ordinal feature descriptors. Therefore, the feature selection aims to achieve an accurate and sparse representation of ordinal measures. And, the optimization subjects to a number of linear inequality constraints, which require that all intra and interclass matching pairs are well separated with a large margin. Ordinal feature selection is formulated as a linear programming (LP) problem so that a solution can be efficiently obtained even on a large-scale feature pool and training database. Extensive experimental results demonstrate that the proposed LP formulation is advantageous over existing feature selection methods, such as mRMR, ReliefF, Boosting, and Lasso for biometric recognition, reporting state-of-the-art accuracy on CASIA and PolyU databases. Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 1 |
| 2013 | Robust Subspace Clustering via Half-Quadratic MinimizationabstractSubspace clustering has important and wide applications in computer vision and pattern recognition. It is a challenging task to learn low-dimensional subspace structures due to the possible errors (e.g., noise and corruptions) existing in high-dimensional data. Recent subspace clustering methods usually assume a sparse representation of corrupted errors and correct the errors iteratively. However large corruptions in real-world applications can not be well addressed by these methods. A novel optimization model for robust subspace clustering is proposed in this paper. The objective function of our model mainly includes two parts. The first part aims to achieve a sparse representation of each high-dimensional data point with other data points. The second part aims to maximize the correntropy between a given data point and its low-dimensional representation with other points. Correntropy is a robust measure so that the influence of large corruptions on subspace clustering can be greatly suppressed. An extension of our method with explicit introduction of representation error terms into the model is also proposed. Half-quadratic minimization is provided as an efficient solution to the proposed robust subspace clustering formulations. Experimental results on Hopkins 155 dataset and Extended Yale Database B demonstrate that our method outperforms state-of-the-art subspace clustering methods. Yingya Zhang, Zhenan Sun, Ran He 0001, Tieniu Tan |
ICCV | 2 |
| 2013 | Face recognition using Histogram of co-occurrence Gabor phase patternsabstractThe fusion of Local Binary Patterns (LBP) and Gabor magnitude features has been demonstrated to be one of the most successful descriptors for face recognition. Recently, several Gabor phase based features like Histogram of Gabor Phase Patterns (HGPP) and Local Gabor XOR Patterns (LGXP) also show competitive results and complementary attributes to Gabor magnitude based features. However, in these two typical Gabor phase based approaches only the binary relationship between neighboring Gabor phases is used, which may lose some discriminative information. To investigate the potential of Gabor phase features for robust face recognition, this paper proposes a novel local descriptor, named Histogram of Co-occurrence Gabor Phase Patterns (HCGPP). In HCGPP, Gabor Phase features are first extracted and quantized into different ranges. Second we estimate the histograms of cooccurrence Gabor phase patterns in each face region. Finally, a nearest-neighbor classifier with the dissimilarity measure χ2is used for classification. Extensive experimental results on FERET and AR databases show the significant advantages of the proposed method over the state-of-the art ones in terms of recognition rate. Zhenhua Chai, Zhenan Sun |
ICIP | 3 |
| 2012 | Semantic Pixel Sets Based Local Binary Patterns for Face Recognition
Zhenhua Chai, Heydi Mendez Vazquez, Ran He 0001, Zhenan Sun, Tieniu Tan |
ACCV (2) | 4 |
| 2012 | Regularization parameter estimation for spectral regression discriminant analysis based on perturbation theory
Jie Gui, Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2012 | Accurate iris localization using contour segments
Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2012 | Robust regularized feature selection for iris recognition via linear programming
Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2012 | Iris image classification based on color information
Hui Zhang 0061, Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2012 | Discriminant sparse neighborhood preserving embedding for face recognition
Jie Gui, Zhenan Sun, Wei Jia 0001, Rong-Xiang Hu, Ying-Ke Lei, Shuiwang Ji |
Pattern Recognit. | 2 |
| 2012 | Noisy iris image matching by using multiple cues
Tieniu Tan, Zhenan Sun, Hui Zhang 0061 |
Pattern Recognit. Lett. | 3 |
| 2011 | Recovery of corrupted low-rank matrices via half-quadratic based nonconvex minimizationabstractRecovering arbitrarily corrupted low-rank matrices arises in computer vision applications, including bioinformatic data analysis and visual tracking. The methods used involve minimizing a combination of nuclear norm and l1norm. We show that by replacing the l1norm on error items with nonconvex M-estimators, exact recovery of densely corrupted low-rank matrices is possible. The robustness of the proposed method is guaranteed by the M-estimator theory. The multiplicative form of half-quadratic optimization is used to simplify the nonconvex optimization problem so that it can be efficiently solved by iterative regularization scheme. Simulation results corroborate our claims and demonstrate the efficiency of our proposed method under tough conditions. Ran He 0001, Zhenan Sun, Tieniu Tan, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2011 | Graph modeling based local descriptor selection via a hierarchical structure for biometric recognitionabstractLocal descriptor based image representation is widely used in biometrics and has achieved promising results. We usually extract the most distinctive local descriptors for image sparse representation due to the large feature space and the redundancy among local descriptors. In this paper, we describe the local descriptor based image representation via a graph model, in which each node is a local descriptor (we call it “atom”) and the edges denote the relationship between atoms. Based on this model, a hierarchical structure is constructed to select the most distinctive local descriptors. Two-layer structure is adopted in our work, including local selection and global selection. In the first layer, L1/Lqregularized least square regression is adopted to reduce the redundancy of local descriptors in local regions. In the second layer, AdaBoost learning is performed for local descriptor selection based on the results of the first layer. We apply this method to long-range personal identification by using binocular regions. Our method can select the distinctive local descriptors and reduce the redundancy among them, and achieve encouraging results on the collected binocular database and CASIA-Iris-Distance. Particularly, our method is about 50 times faster than the traditional AdaBoost learning based method in the experiments. Zhenan Sun, Tieniu Tan |
IJCB | 2 |
| 2011 | Comprehensive assessment of iris image qualityabstractIris image quality critically determines iris recognition performance and the quality metrics of iris images are also useful prior information for adaptive selection of optimal recognition strategy. Iris image quality is jointly determined by multiple factors such as focus, occlusion, off-angle, deformation, etc. So it is a complex problem to assess the overall quality score of an iris image. This paper proposes a novel framework for comprehensive assessment of iris image quality. The contributions of the paper include three aspects: (i) Three novel approaches are proposed to estimate the quality metrics (QM) of defocus, motion blur and off-angle in an iris image respectively, (ii) A fusion method based on likelihood ratio is proposed to combine six quality factors of an iris image into an unified quality score. (iii) A statistical quantization method based on t-test is proposed to adaptively classify the iris images in a database into a number of quality levels. Extensive experiments demonstrate the proposed framework can effectively assess the overall quality of iris images. And the relationship between iris recognition results and the quality level of iris images can be explicitly formulated. Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2011 | Deformable DAISY Matcher for robust iris recognitionabstractIris is rich of texture information for reliable personal identification. However, nonlinear deformation of iris pattern caused by pupil dilation or contraction raises a grand challenge to iris recognition. This paper proposes a novel iris recognition method namely Deformable DAISY Matcher (DDM) for robust iris feature matching. Firstly, dense DAISY descriptors are extracted to represent regional iris features, which are robust against intra-class variations of iris images. Then a set of iris key points are localized on the feature map. Finally deformation tolerant matching strategy is proposed to match corresponding key points of iris images. Experimental results on two iris image databases demonstrate DDM is better than state-of-the-art iris recognition methods. Man Zhang 0005, Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2011 | Iris Matching Based on Personalized Weight MapabstractIris recognition typically involves three steps, namely, iris image preprocessing, feature extraction, and feature matching. The first two steps of iris recognition have been well studied, but the last step is less addressed. Each human iris has its unique visual pattern and local image features also vary from region to region, which leads to significant differences in robustness and distinctiveness among the feature codes derived from different iris regions. However, most state-of-the-art iris recognition methods use a uniform matching strategy, where features extracted from different regions of the same person or the same region for different individuals are considered to be equally important. This paper proposes a personalized iris matching strategy using a class-specific weight map learned from the training images of the same iris class. The weight map can be updated online during the iris recognition procedure when the successfully recognized iris images are regarded as the new training data. The weight map reflects the robustness of an encoding algorithm on different iris regions by assigning an appropriate weight to each feature code for iris matching. Such a weight map trained by sufficient iris templates is convergent and robust against various noise. Extensive and comprehensive experiments demonstrate that the proposed personalized iris matching strategy achieves much better iris recognition performance than uniform strategies, especially for poor quality iris images. Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Texture removal for adaptive level set based iris segmentationabstractLevel set based active contour method has been proposed for iris segmentation in recent years, but it can not converge to iris contours in real applications because of its sensitivity to local gradient extremes due to the complex iris texture. In this paper, a novel scheme is proposed to remove local gradient extremes before using level set directly. Firstly, we use two orthogonal ordinal filters to obtain robust gradient map. Then we localize the iris region on the gradient map by an improved Hough transform. After that, a Semantic Iris Contour Map is generated by combining the spatial information of coarse iris location and the gradient map as the edge indicator for level set segmentation. For robust and accurate segmentation, we propose a convergence criterion and a means of updating the parameters for level set. Finally, the accurate segmentation is obtained by the robust adaptive level set method. Encouraging results on ICE 2005 database and CASIA v3 database show the efficiency and effectiveness of our method. Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2010 | Statistics of local surface curvatures for mis-localized iris detectionabstractEye detection is a hot research topic in computer vision for its wide applications in human-computer interaction, face and iris recognition, etc. However, robust eye detection is still a grand challenge due to the numerous appearance variations of eye images in real-world applications. In this paper, we present a novel local surface curvature analysis method to deal with this problem. Firstly, by regarding an eye image as a 2D surface in 3D space, we propose to use the histogram of local surface curvatures as the general representation of eye pattern. Then, a SVM classifier is employed for eye detection using the histogram vectors of eye and non-eye samples. Extensive experiments are performed and the results show that the proposed method achieves state-of-the-art performance in eye detection. In particular, it is more efficient in mistakenly localized iris detection. Hui Zhang 0061, Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2010 | Hierarchical Fusion of Face and Iris for Personal IdentificationabstractMost existing face and iris fusion schemes are concerned about improving performance on good quality images under controlled environments. In this paper, we propose a hierarchical fusion scheme for low quality images under uncontrolled situations. In the training stage, canonical correlation analysis (CCA) is adopted to construct a statistical mapping from face to iris in pixel level. In the testing stage, firstly the probe face image is used to obtain a subset of candidate gallery samples via regression between the probe face and gallery irises, then ordinal representation and sparse representation are performed on these candidate samples for iris recognition and face recognition respectively. Finally, score level fusion via min-max normalization is performed to make final decision. Experimental results on our low quality database show the outperforming performance of proposed method. Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2010 | Contact Lens Detection Based on Weighted LBPabstractSpoof detection is a critical function for iris recognition because it reduces the risk of iris recognition systems being forged. Despite various counterfeit artifacts, cosmetic contact lens is one of the most common and difficult to detect. In this paper, we proposed a novel fake iris detection algorithm based on improved LBP and statistical features. Firstly, a simplified SIFT descriptor is extracted at each pixel of the image. Secondly, the SIFT descriptor is used to rank the LBP encoding sequence. Then, statistical features are extracted from the weighted LBP map. Lastly, SVM classifier is employed to classify the genuine and counterfeit iris images. Extensive experiments are conducted on a database containing more than 5000 fake iris images by wearing 70 kinds of contact lens, and captured by four iris devices. Experimental results show that the proposed method achieves state-of-the-art performance in contact lens spoof detection. Hui Zhang 0061, Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2010 | Efficient and robust segmentation of noisy iris images for non-cooperative iris recognition
Tieniu Tan, Zhaofeng He 0001, Zhenan Sun |
Image Vis. Comput. | 3 |
| 2010 | Topology modeling for Adaboost-cascade based object detection
Zhaofeng He 0001, Tieniu Tan, Zhenan Sun |
Pattern Recognit. Lett. | 3 |
| 2009 | Hierarchical Shape Primitive Features for Online Text-independent Writer IdentificationabstractThis paper proposes a novel method to text independent writer identification from online handwriting. The main contributions of our method include two parts: shape primitive representation and hierarchical structure. Both shape primitive's features are developed to represent the robust and distinctive characteristics of handwriting in two hierarchies. In first hierarchy, the shape primitives probability distribution function (SPPDF)is defined as the static features, to characterize orientation information of writing style. For each shape primitive, the statistics of pressure is defined as the dynamic shape primitives probability distribution function (DSPPDF) and the second hierarchy we build Gaussian model in dynamic attributes (DA) according to curvature of shape primitives. Experiments were conducted on the NLPR handwriting database collected from 242 persons. The results show that the new method achieves high accuracy, fast speed and low requirement of the amount of characters in handwriting samples. We achieve a writer identification rate of 91.5% with datasets in Chinese text and 93.6% in English text. Bangy Li, Zhenan Sun, Tieniu Tan |
ICDAR | 2 |
| 2009 | Quality-based dynamic threshold for iris matchingabstractCurrent iris recognition systems usually regard poor quality iris images useless since defocused or partially occluded iris images may cause false acceptance. However, such a strategy may lose an opportunity to correctly report a genuine match with poor-quality samples. This paper proposes an adaptive iris matching method to improve the throughput of iris recognition systems. The core idea of the method is to dynamically adjust the decision threshold of iris matching module based on the quality measure of input iris image. So that the poor quality iris images also have a chance to match template database under the controlled false accept rate. Experiment results on the real system demonstrate the effectiveness of the proposed method and the recognition time is expected to be greatly reduced. Zhenan Sun, Tieniu Tan, Zhuoshi Wei |
ICIP | 2 |
| 2009 | Palmprint recognition using coarse-to-fine statistical image representationabstractRecent literatures have revealed that statistics of local texture measures can provide accurate descriptions of palmprint appearances. In this framework, one palmprint image is divided into local blocks with multiple spatial resolutions. The statistical texture descriptions of each block are then concatenated to form a multi-scale image representation. However, resultant high-dimensional statistical features lead to increasing of computational cost. In this paper, we tackle this problem by performing a coarse-to-fine cascade scheme, which makes use of information redundancy of statistical texture descriptions between different spatial scales. In contrast with non-cascade strategies, the proposed method reduces most of computational burden and achieves accurate classification simultaneously. Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2009 | Toward Accurate and Fast Iris Segmentation for Iris BiometricsabstractIris segmentation is an essential module in iris recognition because it defines the effective image region used for subsequent processing such as feature extraction. Traditional iris segmentation methods often involve an exhaustive search of a large parameter space, which is time consuming and sensitive to noise. To address these problems, this paper presents a novel algorithm for accurate and fast iris segmentation. After efficient reflection removal, an Adaboost-cascade iris detector is first built to extract a rough position of the iris center. Edge points of iris boundaries are then detected, and an elastic model named pulling and pushing is established. Under this model, the center and radius of the circular iris boundaries are iteratively refined in a way driven by the restoring forces of Hooke's law. Furthermore, a smoothing spline-based edge fitting scheme is presented to deal with noncircular iris boundaries. After that, eyelids are localized via edge detection followed by curve fitting. The novelty here is the adoption of a rank filter for noise elimination and a histogram filter for tackling the shape irregularity of eyelids. Finally, eyelashes and shadows are detected via a learned prediction model. This model provides an adaptive threshold for eyelash and shadow detection by analyzing the intensity distributions of different iris regions. Experimental results on three challenging iris image databases demonstrate that the proposed algorithm outperforms state-of-the-art methods in both accuracy and speed. Zhaofeng He 0001, Tieniu Tan, Zhenan Sun, Xianchao Qiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Ordinal Measures for Iris RecognitionabstractImages of a human iris contain rich texture information useful for identity authentication. A key and still open issue in iris recognition is how best to represent such textural information using a compact set of features (iris features). In this paper, we propose using ordinal measures for iris feature representation with the objective of characterizing qualitative relationships between iris regions rather than precise measurements of iris image structures. Such a representation may lose some image-specific information, but it achieves a good trade-off between distinctiveness and robustness. We show that ordinal measures are intrinsic features of iris patterns and largely invariant to illumination changes. Moreover, compactness and low computational complexity of ordinal measures enable highly efficient iris recognition. Ordinal measures are a general concept useful for image analysis and many variants can be derived for ordinal feature extraction. In this paper, we develop multilobe differential filters to compute ordinal measures with flexible intralobe and interlobe parameters such as location, scale, orientation, and distance. Experimental results on three public iris image databases demonstrate the effectiveness of the proposed ordinal feature models. Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Boosting ordinal features for accurate and fast iris recognitionabstractIn this paper, we present a novel iris recognition method based on learned ordinal features.Firstly, taking full advantages of the properties of iris textures, a new iris representation method based on regional ordinal measure encoding is presented, which provides an over-complete iris feature set for learning. Secondly, a novel Similarity Oriented Boosting (SOBoost) algorithm is proposed to train an efficient and stable classifier with a small set of features. Compared with Adaboost, SOBoost is advantageous in that it operates on similarity oriented training samples, and therefore provides a better way for boosting strong classifiers. Finally, the well-known cascade architecture is adopted to reorganize the learned SOBoost classifier into a dasiacascadepsila, by which the searching ability of iris recognition towards large-scale deployments is greatly enhanced. Extensive experiments on two challenging iris image databases demonstrate that the proposed method achieves state-of-the-art iris recognition accuracy and speed. In addition, SOBoost outperforms Adaboost (Gentle-Adaboost, JS-Adaboost, etc.) in terms of both accuracy and generalization capability across different iris databases. Zhaofeng He 0001, Zhenan Sun, Tieniu Tan, Xianchao Qiu |
CVPR | 2 |
| 2008 | Robust 3D face recognition in uncontrolled environmentsabstractMost current 3D face recognition algorithms are designed based on the data collected in controlled situations, which leads to the un-guaranteed performance in practical systems. In this paper, we propose a Robust Local Log-Gabor Histograms (RLLGH) method to handle the uncontrolled problems encountered in 3D face recognition. In this challenging topic, large expressions and data noises are two main obstacles. To overcome the large expressions, we choose Log-Gabor features (LGF) to extract the distinctive and robust information embedded in 3D faces, which will be represented as 3D Log-Gabor faces. Data noises are summarized as distorted meshes, hair occlusions and misalignments. To overcome these problems, we introduce a robust local histogram (RLH) strategy, which takes advantage of the robustness of the accurate local statistical information. The combination of LGF and RLH leads to RLLGH. The novelties of this paper come from 1) Our work aims at studying 3D face recognition performance in uncontrolled environments; 2) We find that embedding LGF into the LVC framework leads to robustness in handling large expression variations; 3) The RLH strategy gives a promising way to solve the data noises problem. Our experiments are based on the large expression subset in FRGC2.0 3D face database and the expression subset in CASIA 3D face database. Experimental results show the efficiency, robustness and generalization of our proposed method. Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
CVPR | 2 |
| 2008 | Multispectral palm image fusion for accurate contact-free palmprint recognitionabstractIn this paper, we propose to improve the verification performance of a contract-free palmprint recognition system by means of feature- level image registration and pixel-level fusion of multi-spectral palm images. Our method involves image acquisition via a dedicated device under contact-free and multi-spectral environment, preprocessing to locate region of interest (ROI) from each individual hand images, feature-level registration to align ROIs from different spectral images in one sequence and fusion to combine images from multiple spectra. The advantages of the proposed method include better hygiene and higher verification performance. Given a database composed of images from 330 hands, two out of four state of the art fusion strategies offer significant performance gain and the best equal error rate (EER) is 0.5%. Ying Hao, Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2008 | Enhanced usability of iris recognition via efficient user interface and iris image restorationabstractIn this paper, we investigate the possibility of enhancing the usability of iris recognition via exploration of the specular spots in iris images. Firstly, the spatial configuration of the specular spots in iris images is utilized to estimate the distance between the user and the camera. Based on this a friendly user interface is established to assist users for their range adjustment. Furthermore, the estimated distance is used by an adaptive image restoration scheme to restore the blurred iris image, thereby increasing the depth of field of the iris camera. Experimental results show that the proposed method significantly enhances the usability of iris recognition without noticeable computation cost. Zhaofeng He 0001, Zhenan Sun, Tieniu Tan, Xianchao Qiu |
ICIP | 2 |
| 2008 | Robust eyelid, eyelash and shadow localization for iris recognitionabstractEyelids, eyelashes and shadows are three major challenges for effective iris segmentation, which have not been adequately addressed in the current literature. In this paper, we present a novel method to localize each of them. First, a novel coarse-line to fine-parabola eyelid fitting scheme is developed for accurate and fast eyelid localization. Then, a smart prediction model is established to determine an appropriate threshold for eyelash and shadow detection. Experimental results on the challenging CASIA-IrisV3-Lamp iris image database demonstrate that the proposed method outperforms state-of-the-art methods in both accuracy and speed. Zhaofeng He 0001, Tieniu Tan, Zhenan Sun, Xianchao Qiu |
ICIP | 3 |
| 2008 | Palmprint image synthesis: A preliminary studyabstractIn this paper we present a preliminary study of palmprint image synthesis and propose a framework for synthesizing palmprint texture. We first extract principal lines of real palmprints using edge detection and synthesize wrinkles and ridges of palm using patch-based sampling. Then we incorporate principal lines, wrinkles and ridges to obtain the final synthetic image. After that multiple images are derived from each artificial palm to simulate the intra-class images. Our approach can generate large palmprint databases which preserve inter-class and intra-class variations. Experimental results demonstrate that the synthetic images bear a close resemblance to real palmprints in terms of appearance as well as statistical properties, showing a promising usage in algorithms evaluation and comparison. Zhuoshi Wei, Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2008 | Learning efficient codes for 3D face recognitionabstractFace representation based on the visual codebook becomes popular because of its excellent recognition performance, in which the critical problem is how to learn the most efficient codes to represent the facial characteristics. In this paper, we introduce the quadtree clustering algorithm to learn the facial-codes to boost 3D face recognition performance. The merits of quadtree clustering come from: (1) It is robust to data noises; (2) It can adaptively assign clustering centers according to the density of data distribution. We make a comparison between quadtree and some widely used clustering methods, such as g-means, k-means, normalized-cut and mean-shift. Experimental results show that using the facial- codes learned by quadtree clustering gives the best performance for 3D face recognition. Zhenan Sun, Tieniu Tan |
ICIP | 2 |
| 2008 | How to make iris recognition easier?abstractIris recognition is regarded as the most reliable biometrics and has been widely applied in both public and personal security areas. However users have to highly cooperate with the iris cameras to make his iris images well captured. In this paper, we aim to discuss whether and how we can make iris recognition easier. Firstly the restricting factors of iris image acquisition are analyzed and the optical formulas are derived. Then the solutions of state-of-the-art iris recognition systems are reviewed and summarized. Finally, we propose two novel iris recognition systems with good human-computer-interface but with two different strategies which respectively meet the requirements of low-end and high-end market. Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2008 | Combine hierarchical appearance statistics for accurate palmprint recognitionabstractPalmprint recognition is an active member of biometrics in recent years. State-of-the-art algorithms of palmprint recognition describe appearances of palmprints efficiently through local texture analysis. Following this framework, we propose a novel approach of palmprint recognition in this paper, which represents palmprint images based on statistics and spatial arrangement of appearance descriptors within local image areas. In this method, we firstly design a robust descriptor to encode properties of palmprint appearances of local regions. The whole image is divided into non-overlapped blocks at increasingly fine resolutions successively, so as to describe the spatial layout in hierarchical scales. For a specific spatial resolution, local distributions of the proposed descriptors in the blocks are concatenated to represent structures of palmprint structures. Finally, distribution information of different resolutions is combined to provide complementary descriptive power. Promising experimental results demonstrate that the proposed method achieves even better performances than the state-of-the-art approaches. Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2008 | A hierarchical model for the evaluation of biometric sample qualityabstractThe evaluation of biometric sample quality is of great importance in the evaluation of biometric algorithms. In this paper, we propose a novel hierarchical model to compute the sample quality at three levels. This model is developed on the basis of three types of influencing factors: global factors, subjective factors and variable factors. We adopt different strategies to compute the corresponding three level qualities: database level quality, class level quality and image level quality. The database level quality is estimated by experience. Then, we compute the mean value of variable number of normalized genuine scores, the quantiles of which are used to determine the class level quality. On the image level quality evaluation, a novel concept of subset frequency is proposed. Zhenan Sun, Tieniu Tan |
ICPR | 2 |
| 2008 | Counterfeit iris detection based on texture analysisabstractThis paper addresses the issue of counterfeit iris detection, which is a liveness detection problem in biometrics. Fake iris mentioned here refers to iris wearing color contact lens with textures printed onto them. We propose three measures to detect fake iris: measuring iris edge sharpness, applying Iris-Texton feature for characterizing the visual primitives of iris textures and using selected features based on co-occurrence matrix (CM). Extensive testing is carried out on two datasets containing different types of contact lens with totally 640 fake iris images, which demonstrates that Iris-Texton and CM features are effective and robust in anticounterfeit iris. Detailed comparisons with two state-of-the-art methods are also presented, showing that the proposed iris edge sharpness measure acquires a comparable performance with these two methods, while Iris-Texton and CM features outperform the state-of-the-art. Zhuoshi Wei, Xianchao Qiu, Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2008 | Synthesis of large realistic iris databases using patch-based samplingabstractThis paper presents a framework to synthesize large realistic iris databases, providing an alternative to iris database collection. Firstly, iris patch is used as a basic element to characterize visual primitive of iris texture, and patch-based sampling is applied to create an iris prototype. Then a set of pseudo irises with intra-class variations are derived from the prototype. Qualitative and quantitative studies reveal that synthetic databases are well suited for evaluating iris recognition systems by achieving three goals: (1) the synthetic iris images bear a close resemblance to real iris images in terms of visual appearance; (2) the proposed framework is able to generate databases with large capacity; (3) statistical performance shows that the synthetic iris images hold all the major characteristics of real iris images. Zhuoshi Wei, Tieniu Tan, Zhenan Sun |
ICPR | 3 |
| 2007 | Palmprint Recognition Under Unconstrained Scenes
Zhenan Sun, Tieniu Tan |
ACCV (2) | 2 |
| 2007 | Comparative Studies on Multispectral Palm Image Fusion for Biometrics
Ying Hao, Zhenan Sun, Tieniu Tan |
ACCV (2) | 2 |
| 2007 | Fusion of Face and Palmprint for Personal Identification Based on Ordinal FeaturesabstractIn this paper, we present a face and palmprint multimodal biometric identification method and system to improve the identification performance. Effective classifiers based on ordinal features are constructed for faces and palmprints, respectively. Then, the matching scores from the two classifiers are combined using several fusion strategies. Experimental results on a middle-scale data set have demonstrated the effectiveness of the proposed system. Rufeng Chu, Shengcai Liao, Zhenan Sun, Stan Z. Li, Tieniu Tan |
CVPR | 4 |
| 2007 | Robust 3D Face Recognition Using Learned Visual CodebookabstractIn this paper, we propose a novel learned visual code-book (LVC) for 3D face recognition. In our method, we first extract intrinsic discriminative information embedded in 3D faces using Gabor filters, then K-means clustering is adopted to learn the centers from the filter response vectors. We construct LVC by these learned centers. Finally we represent 3D faces based on LVC and achieve recognition using a nearest neighbor (NN) classifier. The novelty of this paper comes from 1) We first apply textons based methods into 3D face recognition; 2) We encompass the efficiency of Gabor features for face recognition and the robustness of texton strategy for texture classification simultaneously. Our experiments are based on two challenging databases, CASIA 3D face database and FRGC2.0 3D face database. Experimental results show LVC performs better than many commonly used methods. Zhenan Sun, Tieniu Tan |
CVPR | 2 |
| 2007 | Learning Appearance Primitives of Iris Images for Ethnic ClassificationabstractIris pattern is commonly regarded as a kind of phenotypic feature without relation to genes. In our previous work, we argued that iris texture is race related, and its genetic information is illustrated in coarse scale texture features, rather than preserved in the minute local features of state-of-the-art iris recognition algorithms. In this paper, we propose a novel ethnic classification method based on learning appearance primitives of iris images. So we not only confirm that iris texture is race related, but also try to find out which kinds of iris visual primitives make iris images look different between Asian and non-Asian. In our scheme, we learned a small finite vocabulary of micro-structures, which are called iris-textons, to represent visual primitives of iris images. Then we use iris-texton histogram to capture the difference between iris textures. Finally iris images are grouped into two race categories, Asian and non-Asian, by support vector machine (SVM). Based on the proposed method, we get a higher correct classification rate (CCR) of 91.02% than our previous method on a database containing 2400 iris samples. Xianchao Qiu, Zhenan Sun, Tieniu Tan |
ICIP (2) | 2 |
| 2005 | Ordinal Palmprint Represention for Personal IdentificationabstractPalmprint-based personal identification, as a new member in the biometrics family, has become an active research topic in recent years. Although great progress has been made, how to represent palmprint for effective classification is still an open problem. In this paper, we present a novel palmprint representation - ordinal measure, which unifies several major existing palmprint algorithms into a general framework. In this framework, a novel palmprint representation method, namely orthogonal line ordinal features, is proposed. The basic idea of this method is to qualitatively compare two elongated, line-like image regions, which are orthogonal in orientation and generate one bit feature code. A palmprint pattern is represented by thousands of ordinal feature codes. In contrast to the state-of-the-art algorithm reported in the literature, our method achieves higher accuracy, with the equal error rate reduced by 42% for a difficult set, while the complexity of feature extraction is halved. Zhenan Sun, Tieniu Tan, Yunhong Wang 0001, Stan Z. Li |
CVPR (1) | 1 |
| 2005 | Improving iris recognition accuracy via cascaded classifiersabstractAs a reliable approach to human identification, iris recognition has received increasing attention in recent years. The most distinguishing feature of an iris image comes from the fine spatial changes of the image structure. So iris pattern representation must characterize the local intensity variations in iris signals. However, the measurements from minutiae are easily affected by noise, such as occlusions by eyelids and eyelashes, iris localization error, nonlinear iris deformations, etc. This greatly limits the accuracy of iris recognition systems. In this paper, an elastic iris blob matching algorithm is proposed to overcome the limitations of local feature based classifiers (LFC). In addition, in order to recognize various iris images efficiently a novel cascading scheme is proposed to combine the LFC and an iris blob matcher. When the LFC is uncertain of its decision, poor quality iris images are usually involved in intra-class comparison. Then the iris blob matcher is resorted to determine the input iris' identity because it is capable of recognizing noisy images. Extensive experimental results demonstrate that the cascaded classifiers significantly improve the system's accuracy with negligible extra computational cost. Zhenan Sun, Yunhong Wang 0001, Tieniu Tan, Jiali Cui |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2004 | Fast recursive mathematical morphological transformsabstractSince many mathematical morphology operations are recursive transforms of dilation and erosion, this paper proposes fast recursive transforms to reduce computational complexity. The basic idea of the method is to compute the temporary results within a series of adaptive windows and the computing is performed on specific pixels. Each step of the recursive process consists of two parts: 1) computation is limited to the specific pixels (foreground or background pixels) within a window; 2) update the window adoptively and delete those varied pixels. Extensive results show that the time complexity of the method is proportional to the number of the specific pixels. Jiali Cui, Yunhong Wang 0001, Tieniu Tan, Zhenan Sun |
ICIG | 4 |
| 2004 | Cascading statistical and structural classifiers for iris recognitionabstractReliable human identification using iris pattern has recently gained growing interests from pattern recognition researchers. In literature of iris recognition, almost all algorithms are based on statistical information. In this paper, a structural iris image analysis method is proposed, which provides complementary information to statistical classifier. In order to save computational cost, the structural matcher is not consulted unless the statistical classifier is uncertain of its decision. At the second stage, the structural classifier may be combined with statistical classifier with different fusion strategies. The experimental results of decision-level classifiers combination are reported, which demonstrate that the cascaded classification system significantly outperforms single classifier. Zhenan Sun, Yunhong Wang 0001, Tieniu Tan, Jiali Cui |
ICIP | 1 |