EDBT 2026 Demo / reviewers in the wild / expert
Yunlong Wang 0003
dblp:32/8538-3
· DBLP profile ↗
45ranked-venue papers
5as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 16 since 2021Security and privacy · 21 · 1 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks (Abstract Reprint)abstractDeep learning-based face recognition models are vulnerable to adversarial attacks. In contrast to general noises, the presence of imperceptible adversarial noises can lead to catastrophic errors in deep face recognition models. The primary difference between adversarial noise and general noise lies in its specificity. Adversarial attack methods give rise to noises tailored to the characteristics of the individual image and recognition model at hand. Diverse samples and recognition models can engender specific adversarial noise patterns, which pose significant challenges for adversarial defense. Addressing this challenge in the realm of face recognition presents a more formidable endeavor due to the inherent nature of face recognition as an open set task. In order to tackle this challenge, it is imperative to employ customized processing for each individual input sample. Drawing inspiration from the biological immune system, which can identify and respond to various threats, this paper aims to create an artificial immune system to provide adversarial defense for face recognition. The proposed defense model incorporates the principles of antibody cloning, mutation, selection, and memory mechanisms to generate a distinct antibody for each input sample, wherein the term antibody refers to a specialized noise removal manner. Furthermore, we introduce a self-supervised adversarial training mechanism that serves as a simulated rehearsal of immune system invasions. Extensive experimental results demonstrate the efficacy of the proposed method, surpassing state-of-the-art adversarial defense methods. Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
AAAI | 2 |
| 2026 | Symmetric Image-Text Tuning With Entropy-Guided Fusion for Online Continual Learning in Non-Stationary Visual StreamsabstractOnline continual learning studies how models learn from continuous and non-stationary data streams. In this paper, we observe that CLIP models exhibit an asymmetric image-text interaction under online continual learning. Specifically, text features of previously seen classes may introduce unfavorable supervision when paired with visual features of newly observed data, leading to catastrophic forgetting. To alleviate this issue, we propose a simple yet effective symmetric image-text tuning (SIT) strategy that removes such asymmetric text supervision during online learning. We further introduce an entropy-guided fusion (EGF) mechanism that adaptively combines predictions from the pretrained and finetuned branches based on their relative uncertainty. This design allows the model to recover pretrained knowledge when the finetuned branch becomes unreliable, while still preserving plasticity on recently observed classes when confidence is high. In addition, we present MiD-Blurry, an online continual learning benchmark that combines multiple class distribution patterns to better reflect realistic data streams with blurred temporal boundaries. Extensive experiments on standard continual learning benchmarks and the MiD-Blurry setting evaluate inference-at-any-time performance and generalization to future data. The results show that the proposed approach maintains a practical balance between adapting to new data and preserving previously learned information in realistic online learning scenarios. Leyuan Wang, Liuyu Xiang, Yiwei Ru, Yunlong Wang 0003, Zhaofeng He 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Rethinking Class-Incremental Learning From a Dynamic Imbalanced Learning PerspectiveabstractDeep neural networks suffer from catastrophic forgetting when continually learning new concepts. In this paper, we analyze this problem from a data imbalance point of view. We argue that the imbalance between old task and new task data contributes to forgetting of the old tasks. Moreover, the increasing imbalance ratio during incremental learning further aggravates the problem. To address the dynamic imbalance issue, we propose Uniform Prototype Contrastive Learning (UPCL), where uniform and compact features are learned. Specifically, we generate a set of non-learnable uniform prototypes before each task starts. Then we assign these uniform prototypes to each class and guide the feature learning through prototype contrastive learning. We also dynamically adjust the relative margin between old and new classes so that the feature distribution will be maintained balanced and compact. Finally, we demonstrate through extensive experiments that the proposed method achieves state-of-the-art performance on several benchmark including CIFAR-100, ImageNet-100, TinyImageNet, Food-101, and CUB-200. Experimental results show that our approach not only effectively addresses the issue of imbalanced old data in memory but also tackles the problem of imbalanced new data distributions. Leyuan Wang, Liuyu Xiang, Yunlong Wang 0003, Huijia Wu, Huafeng Yang, Jingqian Liu, Zhaofeng He 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action RecognitionabstractIn skeleton-based action recognition, a key challenge is distinguishing between actions with similar trajectories of joints due to the lack of image-level details in skeletal representations. Recognizing that the differentiation of similar actions relies on subtle motion details in specific body parts, we direct our approach to focus on the fine-grained motion of local skeleton components. To this end, we introduce ProtoGCN, a Graph Convolutional Network (GCN)-based model that breaks down the dynamics of entire skeleton sequences into a combination of learnable prototypes representing core motion patterns of action units. By contrasting the reconstruction of prototypes, ProtoGCN can effectively identify and enhance the discriminative representation of similar actions. Without bells and whistles, ProtoGCN achieves state-of-the-art performance on multiple benchmark datasets, including NTU RGB+D, NTU RGB+D 120, Kinetics-Skeleton, and FineGYM, which demonstrates the effectiveness of the proposed method. The code is available at https://github.com/firework8/ProtoGCN. Hongda Liu 0002, Yunfan Liu 0001, Yunlong Wang 0003, Zhenan Sun |
CVPR | 5 |
| 2025 | BiommWave: A Non-Visual Approach for Biometric Recognition Using Millimeter-Wave RadarabstractThis paper explores the application of millimeter-wave (mmWave) radar in biometric recognition. As a non-visual human sensing technology, mmWave radar captures reflection properties and micro-movements, providing a complementary modality to visual appearance. We implemented a complete system pipeline to utilize mmWave sensing for individual recognition. Based on physiological mechanisms, we design preprocessing methods to extract intuitive biometric feature maps, concerning body reflections, cardiopulmonary activity, and micro-motion frequencies. To address the data uncertainty, a dynamic pole-based learning strategy is proposed to construct compact and discriminated feature distributions. In real-world evaluations, the system achieves 95.12% accuracy and an Equal Error Rate (EER) of 1.96%. This work leverages the advantages of mmWave radar for flexible, unconstrained, and private biometric systems. From a non-visual sensing perspective, it explores novel modalities as unique biometric cues, demonstrating significant value of research and applications. Mupei Li, Yunlong Wang 0003, Yiwei Ru, Kunbo Zhang, Zhenan Sun |
IJCB | 2 |
| 2025 | IrisFormer: A Dedicated Transformer Framework for Iris RecognitionabstractWhile Vision Transformer (ViT)-based methods have significantly improved the performance of various vision tasks in natural scenes, progress in iris recognition remains limited. In addition, the human iris contains unique characters that are distinct from natural scenes. To remedy this, this paper investigates a dedicated Transformer framework, termed IrisFormer, for iris recognition and attempts to improve the accuracy by combining the contextual modeling ability of ViT and iris-specific optimization to learn robust, fine-grained, and discriminative features. Specifically, to achieve rotation invariance in iris recognition, we employ relative position encoding instead of regular absolute position encoding for each iris image token, and a horizontal pixel-shifting strategy is utilized during training for data augmentation. Then, to enhance the model's robustness against local distortions such as occlusions and reflections, we randomly mask some tokens during training to force the model to learn representative identity features from only part of the image. Finally, considering that fine-grained features are more discriminative in iris recognition, we retain the entire token sequence for patch-wise feature matching instead of using the standard single classification token. Experiments on three popular datasets demonstrate that the proposed framework achieves competitive performance under both intra- and inter-dataset testing protocols. Xianyun Sun, Caiyong Wang, Yunlong Wang 0003, Jianze Wei, Zhenan Sun |
IEEE Signal Process. Lett. | 3 |
| 2025 | TO-LF: A Texture and Occlusion-Oriented Benchmark Dataset for Light Field Disparity EstimationabstractAccurate disparity estimation in light field (LF) imaging remains challenging due to the narrow baseline between adjacent sub-aperture images (SAIs) and the occlusion effect. Existing learning-based methods suffer from degraded performance in complex scenarios owing to the scarcity of high-quality and diverse training data. To address this limitation, we propose a Texture and Occlusion-oriented Light Field dataset (TO-LF) containing 78 carefully curated images. Unlike the widely used HCI 4D LF benchmark, TO-LF not only provides more training samples but also introduces a more challenging test set with complex occlusions and significant textureless regions. Furthermore, we present a viewpoint-selective sub-pixel cost volume construction method (VS-Sub), which extends disparity labels to the subpixel level for denser cost volumes, and employs dynamic dilated convolutions to differentiate between occluded and non-occluded viewpoints. Comprehensive experiments demonstrate that our framework achieves state-of-the-art (SOTA) performance in disparity estimation. Shubo Zhou, Yunlong Wang 0003, Yingqian Wang 0002, Fei Liu 0031, Xueqin Jiang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Exploring Near-Infrared Iris Image Sequences for High Throughput Iris RecognitionabstractHigh throughput is demanding in real-world iris recognition applications. The challenges mainly originate from the variability in image quality under high-throughput capture conditions. Most of the degraded images are typically filtered out by traditional iris systems through Image Quality Assessment (IQA) module, adversely affecting efficiency and leading to low throughput and poor user experience. Therefore, a better and practical solution is to make the utmost of degraded iris images. In order to investigate the key problems of high-throughput iris recognition, we collect a novel iris sequence dataset under Near-infrared (NIR) illumination. This dataset is specifically constructed for high-throughput evaluation, which faithfully simulates the process of iris sequence acquisition in real-world iris systems. Comprehensive evaluations were conducted to figure out the deficiencies of current iris recognition algorithms. To this end, a testing methodology along with specific evaluation metrics is proposed. It is capable of assessing the throughput performance, e.g., the newly proposed Frame Consumption per Match (FCM). Through performance analysis, several insights were gathered to guide potential directions for developing high-throughput iris recognition algorithms. Furthermore, we consider to leverage iris sequence features for better throughput performance. Continuity sequence criteria and cumulative sequence feature strategy are proposed to enhance the throughput performance of existing algorithms with minimal cost. In summary, this work provides valuable data and rational insights for high-throughput iris recognition studies. The datasets and evaluation toolkit are publicly available on our website1. Mupei Li, Yunlong Wang 0003, Kunbo Zhang, Zhaofeng He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Uncertainty-Aware Bilateral Transformer for Accurate and Reliable Iris SegmentationabstractIris segmentation is a deterministic and critical part of the iris recognition system. However, its performance is usually degraded by data uncertainty in acquisition and annotation, impeding more accurate recognition of the iris recognition system. In the paper, we propose a bilateral self-attention by exploring spatial and visual relationships to effectively distinguish between iris and non-iris regions, then design a bilateral Transformer by enhancing spatial perception and hierarchical feature fusion to mitigate the impact of acquisition uncertainty. Besides, iris segmentation uncertainty learning is developed to estimate the uncertainty map according to prediction discrepancy. With the estimated uncertainty, a weighting scheme and a regularization term are designed to minimize the effect of annotation uncertainty. To investigate data uncertainty, the paper presents a challenging near-infrared iris dataset named UTIris. It comprises 3,690 images with high acquisition uncertainty and provides rich segmentation masks to explore annotation uncertainty. Furthermore, we manually label a large-scale iris dataset, ND-0405 [1], with additional binary maps of iris masks to evaluate segmentation performance. Experimental results on UTIris and four other databases demonstrate the effectiveness of the proposed method in iris segmentation, and its segmentation improvement consequently promotes recognition accuracy. Jianze Wei, Xingyu Gao 0001, Yunlong Wang 0003, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesabstractReconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive progress, they ignore formulating contacts in their frameworks, which results in producing less realistic object meshes. In this work, we explore how to model contacts in an explicit way to benefit the implicit reconstruction of hand-held objects. Our method consists of two components: explicit contact prediction and implicit shape reconstruction. In the first part, we propose a new subtask of directly estimating 3D hand-object contacts from a single image. The part-level and vertex-level graph-based transformers are cascaded and jointly learned in a coarse-to-fine manner for more accurate contact probabilities. In the second part, we introduce a novel method to diffuse estimated contact states from the hand mesh surface to nearby 3D space and leverage diffused contact probabilities to construct the implicit neural representation for the manipulated object. Benefiting from estimating the interaction patterns between the hand and the object, our method can reconstruct more realistic object meshes, especially for object parts that are in contact with hands. Extensive experiments on challenging benchmarks show that the proposed method outperforms the current state of the arts by a great margin. Our code is publicly available at https://junxinghu.github.io/projects/hoi.html. Junxing Hu, Hongwen Zhang 0001, Zerui Chen, Mengcheng Li, Yunlong Wang 0003, Yebin Liu, Zhenan Sun |
AAAI | 5 |
| 2024 | 6-DoF Grasp Detection in Clutter with Enhanced Receptive Field and Graspable Balance Samplingabstract6-DoF grasp detection of small-scale grasps is crucial for robots to perform specific tasks. This paper focuses on enhancing the recognition capability of small-scale grasping, aiming to improve the overall accuracy of grasping prediction results and the generalization ability of the network. We propose an enhanced receptive field method that includes a multi-radii cylinder grouping module and a passive attention module. This method enhances the receptive field area within the graspable space and strengthens the learning of graspable features. Additionally, we design a graspable balance sampling module based on a 3D segmentation network, which enables the network to focus on features of small objects, thereby improving the recognition capability of small-scale grasping. Our network achieves state-of-the-art performance on the GraspNet-1Billion dataset, with an overall improvement of approximately 10% in average precision@k (AP). Furthermore, we deployed our grasp detection model on pybullet grasping platform and in real-world scenarios, which validates the effectiveness of our method. Yunlong Wang 0003 |
IROS | 3 |
| 2024 | A comprehensive research on light field imaging: Theory and applicationabstractAbstract Computational photography is a combination of novel optical designs and processing methods to capture high‐dimensional visual information. As an emerged promising technique, light field (LF) imaging measures the lighting, reflectance, focus, geometry and viewpoint in the free space, which has been widely explored for depth estimation, view synthesis, refocus, rendering, 3D displays, microscopy and other applications in computer vision in the past decades. In this paper, the authors present a comprehensive research survey on the LF imaging theory, technology and application. Firstly, the LF imaging process based on a MicroLens Array structure is derived, that is MLA‐LF. Subsequently, the innovations of LF imaging technology are presented in terms of the imaging prototype, consumer LF camera and LF displays in Virtual Reality (VR) and Augmented Reality (AR). Finally the applications and challenges of LF imaging integrating with deep learning models are analysed, which consist of depth estimation, saliency detection, semantic segmentation, de‐occlusion and defocus deblurring in recent years. It is believed that this paper will be a good reference for the future research on LF imaging technology in Artificial Intelligence era. Fei Liu 0031, Yunlong Wang 0003, Shubo Zhou, Kunbo Zhang |
IET Comput. Vis. | 2 |
| 2024 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
Int. J. Comput. Vis. | 2 |
| 2024 | Personalized Graph Generation for Monocular 3D Human Pose and Shape Estimationabstract3D human pose and shape estimation from a single RGB image is an appealing yet challenging task. Due to the graph-like nature of human parametric models, a growing number of graph neural network-based approaches have been proposed and achieved promising results. However, existing methods build graphs for different instances based on the same template SMPL mesh, neglecting the geometric perception of individual properties. In this work, we propose an end-to-end method named Personalized Graph Generation (PGG) to construct the geometry-aware graph from an intermediate predicted human mesh. Specifically, a convolutional module initially regresses a coarse SMPL mesh tailored for each sample. Guided by the 3D structure of this personalized mesh, PGG extracts the local features from the 2D feature map. Then, these geometry-aware features are integrated with the specific coarse SMPL parameters as vertex features. Furthermore, a body-oriented adjacency matrix is adaptively generated according to the coarse mesh. It considers individual full-body relations between vertices, enhancing the perception of body geometry. Finally, a graph attentional module is utilized to predict the residuals to get the final results. Quantitative experiments across four benchmarks and qualitative comparisons on more datasets show that the proposed method outperforms state-of-the-art approaches for 3D human pose and shape estimation. Junxing Hu, Hongwen Zhang 0001, Yunlong Wang 0003, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Understanding Deep Face Representation via Attribute RecoveryabstractDeep neural networks have proven to be highly effective in the face recognition task, as they can map raw samples into a discriminative high-dimensional representation space. However, understanding this complex space proves to be challenging for human observers. In this paper, we propose a novel approach that interprets deep face recognition models via facial attributes. To achieve this, we introduce a two-stage framework that recovers attributes from the deep face representations. This framework allows us to quantitatively measure the significance of facial attributes in relation to the recognition model. Moreover, this framework enables us to generate sample-specific explanations through counterfactual methodology. These explanations are not only understandable but also quantitative. Through the proposed approach, we are able to acquire a deeper understanding of how the recognition model conceptualizes the notion of “identity” and understand the reasons behind the error decisions made by the deep models. By utilizing attributes as an interpretable interface, the proposed method marks a paradigm shift in our comprehension of deep face recognition models. It allows a complex model, obtained through gradient backpropagation, to effectively “communicate” with humans. The source code is available here, or you can visit this website:https://github.com/RenMin1991/Facial-Attribute-Recovery. Yuhao Zhu 0003, Yunlong Wang 0003, Yongzhen Huang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Multi-Faceted Knowledge-Driven Graph Neural Network for Iris SegmentationabstractAccurate iris segmentation, especially around the iris inner and outer boundaries, is still a formidable challenge. Pixels within these areas are difficult to semantically distinguish since they have similar visual characteristics and close spatial positions. To tackle this problem, the paper proposes an iris segmentation graph neural network (ISeGraph) for accurate segmentation. ISeGraph regards individual pixels as nodes within the graph and constructs self-adaptive edges according to multi-faceted knowledge, including visual similarity, positional correlation, and semantic consistency for feature aggregation. Specifically, visual similarity strengthens the connections between nodes sharing similar visual characteristics, while positional correlation assigns weights according to the spatial distance between nodes. In contrast to the above knowledge, semantic consistency maps nodes into a semantic space and learns pseudo-labels to define relationships based on label consistency. ISeGraph leverages multi-faceted knowledge to generate self-adaptive relationships for accurate iris segmentation. Furthermore, a pixel-wise adaptive normalization module is developed to increase the feature discriminability. It takes informative features in the shallow layer as a reference to improve the segmentation features from a statistical perspective. Experimental results on three iris datasets illustrate that the proposed method achieves superior performance in iris segmentation, increasing the segmentation accuracy in areas near the iris boundaries. Jianze Wei, Yunlong Wang 0003, Xingyu Gao 0001, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Sclera-TransFuse: Fusing Swin Transformer and CNN for Accurate Sclera SegmentationabstractSclera segmentation is a crucial step in sclera recognition, which has been greatly advanced by Convolutional Neural Networks (CNNs). However, when dealing with non-ideal eye images, many existing CNN-based approaches are still prone to failure. One major reason is that due to the limited range of receptive fields, CNNs are difficult to effectively model global semantic relevance and thus robustly resist noise interference. To solve this problem, this paper proposes a novel two-stream hybrid model, named Sclera-TransFuse, to integrate classical ResNet-34 and recently emerging Swin Transformer encoders. Specially, the self-attentive Swin Transformer has shown a strong ability in capturing long-range spatial dependencies and has a hierarchical structure similar to CNNs. The dual encoders firstly extract coarse- and fine-grained feature representations at hierarchical stages, separately. Then a novel Cross-Domain Fusion (CDF) module based on information interaction and self-attention mechanism is introduced to efficiently fuse the multi-scale features extracted from dual encoders. Finally, the fused features are progressively upsampled and aggregated to predict the sclera masks in the decoder meanwhile deep supervision strategies are employed to learn intermediate feature representations better and faster. Experimental results show that Sclera-TransFuse achieves state-of-the-art performance on various sclera segmentation benchmarks. Additionally, a UBIRIS.v2 subset of 683 eye images with manually labeled sclera masks, and our codes are publicly available to the community through https://github.com/Ihqqq/Sclera-TransFuse. Caiyong Wang, Guangzhe Zhao, Zhaofeng He 0001, Yunlong Wang 0003, Zhenan Sun |
IJCB | 5 |
| 2023 | Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological IntelligenceabstractAffective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP. Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun |
ACM Multimedia | 4 |
| 2023 | Multiscale Dynamic Graph Representation for Biometric Recognition With OcclusionsabstractOcclusion is a common problem with biometric recognition in the wild. The generalization ability of CNNs greatly decreases due to the adverse effects of various occlusions. To this end, we propose a novel unified framework integrating the merits of both CNNs and graph models to overcome occlusion problems in biometric recognition, called multiscale dynamic graph representation (MS-DGR). More specifically, a group of deep features reflected on certain subregions is recrafted into a feature graph (FG). Each node inside the FG is deemed to characterize a specific local region of the input sample, and the edges imply the co-occurrence of non-occluded regions. By analyzing the similarities of the node representations and measuring the topological structures stored in the adjacent matrix, the proposed framework leverages dynamic graph matching to judiciously discard the nodes corresponding to the occluded parts. The multiscale strategy is further incorporated to attain more diverse nodes representing regions of various sizes. Furthermore, the proposed framework exhibits a more illustrative and reasonable inference by showing the paired nodes. Extensive experiments demonstrate the superiority of the proposed framework, which boosts the accuracy in both natural and occlusion-simulated cases by a large margin compared with that of baseline methods. Yunlong Wang 0003, Yuhao Zhu 0003, Kunbo Zhang, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | AIF-LFNet: All-in-Focus Light Field Super-Resolution Method Considering the Depth-Varying DefocusabstractAs an aperture-divided computational imaging system, microlens array (MLA) -based light field (LF) imaging is playing an increasingly important role in computer vision. As the trade-off between the spatial and angular resolutions, deep learning (DL) -based image super-resolution (SR) methods have been applied to enhance the spatial resolution. However, in existing DL-based methods, the depth-varying defocus is not considered both in dataset development and algorithm design, which restricts many applications such as depth estimation and object recognition. To overcome this shortcoming, a super-resolution task that reconstructs all-in-focus high-resolution (HR) LF images from low-resolution (LR) LF images is proposed by designing a large dataset and proposing a convolutional neural network (CNN) -based SR method. The dataset is constructed by using Blender software, consisting of 150 light field images used as training data, and 15 light field images used as validation and testing data. The proposed network is designed by proposing the dilated deformable convolutional network (DCN) -based feature extraction block and the LF subaperture image (SAI) Deblur-SR block. The experimental results demonstrate that the proposed method achieves more appealing results both quantitatively and qualitatively. Shubo Zhou, Yunlong Wang 0003, Zhenan Sun, Kunbo Zhang, Xueqin Jiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | IrisGuideNet: Guided Localization and Segmentation Network for Unconstrained Iris BiometricsabstractIn recent years, unconstraint iris biometric is becoming more prevalent due to its wide range of user applications. But since it allows less user co-operation, it presents numerous challenges to the iris preprocessing task of localisation and segmentation (ILS). To address these challenges, many ILS techniques have been proposed with the deep learning CNN based approaches been the most effective. Training the CNN is data intensive and most of the existing CNN based ILS adopt general purpose CNN without any iris specific guidance. However, the available iris dataset comprises of small subsets with labelled images. As such, the existing CNN models can be less effective as they are trained with these dataset. Hence, in this paper, we propose a guided CNN based ILS technique termed IrisGuideNet by incorporating known iris specific heuristics into the network pipeline. IrisGuideNet has an encoder-decoder structure designed to be invariant to translation and rotation and can capture iris at multiple scales. To address the iris limited data problem, unlike the existing CNN based ILS, during the training process, we adopt the deep supervision technique, employ hybrid losses and introduce a novel iris specific heuristics named Iris Regularization Term (IRT) in other to effectively train the network. At inference, we introduce a novel Iris Infusion Module (IIM) that utilise the geometrical relationships between the ILS outputs to refine the predicted outputs through logical operations. Our models were trained and evaluated with the recently published NIR-ISL Challange * datasets and has proven to be effective as it has outperformed most of the participating models across all the database categories in the competition. Jawad Muhammad, Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Exploring Bias in Sclera Segmentation Models: A Group Evaluation ApproachabstractBias and fairness of biometric algorithms have been key topics of research in recent years, mainly due to the societal, legal and ethical implications of potentially unfair decisions made by automated decision-making models. A considerable amount of work has been done on this topic across different biometric modalities, aiming at better understanding the main sources of algorithmic bias or devising mitigation measures. In this work, we contribute to these efforts and present the first study investigating bias and fairness of sclera segmentation models. Although sclera segmentation techniques represent a key component of sclera-based biometric systems with a considerable impact on the overall recognition performance, the presence of different types of biases in sclera segmentation methods is still underexplored. To address this limitation, we describe the results of a group evaluation effort (involving seven research groups), organized to explore the performance of recent sclera segmentation models within a common experimental framework and study performance differences (and bias), originating from various demographic as well as environmental factors. Using five diverse datasets, we analyze seven independently developed sclera segmentation models in different experimental configurations. The results of our experiments suggest that there are significant differences in the overall segmentation performance across the seven models and that among the considered factors, ethnicity appears to be the biggest cause of bias. Additionally, we observe that training with representative and balanced data does not necessarily lead to less biased results. Finally, we find that in general there appears to be a negative correlation between the amount of bias observed (due to eye color, ethnicity and acquisition device) and the overall segmentation performance, suggesting that advances in the field of semantic segmentation may also help with mitigating bias. Matej Vitek, Abhijit Das 0001, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Jalil Nourmohammadi-Khiarak, Mohsen Akbari Shahpar, Meysam Asgari-Chenaghlu, Farhang Jaryani, Juan E. Tapia, Andres Valenzuela, Caiyong Wang, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Kiran B. Raja, Gourav Gupta, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, S. V. Aruna Kumar, B. S. Harish, Umapada Pal 0001, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 13 |
| 2023 | Contextual Measures for Iris RecognitionabstractThe iris patterns of the human contain a large amount of randomly distributed and irregularly shaped microstructures. These microstructures make the human iris informative biometric traits. To learn identity representation from them, this paper regards each iris region as a potential microstructure and proposes contextual measures (CM) to model the correlations between them. CM adopts two parallel branches to learn global and local contexts in iris image. The first one is the globally contextual measure branch. It measures the global context involving the relationships between all regions for feature aggregation and is robust to local occlusions. Besides, we improve its spatial perception considering the positional randomness of the microstructures. The other one is the locally contextual measure branch. This branch considers the role of local details in the phenotypic distinctiveness of iris patterns and learns a series of relationship atoms to capture contextual information from a local perspective. In addition, we develop the perturbation bottleneck to make sure that the two branches learn divergent contexts. It introduces perturbation to limit the information flow from input images to identity features, forcing CM to learn discriminative contextual information for iris recognition. Experimental results suggest that global and local contexts are two different clues critical for accurate iris recognition. The superior performance on four benchmark iris datasets demonstrates the effectiveness of the proposed approach in within-database and cross-database scenarios. Jianze Wei, Yunlong Wang 0003, Huaibo Huang, Ran He 0001, Zhenan Sun, Xingyu Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | PDVN: A Patch-based Dual-view Network for Face Liveness Detection using Light Field Focal StackabstractLight Field Focal Stack (LFFS) can be efficiently rendered from a light field (LF) image captured by plenoptic cameras. Differences in the 3D surface and texture of biometric samples are internally reflected in the defocus blur and local patterns between the rendered slices of LFFS. This unique property makes LFFS quite appropriate to differentiate presentation attack instruments (PAIs) from bona fide samples. A patch-based dual-view network (PDVN) is proposed in this paper to leverage the merits of LFFS for face presentation attack detection (PAD). First, original LFFS data are divided into various local patches along spatial dimensions, which distracts the model from learning the useless facial semantics and greatly relieve the problem of insufficient samples. The strategy of dual-view branches is innovatively proposed, wherein the original view and microscopic view can simultaneously contribute to liveness detection. Separable 3D convolution on the focal dimension is verified to be more effective than vanilla 3D convolution for extracting discriminative features from LFFS data. The voting mechanism on predictions of patch LFFS samples further strengthens the robustness of the proposed framework. PDVN is compared with other face PAD methods on IST LLFFSD dataset and achieves perfect performance, i.e., ACER drops to 0. Yunlong Wang 0003, Mupei Li, Zhengquan Luo, Zhenan Sun |
IJCB | 1 |
| 2022 | D-ESRGAN: A Dual-Encoder GAN with Residual CNN and Vision Transformer for Iris Image Super-ResolutionabstractIris images captured in less-constrained environments, especially at long distances often suffer from the interference of low resolution, resulting in the loss of much valid iris texture information for iris recognition. In this paper, we propose a dual-encoder super-resolution generative adversarial network (D-ESRGAN) for compensating texture lost of the raw image meanwhile maintaining the newly generated textures more natural. Specifically, the proposed D-ESRGAN not only integrates the residual CNN encoder to extract local features, but also employs an emerging vision transformer encoder to capture global associative information. The local and global features from two encoders are further fused for the subsequent reconstruction of high-resolution features. During the training, we develop a three-stage strategy to alleviate the problem that generative adversarial networks are prone to collapse. Moreover, to boost the iris recognition performance, we introduce a triplet loss to push away the distance of super-resolved iris images with different IDs, and pull the distance of super-resolved iris images with the same ID much closer. Experimental results on the public CASIA-Iris-distance and CASIA-Iris-M1 datasets show that D-ESRGAN archives better performance than state-of-the-art baselines in terms of both super-resolution image quality metrics and iris recognition metric. Caiyong Wang, Gaosheng Wu, Yunlong Wang 0003, Zhenan Sun |
IJCB | 4 |
| 2022 | Disentangled Federated Learning for Tackling Attributes Skew via Invariant Aggregation and Diversity TransferringabstractAttributes skew hinders the current federated learning (FL) frameworks from consistent optimization directions among the clients, which inevitably leads to performance reduction and unstable convergence. The core problems lie in that: 1) Domain-specific attributes, which are non-causal and only locally valid, are indeliberately mixed into global aggregation. 2) The one-stage optimizations of entangled attributes cannot simultaneously satisfy two conflicting objectives, i.e., generalization and personalization. To cope with these, we proposed disentangled federated learning (DFL) to disentangle the domain-specific and cross-invariant attributes into two complementary branches, which are trained by the proposed alternating local-global optimization independently. Importantly, convergence analysis proves that the FL system can be stably converged even if incomplete client models participate in the global aggregation, which greatly expands the application scope of FL. Extensive experiments verify that DFL facilitates FL with higher performance, better interpretability, and faster convergence rate, compared with SOTA FL methods on both manually synthesized and realistic attributes skew datasets. Zhengquan Luo, Yunlong Wang 0003, Zilei Wang, Zhenan Sun, Tieniu Tan |
ICML | 2 |
| 2022 | Cross-Spectral Iris Recognition by Learning Device-Specific BandabstractCross-spectral recognition is still an open challenge in iris recognition. In cross-spectral iris recognition, there exist distinct device-specific bands between near-infrared (NIR) and visible (VIS) images, resulting in the distribution gap between samples from different spectra and thus severe degradation in recognition performance. To tackle this problem, we propose a new cross-spectral iris recognition method to learn spectral-invariant features by estimating device-specific bands. In the proposed method,GaborTridentNetwork (GTN) first utilizes the Gabor function’s priors to perceive iris textures under different spectra, and then codes the device-specific band as the residual component to assist the generation of spectral-invariant features. By investigating the device-specific band, GTN effectively reduces the impact of device-specific bands on identity features. Besides, we make three efforts to further reduce the distribution gap. First,SpectralAdversarialNetwork (SAN) adopts a class-level adversarial strategy to align feature distributions. Second,Sample-Anchor (SA) loss upgrades triplet loss by pulling samples to their class center and pushing away from other class centers. Third, we develop a higher-order alignment loss to measures the distribution gap according to space bases and distribution shapes. Extensive experiments on five iris datasets demonstrate the efficacy of our proposed method for cross-spectral iris recognition. Jianze Wei, Yunlong Wang 0003, Yi Li 0018, Ran He 0001, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Perturbation Inactivation Based Adversarial Defense for Face RecognitionabstractDeep learning-based face recognition models are vulnerable to adversarial attacks. To curb these attacks, most defense methods aim to improve the robustness of recognition models against adversarial perturbations. However, the generalization capacities of these methods are quite limited. In practice, they are still vulnerable to unseen adversarial attacks. Deep learning models are fairly robust to general perturbations, such as Gaussian noises. A straightforward approach is to inactivate the adversarial perturbations so that they can be easily handled as general perturbations. In this paper, a plug-and-play adversarial defense method, named perturbation inactivation (PIN), is proposed to inactivate adversarial perturbations for adversarial defense. We discover that the perturbations in different subspaces have different influences on the recognition model. There should be a subspace, called the immune space, in which the perturbations have fewer adverse impacts on the recognition model than in other subspaces. Hence, our method estimates the immune space and inactivates the adversarial perturbations by restricting them to this subspace. The proposed method can be generalized to unseen adversarial perturbations since it does not rely on a specific kind of adversarial attack method. This approach not only outperforms several state-of-the-art adversarial defense methods but also demonstrates a superior generalization capacity through exhaustive experiments. Moreover, the proposed method can be successfully applied to four commercial APIs without additional training, indicating that it can be easily generalized to existing face recognition systems. The source code is available at https://github.com/RenMin1991/Perturbation-Inactivate.. Yuhao Zhu 0003, Yunlong Wang 0003, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Towards More Discriminative and Robust Iris Recognition by Learning Uncertain FactorsabstractThe uncontrollable acquisition process limits the performance of iris recognition. In the acquisition process, various inevitable factors, including eyes, devices, and environment, hinder the iris recognition system from learning a discriminative identity representation. This leads to severe performance degradation. In this paper, we explore uncertain acquisition factors and propose uncertainty embedding (UE) and uncertainty-guided curriculum learning (UGCL) to mitigate the influence of acquisition factors. UE represents an iris image using a probabilistic distribution rather than a deterministic point (binary template or feature vector) that is widely adopted in iris recognition methods. Specifically, UE learns identity and uncertainty features from the input image, and encodes them as two independent components of the distribution, mean and variance. Based on this representation, an input image can be regarded as an instantiated feature sampled from the UE, and we can also generate various virtual features through sampling. UGCL is constructed by imitating the progressive learning process of newborns. Particularly, it selects virtual features to train the model in an easy-to-hard order at different training stages according to their uncertainty. In addition, an instance-level enhancement method is developed by utilizing local and global statistics to mitigate the data uncertainty from image noise and acquisition conditions in the pixel-level space. The experimental results on six benchmark iris datasets verify the effectiveness and generalization ability of the proposed method on same-sensor and cross-sensor recognition. Jianze Wei, Huaibo Huang, Yunlong Wang 0003, Ran He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | A Large-scale Database for Less Cooperative Iris RecognitionabstractSince the outbreak of the COVID-19 pandemic, iris recognition has been used increasingly as contactless and unaffected by face masks. Although less user cooperation is an urgent demand for existing systems, corresponding manually annotated databases could hardly be obtained. This paper presents a large-scale database of near-infrared iris images named CASIA-Iris-Degradation Version 1.0 (DV1), which consists of 15 subsets of various degraded images, simulating less cooperative situations such as illumination, off-angle, occlusion, and nonideal eye state. A lot of open-source segmentation and recognition methods are compared comprehensively on the DV1 using multiple evaluations, and the best among them are exploited to conduct ablation studies on each subset. Experimental results show that even the best deep learning frameworks are not robust enough on the database, and further improvements are recommended for challenging factors such as half-open eyes, off-angle, and pupil dilation. Therefore, we publish the DV1 with manual annotations online to promote iris recognition. (http://www.cripacsir.cn/dataset/) Junxing Hu, Leyuan Wang, Zhengquan Luo, Yunlong Wang 0003, Zhenan Sun |
IJCB | 4 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 2 |
| 2021 | An End-to-End Autofocus Camera for Iris on the MoveabstractFor distant iris recognition, a long focal length lens is generally used to ensure the resolution of iris images, which reduces the depth of field and leads to potential defocus blur. To accommodate users standing statically at different distances, it is necessary to control focus quickly and accurately. And for users in motion, it is also expected to acquire a sufficient amount of accurately focused iris images. In this paper, we introduced a novel rapid auto-focus camera for active refocusing of the iris area of the moving objects with a focus-tunable lens. Our end-to-end computational algorithm can predict the best focus position from one single blurred image and generate the proper lens diopter control signal automatically. This scene-based active manipulation method enables real-time focus tracking of the iris area of a moving object. We built a testing bench to collect real-world focal stacks for evaluation of the autofocus methods. Our camera has reached an autofocus speed of over 50 fps. The results demonstrate the advantages of our proposed camera for biometric perception in static and dynamic scenes. The code is available at https://github.com/Debatrix/AquulaCam. Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun |
IJCB | 3 |
| 2021 | Learning Instance-level Spatial-Temporal Patterns for Person Re-identificationabstractPerson re-identification (Re-ID) aims to match pedestrians under dis-joint cameras. Most Re-ID methods formulate it as visual representation learning and image search, and its accuracy is consequently affected greatly by the search space. Spatial-temporal information has been proven to be efficient to filter irrelevant negative samples and significantly improve Re-ID accuracy. However, existing spatial-temporal person Re-ID methods are still rough and do not exploit spatial-temporal information sufficiently. In this paper, we propose a novel Instance-level and Spatial-Temporal Disentangled Re-ID method (InSTD), to improve Re-ID accuracy. In our proposed framework, personalized information such as moving direction is explicitly considered to further narrow down the search space. Besides, the spatial-temporal transferring probability is disentangled from joint distribution to marginal distribution, so that outliers can also be well modeled. Abundant experimental analyses are presented, which demonstrates the superiority and provides more insights into our method. The proposed method achieves mAP of 90.8% on Market-1501 and 89.1% on DukeMTMC-reID, improving from the baseline 82.2% and 72.7%, respectively. Besides, in order to provide a better benchmark for person re-identification, we release a cleaned data list of DukeMTMC-reID with this paper: https://github.com/RenMin1991/cleaned-DukeMTMC-reID/ Lingxiao He, Xingyu Liao, Wu Liu 0005, Yunlong Wang 0003, Tieniu Tan |
ICCV | 5 |
| 2021 | CASIA-Face-Africa: A Large-Scale African Face Image DatabaseabstractFace recognition is a popular and well-studied area with wide applications in our society. However, racial bias had been proven to be inherent in most State Of The Art (SOTA) face recognition systems. Many investigative studies on face recognition algorithms have reported higher false positive rates of African subjects cohorts than the other cohorts. Lack of large-scale African face image databases in public domain is one of the main restrictions in studying the racial bias problem of face recognition. To this end, we collect a face image database namely CASIA-Face-Africa which contains 38,546 images of 1,183 African subjects. Multi-spectral cameras are utilized to capture the face images under various illumination settings. Demographic attributes and facial expressions of the subjects are also carefully recorded. For landmark detection, each face image in the database is manually labeled with 68 facial keypoints. A group of evaluation protocols are constructed according to different applications, tasks, partitions and scenarios. The performances of SOTA face recognition algorithms without re-training are reported as baselines. The proposed database along with its face landmark annotations, evaluation protocols and preliminary results form a good benchmark to study the essential aspects of face biometrics for African subjects, especially face image preprocessing, face feature analysis and matching, facial expression recognition, sex/age estimation, ethnic classification, face image generation, etc. The database can be downloaded from our website. Jawad Muhammad, Yunlong Wang 0003, Caiyong Wang, Kunbo Zhang, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Dynamic Graph Representation for Occlusion Handling in BiometricsabstractThe generalization ability of Convolutional neural networks (CNNs) for biometrics drops greatly due to the adverse effects of various occlusions. To this end, we propose a novel unified framework integrated the merits of both CNNs and graphical models to learn dynamic graph representations for occlusion problems in biometrics, called Dynamic Graph Representation (DGR). Convolutional features onto certain regions are re-crafted by a graph generator to establish the connections among the spatial parts of biometrics and build Feature Graphs based on these node representations. Each node of Feature Graphs corresponds to a specific part of the input image and the edges express the spatial relationships between parts. By analyzing the similarities between the nodes, the framework is able to adaptively remove the nodes representing the occluded parts. During dynamic graph matching, we propose a novel strategy to measure the distances of both nodes and adjacent matrixes. In this way, the proposed method is more convincing than CNNs-based methods because the dynamic graph method implies a more illustrative and reasonable inference of the biometrics decision. Experiments conducted on iris and face demonstrate the superiority of the proposed framework, which boosts the accuracy of occluded biometrics recognition by a large margin comparing with baseline methods. Yunlong Wang 0003, Zhenan Sun, Tieniu Tan |
AAAI | 2 |
| 2020 | A Lightweight Multi-Label Segmentation Network for Mobile Iris BiometricsabstractThis paper proposes a novel, lightweight deep convolutional neural network specifically designed for iris segmentation of noisy images acquired by mobile devices. Unlike previous studies, which only focused on improving the accuracy of segmentation mask using the popular CNN technology, our method is a complete end-to-end iris segmentation solution, i.e., segmentation mask and parameterized pupillary and limbic boundaries of the iris are obtained simultaneously, which further enables CNN-based iris segmentation to be applied in any regular iris recognition systems. By introducing an intermediate pictorial boundary representation, predictions of iris boundaries and segmentation mask have collectively formed a multi-label semantic segmentation problem, which could be well solved by a carefully adapted stacked hourglass network. Experimental results show that our method achieves competitive or state-of-the-art performance in both iris segmentation and localization on two challenging mobile iris databases. Caiyong Wang, Yunlong Wang 0003, Boqiang Xu, Yong He 0009, Zhiwei Dong, Zhenan Sun |
ICASSP | 2 |
| 2020 | SSBC 2020: Sclera Segmentation Benchmarking Competition in the Mobile EnvironmentabstractThe paper presents a summary of the 2020 Sclera Segmentation Benchmarking Competition (SSBC), the 7th in the series of group benchmarking efforts centred around the problem of sclera segmentation. Different from previous editions, the goal of SSBC 2020 was to evaluate the performance of sclera-segmentation models on images captured with mobile devices. The competition was used as a platform to assess the sensitivity of existing models to i) differences in mobile devices used for image capture and ii) changes in the ambient acquisition conditions. 26 research groups registered for SSBC 2020, out of which 13 took part in the final round and submitted a total of 16 segmentation models for scoring. These included a wide variety of deep-learning solutions as well as one approach based on standard image processing techniques. Experiments were conducted with three recent datasets. Most of the segmentation models achieved relatively consistent performance across images captured with different mobile devices (with slight differences across devices), but struggled most with low-quality images captured in challenging ambient conditions, i.e., in an indoor environment and with poor lighting. Matej Vitek, Abhijit Das 0001, Yann Pourcenoux, Alexandre Missler, C. Paumier, Sumanta Das, Ishita De Ghosh, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Junxing Hu, Yong He 0009, Caiyong Wang, Yunlong Wang 0003, Zhenan Sun, Dailé Osorio Roig, Christian Rathgeb, Christoph Busch 0001, Juan E. Tapia, Andres Valenzuela, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, Sabari Nathan, R. Suganya 0001, Vineet Mehta, Abhinav Dhall, Kiran B. Raja, Gourav Gupta, Jalil Nourmohammadi-Khiarak, Mohsen Akbari-Shahper, Farhang Jaryani, Meysam Asgari-Chenaghlu, Ritesh Vyas, Sristi Dakshit, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 19 |
| 2020 | Recognition Oriented Iris Image Quality Assessment in the Feature SpaceabstractA large portion of iris images captured in real world scenarios are poor quality due to the uncontrolled environment and the non-cooperative subject. To ensure that the recognition algorithm is not affected by low-quality images, traditional hand-crafted factors based methods discard most images, which will cause system timeout and disrupt user experience. In this paper, we propose a recognition-oriented quality metric and assessment method for iris image to deal with the problem. The method regards the iris image em-beddings Distance in Feature Space (DFS) as the quality metric and the prediction is based on deep neural networks with the attention mechanism. The quality metric proposed in this paper can significantly improve the performance of the recognition algorithm while reducing the number of images discarded for recognition, which is advantageous over hand-crafted factors based iris quality assessment methods. The relationship between Image Rejection Rate (IRR) and Equal Error Rate (EER) is proposed to evaluate the performance of the quality assessment algorithm under the same image quality distribution and the same recognition algorithm. Compared with hand-crafted factors based methods, the proposed method is a trial to bridge the gap between the image quality assessment and biometric recognition. Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun |
IJCB | 4 |
| 2020 | All-in-Focus Iris Camera With a Great Capture VolumeabstractImaging volume of an iris recognition system has been restricting the throughput and cooperation convenience in biometric applications. Numerous improvement trials are still impractical to supersede the dominant fixed-focus lens in stand-off iris recognition due to incremental performance increase and complicated optical design. In this study, we develop a novel all-in-focus iris imaging system using a focus-tunable lens and a 2D steering mirror to greatly extend capture volume by spatiotemporal multiplexing method. Our iris imaging depth of field extension system requires no mechanical motion and is capable to adjust the focal plane at extremely high speed. In addition, the motorized reflection mirror adaptively steers the light beam to extend the horizontal and vertical field of views in an active manner. The proposed all-in-focus iris camera increases the depth of field up to 3.9 m which is afactor of 37.5 compared with conventional long focal lens. We also experimentally demonstrate the capability of this 3D light beam steering imaging system in real-time multi-person iris refocusing using dynamic focal stacks and the potential of continuous iris recognition for moving participants. Kunbo Zhang, Zhenteng Shen, Yunlong Wang 0003, Zhenan Sun |
IJCB | 3 |
| 2020 | A Novel Deep-learning Pipeline for Light Field Image Based Material RecognitionabstractThe primitive basis of image based material recognition builds upon the fact that discrepancies in the reflectances of distinct materials lead to imaging differences under multiple viewpoints. LF cameras possess coherent abilities to capture multiple sub-aperture views (SAIs) within one exposure, which can provide appropriate multi-view sources for material recognition. In this paper, a unified “Factorize-Connect-Merge” (FCM) deep-learning pipeline is proposed to solve problems of light field image based material recognition. 4D light-field data as input is initially decomposed into consecutive 3D light-field slices. Shallow CNN is leveraged to extract low-level visual features of each view inside these slices. As to establish correspondences between these SAIs, Bidirectional Long-Short Term Memory (Bi-LSTM) network is built upon these low-level features to model the imaging differences. After feature selection including concatenation and dimension reduction, effective and robust feature representations for material recognition can be extracted from 4D light-field data. Experimental results indicate that the proposed pipeline can obtain remarkable performances on both tasks of single-pixel material classification and full-image material segmentation. In addition, the proposed pipeline can potentially benefit and inspire other researchers who may also take LF images as input and need to extract 4D light-field representations for computer vision tasks such as object classification, semantic segmentation and edge detection. Yunlong Wang 0003, Kunbo Zhang, Zhenan Sun |
ICPR | 1 |
| 2020 | Towards Complete and Accurate Iris Segmentation Using Deep Multi-Task Attention Network for Non-Cooperative Iris RecognitionabstractIris images captured in non-cooperative environments often suffer from adverse noise, which challenges many existing iris segmentation methods. To address this problem, this paper proposes a high-efficiency deep learning based iris segmentation approach, named IrisParseNet. Different from many previous CNN-based iris segmentation methods, which only focus on predicting accurate iris masks by following popular semantic segmentation frameworks, the proposed approach is a complete iris segmentation solution, i.e., iris mask and parameterized inner and outer iris boundaries are jointly achieved by actively modeling them into a unified multi-task network. Moreover, an elaborately designed attention module is incorporated into it to improve the segmentation performance. To train and evaluate the proposed approach, we manually label three representative and challenging iris databases, i.e., CASIA.v4-distance, UBIRIS.v2, and MICHE-I, which involve multiple illumination (NIR, VIS) and imaging sensors (long-range and mobile iris cameras), along with various types of noises. Additionally, several unified evaluation protocols are built for fair comparisons. Extensive experiments are conducted on these newly annotated databases, and results show that the proposed approach achieves state-of-the-art performance on various benchmarks. Further, as a general drop-in replacement, the proposed iris segmentation method can be used for any iris recognition methodology, and would significantly improve the performance of non-cooperative iris recognition. Caiyong Wang, Jawad Muhammad, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Binocular Light-Field: Imaging Theory and Occlusion-Robust Depth Perception ApplicationabstractBinocular stereo vision (SV) has been widely used to reconstruct the depth information, but it is quite vulnerable to scenes with strong occlusions. As an emerging computational photography technology, light-field (LF) imaging brings about a novel solution to passive depth perception by recording multiple angular views in a single exposure. In this paper, we explore binocular SV and LF imaging to form the binocular-LF imaging system. An imaging theory is derived by modeling the imaging process and analyzing disparity properties based on the geometrical optics theory. Then an accurate occlusion-robust depth estimation algorithm is proposed by exploiting multibaseline stereo matching cues and defocus cues. The occlusions caused by binocular SV and LF imaging are detected and handled to eliminate the matching ambiguities and outliers. Finally, we develop a binocular-LF database and capture realworld scenes by our binocular-LF system to test the accuracy and robustness. The experimental results demonstrate that the proposed algorithm definitely recovers high quality depth maps with smooth surfaces and precise geometric shapes, which tackles the drawbacks of binocular SV and LF imaging simultaneously. Fei Liu 0031, Shubo Zhou, Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 3 |
| 2018 | End-to-End View Synthesis for Light Field Imaging with Pseudo 4DCNN
Yunlong Wang 0003, Fei Liu 0031, Zilei Wang, Guangqi Hou, Zhenan Sun, Tieniu Tan |
ECCV (2) | 1 |
| 2018 | LFNet: A Novel Bidirectional Recurrent Convolutional Neural Network for Light-Field Image Super-ResolutionabstractThe low spatial resolution of light-field image poses significant difficulties in exploiting its advantage. To mitigate the dependency of accurate depth or disparity information as priors for light-field image super-resolution, we propose an implicitly multi-scale fusion scheme to accumulate contextual information from multiple scales for super-resolution reconstruction. The implicitly multi-scale fusion scheme is then incorporated into bidirectional recurrent convolutional neural network, which aims to iteratively model spatial relations between horizontally or vertically adjacent sub-aperture images of light-field data. Within the network, the recurrent convolutions are modified to be more effective and flexible in modeling the spatial correlations between neighboring views. A horizontal sub-network and a vertical sub-network of the same network structure are ensembled for final outputs via stacked generalization. Experimental results on synthetic and real-world data sets demonstrate that the proposed method outperforms other state-of-the-art methods by a large margin in peak signal-to-noise ratio and gray-scale structural similarity indexes, which also achieves superior quality for human visual systems. Furthermore, the proposed method can enhance the performance of light field applications such as depth estimation. Yunlong Wang 0003, Fei Liu 0031, Kunbo Zhang, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 1 |
| 2016 | A simple and robust super resolution method for light field imagesabstractLight field cameras generate low-resolution images due to the tradeoff between spatial and angular resolution. Traditional light field super-resolution (LFSR) methods depend on prior knowledge of depth information. This paper presents a projection-based LFSR solution without prior information based on redefinition of the mapping function between disparity and shearing shift. Moreover, simplified variational regularization is imposed in global optimization formulation to the rendered high-resolution images. Both a synthetic dataset and a real-world dataset of light field images captured by a self-developed light field camera are used to demonstrate the state-of-the-art performance of the proposed method. Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Zilei Wang, Tieniu Tan |
ICIP | 1 |