Kunbo Zhang

dblp:222/3947 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 10 since 2021Security and privacy · 13 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CLower: Detecting Compiler Pessimization Bugs through Redundant Memory Accesses
abstract
Compilers are expected to generate optimized code, but they sometimes introduce pessimizations, quality-degrading redundant instructions. These bugs not only incur performance overhead but also, critically, expand the attack surface by introducing unexpected side effects (e.g., redundant memory accesses) without breaking compilation correctness. Existing bug-finding methods are neither designed for nor effective at identifying such security-sensitive pessimizations. This paper presents CLower, a novel, black-box approach for automatically detecting compiler pessimizations via redundant memory accesses. CLower’s core insight is that any extra global memory accesses in a fully optimized binary, compared to the source, indicate a pessimization. To reliably distinguish compiler-introduced redundancy from source-level redundancy, we generate random C programs in which each global variable has a predetermined, controlled number of memory accesses. CLower then executes the instrumented binary and verifies whether superfluous accesses have been introduced during compilation. We applied CLower to GCC and LLVM, reporting 23 unique bugs (21 in GCC, 2 in Clang), with 16 confirmed as new pessimization bugs. Our evaluation shows that CLower accurately detects diverse, impactful pes-simization bugs, the majority of which (75%) also manifest for heap-allocated objects, demonstrating that the underlying compiler flaws are general and not limited to global memory. Furthermore, we identify a systematic conflict between compiler optimizations and pessimization bugs, which causes many such bugs to remain hidden in compiler versions. This study sheds light on the under-explored area of compiler pessimization and provides a practical tool for improving compiler quality.
Jianhao Xu, Kunbo Zhang, Mathias Payer, Kangjie Lu, Bing Mao 0001
Proc. ACM Program. Lang.2
2025 BiommWave: A Non-Visual Approach for Biometric Recognition Using Millimeter-Wave Radar
abstract
This paper explores the application of millimeter-wave (mmWave) radar in biometric recognition. As a non-visual human sensing technology, mmWave radar captures reflection properties and micro-movements, providing a complementary modality to visual appearance. We implemented a complete system pipeline to utilize mmWave sensing for individual recognition. Based on physiological mechanisms, we design preprocessing methods to extract intuitive biometric feature maps, concerning body reflections, cardiopulmonary activity, and micro-motion frequencies. To address the data uncertainty, a dynamic pole-based learning strategy is proposed to construct compact and discriminated feature distributions. In real-world evaluations, the system achieves 95.12% accuracy and an Equal Error Rate (EER) of 1.96%. This work leverages the advantages of mmWave radar for flexible, unconstrained, and private biometric systems. From a non-visual sensing perspective, it explores novel modalities as unique biometric cues, demonstrating significant value of research and applications.
Mupei Li, Yunlong Wang 0003, Yiwei Ru, Kunbo Zhang, Zhenan Sun
IJCB4
2025 Hierarchical Emotion-Guided Masked Transformer for Long-Sequence Co-Speech Gestures with Partial Supervision
abstract
This work proposes a method for co-speech gesture generation that produces natural, emotionally expressive body movements synchronized with speech. Existing methods typically assume complete motion supervision and static emotional states, limiting training data leverage and impairing fine-grained emotion transitions, ultimately leading to incoherent long-sequence gesture generation. In contrast, we propose a Hierarchical Emotion-Guided Masked Transformer that tackles these limitations. We introduce a Mask-Based Motion Modeling Strategy in a discrete latent space, enabling learning from partially annotated data and ensuring physically plausible motions under incomplete supervision, alongside a hierarchical Emotion Guidance Adaptor that injects time-varying emotional cues at multiple transformer levels, capturing both global emotion and subtle local nuances. An Alternating Optimization Mechanism is employed to decouple semantic alignment from emotion modulation during training, stabilizing learning and improving expression fidelity. At inference, a Mask-Guided Inference strategy seamlessly extends gestures over long sequences, mitigating boundary discontinuities and drift. Evaluations on a partially annotated dataset featuring natural emotional transitions show our method surpasses existing approaches in long-sequence co-speech gesture generation, yielding improved gestures-peech synchrony, enhanced motion diversity, and more faithful emotion-conveying ability.
Kunbo Zhang, Zhenan Sun
IJCB4
2025 Correction: Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan
Int. J. Comput. Vis.4
2025 Exploring Near-Infrared Iris Image Sequences for High Throughput Iris Recognition
abstract
High throughput is demanding in real-world iris recognition applications. The challenges mainly originate from the variability in image quality under high-throughput capture conditions. Most of the degraded images are typically filtered out by traditional iris systems through Image Quality Assessment (IQA) module, adversely affecting efficiency and leading to low throughput and poor user experience. Therefore, a better and practical solution is to make the utmost of degraded iris images. In order to investigate the key problems of high-throughput iris recognition, we collect a novel iris sequence dataset under Near-infrared (NIR) illumination. This dataset is specifically constructed for high-throughput evaluation, which faithfully simulates the process of iris sequence acquisition in real-world iris systems. Comprehensive evaluations were conducted to figure out the deficiencies of current iris recognition algorithms. To this end, a testing methodology along with specific evaluation metrics is proposed. It is capable of assessing the throughput performance, e.g., the newly proposed Frame Consumption per Match (FCM). Through performance analysis, several insights were gathered to guide potential directions for developing high-throughput iris recognition algorithms. Furthermore, we consider to leverage iris sequence features for better throughput performance. Continuity sequence criteria and cumulative sequence feature strategy are proposed to enhance the throughput performance of existing algorithms with minimal cost. In summary, this work provides valuable data and rational insights for high-throughput iris recognition studies. The datasets and evaluation toolkit are publicly available on our website1.
Mupei Li, Yunlong Wang 0003, Kunbo Zhang, Zhaofeng He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2025 Cross-Optical Property Image Translation for Face Anti-Spoofing: From Visible to Polarization
abstract
Despite the development of spectral sensors and spectral data-driven learning methods which have led to significant advances in face anti-spoofing (FAS), the singular dimensionality of spectral information often results in poor robustness and weak generalization. Polarization, another fundamental property of light, can reveal intrinsic differences between genuine and fake faces with advantaged performance in precision, robustness, and generalizability. In this paper, we propose a facial image translation method from visible light (VIS) to polarization (VPT), capable of generating valuable polarimetric optical characteristics for facial presentation attack detection using VIS spectrum information input only. Specifically, the VPT method adopts a multi-stream network structure, comprising a main network and two branch networks, to translate VIS images into degree of polarization (DoP) images and Stokes polarization parameters${S}_{1}$and${S}_{2}$. To further improve image translation quality, we introduce a frequency-domain consistency loss as a complement to the existing spatial losses to narrow the gap in the frequency domain. The physical mapping relations for the DoP and Stokes parameters are employed, and the Stokes loss is designed to ensure that the generated polarization modalities conform to objective physical laws. Extensive experiments on the CASIA-Polar and CASIA-SURF datasets demonstrate the superiority of VPT over other baseline methods in terms of polarization image quality and its remarkable performance in the FAS task. This work leverages the inherent physical advantages of polarization information in material discrimination tasks while addressing hardware limitations in polarization image collection, proposing a novel solution for face recognition system security control.
Yu Tian 0017, Kunbo Zhang, Yalin Huang, Leyuan Wang, Yue Liu 0005, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.2
2024 EditHuman: Fine-Grained Text-Driven Human Video Editing
abstract
Recently, video editing has made significant advances. Human character, as one of the core elements in video editing, has attracted great research attention. However, when editing characters with strong structural information, previous methods generally encounter blurring and distortion in the limbs. In this paper, we present EditHuman, a model to realize fine-grained text-driven human video editing tasks, which achieves continuous pose movements and high-quality limb expression. Considering complex body structures and continuity of motion, more precise designs are needed to obtain practical performance. Specifically, we propose a Cascaded UNet (CAU) to realize a coarse-to-fine denoising process and refined noise estimation. Meanwhile, we introduce two Heatmap-Centric Attention Modules called Key-Element Attention (KEA) and Key-Temporal Attention (KTA) to enhance the quality of human limb expression and inter-frame continuity. Moreover, we utilize the estimated heatmap to guide the noise prediction, which further refines the video quality. Extensive experiments show that EditHuman has achieved the SOTA performance.
Kaiduo Zhang, Muyi Sun, Junxing Hu, Kunbo Zhang, Zhenan Sun
IJCB4
2024 A comprehensive research on light field imaging: Theory and application
abstract
Abstract Computational photography is a combination of novel optical designs and processing methods to capture high‐dimensional visual information. As an emerged promising technique, light field (LF) imaging measures the lighting, reflectance, focus, geometry and viewpoint in the free space, which has been widely explored for depth estimation, view synthesis, refocus, rendering, 3D displays, microscopy and other applications in computer vision in the past decades. In this paper, the authors present a comprehensive research survey on the LF imaging theory, technology and application. Firstly, the LF imaging process based on a MicroLens Array structure is derived, that is MLA‐LF. Subsequently, the innovations of LF imaging technology are presented in terms of the imaging prototype, consumer LF camera and LF displays in Virtual Reality (VR) and Augmented Reality (AR). Finally the applications and challenges of LF imaging integrating with deep learning models are analysed, which consist of depth estimation, saliency detection, semantic segmentation, de‐occlusion and defocus deblurring in recent years. It is believed that this paper will be a good reference for the future research on LF imaging technology in Artificial Intelligence era.
Fei Liu 0031, Yunlong Wang 0003, Shubo Zhou, Kunbo Zhang
IET Comput. Vis.5
2024 Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan
Int. J. Comput. Vis.4
2023 Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological Intelligence
abstract
Affective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP.
Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun
ACM Multimedia5
2023 Multiscale Dynamic Graph Representation for Biometric Recognition With Occlusions
abstract
Occlusion is a common problem with biometric recognition in the wild. The generalization ability of CNNs greatly decreases due to the adverse effects of various occlusions. To this end, we propose a novel unified framework integrating the merits of both CNNs and graph models to overcome occlusion problems in biometric recognition, called multiscale dynamic graph representation (MS-DGR). More specifically, a group of deep features reflected on certain subregions is recrafted into a feature graph (FG). Each node inside the FG is deemed to characterize a specific local region of the input sample, and the edges imply the co-occurrence of non-occluded regions. By analyzing the similarities of the node representations and measuring the topological structures stored in the adjacent matrix, the proposed framework leverages dynamic graph matching to judiciously discard the nodes corresponding to the occluded parts. The multiscale strategy is further incorporated to attain more diverse nodes representing regions of various sizes. Furthermore, the proposed framework exhibits a more illustrative and reasonable inference by showing the paired nodes. Extensive experiments demonstrate the superiority of the proposed framework, which boosts the accuracy in both natural and occlusion-simulated cases by a large margin compared with that of baseline methods.
Yunlong Wang 0003, Yuhao Zhu 0003, Kunbo Zhang, Zhenan Sun
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 AIF-LFNet: All-in-Focus Light Field Super-Resolution Method Considering the Depth-Varying Defocus
abstract
As an aperture-divided computational imaging system, microlens array (MLA) -based light field (LF) imaging is playing an increasingly important role in computer vision. As the trade-off between the spatial and angular resolutions, deep learning (DL) -based image super-resolution (SR) methods have been applied to enhance the spatial resolution. However, in existing DL-based methods, the depth-varying defocus is not considered both in dataset development and algorithm design, which restricts many applications such as depth estimation and object recognition. To overcome this shortcoming, a super-resolution task that reconstructs all-in-focus high-resolution (HR) LF images from low-resolution (LR) LF images is proposed by designing a large dataset and proposing a convolutional neural network (CNN) -based SR method. The dataset is constructed by using Blender software, consisting of 150 light field images used as training data, and 15 light field images used as validation and testing data. The proposed network is designed by proposing the dilated deformable convolutional network (DCN) -based feature extraction block and the LF subaperture image (SAI) Deblur-SR block. The experimental results demonstrate that the proposed method achieves more appealing results both quantitatively and qualitatively.
Shubo Zhou, Yunlong Wang 0003, Zhenan Sun, Kunbo Zhang, Xueqin Jiang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 IrisGuideNet: Guided Localization and Segmentation Network for Unconstrained Iris Biometrics
abstract
In recent years, unconstraint iris biometric is becoming more prevalent due to its wide range of user applications. But since it allows less user co-operation, it presents numerous challenges to the iris preprocessing task of localisation and segmentation (ILS). To address these challenges, many ILS techniques have been proposed with the deep learning CNN based approaches been the most effective. Training the CNN is data intensive and most of the existing CNN based ILS adopt general purpose CNN without any iris specific guidance. However, the available iris dataset comprises of small subsets with labelled images. As such, the existing CNN models can be less effective as they are trained with these dataset. Hence, in this paper, we propose a guided CNN based ILS technique termed IrisGuideNet by incorporating known iris specific heuristics into the network pipeline. IrisGuideNet has an encoder-decoder structure designed to be invariant to translation and rotation and can capture iris at multiple scales. To address the iris limited data problem, unlike the existing CNN based ILS, during the training process, we adopt the deep supervision technique, employ hybrid losses and introduce a novel iris specific heuristics named Iris Regularization Term (IRT) in other to effectively train the network. At inference, we introduce a novel Iris Infusion Module (IIM) that utilise the geometrical relationships between the ILS outputs to refine the predicted outputs through logical operations. Our models were trained and evaluated with the recently published NIR-ISL Challange * datasets and has proven to be effective as it has outperformed most of the participating models across all the database categories in the competition.
Jawad Muhammad, Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2023 Polarized Image Translation From Nonpolarized Cameras for Multimodal Face Anti-Spoofing
abstract
In face antispoofing, it is desirable to have multimodal images to demonstrate liveness cues from various perspectives. However, in most face recognition scenarios, only a single modality, namely visible lighting (VIS) facial images is available. This paper first investigates the possibility of generating polarized (Polar) images from VIS cameras without changing the existing recognition devices to improve the accuracy and robustness of Presentation Attack Detection (PAD) in face biometrics. A novel multimodal face antispoofing framework is proposed based on the machine-learning relationship between VIS and Polar images of genuine faces. Specifically, a dual-modal central differential convolutional network (CDCN) is developed to capture the inherent spoofing features between the VIS and the generated Polar modalities. Quantitative and qualitative experimental results show that our proposed framework not only generates realistic Polar face images but also improves the state-of-the-art face anti-spoofing results on the VIS modal database (i.e. CASIA-SURF). Moreover, a polar face database, CASIA-Polar, has been constructed and will be shared with the public at http://biometrics.idealtest.org to inspire future applications within the biometric anti-spoofing field.
Yu Tian 0017, Yalin Huang, Kunbo Zhang, Yue Liu 0005, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2023 Pose-Appearance Relational Modeling for Video Action Recognition
abstract
Recent studies of video action recognition can be classified into two categories: the appearance-based methods and the pose-based methods. The appearance-based methods generally cannot model temporal dynamics of large motion well by virtue of optical flow estimation, while the pose-based methods ignore the visual context information such as typical scenes and objects, which are also important cues for action understanding. In this paper, we tackle these problems by proposing a Pose-Appearance Relational Network (PARNet), which models the correlation between human pose and image appearance, and combines the benefits of these two modalities to improve the robustness towards unconstrained real-world videos. There are three network streams in our model, namely pose stream, appearance stream and relation stream. For the pose stream, a Temporal Multi-Pose RNN module is constructed to obtain the dynamic representations through temporal modeling of 2D poses. For the appearance stream, a Spatial Appearance CNN module is employed to extract the global appearance representation of the video sequence. For the relation stream, a Pose-Aware RNN module is built to connect pose and appearance streams by modeling action-sensitive visual context information. Through jointly optimizing the three modules, PARNet achieves superior performances compared with the state-of-the-arts on both the pose-complete datasets (KTH, Penn-Action, UCF11) and the challenging pose-incomplete datasets (UCF101, HMDB51, JHMDB), demonstrating its robustness towards complex environments and noisy skeletons. Its effectiveness on NTU-RGBD dataset is also validated even compared with 3D skeleton-based methods. Furthermore, an appearance-enhanced PARNet equipped with a RGB-based I3D stream is proposed, which outperforms the Kinetics pre-trained competitors on UCF101 and HMDB51. The better experimental results verify the potentials of our framework by integrating various modules.
Mengmeng Cui, Wei Wang 0115, Kunbo Zhang, Zhenan Sun, Liang Wang 0001
IEEE Trans. Image Process.3
2021 Avoiding Spectacles Reflections on Iris Images Using A Ray-tracing Method
abstract
Spectacles reflection removal is a challenging problem in iris recognition research. The reflection of the spectacles usually contaminates the iris image acquired under infrared illumination. The intense light reflection caused by the active light source makes reflection removal more challenging than normal scenes since important iris texture features are entirely obscured. Eliminating unnecessary reflections can effectively improve iris recognition system performance. This paper proposes a spectacle reflection removal algorithm based on ray coding and ray tracking to remove spectacle reflection in iris images. By decoding the light source’s encoded light beam, the iris imaging device eliminates most of the stray light. Our binocular imaging device tracks the light path to obtain parallax information and realizes reflected light spot removal through image fusion. We designed a prototype system to verify our proposed method in this paper. This method can effectively eliminate reflections without changing iris texture and improve iris recognition in complex scenarios.
Kunbo Zhang, Leyuan Wang
IJCB2
2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization
abstract
For iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research.
Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad
IJCB3
2021 An End-to-End Autofocus Camera for Iris on the Move
abstract
For distant iris recognition, a long focal length lens is generally used to ensure the resolution of iris images, which reduces the depth of field and leads to potential defocus blur. To accommodate users standing statically at different distances, it is necessary to control focus quickly and accurately. And for users in motion, it is also expected to acquire a sufficient amount of accurately focused iris images. In this paper, we introduced a novel rapid auto-focus camera for active refocusing of the iris area of the moving objects with a focus-tunable lens. Our end-to-end computational algorithm can predict the best focus position from one single blurred image and generate the proper lens diopter control signal automatically. This scene-based active manipulation method enables real-time focus tracking of the iris area of a moving object. We built a testing bench to collect real-world focal stacks for evaluation of the autofocus methods. Our camera has reached an autofocus speed of over 50 fps. The results demonstrate the advantages of our proposed camera for biometric perception in static and dynamic scenes. The code is available at https://github.com/Debatrix/AquulaCam.
Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun
IJCB2
2021 CASIA-Face-Africa: A Large-Scale African Face Image Database
abstract
Face recognition is a popular and well-studied area with wide applications in our society. However, racial bias had been proven to be inherent in most State Of The Art (SOTA) face recognition systems. Many investigative studies on face recognition algorithms have reported higher false positive rates of African subjects cohorts than the other cohorts. Lack of large-scale African face image databases in public domain is one of the main restrictions in studying the racial bias problem of face recognition. To this end, we collect a face image database namely CASIA-Face-Africa which contains 38,546 images of 1,183 African subjects. Multi-spectral cameras are utilized to capture the face images under various illumination settings. Demographic attributes and facial expressions of the subjects are also carefully recorded. For landmark detection, each face image in the database is manually labeled with 68 facial keypoints. A group of evaluation protocols are constructed according to different applications, tasks, partitions and scenarios. The performances of SOTA face recognition algorithms without re-training are reported as baselines. The proposed database along with its face landmark annotations, evaluation protocols and preliminary results form a good benchmark to study the essential aspects of face biometrics for African subjects, especially face image preprocessing, face feature analysis and matching, facial expression recognition, sex/age estimation, ethnic classification, face image generation, etc. The database can be downloaded from our website.
Jawad Muhammad, Yunlong Wang 0003, Caiyong Wang, Kunbo Zhang, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.4
2020 Recognition Oriented Iris Image Quality Assessment in the Feature Space
abstract
A large portion of iris images captured in real world scenarios are poor quality due to the uncontrolled environment and the non-cooperative subject. To ensure that the recognition algorithm is not affected by low-quality images, traditional hand-crafted factors based methods discard most images, which will cause system timeout and disrupt user experience. In this paper, we propose a recognition-oriented quality metric and assessment method for iris image to deal with the problem. The method regards the iris image em-beddings Distance in Feature Space (DFS) as the quality metric and the prediction is based on deep neural networks with the attention mechanism. The quality metric proposed in this paper can significantly improve the performance of the recognition algorithm while reducing the number of images discarded for recognition, which is advantageous over hand-crafted factors based iris quality assessment methods. The relationship between Image Rejection Rate (IRR) and Equal Error Rate (EER) is proposed to evaluate the performance of the quality assessment algorithm under the same image quality distribution and the same recognition algorithm. Compared with hand-crafted factors based methods, the proposed method is a trial to bridge the gap between the image quality assessment and biometric recognition.
Leyuan Wang, Kunbo Zhang, Yunlong Wang 0003, Zhenan Sun
IJCB2
2020 All-in-Focus Iris Camera With a Great Capture Volume
abstract
Imaging volume of an iris recognition system has been restricting the throughput and cooperation convenience in biometric applications. Numerous improvement trials are still impractical to supersede the dominant fixed-focus lens in stand-off iris recognition due to incremental performance increase and complicated optical design. In this study, we develop a novel all-in-focus iris imaging system using a focus-tunable lens and a 2D steering mirror to greatly extend capture volume by spatiotemporal multiplexing method. Our iris imaging depth of field extension system requires no mechanical motion and is capable to adjust the focal plane at extremely high speed. In addition, the motorized reflection mirror adaptively steers the light beam to extend the horizontal and vertical field of views in an active manner. The proposed all-in-focus iris camera increases the depth of field up to 3.9 m which is afactor of 37.5 compared with conventional long focal lens. We also experimentally demonstrate the capability of this 3D light beam steering imaging system in real-time multi-person iris refocusing using dynamic focal stacks and the potential of continuous iris recognition for moving participants.
Kunbo Zhang, Zhenteng Shen, Yunlong Wang 0003, Zhenan Sun
IJCB1
2020 A Novel Deep-learning Pipeline for Light Field Image Based Material Recognition
abstract
The primitive basis of image based material recognition builds upon the fact that discrepancies in the reflectances of distinct materials lead to imaging differences under multiple viewpoints. LF cameras possess coherent abilities to capture multiple sub-aperture views (SAIs) within one exposure, which can provide appropriate multi-view sources for material recognition. In this paper, a unified “Factorize-Connect-Merge” (FCM) deep-learning pipeline is proposed to solve problems of light field image based material recognition. 4D light-field data as input is initially decomposed into consecutive 3D light-field slices. Shallow CNN is leveraged to extract low-level visual features of each view inside these slices. As to establish correspondences between these SAIs, Bidirectional Long-Short Term Memory (Bi-LSTM) network is built upon these low-level features to model the imaging differences. After feature selection including concatenation and dimension reduction, effective and robust feature representations for material recognition can be extracted from 4D light-field data. Experimental results indicate that the proposed pipeline can obtain remarkable performances on both tasks of single-pixel material classification and full-image material segmentation. In addition, the proposed pipeline can potentially benefit and inspire other researchers who may also take LF images as input and need to extract 4D light-field representations for computer vision tasks such as object classification, semantic segmentation and edge detection.
Yunlong Wang 0003, Kunbo Zhang, Zhenan Sun
ICPR2
2018 LFNet: A Novel Bidirectional Recurrent Convolutional Neural Network for Light-Field Image Super-Resolution
abstract
The low spatial resolution of light-field image poses significant difficulties in exploiting its advantage. To mitigate the dependency of accurate depth or disparity information as priors for light-field image super-resolution, we propose an implicitly multi-scale fusion scheme to accumulate contextual information from multiple scales for super-resolution reconstruction. The implicitly multi-scale fusion scheme is then incorporated into bidirectional recurrent convolutional neural network, which aims to iteratively model spatial relations between horizontally or vertically adjacent sub-aperture images of light-field data. Within the network, the recurrent convolutions are modified to be more effective and flexible in modeling the spatial correlations between neighboring views. A horizontal sub-network and a vertical sub-network of the same network structure are ensembled for final outputs via stacked generalization. Experimental results on synthetic and real-world data sets demonstrate that the proposed method outperforms other state-of-the-art methods by a large margin in peak signal-to-noise ratio and gray-scale structural similarity indexes, which also achieves superior quality for human visual systems. Furthermore, the proposed method can enhance the performance of light field applications such as depth estimation.
Yunlong Wang 0003, Fei Liu 0031, Kunbo Zhang, Guangqi Hou, Zhenan Sun, Tieniu Tan
IEEE Trans. Image Process.3