VLDB 2026 Research / reviewers in the wild / expert
Hu Han 0001
dblp:03/6451-1
· DBLP profile ↗
78ranked-venue papers
10as first author
36since 2021 · last 2026
0000-0001-6010-1792ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 43 · 7 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021Security and privacy · 11 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal RegionsabstractFacial analysis is a fundamental problem in vision–language research, with important applications in affective computing. However, existing methods primarily focus on global facial attributes or single-dimension analysis, lacking fine-grained, interpretable multi-attribute modeling of arbitrary local facial regions. We introduce FaceFocalDesc, a new problem that aims to generate and recognize multi-attribute natural language descriptions for arbitrarily selected facial regions. The target attributes include facial action units, emotional states, and age. We argue that explicit region-level modeling enables more controllable and interpretable facial understanding. To support this task, we construct a new dataset with region-level annotations and corresponding language descriptions. We further propose Focal-RegionFace, a vision–language model fine-tuned from Qwen2.5-VL, which progressively refines its focus on localized facial features through multi-stage training. Experiments show that Focal-RegionFace achieves state-of-the-art performance on the proposed benchmark under both standard and newly introduced metrics, demonstrating its effectiveness in fine-grained region-focused facial analysis. Kaiwen Zheng 0002, Junchen Fu, Songpei Xu, Yaoqin He, Joemon M. Jose, Hu Han 0001, Xuri Ge |
ICMR | 6 |
| 2026 | MoMBS: Mixed-order sampling improves training on heterogeneous-quality data for universal lesion detection
Jingsong Liu, Peter J. Schüffler, Hu Han 0001, Shaohua Kevin Zhou |
Medical Image Anal. | 4 |
| 2026 | Distillation-SAM: Knowledge Distillation-Based Auto-Prompt Embedding Learning for Surgical Image SegmentationabstractSurgical image segmentation is vital for various stages of surgical procedures, from preoperative planning to real-time navigation and postoperative assessment. Despite advances in deep learning, current surgical image segmentation methods remain limited. They primarily target instrument segmentation and show poor generalizability across different surgical settings. While the Segment Anything Model (SAM) shows robust generalization capabilities in the segmentation of natural images, adapting SAM to surgical and medical images faces challenges because of its reliance on high-quality user-provided prompts and inherent lack of design for multi-class semantic segmentation. To address these limitations, we propose Distillation-SAM, an effective method that adapts SAM for accurate surgical image segmentation without user-provided prompts while freezing its encoder and decoder. Distillation-SAM introduces a trainable adapter branch that learns both sparse auto-prompt embeddings and enriched image features with dense auto-prompt embeddings, enabling the segmentation of surgical objects such as vessels, instruments, and tissues. We propose a direct knowledge distillation constraint for these auto-prompt embedding learnings by using embeddings derived from ground-truth masks as guidance. To enable multi-class semantic segmentation using SAM, we revise the mask score regression branch in SAM's decoder by incorporating a trainable Multilayer Perceptron to predict mask categories while keeping other parameters frozen. Our experiments in multiple surgical datasets, including IVIS, EndoVis2017, and Cholecseg8k, demonstrate that distillation-SAM outperforms existing methods in vessel, tissue, and instrument segmentation. Jiyang Tang, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-To-Video Diffusion ModelsabstractThe progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing motion transfer methods explored the motion representations of reference videos to guide generation. Nevertheless, these methods typically rely on sample-specific optimization strategy, resulting in high computational burdens. In this paper, we propose EfficientMT, a novel and efficient end-to-end framework for video motion transfer. By leveraging a small set of synthetic paired motion transfer samples, EfficientMT effectively adapts a pretrained T2V model into a general motion transfer framework that can accurately capture and reproduce diverse motion patterns. Specifically, we repurpose the backbone of the T2V model to extract temporal information from reference videos, and further propose a scaler module to distill motion-related information. Subsequently, we introduce a temporal integration mechanism that seamlessly incorporates reference motion features into the video generation process. After training on our self-collected synthetic paired samples, EfficientMT enables general video motion transfer without requiring test-time optimization. Extensive experiments demonstrate that our EfficientMT outperforms existing methods in efficiency while maintaining flexible motion controllability. Our code will be available https://github.com/PrototypeNx/EfficientMT. Yufei Cai, Hu Han 0001, Yuxiang Wei 0001, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2025 | Natural Adversarial Mask for Face Identity Protection in Physical WorldabstractFacial recognition (FR) technology offers convenience in our daily lives, but it also raises serious privacy issues due to unauthorized FR applications. To protect facial privacy, existing methods have proposed adversarial face examples that can fool FR systems. However, most of these methods work only in the digital domain and do not consider natural physical protections. In this paper, we present NatMask, a 3D-based method for creating natural and realistic adversarial face masks that can preserve facial identity in the physical world. Our method utilizes 3D face reconstruction and differentiable rendering to generate 2D face images with natural-looking facial masks. Moreover, we propose an identity-aware style injection (IASI) method to improve the naturalness and transferability of the mask texture. We evaluate our method on two face datasets to verify its effectiveness in protecting face identity against four state-of-the-art (SOTA) FR models and three commercial FR APIs in both digital and physical domains under black-box impersonation and dodging strategies. Experiments show that our method can generate adversarial masks with superior naturalness and physical realizability to safeguard face identity, outperforming SOTA methods by a large margin. Tianxin Xie, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Leveraging face-prior knowledge for general face representation learning
Haomiao Sun, Shiguang Shan, Hu Han 0001 |
Pattern Recognit. | 4 |
| 2024 | Decoupled Textual Embeddings for Customized Image GenerationabstractCustomized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned concept, limiting the potential to compose concept into new scenes. To address these issues, we propose the DETEX, a novel approach that learns the disentangled concept embedding for flexible customized text-to-image generation. Unlike conventional methods that learn a single concept embedding from the given images, our DETEX represents each image using multiple word embeddings during training, i.e., a learnable image-shared subject embedding and several image-specific subject-unrelated embeddings. To decouple irrelevant attributes (i.e., background and pose) from the subject embedding, we further present several attribute mappers that encode each image as several image-specific subject-unrelated embeddings. To encourage these unrelated embeddings to capture the irrelevant information, we incorporate them with corresponding attribute words and propose a joint training strategy to facilitate the disentanglement. During inference, we only use the subject embedding for image generation, while selectively using image-specific embeddings to retain image-specified attributes. Extensive experiments demonstrate that the subject embedding obtained by our method can faithfully represent the target concept, while showing superior editability compared to the state-of-the-art methods. Our code will be available at https://github.com/PrototypeNx/DETEX. Yufei Cai, Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Hu Han 0001, Wangmeng Zuo |
AAAI | 5 |
| 2024 | Multi-View Consistent 3D GAN Inversion via Bidirectional Encoderabstract3D GAN inversion enables not only 3D reconstruction from a 2D image, but also novel view synthesis and image editing. Existing works ensure the novel view synthesis quality by constraining the synthesized views to conform to the real image distribution. However, most of the methods did not consider the multi-view consistency, i.e., different photos of the same 3D scene via 3D GAN inversion can be inverted to the same 3D scene. In this paper, we propose a bidirectional encoder (BiDiE) for 3D GAN inversion that can improve the multi-view consistency and alleviate the interference of camera parameter prediction errors. On the one hand, the bidirectional encoder takes real images as input, estimates the camera parameters, and performs 3D reconstruction. On the other hand, the bidirectional encoder takes randomly sampled latent code and camera parameters as input, and generates synthesized images to assist in the latent code learning process. In addition, we extend the latent space from W+ to W++ to improve its reconstruction and editing capabilities. Experiments on the FFHQ, CelebA-HQ and Multi-PIE datasets prove that our proposed method outperforms state-of-the-art methods in multi-view consistent reconstruction as well as editing capability.11Code and datasets are available at https://github.com/WHZMM/BiDiE Haozhan Wu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2024 | Deep Subdomain Alignment for Cross-domain Image ClassificationabstractUnsupervised domain adaptation (UDA), which aims to transfer knowledge learned from a labeled source domain to an unlabeled target domain, is useful for various cross-domain image classification scenarios. A commonly used approach for UDA is to minimize the distribution differences between two domains, and subdomain alignment is found to be an effective method. However, most of the existing subdomain alignment methods are based on adversarial learning and focus on subdomain alignment procedures without considering the discriminability among individual subdomains, resulting in slow convergence and unsatisfactory adaptation results. To address these issues, we propose a novel deep subdomain alignment method for UDA in image classification, which consists of a Union Subdo-main Contrastive Learning (USCL) module and a Multi-view Subdomain Alignment (MvSA) strategy. USCL can create discriminative and dispersed subdomains by bringing samples from the same subdomain closer while pushing away samples from different subdomains. MvSA makes use of labeled source domain data and easy target domain data to perform target-to-source and target-to-target alignment. Experimental results on three image classifi-cation datasets (Office-31, Office-Home, Visda-17) demonstrate that our proposed method is effective for UDA and achieves promising results in several cross-domain image classification tasks. Our code will be available: https://github.com/zhaoyewei/DSACDIC. Yewei Zhao, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
WACV | 2 |
| 2024 | Fine-Grained Open-Set Deepfake Detection via Unsupervised Domain AdaptationabstractDeepfake represented by face swapping and face reenactment can transfer the appearance and behavioral expressions of a face in one video image to another face in a different video. In recent years, with the advancement of deep learning techniques, deepfake technology has developed rapidly, achieving increasingly realistic effects. Therefore, many researchers have begun to study deepfake detection research. However, most existing studies on deepfake detection are mainly limited to binary classification of real and fake images, rather than identifying different methods in an open-world scenario, leading to failures in dealing with unknown deepfake categories in practice. In this paper, we propose an unsupervised domain adaptation method for fine-grained open-set deepfake detection. Our method first uses labeled data from the source domain for model pre-training to establish the ability of recognizing different deepfake methods in the source domain. Then, the method uses a Network Memorization based Adaptive Clustering (NMAC) approach to cluster unlabeled images in the target domain and designs a Pseudo-Label Generation (PLG) to generate virtual class labels for unknown deepfake categories by matching the adaptive clustering results with the known deepfake categories in the source domain. Finally, we retrain the initial multi-class deepfake detection model using labeled data of the source domain and pseudo-labeled data of the target domain to improve its generalization ability to unknown deepfake classes presented in the target domain. We validate the effectiveness of the proposed method under multiple open-set fine-grained deepfake detection tasks based on three deepfake datasets (ForgerNet, FaceForensics++, and FakeAVCeleb). Experimental results show that our method has better domain generalization ability than the state-of-the-art methods, and achieves promising performance in fine-grained open-set deepfake detection. Xinye Zhou, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Consensus-Agent Deep Reinforcement Learning for Face AgingabstractFace aging tasks aim to simulate changes in the appearance of faces over time. However, due to the lack of data on different ages under the same identity, existing models are commonly trained using mapping between age groups. This makes it difficult for most existing aging methods to accurately capture the correspondence between individual identities and aging features, leading to generating faces that do not match the real aging appearance. In this paper, we re-annotate the CACD2000 dataset and propose a consensus-agent deep reinforcement learning method to solve the aforementioned problem. Specifically, we define two agents, the aging process agent and the aging personalization agent, and model the task of matching aging features as a Markov decision process. The aging process agent simulates the aging process of an individual, while the aging personalization agent calculates the difference between the aging appearance of an individual and the average aging appearance. The two agents iteratively adjust the matching degree between the target aging feature and the current identity through a form of synergistic cooperation. Extensive experimental results on four face aging datasets show that our model achieves convincing performance compared to the current state-of-the-art methods. Ling Lin 0002, Hao Liu 0019, Jinqiao Liang, Jiao Feng, Hu Han 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Unit DetectionabstractThe Facial Action Coding System (FACS) encodes the action units (AUs) in facial images, which has attracted extensive research attention due to its wide use in facial expression analysis. Many methods that perform well on automatic facial action unit (AU) detection primarily focus on modeling various AU relations between corresponding local muscle areas or mining global attention–aware facial features; however, they neglect the dynamic interactions among local-global features. We argue that encoding AU features just from one perspective may not capture the rich contextual information between regional and global face features, as well as the detailed variability across AUs, because of the diversity in expression and individual characteristics. In this article, we propose a novel Multi-level Graph Relational Reasoning Network (termed MGRR-Net ) for facial AU detection. Each layer of MGRR-Net performs a multi-level (i.e., region-level, pixel-wise, and channel-wise level) feature learning. On the one hand, the region-level feature learning from the local face patch features via graph neural network can encode the correlation across different AUs. On the other hand, pixel-wise and channel-wise feature learning via graph attention networks (GAT) enhance the discrimination ability of AU features by adaptively recalibrating feature responses of pixels and channels from global face features. The hierarchical fusion strategy combines features from the three levels with gated fusion cells to improve AU discriminative ability. Extensive experiments on DISFA and BP4D AU datasets show that the proposed approach achieves superior performance than the state-of-the-art methods. Xuri Ge, Joemon M. Jose, Songpei Xu, Xiao Liu 0040, Hu Han 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | SurgNet: Self-Supervised Pretraining With Semantic Consistency for Vessel and Instrument Segmentation in Surgical ImagesabstractBlood vessel and surgical instrument segmentation is a fundamental technique for robot-assisted surgical navigation. Despite the significant progress in natural image segmentation, surgical image-based vessel and instrument segmentation are rarely studied. In this work, we propose a novel self-supervised pretraining method (SurgNet) that can effectively learn representative vessel and instrument features from unlabeled surgical images. As a result, it allows for precise and efficient segmentation of vessels and instruments with only a small amount of labeled data. Specifically, we first construct a region adjacency graph (RAG) based on local semantic consistency in unlabeled surgical images and use it as a self-supervision signal for pseudo-mask segmentation. We then use the pseudo-mask to perform guided masked image modeling (GMIM) to learn representations that integrate structural information of intraoperative objectives more effectively. Our pretrained model, paired with various segmentation methods, can be applied to perform vessel and instrument segmentation accurately using limited labeled data for fine-tuning. We build an Intraoperative Vessel and Instrument Segmentation (IVIS) dataset, comprised of ~3 million unlabeled images and over 4,000 labeled images with manual vessel and instrument annotations to evaluate the effectiveness of our self-supervised pretraining method. We also evaluated the generalizability of our method to similar tasks using two public datasets. The results demonstrate that our approach outperforms the current state-of-the-art (SOTA) self-supervised representation learning methods in various surgical image segmentation tasks. Hu Han 0001, Zhiming Zhao, Xilin Chen 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | ReCoT: Regularized Co-Training for Facial Action Unit Recognition with Noisy Labels
Hu Han 0001, Shiguang Shan, Zhilong Ji, Jinfeng Bai, Xilin Chen 0001 |
BMVC | 2 |
| 2023 | DISC: Learning from Noisy Labels via Dynamic Instance-Specific Selection and CorrectionabstractExisting studies indicate that deep neural networks (DNNs) can eventually memorize the label noise. We observe that the memorization strength of DNNs towards each instance is different and can be represented by the confidence value, which becomes larger and larger during the training process. Based on this, we propose a Dynamic Instance-specific Selection and Correction method (DISC) for learning from noisy labels (LNL). We first use a two- view-based backbone for image classification, obtaining confidence for each image from two views. Then we propose a dynamic threshold strategy for each instance, based on the momentum of each instance's memorization strength in previous epochs to select and correct noisy labeled data. Benefiting from the dynamic threshold strategy and two-view learning, we can effectively group each instance into one of the three subsets (i.e., clean, hard, and purified) based on the prediction consistency and discrepancy by two views at each epoch. Finally, we employ different regularization strategies to conquer subsets with different degrees of label noise, improving the whole network's robustness. Comprehensive evaluations on three controllable and four real-world LNL benchmarks show that our method outperforms the state-of-the-art (SOTA) methods to leverage useful information in noisy data while alleviating the pollution of label noise. Code is available at https://github.com/JackYFL/DISC. Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2023 | Data-Free Knowledge Distillation via Feature Exchange and Activation Region ConstraintabstractDespite the tremendous progress on data-free knowledge distillation (DFKD) based on synthetic data generation, there are still limitations in diverse and efficient data synthesis. It is naive to expect that a simple combination of generative network-based data synthesis and data augmentation will solve these issues. Therefore, this paper proposes a novel data-free knowledge distillation method (Spaceship-Net) based on channel-wise feature exchange (CFE) and multi-scale spatial activation region consistency (mSARC) constraint. Specifically, CFE allows our generative network to better sample from the feature space and efficiently synthesize diverse images for learning the student network. However, using CFE alone can severely amplify the unwanted noises in the synthesized images, which may result in failure to improve distillation learning and even have negative effects. Therefore, we propose mSARC to assure the student network can imitate not only the logit output but also the spatial activation region of the teacher network in order to alleviate the influence of unwanted noises in diverse synthetic images on distillation learning. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, Imagenette, and ImageNet100 show that our method can work well with different backbone networks, and outperform the state-of-the-art DFKD methods. Code will be available at: https://github.com/skgyu/Spaceship-Net. Shikang Yu, Hu Han 0001, Shuqiang Jiang |
CVPR | 3 |
| 2023 | Intrinsic Imaging Model Enhanced Contrastive Face Representation LearningabstractHumans can easily perceive numerous information from faces, only part of which has been achieved by a machine, thanks to the availability of large-scale face images with supervision signals of those specific tasks. More face perception tasks, like rare expression or attribute recognition, and genetic syndrome diagnosis, are not solved due to a critical shortage of supervised data. One possible way to solve these tasks is leveraging ubiquitous large-scale unsupervised face images and building a foundation face model via methods like contrastive learning (CL), which is, however, not aware of the intrinsic physics of the human face. In consideration of this shortcoming, this paper proposes to enhance contrastive face representation learning by the physical imaging model. Specifically, besides the CL-backbone network, we also design an auxiliary bypass pathway to constrain the CL-backbone to support the ability of accurately re-rendering the face with a differentiable physical imaging model after decomposing an input face image into intrinsic 3D imaging factors. With this design, the CL network is endowed the capacity of implicitly “knowing” the 3D of the face rather than the 2D pixels only. In experiments, we learn face representations from the CelebA and WebFace-42M datasets in unsupervised mode and evaluate the generalization capability of the representations with three different downstream tasks in case of limited supervised data. The experimental results clearly justify the effectiveness of the proposed method. Haomiao Sun, Shiguang Shan, Hu Han 0001 |
FG | 3 |
| 2023 | Modeling the Relative Visual Tempo for Self-supervised Skeleton-based Action RecognitionabstractVisual tempo characterizes the dynamics and the temporal evolution, which helps describe actions. Recent approaches directly perform visual tempo prediction on skeleton sequences, which may suffer from insufficient feature representation issue. In this paper, we observe that relative visual tempo is more in line with human intuition, and thus providing more effective supervision signals. Based on this, we propose a novel Relative Visual Tempo Contrastive Learning framework for skeleton action Representation (RVTCLR). Specifically, we design a Relative Visual Tempo Learning (RVTL) task to explore the motion information in intra-video clips, and an Appearance-Consistency (AC) task to learn appearance information simultaneously, resulting in more representative spatiotemporal features. Furthermore, skeleton sequence data is much sparser than RGB data, making the network learn shortcuts, and overfit to low-level information such as skeleton scales. To learn high-order semantics, we further design a new Distribution-Consistency (DC) branch, containing three components: Skeleton-specific Data Augmentation (S-DA), Fine-grained Skeleton Encoding Module (FSEM), and Distribution-aware Diversity (DD) Loss. We term our entire method (RVTCLR with DC) as RVTCLR+. Extensive experiments on NTU RGB+D 60 and NTU RGB+D 120 datasets demonstrate that our RVTCLR+ can achieve competitive results over the state-of-the-art methods. Code is available at https://github.com/Zhuysheng/RVTCLR. Yisheng Zhu, Hu Han 0001, Zhengtao Yu 0001, Guangcan Liu |
ICCV | 2 |
| 2023 | CMOS-GAN: Semi-Supervised Generative Adversarial Model for Cross-Modality Face Image SynthesisabstractCross-modality face image synthesis such as sketch-to-photo, NIR-to-RGB, and RGB-to-depth has wide applications in face recognition, face animation, and digital entertainment. Conventional cross-modality synthesis methods usually require paired training data, i.e., each subject has images of both modalities. However, paired data can be difficult to acquire, while unpaired data commonly exist. In this paper, we propose a novel semi-supervised cross-modality synthesis method (namely CMOS-GAN), which can leverage both paired and unpaired face images to learn a robust cross-modality synthesis model. Specifically, CMOS-GAN uses a generator of encoder-decoder architecture for new modality synthesis. We leverage pixel-wise loss, adversarial loss, classification loss, and face feature loss to exploit the information from both paired multi-modality face images and unpaired face images for model learning. In addition, since we expect the synthetic new modality can also be helpful for improving face recognition accuracy, we further use a modified triplet loss to retain the discriminative features of the subject in the synthetic modality. Experiments on three cross-modality face synthesis tasks (NIR-to-VIS, RGB-to-depth, and sketch-to-photo) show the effectiveness of the proposed approach compared with the state-of-the-art. In addition, we also collect a large-scale RGB-D dataset (VIPL-MumoFace-3K) for the RGB-to-depth synthesis task. We plan to open-source our code and VIPL-MumoFace-3K dataset to the community (https://github.com/skgyu/CMOS-GAN). Shikang Yu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Semi-Supervised CT Lesion Segmentation Using Uncertainty-Based Data Pairing and SwapMixabstractSemi-supervised learning (SSL) methods show their powerful performance to deal with the issue of data shortage in the field of medical image segmentation. However, existing SSL methods still suffer from the problem of unreliable predictions on unannotated data due to the lack of manual annotations for them. In this paper, we propose an unreliability-diluted consistency training (UDiCT) mechanism to dilute the unreliability in SSL by assembling reliable annotated data into unreliable unannotated data. Specifically, we first propose an uncertainty-based data pairing module to pair annotated data with unannotated data based on a complementary uncertainty pairing rule, which avoids two hard samples being paired off. Secondly, we develop SwapMix, a mixed sample data augmentation method, to integrate annotated data into unannotated data for training our model in a low-unreliability manner. Finally, UDiCT is trained by minimizing a supervised loss and an unreliability-diluted consistency loss, which makes our model robust to diverse backgrounds. Extensive experiments on three chest CT datasets show the effectiveness of our method for semi-supervised CT lesion segmentation. Pengchong Qiao, Guoli Song, Hu Han 0001, Yonghong Tian 0001, Yongsheng Liang 0001, Xi Li 0011, Shaohua Kevin Zhou, Jie Chen 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Improving Face Anti-spoofing via Advanced Multi-perspective Feature LearningabstractFace anti-spoofing (FAS) plays a vital role in securing face recognition systems. Previous approaches usually learn spoofing features from a single perspective, in which only universal cues shared by all attack types are explored. However, such single-perspective-based approaches ignore the differences among various attacks and commonness between certain attacks and bona fides, thus tending to neglect some non-universal cues that contain strong discernibility against certain types. As a result, when dealing with multiple types of attacks, the above approaches may suffer from the uncomprehensive representation of bona fides and spoof faces. In this work, we propose a novel Advanced Multi-Perspective Feature Learning network (AMPFL), in which multiple perspectives are adopted to learn discriminative features, to improve the performance of FAS. Specifically, the proposed network first learns universal cues and several perspective-specific cues from multiple perspectives, then aggregates the above features and further enhances them to perform face anti-spoofing. In this way, AMPFL obtains features that are difficult to be captured by single-perspective-based methods and provides more comprehensive information on bona fides and spoof faces, thus achieving better performance for FAS. Experimental results show that our AMPFL achieves promising results in public databases, and it effectively solves the issues of single-perspective-based approaches. Zhuming Wang, Yaowen Xu, Lifang Wu, Hu Han 0001, Zun Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Towards High-Fidelity Face Self-Occlusion Recovery via Multi-View Residual-Based GAN InversionabstractFace self-occlusions are inevitable due to the 3D nature of the human face and the loss of information in the projection process from 3D to 2D images. While recovering face self-occlusions based on 3D face reconstruction, e.g., 3D Morphable Model (3DMM) and its variants provides an effective solution, most of the existing methods show apparent limitations in expressing high-fidelity, natural, and diverse facial details. To overcome these limitations, we propose in this paper a new generative adversarial network (MvInvert) for natural face self-occlusion recovery without using paired image-texture data. We design a coarse-to-fine generator for photorealistic texture generation. A coarse texture is computed by inpainting the invisible areas in the photorealistic but incomplete texture sampled directly from the 2D image using the unrealistic but complete statistical texture from 3DMM. Then, we design a multi-view Residual-based GAN Inversion, which re-renders and refines multi-view 2D images, which are used for extracting multiple high-fidelity textures. Finally, these high-fidelity textures are fused based on their visibility maps via Poisson blending. To perform adversarial learning to assure the quality of the recovered texture, we design a discriminator consisting of two heads, i.e., one for global and local discrimination between the recovered texture and a small set of real textures in UV space, and the other for discrimination between the input image and the re-rendered 2D face images via pixel-wise, identity, and adversarial losses. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods in face self-occlusion recovery under unconstrained scenarios. Hu Han 0001, Shiguang Shan |
AAAI | 2 |
| 2022 | SATr: Slice Attention with Transformer for Universal Lesion Detection
Hu Han 0001, Shaohua Kevin Zhou |
MICCAI (3) | 3 |
| 2022 | A Spatio-Temporal Approach for Apathy ClassificationabstractApathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient’s presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both, patients and clinical staff, hindering among other large-scale diagnostics. In this work, we propose a novel spatio-temporal framework for apathy classification, which is streamlined to analyze facial dynamics and emotion in videos. Specifically, we divide the videos into smaller clips, and proceed to extract associated facial dynamics and emotion-based features. Statistical representations/descriptors based on each feature and clip serve as input of the proposed Gated Recurrent Unit (GRU)-architecture. Temporal representations of individual features at the lower level of the proposed architecture are combined at deeper layers of the proposed GRU architecture, in order to obtain the final feature-set for apathy classification. Based on extensive experiments, we show that fusion of characteristics such as emotion and facial dynamics in proposed deep-bi-directional GRU obtains an accuracy of 95.34% in apathy classification. Abhijit Das 0001, Xuesong Niu, Antitza Dantcheva, S. L. Happy, Hu Han 0001, Radia Zeghari, Philippe Robert, Shiguang Shan, François Brémond, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Zero-Shot Embedding via Regularization-Based Recollection and Residual Familiarity ProcessesabstractThe goal of zero-shot learning (ZSL) is to transfer knowledge learned from seen classes during training to unseen classes for testing, with the help of auxiliary information, such as attributes and descriptions. Most of the existing methods view ZSL as a label-embedding problem, in which class and image representations are embedded to a common space. However, many methods either show a bias toward seen classes caused by the projection domain-shift problem, or sacrifice the performance of seen classes to generalize to unseen ones. In this article, we present an embedding approach for ZSL, which is motivated by human recognition memory, namely, recollection and familiarity (R&F). We propose a decoder to regularize the nonlinear mapping between the semantic space and the visual space, which represents the reasonable recollection process, and use a residual block to refine the recognition ability for seen classes, which indicates the familiarity process. R&F can generalize well to unseen classes, while retaining the discriminative ability for the seen classes. Extensive experiments are conducted on Animals with Attribute (AwA1), Animals with Attributes 2 (AwA2), Attribute Pascal&Yahoo (aPY), SUN Attribute (SUN), Caltech-UCSD-Birds 200-2011 (CUB), and ImageNet databases. As qualitative and quantitative results show, the proposed approach outperforms state-of-the-art embedding-based methods by a large margin and significantly alleviates the projection domain-shift problem. Mengyao Lyu, Hu Han 0001, Xiangzhi Bai |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | Dual-GAN: Joint BVP and Noise Modeling for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) based physiological measurement has great application values in health monitoring, emotion analysis, etc. Existing methods mainly focus on how to enhance or extract the very weak blood volume pulse (BVP) signals from face videos, but seldom explicitly model the noises that dominate face video content. Thus, they may suffer from poor generalization ability in unseen scenarios. This paper proposes a novel adversarial learning approach for rPPG based physiological measurement by using Dual Generative Adversarial Networks (Dual-GAN) to model the BVP predictor and noise distribution jointly. The BVP-GAN aims to learn a noise-resistant mapping from input to ground-truth BVP, and the Noise-GAN aims to learn the noise distribution. The two GANs can promote each other’s capability, leading to improved feature disentanglement between BVP and noises. Besides, a plug-and-play block named ROI alignment and fusion (ROI-AF) block is proposed to alleviate the inconsistencies between different ROIs and exploit informative features from a wider receptive field in terms of ROIs. In comparison to state-of-the-art methods, our approach achieves better performance in heart rate, heart rate variability, and respiration frequency estimation from face videos. Hao Lu 0009, Hu Han 0001, Shaohua Kevin Zhou |
CVPR | 2 |
| 2021 | BVPNet: Video-to-BVP Signal Prediction for Remote Heart Rate EstimationabstractIn this paper, we propose a new method for remote photoplethysmography (rPPG) based heart rate (HR) estimation. In particular, our proposed method BVPNet is streamlined to predict the blood volume pulse (BVP) signals from face videos. Towards this, we firstly define ROIs based on facial landmarks and then extract the raw temporal signal from each ROI. Then the extracted signals are pre-processed via first-order difference and Butterworth filter and combined to form a Spatial-Temporal map (STMap). We then propose to revise U-Net, in order to predict BVP signals from the STMap. BVPNet takes into account both temporal and frequency domain losses in order to learn better than conventional models. Our experimental results suggest that our BVPNet outperforms the state-of-the-art methods on two publicly available datasets (MMSE-HR and VIPL-HR). Abhijit Das 0001, Hao Lu 0009, Hu Han 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001 |
FG | 3 |
| 2021 | Local Global Relational Network for Facial Action Units RecognitionabstractMany existing facial action units (AUs) recognition approaches often enhance the AU representation by combining local features from multiple independent branches, each corresponding to a different AU. However, such multi-branch combination-based methods usually neglect potential mutual assistance and exclusion relationship between AU branches or simply employ a pre-defined and fixed knowledge-graph as a prior. In addition, extracting features from pre-defined AU regions of regular shapes limits the representation ability. In this paper, we propose a novel Local Global Relational Network (LGRNet) for facial AU recognition. LGRNet mainly consists of two novel structures, i.e., a skip-BiLSTM module which models the latent mutual assistance and exclusion relationship among local AU features from multiple branches to enhance the feature robustness, and a feature fusion&refining module which explores the complementarity between local AUs and the whole face in order to refine the local AU features to improve the discriminability. Experiments on the BP4D and DISFA AU datasets show that the proposed approach outperforms the state-of-the-art methods by a large margin. Xuri Ge, Hu Han 0001, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, Xiao Liu 0040 |
FG | 3 |
| 2021 | Exploiting Non-uniform Inherent Cues to Improve Presentation Attack DetectionabstractFace anti-spoofing plays a vital role in face recognition systems. The existed deep learning approaches have effectively improved the performance of presentation attack detection (PAD). However, they learn a uniform feature for different types of presentation attacks, which ignore the diversity of the inherent cues presented in different spoofing types. As a result, they can not effectively represent the intrinsic difference between different spoof faces and live faces, and the performance drops on the cross-domain databases. In this paper, we introduce the inherent cues of different spoofing types by non-uniform learning as complements to uniform features. Two lightweight sub-networks are designed to learn inherent motion patterns from photo attacks and the inherent texture cues from video attacks. Furthermore, an element-wise weighting fusion strategy is proposed to integrate the non-uniform inherent cues and uniform features. Extensive experiments on four public databases demonstrate that our approach outperforms the state-of-the-art methods and achieves a superior performance of 3.7% ACER in the cross-domain Protocol 4 of the Oulu-NPU database. Code is available at https://github.com/BJUT-VIP/Non-uniform-cues. Yaowen Xu, Zhuming Wang, Hu Han 0001, Lifang Wu, Yongluo Liu |
IJCB | 3 |
| 2021 | Conditional Training with Bounding Map for Universal Lesion Detection
Hu Han 0001, Ying Chi, Shaohua Kevin Zhou |
MICCAI (5) | 3 |
| 2021 | M-SEAM-NAM: Multi-instance Self-supervised Equivalent Attention Mechanism with Neighborhood Affinity Module for Double Weakly Supervised Segmentation of COVID-19
Wen Tang 0005, Han Kang, Pengxin Yu, Hu Han 0001, Rongguo Zhang, Kuan Chen |
MICCAI (7) | 5 |
| 2021 | Semi-Supervised Natural Face De-OcclusionabstractOcclusions are often present in face images in the wild, e.g., under video surveillance and forensic scenarios. Existing face de-occlusion methods are limited as they require the knowledge of an occlusion mask. To overcome this limitation, we propose in this paper a new generative adversarial network (named OA-GAN) for natural face de-occlusion without an occlusion mask, enabled by learning in a semi-supervised fashion using (i) paired images with known masks of artificial occlusions and (ii) natural images without occlusion masks. The generator of our approach first predicts an occlusion mask, which is used for filtering the feature maps of the input image as a semantic cue for de-occlusion. The filtered feature maps are then used for face completion to recover a non-occluded face image. The initial occlusion mask prediction might not be accurate enough, but it gradually converges to the accurate one because of the adversarial loss we use to perceive which regions in a face image need to be recovered. The discriminator of our approach consists of an adversarial loss, distinguishing the recovered face images from natural face images, and an attribute preserving loss, ensuring that the face image after de-occlusion can retain the attributes of the input face image. Experimental evaluations on the widely used CelebA dataset and a dataset with natural occlusions we collected show that the proposed approach can outperform the state of the art methods in natural face de-occlusion. Jiancheng Cai, Hu Han 0001, Jiyun Cui, Jie Chen 0001, Li Liu 0002, Shaohua Kevin Zhou |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Deep Conditional Distribution Learning for Age EstimationabstractAge estimation is a challenging task not only because face appearance is affected by illumination, pose, and expression, but also because there exists age label ambiguity among different demographic groups. In this work, we first revisit different label distribution learning (LDL) based age estimation methods and propose a more general formulation, which can unify individual LDL-based age estimation methods, as well as the traditional regression, classification, and ranking based age estimation methods. Based on such a general formulation, we propose a novel deep conditional distribution learning (DCDL) method, which can flexibly leverage a varying number of auxiliary face attributes to achieve adaptive age-related feature learning and improve age estimation robustness against the challenges above. Experimental results on multiple age estimation datasets (MORPH II, AgeDB, FG-NET, MegaAge-Asian, CLAP2016, UTK-Face, and LFW+) show that the proposed approach outperforms the state-of-the-art age estimation methods by a large margin. In addition, the proposed approach can generalize well to other human attributes estimation tasks, like height, weight, and body mass index (BMI) estimation. Haomiao Sun, Hongyu Pan, Hu Han 0001, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Unsupervised Adversarial Domain Adaptation for Cross-Domain Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is essential for securing the widely used face recognition systems. Most of the existing PAD methods do not generalize well to unseen scenarios because labeled training data of the new domain is usually not available. In light of this, we propose an unsupervised domain adaptation with disentangled representation (DR-UDA) approach to improve the generalization capability of PAD into new scenarios. DR-UDA consists of three modules, i.e., ML-Net, UDA-Net and DR-Net. ML-Net aims to learn a discriminative feature representation using the labeled source domain face images via metric learning. UDA-Net performs unsupervised adversarial domain adaptation in order to optimize the source domain and target domain encoders jointly, and obtain a common feature space shared by both domains. As a result, the source domain PAD model can be effectively transferred to the unlabeled target domain for PAD. DR-Net further disentangles the features irrelevant to specific domains by reconstructing the source and target domain face images from the common feature space. Therefore, DR-UDA can learn a disentangled representation space which is generative for face images in both domains and discriminative for live vs. spoof classification. The proposed approach shows promising generalization capability in several public-domain face PAD databases. Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Collaborative Local-Global Learning for Temporal Action ProposalabstractTemporal action proposal generation is an essential and challenging task in video understanding, which aims to locate the temporal intervals that likely contain the actions of interest. Although great progress has been made, the problem is still far from being well solved. In particular, prevalent methods can handle well only the local dependencies (i.e., short-term dependencies) among adjacent frames but are generally powerless in dealing with the global dependencies (i.e., long-term dependencies) between distant frames. To tackle this issue, we propose CLGNet, a novel Collaborative Local-Global Learning Network for temporal action proposal. The majority of CLGNet is an integration of Temporal Convolution Network and Bidirectional Long Short-Term Memory, in which Temporal Convolution Network is responsible for local dependencies while Bidirectional Long Short-Term Memory takes charge of handling the global dependencies. Furthermore, an attention mechanism called the background suppression module is designed to guide our model to focus more on the actions. Extensive experiments on two benchmark datasets, THUMOS’14 and ActivityNet-1.3, show that the proposed method can outperform state-of-the-art methods, demonstrating the strong capability of modeling the actions with varying temporal durations. Yisheng Zhu, Hu Han 0001, Guangcan Liu, Qingshan Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | NAS-HR: Neural architecture search for heart rate estimation from face videosabstractIn anticipation of its great potential application to natural human-computer interaction and health monitoring, heart-rate (HR) estimation based on remote photoplethysmography has recently attracted increasing research attention. Whereas the recent deep-learning-based HR estimation methods have achieved promising performance, their computational costs remain high, particularly in mobile-computing scenarios. We propose a neural architecture search approach for HR estimation to automatically search a lightweight network that can achieve even higher accuracy than a complex network while reducing the computational cost. First, we define the regions of interests based on face landmarks and then extract the raw temporal pulse signals from the R, G, and B channels in each ROI. Then, pulse-related signals are extracted using a plane-orthogonal-to-skin algorithm, which are combined with the R and G channel signals to create a spatial-temporal map. Finally, a differentiable architecture search approach is used for the network-structure search. Compared with the state-of-the-art methods on the public-domain VIPL-HR and PURE databases, our method achieves better HR estimation performance in terms of several evaluation metrics while requiring a much lower computational cost1. Hao Lu 0009, Hu Han 0001 |
Virtual Real. Intell. Hardw. | 2 |
| 2020 | Cross-Domain Face Presentation Attack Detection via Multi-Domain Disentangled Representation LearningabstractFace presentation attack detection (PAD) has been an urgent problem to be solved in the face recognition systems. Conventional approaches usually assume the testing and training are within the same domain; as a result, they may not generalize well into unseen scenarios because the representations learned for PAD may overfit to the subjects in the training set. In light of this, we propose an efficient disentangled representation learning for cross-domain face PAD. Our approach consists of disentangled representation learning (DR-Net) and multi-domain learning (MD-Net). DR-Net learns a pair of encoders via generative models that can disentangle PAD informative features from subject discriminative features. The disentangled features from different domains are fed to MD-Net which learns domain-independent features for the final cross-domain face PAD task. Extensive experiments on several public datasets validate the effectiveness of the proposed approach for cross-domain PAD. Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2020 | Video-Based Remote Physiological Measurement via Cross-Verified Feature Disentangling
Xuesong Niu, Zitong Yu, Hu Han 0001, Shiguang Shan, Guoying Zhao 0001 |
ECCV (2) | 3 |
| 2020 | Bounding Maps for Universal Lesion Detection
Hu Han 0001, Shaohua Kevin Zhou |
MICCAI (4) | 2 |
| 2020 | Miss the Point: Targeted Adversarial Attack on Multiple Landmark Detection
Qingsong Yao, Zecheng He, Hu Han 0001, Shaohua Kevin Zhou |
MICCAI (4) | 3 |
| 2020 | RhythmNet: End-to-End Heart Rate Estimation From Face via Spatial-Temporal RepresentationabstractHeart rate (HR) is an important physiological signal that reflects the physical and emotional status of a person. Traditional HR measurements usually rely on contact monitors, which may cause inconvenience and discomfort. Recently, some methods have been proposed for remote HR estimation from face videos; however, most of them focus on well-controlled scenarios, their generalization ability into less-constrained scenarios (e.g., with head movement, and bad illumination) are not known. At the same time, lacking large-scale HR databases has limited the use of deep models for remote HR estimation. In this paper, we propose an end-to-end RhythmNet for remote HR estimation from the face. In RyhthmNet, we use a spatial-temporal representation encoding the HR signals from multiple ROI volumes as its input. Then the spatial-temporal representations are fed into a convolutional network for HR estimation. We also take into account the relationship of adjacent HR measurements from a video sequence via Gated Recurrent Unit (GRU) and achieves efficient HR measurement. In addition, we build a large-scale multi-modal HR database (named as VIPL-HRVIPL-HR is available at: ), which contains 2,378 visible light videos (VIS) and 752 near-infrared (NIR) videos of 107 subjects. Our VIPL-HR database contains various variations such as head movements, illumination variations, and acquisition device changes, replicating a less-constrained scenario for HR estimation. The proposed approach outperforms the state-of-the-art methods on both the public-domain and our VIPL-HR databases. Xuesong Niu, Shiguang Shan, Hu Han 0001, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | High-Resolution Chest X-Ray Bone Suppression Using Unpaired CT Structural PriorsabstractThere is clinical evidence that suppressing the bone structures in Chest X-rays (CXRs) improves diagnostic value, either for radiologists or computer-aided diagnosis. However, bone-free CXRs are not always accessible. We hereby propose a coarse-to-fine CXR bone suppression approach by using structural priors derived from unpaired computed tomography (CT) images. In the low-resolution stage, we use the digitally reconstructed radiograph (DRR) image that is computed from CT as a bridge to connect CT and CXR. We then perform CXR bone decomposition by leveraging the DRR bone decomposition model learned from unpaired CTs and domain adaptation between CXR and DRR. To further mitigate the domain differences between CXRs and DRRs and speed up the learning convergence, we perform all the aboved operations in Laplacian of Gaussian (LoG) domain. After obtaining the bone decomposition result in DRR, we upsample it to a high resolution, based on which the bone region in the original high-resolution CXR is cropped and processed to produce a high-resolution bone decomposition result. Finally, such a produced bone image is subtracted from the original high-resolution CXR to obtain the bone suppression result. We conduct experiments and clinical evaluations based on two benchmarking CXR databases to show that (i) the proposed method outperforms the state-of-the-art unsupervised CXR bone suppression approaches; (ii) the CXRs with bone suppression are instrumental to radiologists for reducing their false-negative rate of lung diseases from 15% to 8%; and (iii) state-of-the-art disease classification performances are achieved by learning a deep network that takes the original CXR and its bone-suppressed image as inputs. Hu Han 0001, Zeju Li, Jingjing Lu, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Weakly Supervised Image Classification Through Noise RegularizationabstractWeakly supervised learning is an essential problem in computer vision tasks, such as image classification, object recognition, etc., because it is expected to work in the scenarios where a large dataset with clean labels is not available. While there are a number of studies on weakly supervised image classification, they usually limited to either single-label or multi-label scenarios. In this work, we propose an effective approach for weakly supervised image classification utilizing massive noisy labeled data with only a small set of clean labels (e.g., 5%). The proposed approach consists of a clean net and a residual net, which aim to learn a mapping from feature space to clean label space and a residual mapping from feature space to the residual between clean labels and noisy labels, respectively, in a multi-task learning manner. Thus, the residual net works as a regularization term to improve the clean net training. We evaluate the proposed approach on two multi-label datasets (OpenImage and MS COCO2014) and a single-label dataset (Clothing1M). Experimental results show that the proposed approach outperforms the state-of-the-art methods, and generalizes well to both single-label and multi-label scenarios. Mengying Hu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2019 | Local Relationship Learning With Person-Specific Shape Regularization for Facial Action Unit DetectionabstractEncoding individual facial expressions via action units (AUs) coded by the Facial Action Coding System (FACS) has been found to be an effective approach in resolving the ambiguity issue among different expressions. While a number of methods have been proposed for AU detection, robust AU detection in the wild remains a challenging problem because of the diverse baseline AU intensities across individual subjects, and the weakness of appearance signal of AUs. To resolve these issues, in this work, we propose a novel AU detection method by utilizing local information and the relationship of individual local face regions. Through such a local relationship learning, we expect to utilize rich local information to improve the AU detection robustness against the potential perceptual inconsistency of individual local regions. In addition, considering the diversity in the baseline AU intensities of individual subjects, we further regularize local relationship learning via person-specific face shape information, i.e., reducing the influence of person-specific shape information, and obtaining more AU discriminative features. The proposed approach outperforms the state-of-the-art methods on two widely used AU detection datasets in the public domain (BP4D and DISFA). Xuesong Niu, Hu Han 0001, Songfan Yang, Yan Huang 0008, Shiguang Shan |
CVPR | 2 |
| 2019 | FCSR-GAN: End-to-end Learning for Joint Face Completion and Super-resolutionabstractCombined variations such as low-resolution and occlusion often present in face images in the wild, e.g., under the scenario of video surveillance. While most of the existing face enhancement approaches only handle one type of variation per model, in this paper, we propose a deep generative adversarial network (FCSR-GAN) for joint face completion and face super-resolution via one model. The generator of FCSR-GAN aims to recover a high-resolution face image without occlusion given an input low-resolution face image with partial occlusions. The discriminator of FCSR-GAN consists of two adversarial losses, a perceptual loss, and a face parsing loss, which assure the high quality of the recovered face images. Experimental results on several public-domain databases (CelebA and Helen) show that the proposed approach outperforms the state-of-the-art methods in jointly doing face super-resolution (up to 4×) and face completion from low-resolution face images with occlusions. Jiancheng Cai, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2019 | Robust Remote Heart Rate Estimation from Face Utilizing Spatial-temporal AttentionabstractIn this work, we propose an end-to-end approach for robust remote heart rate (HR) measurement gleaned from facial videos. Specifically the approach is based on remote photoplethysmography (rPPG), which constitutes a pulse triggered perceivable chromatic variation, sensed in RGB-face videos. Consequently, rPPGs can be affected in less-constrained settings. To unpin the shortcoming, the proposed algorithm utilizes a spatio-temporal attention mechanism, which places focus on the salient features included in rPPG-signals. In addition, we propose an effective rPPG augmentation approach, generating multiple rPPG signals with varying HRs from a single face video. Experimental results on the public datasets VIPL-HR and MMSE-HR show that the proposed method outperforms state-of-the-art algorithms in remote HR estimation. Xuesong Niu, Xingyuan Zhao, Hu Han 0001, Abhijit Das 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001 |
FG | 3 |
| 2019 | Improving Face Sketch Recognition via Adversarial Sketch-Photo TransformationabstractFace sketch-photo transformation has broad applications in forensics, law enforcement, and digital entertainment, particular for face recognition systems that are designed for photo-to-photo matching. While there are a number of methods for face photo-to-sketch transformation, studies on sketch-to-photo transformation remain limited. In this paper, we propose a novel conditional CycleGAN for face sketch-to-photo transformation. Specifically, we leverage the advantages of CycleGAN and conditional GANs and design a feature-level loss to assure the high quality of the generated face photos from sketches. The generated face photos are used, as a replacement of face sketches, and particularly for face identification against a gallery set of mugshot photos. Experimental results on the public-domain database CUFSF show that the proposed approach is able to generate realistic photos from sketches, and the generated photos are instrumental in improving the sketch identification accuracy against a large gallery set. Shikang Yu, Hu Han 0001, Shiguang Shan, Antitza Dantcheva, Xilin Chen 0001 |
FG | 2 |
| 2019 | 3D U2-Net: A 3D Universal U-Net for Multi-domain Medical Image Segmentation
Chao Huang 0009, Hu Han 0001, Qingsong Yao, Shankuan Zhu, Shaohua Kevin Zhou |
MICCAI (2) | 2 |
| 2019 | Encoding CT Anatomy Knowledge for Unpaired Chest X-ray Image Decomposition
Zeju Li, Hu Han 0001, Gonglei Shi, Jiannan Wang 0005, Shaohua Kevin Zhou |
MICCAI (6) | 3 |
| 2019 | Multi-label Co-regularization for Semi-supervised Facial Action Unit RecognitionabstractFacial action units (AUs) recognition is essential for emotion analysis and has been widely applied in mental state analysis. Existing work on AU recognition usually requires big face dataset with accurate AU labels. However, manual AU annotation requires expertise and can be time-consuming. In this work, we propose a semi-supervised approach for AU recognition utilizing a large number of web face images without AU labels and a small face dataset with AU labels inspired by the co-training methods. Unlike traditional co-training methods that require provided multi-view features and model re-training, we propose a novel co-training method, namely multi-label co-regularization, for semi-supervised facial AU recognition. Two deep neural networks are used to generate multi-view features for both labeled and unlabeled face images, and a multi-view loss is designed to enforce the generated features from the two views to be conditionally independent representations. In order to obtain consistent predictions from the two views, we further design a multi-label co-regularization loss aiming to minimize the distance between the predicted AU probability distributions of the two views. In addition, prior knowledge of the relationship between individual AUs is embedded through a graph convolutional network (GCN) for exploiting useful information from the big unlabeled dataset. Experiments on several benchmarks show that the proposed approach can effectively leverage large datasets of unlabeled face images to improve the AU recognition robustness and outperform the state-of-the-art semi-supervised AU recognition methods. Xuesong Niu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
NeurIPS | 2 |
| 2019 | Tattoo Image Search at Scale: Joint Detection and Compact Representation LearningabstractThe explosive growth of digital images in video surveillance and social media has led to the significant need for efficient search of persons of interest in law enforcement and forensic applications. Despite tremendous progress in primary biometric traits (e.g., face and fingerprint) based person identification, a single biometric trait alone can not meet the desired recognition accuracy in forensic scenarios. Tattoos, as one of the important soft biometric traits, have been found to be valuable for assisting in person identification. However, tattoo search in a large collection of unconstrained images remains a difficult problem, and existing tattoo search methods mainly focus on matching cropped tattoos, which is different from real application scenarios. To close the gap, we propose an efficient tattoo search approach that is able to learn tattoo detection and compact representation jointly in a single convolutional neural network (CNN) via multi-task learning. While the features in the backbone network are shared by both tattoo detection and compact representation learning, individual latent layers of each sub-network optimize the shared features toward the detection and feature learning tasks, respectively. We resolve the small batch size issue inside the joint tattoo detection and compact representation learning network via random image stitch and preceding feature buffering. We evaluate the proposed tattoo search system using multiple public-domain tattoo benchmarks, and a gallery set with about 300K distracter tattoo images compiled from these datasets and images from the Internet. In addition, we also introduce a tattoo sketch dataset containing 300 tattoos for sketch-based tattoo search. Experimental results show that the proposed approach has superior performance in tattoo detection and tattoo search at scale compared to several state-of-the-art tattoo retrieval algorithms. Hu Han 0001, Anil K. Jain 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | TKDN: Scene Text Detection via Keypoints Detection
Yuanshun Cui, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (5) | 3 |
| 2018 | Multi-label Learning from Noisy Labels with Non-linear Feature Transformation
Mengying Hu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (5) | 2 |
| 2018 | VIPL-HR: A Multi-modal Database for Pulse Estimation from Less-Constrained Face Video
Xuesong Niu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (5) | 2 |
| 2018 | Mean-Variance Loss for Deep Age Estimation From a FaceabstractAge estimation has wide applications in video surveillance, social networking, and human-computer interaction. Many of the published approaches simply treat age estimation as an exact age regression problem, and thus do not leverage a distribution's robustness in representing labels with ambiguity such as ages. In this paper, we propose a new loss function, called mean-variance loss, for robust age estimation via distribution learning. Specifically, the mean-variance loss consists of a mean loss, which penalizes difference between the mean of the estimated age distribution and the ground-truth age, and a variance loss, which penalizes the variance of the estimated age distribution to ensure a concentrated distribution. The proposed mean-variance loss and softmax loss are jointly embedded into Convolutional Neural Networks (CNNs) for age estimation. Experimental results on the FG-NET, MORPH Album II, CLAP2016, and AADB databases show that the proposed approach outperforms the state-of-the-art age estimation methods by a large margin, and generalizes well to image aesthetics assessment. Hongyu Pan, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2018 | RGB-D Face Recognition via Deep Complementary and Common Feature LearningabstractRGB-D face recognition has attracted increasing attentions in recent years because of its robustness in unconstrained environment. However, existing approaches either handle individual modalities using completely separate pipelines or treat all the modalities equally using the same pipeline. Such approaches did not adequately consider the modality differences and exploit the modality correlations. We propose a novel approach for RGB-D face recognition that is able to learn complementary features from multiple modalities and common features between different modalities. Specifically, we introduce a joint loss taking activation from both modality-specific feature learning networks, and enforcing the features to be learned in a complementary way. We further extend the capability of this multi-modality (e.g., RGB-D vs. RGB-D) matcher into cross-modality (e.g., RGB vs. RGB-D) scenarios by learning a common feature transformation mapping different modalities into the same feature space. Experimental results on a number of public RGB-D face databases (e.g., EURECOM, VAP, IIIT-D, and BUAA), and a large RGB-D database we collected, show the impressive performance of the proposed approach. Hao Zhang 0203, Hu Han 0001, Jiyun Cui, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2018 | HeadNet: Pedestrian Head Detection Utilizing Body in ContextabstractPedestrian head with arbitrary poses and size is prohibitively difficult to detect in many real world applications. An appealing alternative is to utilize object detection technologies, which tend to be more and more mature and faster. However, general object detection technologies can hardly work in complicated scenarios where many heads are often too small to detect. In this paper, we present a novel approach that learns a semantic connection between pedestrian head and other body parts for head detection. Specifically, the proposed model, named as HeadNet, is based on PVANet backbone and also introduces beneficial strategies including online hard example mining (OHEM), fine-grained feature maps, RoI Align and Body in Context (BiC). Experiments demonstrate that our approach is able to utilize spatial semantics of the entire body effectively, and gains inspiring performance for pedestrian head detection. Xufen Cai, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
FG | 3 |
| 2018 | Face Alignment across Large Pose via MT-CNN Based 3D Shape ReconstructionabstractFace alignment plays an important role for robust face recognition and analysis applications in the wild. While a number of face alignment methods are available, large-pose face alignment remains a very challenging problem due to the ambiguity of facial keypoints in 2D face images. Recent attempts to solve this problem via 3D model fitting show more robustness against large poses and 2D ambiguity, but their accuracy and speed are still limited. We propose a 3D reconstruction based method to quickly and accurately detect 2D facial landmarks and estimate their visibilities. By designing a cascaded multi-task CNN model, we can efficiently reconstruct the 3D face shape, together with pose estimation as an auxiliary task. Finally, the landmarks on 3D shape are projected to the 2D face image to get the 2D landmarks and their visibilities. Experimental results on the challenging 300W-LP, AFLW2000-3D, and AFLW databases show that the proposed approach can be comparable with the state-of-the-art methods and is able to run in real time (32ms per image) on 3.4 GHz CPU. Gang Zhang 0005, Hu Han 0001, Shiguang Shan, Xingguang Song, Xilin Chen 0001 |
FG | 2 |
| 2018 | Automatic Engagement Prediction with GAP FeatureabstractIn this paper, we propose an automatic engagement prediction method for the Engagement in the Wild sub-challenge of EmotiW 2018. We first design a novel Gaze-AU-Pose (GAP) feature taking into account the information of gaze, action units and head pose of a subject. The GAP feature is then used for the subsequent engagement level prediction. To efficiently predict the engagement level for a long-time video, we divide the long-time video into multiple overlapped video clips and extract GAP feature for each clip. A deep model consisting of a Gated Recurrent Unit (GRU) layer and a fully connected layer is used as the engagement predictor. Finally, a mean pooling layer is applied to the per-clip estimation to get the final engagement level of the whole video. Experimental results on the validation set and test set show the effectiveness of the proposed approach. In particular, our approach achieves a promising result with an MSE of 0.0724 on the test set of Engagement Prediction Challenge of EmotiW 2018.t with an MSE of 0.072391 on the test set of Engagement Prediction Challenge of EmotiW 2018. Xuesong Niu, Hu Han 0001, Jiabei Zeng, Xuran Sun, Shiguang Shan, Yan Huang 0008, Songfan Yang, Xilin Chen 0001 |
ICMI | 2 |
| 2018 | SynRhythm: Learning a Deep Heart Rate Estimator from General to SpecificabstractRemote photoplethysmography (rPPG) based noncontact heart rate (HR) measurement from a face video has drawn increasing attention recently because of its potential applications in many scenarios such as training aid, health monitoring, and nursing care. Although a number of methods have been proposed, most of them are designed under certain assumptions and could fail when such assumptions do not hold. At the same time, while deep learning based methods have been reported to achieve promising results in many computer vision tasks, their use in rPPG-based heart rate estimation has been limited due to the very limited data available in public domain. To overcome this limitation and leverage the strong modeling ability of deep neural networks, in this paper, we propose a novel spatial-temporal representation for the HR signal and design a general-to-specific transfer learning strategy to train a deep heart rate estimator from a large volume of synthetic rhythm signals and a limited number of available face video data. Experiment results on the public-domain databases show the effectiveness of the proposed approach. Xuesong Niu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ICPR | 2 |
| 2018 | Revised Contrastive Loss for Robust Age Estimation from FaceabstractAge estimation has broad applications in many fields, such as video surveillance, social networking, and human-computer interaction. Many of the existing approaches treat age estimation as a classification problem; however, the individual age values are not independent classes; they have an ordinal relationship. Classification loss such as softmax is not able to model such kind of relationship. In this paper, we propose a new loss, called revised contrastive loss, to model the ordinal relationship of individual ages. Specifically, the revised contrastive loss is proposed to penalize the distance between two face images in the feature space according to their age difference, which makes the learned features more discriminative for the age estimation task. We embed the proposed revised contrastive loss and softmax loss into a Convolutional Neural Network (CNN), and optimize the networks via Stochastic Gradient Descent (SGD) in an end-to-end fashion. Experimental results on a number of challenging face aging databases (FG-NET, MORPH Album II, and CLAP2016) show that the proposed approach outperforms the state-of-the-art methods by a large margin using a single model. Hongyu Pan, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ICPR | 2 |
| 2018 | Scene Text Detection via Deep Semantic Feature Fusion and Attention-based RefinementabstractDespite tremendous progress in scene text detection in the past few years, efficient text detection in the wild remains challenging, particularly for the texts have large rotations, and the complicated background areas that are easily confused with text. In this paper, we propose an effective approach for scene text detection, which consists of initial text detection using the proposed deep semantic feature fusion of a fully convolutional network (FCN), and text detection refinement by our attention based text vs. non-text classifier learned in a fine-to-coarse fashion. The proposed approach outperforms the state-of-the-art scene text detection algorithms on the public-domain ICDAR2015 dataset, achieving an accuracy of 0.83 in terms of F-measure. Yuanshun Cui, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
ICPR | 3 |
| 2018 | Heterogeneous Face Attribute Estimation: A Deep Multi-Task Learning ApproachabstractFace attribute estimation has many potential applications in video surveillance, face retrieval, and social media. While a number of methods have been proposed for face attribute estimation, most of them did not explicitly consider the attribute correlation and heterogeneity (e.g., ordinal versus nominal and holistic versus local) during feature representation learning. In this paper, we present a Deep Multi-Task Learning (DMTL) approach to jointly estimate multiple heterogeneous attributes from a single face image. In DMTL, we tackle attribute correlation and heterogeneity with convolutional neural networks (CNNs) consisting of shared feature learning for all the attributes, and category-specific feature learning for heterogeneous attributes. We also introduce an unconstrained face database (LFW+), an extension of public-domain LFW, with heterogeneous demographic attributes (age, gender, and race) obtained via crowdsourcing. Experimental results on benchmarks with multiple face attributes (MORPH II, LFW+, CelebA, LFWA, and FotW) show that the proposed approach has superior performance compared to state of the art. Finally, evaluations on a public-domain face database (LAP) with a single attribute show that the proposed approach has excellent generalization ability. Hu Han 0001, Anil K. Jain 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Deep Multi-Task Learning for Joint Prediction of Heterogeneous Face AttributesabstractFace attribute prediction has important applications in video surveillance, face retrieval, and social media. While a number of methods have been proposed for face attribute prediction, most of them did not explicitly consider the attribute correlation and heterogeneity during feature learning. In this paper, we propose a Deep Multi-Task Learning (DMTL) network to jointly learn multiple models; each addresses the prediction of one category of homogenous attributes. Specifically, we group the heterogeneous face attributes into two categories (i.e., nominal and ordinal), and design corresponding prediction models. At the same time, we use a convolutional neural network (CNN) for early stage feature learning, which is shared by all the attributes. Experiments on the public-domain MORPH II, CelebA, and LFWA databases show that the proposed approach outperforms the state of the art in joint face attribute prediction, and has good generalization ability. Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2017 | Continuous heart rate measurement from face: A robust rPPG approach with distribution learningabstractNon-contact heart rate (HR) measurement via remote photoplethysmography (rPPG) has drawn increasing attention. While a number of methods have been reported, most of them did not take into account the continuous HR measurement problem, which is more challenging due to limited observed video frames and the requirement of speed. In this paper, we present a real-time rPPG method for continuous HR measurement from face videos. We use a multi-patch ROI strategy to remove outlier signals. Chrominance feature is then generated from each ROI to reduce the color channel magnitude differences, which is followed by temporal filtering to suppress the artifacts. In addition, considering the temporal relationship of neighboring HR rhythms, we learn a HR distribution based on historical HR measurements, and apply it to the succeeding HR estimations. Experiment results on the public-domain MAHNOB-HCI database and user tests with commodity webcams show the effectiveness of the proposed approach. Xuesong Niu, Hu Han 0001, Shiguang Shan, Xilin Chen 0001 |
IJCB | 2 |
| 2016 | Secure Face Unlock: Spoof Detection on SmartphonesabstractWith the wide deployment of the face recognition systems in applications from deduplication to mobile device unlocking, security against the face spoofing attacks requires increased attention; such attacks can be easily launched via printed photos, video replays, and 3D masks of a face. We address the problem of face spoof detection against the print (photo) and replay (photo or video) attacks based on the analysis of image distortion (e.g., surface reflection, moiré pattern, color distortion, and shape deformation) in spoof face images (or video frames). The application domain of interest is smartphone unlock, given that the growing number of smartphones have the face unlock and mobile payment capabilities. We build an unconstrained smartphone spoof attack database (MSU USSA) containing more than 1000 subjects. Both the print and replay attacks are captured using the front and rear cameras of a Nexus 5 smartphone. We analyze the image distortion of the print and replay attacks using different: 1) intensity channels (R, G, B, and grayscale); 2) image regions (entire image, detected face, and facial component between nose and chin); and 3) feature descriptors. We develop an efficient face spoof detection system on an Android smartphone. Experimental results on the public-domain Idiap Replay-Attack, CASIA FASD, and MSU-MFSD databases, and the MSU USSA database show that the proposed approach is effective in face spoof detection for both the cross-database and intra-database testing scenarios. User studies of our Android face spoof detection system involving 20 participants show that the proposed approach works very well in real application scenarios. Keyurkumar Patel, Hu Han 0001, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | Demographic Estimation from Face Images: Human vs. Machine PerformanceabstractDemographic estimation entails automatic estimation of age, gender and race of a person from his face image, which has many potential applications ranging from forensics to social media. Automatic demographic estimation, particularly age estimation, remains a challenging problem because persons belonging to the same demographic group can be vastly different in their facial appearances due to intrinsic and extrinsic factors. In this paper, we present a generic framework for automatic demographic (age, gender and race) estimation. Given a face image, we first extract demographic informative features via a boosting algorithm, and then employ a hierarchical approach consisting of between-group classification, and within-group regression. Quality assessment is also developed to identify low-quality face images that are difficult to obtain reliable demographic estimates. Experimental results on a diverse set of face image databases, FG-NET (1K images), FERET (3K images), MORPH II (75K images), PCSO (100K images), and a subset of LFW (4K images), show that the proposed approach has superior performance compared to the state of the art. Finally, we use crowdsourcing to study the human perception ability of estimating demographics from face images. A side-by-side comparison of the demographic estimates from crowdsourced data and the proposed algorithm provides a number of insights into this challenging problem. Hu Han 0001, Charles Otto, Xiaoming Liu 0002, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Face Spoof Detection With Image Distortion AnalysisabstractAutomatic face recognition is now widely used in applications ranging from deduplication of identity to authentication of mobile payment. This popularity of face recognition has raised concerns about face spoof attacks (also known as biometric sensor presentation attacks), where a photo or video of an authorized person's face could be used to gain access to facilities or services. While a number of face spoof detection techniques have been proposed, their generalization ability has not been adequately addressed. We propose an efficient and rather robust face spoof detection algorithm based on image distortion analysis (IDA). Four different features (specular reflection, blurriness, chromatic moment, and color diversity) are extracted to form the IDA feature vector. An ensemble classifier, consisting of multiple SVM classifiers trained for different face spoof attacks (e.g., printed photo and replayed video), is used to distinguish between genuine (live) and spoof faces. The proposed approach is extended to multiframe face spoof detection in videos using a voting-based scheme. We also collect a face spoof database, MSU mobile face spoofing database (MSU MFSD), using two mobile devices (Google Nexus 5 and MacBook Air) with three types of spoof attacks (printed photo, replayed video with iPhone 5S, and replayed video with iPad Air). Experimental results on two public-domain face spoof databases (Idiap REPLAY-ATTACK and CASIA FASD), and the MSU MFSD database show that the proposed approach outperforms the state-of-the-art methods in spoof detection. Our results also highlight the difficulty in separating genuine and spoof faces, especially in cross-database and cross-device scenarios. Di Wen 0001, Hu Han 0001, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Nighttime face recognition at large standoff: Cross-distance and cross-spectral matching
Dongoh Kang, Hu Han 0001, Anil K. Jain 0001, Seong-Whan Lee |
Pattern Recognit. | 2 |
| 2014 | Unconstrained Face Recognition: Identifying a Person of Interest From a Media CollectionabstractAs face recognition applications progress from constrained sensing and cooperative subjects scenarios (e.g., driver's license and passport photos) to unconstrained scenarios with uncooperative subjects (e.g., video surveillance), new challenges are encountered. These challenges are due to variations in ambient illumination, image resolution, background clutter, facial pose, expression, and occlusion. In forensic investigations where the goal is to identify a person of interest, often based on low quality face images and videos, we need to utilize whatever source of information is available about the person. This could include one or more video tracks, multiple still images captured by bystanders (using, for example, their mobile phones), 3-D face models constructed from image(s) and video(s), and verbal descriptions of the subject provided by witnesses. These verbal descriptions can be used to generate a face sketch and provide ancillary information about the person of interest (e.g., gender, race, and age). While traditional face matching methods generally take a single media (i.e., a still face image, video track, or face sketch) as input, this paper considers using the entire gamut of media as a probe to generate a single candidate list for the person of interest. We show that the proposed approach boosts the likelihood of correctly identifying the person of interest through the use of different fusion schemes, 3-D face models, and incorporation of quality measures for fusion and video frame selection. Lacey Best-Rowden, Hu Han 0001, Charles Otto, Brendan Klare, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | The FaceSketchID System: Matching Facial Composites to MugshotsabstractFacial composites are widely used by law enforcement agencies to assist in the identification and apprehension of suspects involved in criminal activities. These composites, generated from witness descriptions, are posted in public places and media with the hope that some viewers will provide tips about the identity of the suspect. This method of identifying suspects is slow, tedious, and may not lead to the timely apprehension of a suspect. Hence, there is a need for a method that can automatically and efficiently match facial composites to large police mugshot databases. Because of this requirement, facial composite recognition is an important topic for biometrics researchers. While substantial progress has been made in nonforensic facial composite (or viewed composite) recognition over the past decade, very little work has been done using operational composites relevant to law enforcement agencies. Furthermore, no facial composite to mugshot matching systems have been documented that are readily deployable as standalone software. Thus, the contributions of this paper include: 1) an exploration of composite recognition use cases involving multiple forms of facial composites; 2) the FaceSketchID System, a scalable, and operationally deployable software system that achieves state-of-the-art matching accuracy on facial composites using two algorithms (holistic and component based); and 3) a study of the effects of training data on algorithm performance. We present experimental results using a large mugshot gallery that is representative of a law enforcement agency’s mugshot database. All results are compared against three state-of-the-art commercial-off-the-shelf face recognition systems. Scott Klum, Hu Han 0001, Brendan Klare, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | A comparative study on illumination preprocessing in face recognition
Hu Han 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
Pattern Recognit. | 1 |
| 2013 | Matching Composite Sketches to Face Photos: A Component-Based ApproachabstractThe problem of automatically matching composite sketches to facial photographs is addressed in this paper. Previous research on sketch recognition focused on matching sketches drawn by professional artists who either looked directly at the subjects (viewed sketches) or used a verbal description of the subject's appearance as provided by an eyewitness (forensic sketches). Unlike sketches hand drawn by artists, composite sketches are synthesized using one of the several facial composite software systems available to law enforcement agencies. We propose a component-based representation (CBR) approach to measure the similarity between a composite sketch and mugshot photograph. Specifically, we first automatically detect facial landmarks in composite sketches and face photos using an active shape model (ASM). Features are then extracted for each facial component using multiscale local binary patterns (MLBPs), and per component similarity is calculated. Finally, the similarity scores obtained from individual facial components are fused together, yielding a similarity score between a composite sketch and a face photo. Matching performance is further improved by filtering the large gallery of mugshot images using gender information. Experimental results on matching 123 composite sketches against two galleries with 10,123 and 1,316 mugshots show that the proposed method achieves promising performance (rank-100 accuracies of 77.2% and 89.4%, respectively) compared to a leading commercial face recognition system (rank-100 accuracies of 22.8% and 52.0%) and densely sampled MLBP on holistic faces (rank-100 accuracies of 27.6% and 10.6%). We believe our prototype system will be of great value to law enforcement agencies in apprehending suspects in a timely fashion. Hu Han 0001, Brendan Klare, Kathryn Bonnen, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Separability Oriented Preprocessing for Illumination-Insensitive Face Recognition
Hu Han 0001, Shiguang Shan, Xilin Chen 0001, Shihong Lao, Wen Gao 0001 |
ECCV (7) | 1 |
| 2010 | Lighting Aware Preprocessing for Face Recognition across Varying Illumination
Hu Han 0001, Shiguang Shan, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001 |
ECCV (2) | 1 |
| 2010 | Gray-scale super-resolution for face recognition from low Gray-scale resolution face imagesabstractToday's camera sensors usually have a high gray-scale resolution, e.g. 256, however, due to the dramatic lighting variations, the gray-scales distributed to the face region might be far less than 256. Therefore, besides low spatial resolution, a practical face recognition system must also handle degraded face images of low gray-scale resolution (LGR). In the last decade, low spatial resolution problem has been studied prevalently, but LGR problem was rarely studied. Aiming at robust face recognition, this paper makes a first primary attempt to investigate explicitly the LGR problem and empirically reveals that LGR indeed degrades face recognition method significantly. Possible solutions to the problem are discussed and grouped into three categories: gray-scale resolution invariant features, gray-scale degradation modeling and Gray-scale Super-Resolution (GSR). Then, we propose a Coupled Subspace Analysis (CSA) based GSR method to recover the high gray-scale resolution image from a single input LGR image. Extensive experiments on FERET and CMU-PIE face databases show that the proposed method can not only dramatically increase the gray-scale resolution and visualization quality, but also impressively improve the accuracy of face recognition. Hu Han 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
ICIP | 1 |
| 2010 | Maximizing intra-individual correlations for illumination-insensitive face recognitionabstractIllumination variation has been one of the most intractable problems in face recognition and many approaches have been proposed to handle illumination problem in the last decades of years. The key problem is how to get stable similarity measurements between two face images of the same individual but captured under dramatically different lighting conditions. We propose a framework to optimize the illumination normalization for a pair of gallery and probe face images by maximizing a correlation (MAC) between them. The illumination normalization in the proposed framework tends to maximize the intra-individual correlations instead of both the inter- and intra-individual correlations. Experiments on Extended YaleB and CMU-PIE face databases show the effectiveness of our proposed approach in face recognition across varying lighting conditions. Hu Han 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
ICIP | 1 |
| 2008 | Illumination transfer using homomorphic wavelet filtering and its application to light-insensitive face recognitionabstractIn this paper, we propose a novel homomorphic wavelet filtering based illumination transfer technique to change the dominant lighting of one face image (source face image) to another face image (reference face image ). Specifically, in the proposed method, based on the ldquoreflectance-illuminationrdquo imaging model, we first obtain an approximate estimate of the illumination component of the face image through a wavelet-based Multiresolution Analysis (MRA) in the logarithm domain of the input image. Then, a homomorphic filtering procedure is applied to improve the accuracy of the illumination component estimation. Finally, the source face image is re-lighted by substituting the estimated illumination component of the reference image for that of the source image. The proposed method is entirely an image processing based method without any 3D geometry modeling steps, so it is simple but effective. The method is also applied easily to illumination invariant face recognition by transferring a standard illumination to all the face images. Experimental results show that our method is quite effective for both illumination transfer and illumination-insensitive face recognition. Hu Han 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
FG | 1 |