EDBT 2026 Demo / reviewers in the wild / expert
Jiani Hu
dblp:20/5087
· DBLP profile ↗
72ranked-venue papers
5as first author
21since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FASTER: Face Attribute Sliders with Semantic RewardsabstractLarge-scale text-to-image generative models have demonstrated remarkable success in generating diverse and high-quality faces. However, current methods for face editing often unintentionally modify facial features that are intended to be preserved. Multi-step denoising methods necessitate storing multi-step gradients, leading to considerable time and memory consumption. In this study, we propose FASTER(Face Attribute Sliders wiTh sEmantic Rewards), an effective method that employs stable diffusion models for face attribute editing. The key idea is to identify a low-rank attribute editing direction by leveraging attribute reward and S-CLIP reward between the original face and the edited face. This process helps to establish the desired face attribute slider. To acquire the edited face, we introduce an efficient one-step reward technique by utilizing denoised results at random timesteps for learning. This technique reduces training time by 6x. FASTER achieves 98.67% editing accuracy while simultaneously improving attribute preservation by nearly 10% compared to other methods on the CelebA-HQ dataset, all without compromising identity information. Jingyan Chen, Lanxiang Zhou, Han Fang 0002, Zerun Feng, Chao Ban, Hao Sun 0038, Jiani Hu |
ICASSP | 8 |
| 2025 | Towards Interactive Deepfake AnalysisabstractExisting deepfake analysis methods are primarily based on discriminative models, which significantly limit their application scenarios. This paper aims to explore interactive deepfake analysis by performing instruction tuning on multi-modal large language models (MLLMs). This will face challenges such as the lack of datasets and benchmarks, and low training efficiency. To address these issues, we introduce (1) a GPT-assisted data construction process resulting in an instruction-following dataset called DFA-Instruct, (2) a benchmark named DFA-Bench, designed to comprehensively evaluate the capabilities of MLLMs in deepfake detection, deepfake classification, and artifact description, and (3) construct an interactive deepfake analysis system called DFA-GPT, as a strong baseline for the community, with the Low-Rank Adaptation (LoRA) module. The dataset and code will be made available at https://github.com/lxq1000/DFA-Instruct to facilitate further research. Lixiong Qin, Yuhan Qiu, Dingheng Zeng, Jiani Hu, Weihong Deng |
ICASSP | 6 |
| 2025 | Variational Vision Transformer with Anti-Over-Smoothing Strategy for Robust Face Recognition
Yuying Zhao, Jiani Hu, Chun-Guang Li |
PRCV (7) | 2 |
| 2025 | Marginal debiased network for fair visual recognition
Mei Wang 0001, Weihong Deng, Jiani Hu, Sen Su |
Pattern Recognit. | 3 |
| 2025 | Unsupervised evaluation for out-of-distribution detection
Jiani Hu, Dongchao Wen, Weihong Deng |
Pattern Recognit. | 2 |
| 2025 | A perturbed match filtering approach for face image quality assessment
Yuying Zhao, Mei Wang 0001, Jiani Hu, Weihong Deng, Chun-Guang Li |
Pattern Recognit. | 3 |
| 2025 | DDL: Dynamic Direction Learning for Semi-Supervised Facial Expression RecognitionabstractMost semi-supervised facial expression recognition (FER) algorithms leverage pseudo-labeling to mine additional information from unlabeled samples. Despite its good performance, two critical issues persist: class imbalance and domain shift. The former is a typical challenge due to the significant variation in sample numbers across different FER classes, resulting in highly imbalanced pseudo labels in existing semi-supervised methods. For the latter, given that labeled and unlabeled data usually come from different sources, a considerable domain gap might exist, leading the model to generate low-quality pseudo labels. To tackle these issues, we introduce a novel semi-supervised FER algorithm called Dynamic Direction Learning (DDL), which consists of adaptive balance learning (ABL) and adaptive alignment learning (AAL). ABL allows a balanced training process by dynamically adjusting the constraints of self-training based on the performance of a balanced validation dataset. Moreover, AAL adaptively aligns the feature distribution of labeled and unlabeled data by minimizing their distance in feature space. Additionally, a role rotation mechanism (RRM) is proposed to avoid confirmation bias, which further improves self-training. Extensive experiments demonstrate that DDL achieves state-of-the-art performance on different FER datasets. Yuhang Zhang 0016, Han Fang 0002, Jiani Hu, Weihong Deng |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Generalizable Facial Expression Recognition
Yuhang Zhang 0016, Xiuqi Zheng, Chenyi Liang, Jiani Hu, Weihong Deng |
ECCV (14) | 4 |
| 2024 | FedSC: Federated Generalized Face Anti-Spoofing via Shuffled Codebook
Mei Wang 0001, Weihong Deng, Jiani Hu |
ICPR (3) | 4 |
| 2024 | Efficient Face Super-Resolution via Wavelet-based Feature Enhancement NetworkabstractFace super-resolution aims to reconstruct a high-resolution face image from a low-resolution face image. Previous methods typically employ an encoder-decoder structure to extract facial structural features, where the direct downsampling inevitably introduces distortions, especially to high-frequency features such as edges. To address this issue, we propose a wavelet-based feature enhancement network, which mitigates feature distortion by losslessly decomposing the input feature into high and low-frequency components using the wavelet transform and processing them separately. To improve the efficiency of facial feature extraction, a full domain Transformer is further proposed to enhance local, regional, and global facial features. Such designs allow our method to perform better without stacking many modules as previous methods did. Experiments show that our method effectively balances performance, model size, and speed. Code link: https://github.com/PRIS-CV/WFEN. Heng Guo 0003, Xuannan Liu, Kongming Liang, Jiani Hu, Zhanyu Ma, Jun Guo 0002 |
ACM Multimedia | 5 |
| 2024 | GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningabstractRecent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality. Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038 |
ACM Multimedia | 10 |
| 2024 | Joint recognition of basic and compound facial expressions by mining latent soft labels
Mei Wang 0001, Bo Xiao 0006, Jiani Hu, Weihong Deng |
Pattern Recognit. | 4 |
| 2024 | SwinFace: A Multi-Task Transformer for Face Recognition, Expression Recognition, Age Estimation and Attribute EstimationabstractIn recent years, vision transformers have been introduced into face recognition and analysis and have achieved performance breakthroughs. However, most previous methods generally train a single model or an ensemble of models to perform the desired task, which ignores the synergy among different tasks and fails to achieve improved prediction accuracy, increased data efficiency, and reduced training time. This paper presents a multi-purpose algorithm for simultaneous face recognition, facial expression recognition, age estimation, and face attribute estimation (40 attributes including gender) based on a single Swin Transformer. Our design, the SwinFace, consists of a single shared backbone together with a subnet for each set of related tasks. To address the conflicts among multiple tasks and meet the different demands of tasks, a Multi-Level Channel Attention (MLCA) module is integrated into each task-specific analysis subnet, which can adaptively select the features from optimal levels and channels to perform the desired tasks. Extensive experiments show that the proposed model has a better understanding of the face and achieves excellent performance for all tasks. Especially, it achieves 90.97% accuracy on RAF-DB and 0.22 ϵ-error on CLAP2015, which are state-of-the-art results on facial expression recognition and age estimation respectively. The code and models will be made publicly available at https://github.com/lxq1000/SwinFace. Lixiong Qin, Mei Wang 0001, Chao Deng 0002, Jiani Hu, Weihong Deng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Momentum Distillation Improves Multimodal Sentiment Analysis
Weihong Deng, Jiani Hu |
PRCV (1) | 3 |
| 2022 | A universal approach for integrating super large-scale single-cell transcriptomes by exploring gene rankingsabstractAdvancement in single-cell RNA sequencing leads to exponential accumulation of single-cell expression data. However, there is still lack of tools that could integrate these unlimited accumulations of single-cell expression data. Here, we presented a universal approach iSEEEK for integrating super large-scale single-cell expression via exploring expression rankings of top-expressing genes. We developed iSEEEK with 11.9 million single cells. We demonstrated the efficiency of iSEEEK with canonical single-cell downstream tasks on five heterogenous datasets encompassing human and mouse samples. iSEEEK achieved good clustering performance benchmarked against well-annotated cell labels. In addition, iSEEEK could transfer its knowledge learned from large-scale expression data on new dataset that was not involved in its development. iSEEEK enables identification of gene-gene interaction networks that are characteristic of specific cell types. Our study presents a simple and yet effective method to integrate super large-scale single-cell transcriptomes and would facilitate translational single-cell research from bench to bedside. Hongru Shen, Xilin Shen, Mengyao Feng, Yichen Yang 0004, Jiani Hu, Jilei Liu, Jilong Yang, Xiangchun Li |
Briefings Bioinform. | 8 |
| 2022 | Scalable batch-correction approach for integrating large-scale single-cell transcriptomesabstractIntegration of accumulative large-scale single-cell transcriptomes requires scalable batch-correction approaches. Here we propose Fugue, a simple and efficient batch-correction method that is scalable for integrating super large-scale single-cell transcriptomes from diverse sources. The core idea of the method is to encode batch information as trainable parameters and add it to single-cell expression profile; subsequently, a contrastive learning approach is used to learn feature representation of the additive expression profile. We demonstrate the scalability of Fugue by integrating all single cells obtained from the Human Cell Atlas. We benchmark Fugue against current state-of-the-art methods and show that Fugue consistently achieves improved performance in terms of data alignment and clustering preservation. Our study will facilitate the integration of single-cell transcriptomes at increasingly large scale. Xilin Shen, Hongru Shen, Mengyao Feng, Jiani Hu, Jilei Liu, Yichen Yang 0004, Xiangchun Li |
Briefings Bioinform. | 5 |
| 2022 | Dynamic Training Data Dropout for Robust Deep Face RecognitionabstractLearning with noise is a practically challenging problem in deep face recognition. Despite the success of large margin softmax loss functions, these methods are designed for clean face databases. Considering the inevitable noise in the large scale databases, we first analyze the performance of noise in the training databases. For noise-robust deep face recognition, we propose a dynamic training data dropout (DTDD) method to dynamically filter the noise in the training database and gradually form a stable refined database for model learning. Specifically, we leverage the information provided by the model predictions of accumulated training epochs, which can distinguish regular samples and noise effectively and accurately. The proposed DTDD method is easy and stable for implementation, and can be combined with existing state-of-the-art loss functions and network architectures. Extensive experiments on CASIA-WebFace, VGGFace2, and MS-Celeb-1 M databases empirically demonstrate that our proposed method can robustly train deep face recognition models in the presence of label noise and low quality images. Yaoyao Zhong, Weihong Deng, Han Fang 0002, Jiani Hu, Dongyue Zhao, Dongchao Wen |
IEEE Trans. Multim. | 4 |
| 2021 | Augmented Face Representation Learning via Transitive DistillationabstractThe wild face of large variations is hard to recognize in unconstrained scenarios. To tackle this issue, existing works synthesize and augment the variation-specific faces for recognition. However, directly feeding generated samples results in negative transfer, because the feature spaces are shifted compared with normal samples. Instead, we propose a transitive distillation network (TDNet) that introduces a transitive domain to transfer cross-variation representations, which alleviates the negative influence of synthesized data. Specifically, data of diverse variations are firstly synthesized. Then we construct distributions from different variations as teachers to distill student. The negative transfer is mitigated by adopting adaptor as a bridge to break large domain distance. To handle faces of different quality, we propose a novel strategy to define easy and hard samples, which are utilized to select specific transitive status. Meanwhile, bilateral classification with curriculum learning is proposed to improve confidence of synthesized data gradually, enhancing the robustness of representation learning. Experiments show that our method achieves superiority on unconstrained face benchmarks such as IJB-C and SCface, while maintaining competence on general test sets. Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen |
FG | 4 |
| 2021 | Adaptive Label Noise Cleaning with Meta-Supervision for Deep Face RecognitionabstractThe training of a deep face recognition system usually faces the interference of label noise in the training data. However, it is difficult to obtain a high-precision cleaning model to remove these noises. In this paper, we propose an adaptive label noise cleaning algorithm based on meta-learning for face recognition datasets, which can learn the distribution of the data to be cleaned and make automatic adjustments based on class differences. It first learns re-liable cleaning knowledge from well-labeled noisy data, then gradually transfers it to the target data with meta-supervision to improve performance. A threshold adapter module is also proposed to address the drift problem in transfer learning methods. Extensive experiments clean two noisy in-the-wild face recognition datasets and show the effectiveness of the proposed method to reach state-of-the-art performance on the IJB-C face recognition benchmark. Yaobin Zhang, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen |
ICCV | 4 |
| 2021 | Orthogonality Loss: Learning Discriminative Representations for Face RecognitionabstractConvolutional neural networks have achieved excellent performance on face recognition (FR) by learning the high discriminative features with advanced loss functions. These improved loss functions share the similar idea for maximizing inter-class variance or minimizing intra-class variance. In this article, from a different perspective, we consider enlarging the inter-class variance by directly penalizing weight vectors of last fully connected layer, which represent the center of classes. To the end, we propose Orthogonality loss as an elegant penalty item appends to common classification loss to learn the discriminative representations. The main idea is that in order for weight vectors to be discriminative, it should be as close as possible to be orthogonal to each other in the vector space. More specifically, the optimization objective of Orthogonality loss is the first moment and second moment of cosine similarity of weight vectors. We performed the empirical studies through simulating the long-tail datasets to show the generalization ability of the proposed approach on long-tail distribution datasets. Further, extensive experiments on large-scale face recognition benchmarks including the Labeled Face in the Wild (LFW), the IARPA Janus Benchmark A (IJB-A), IJB-B, IJB-C, MegaFace Challenge 1 (MF1) and MS-Celeb-1M Low-shot Learning demonstrated that Orthogonality loss outperforms strong baselines, which showcases the extensive suitability and effectiveness of Orthogonality loss. Shan-Ming Yang, Weihong Deng, Mei Wang 0001, Junping Du 0001, Jiani Hu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face RecognitionabstractDeep face recognition has achieved great success due to large-scale training databases and rapidly developing loss functions. The existing algorithms devote to realizing an ideal idea: minimizing the intra-class distance and maximizing the inter-class distance. However, they may neglect that there are also low quality training images which should not be optimized in this strict way. Considering the imperfection of training databases, we propose that intra-class and inter-class objectives can be optimized in a moderate way to mitigate overfitting problem, and further propose a novel loss function, named sigmoid-constrained hypersphere loss (SFace). Specifically, SFace imposes intra-class and inter-class constraints on a hypersphere manifold, which are controlled by two sigmoid gradient re-scale functions respectively. The sigmoid curves precisely re-scale the intra-class and inter-class gradients so that training samples can be optimized to some degree. Therefore, SFace can make a better balance between decreasing the intra-class distances for clean examples and preventing overfitting to the label noise, and contributes more robust deep face recognition models. Extensive experiments of models trained on CASIA-WebFace, VGGFace2, and MS-Celeb-1M databases, and evaluated on several face recognition benchmarks, such as LFW, MegaFace and IJB-C databases, have demonstrated the superiority of SFace. Yaoyao Zhong, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen |
IEEE Trans. Image Process. | 3 |
| 2020 | Global-Local GCN: Large-Scale Label Noise Cleansing for Face RecognitionabstractIn the field of face recognition, large-scale web-collected datasets are essential for learning discriminative representations, but they suffer from noisy identity labels, such as outliers and label flips. It is beneficial to automatically cleanse their label noise for improving recognition accuracy. Unfortunately, existing cleansing methods cannot accurately identify noise in the wild. To solve this problem, we propose an effective automatic label noise cleansing framework for face recognition datasets, FaceGraph. Using two cascaded graph convolutional networks, FaceGraph performs global-to-local discrimination to select useful data in a noisy environment. Extensive experiments show that cleansing widely used datasets, such as CASIA-WebFace, VGGFace2, MegaFace2, and MS-Celeb-1M, using the proposed method can improve the recognition performance of state-of-the-art representation learning methods like Arcface. Further, we cleanse massive self-collected celebrity data, namely MillionCelebs, to provide 18.8M images of 636K identities. Training with the new data, Arcface surpasses state-of-the-art performance by a notable margin to reach 95.62% TPR at 1e-5 FPR on the IJB-C benchmark. Yaobin Zhang, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen |
CVPR | 4 |
| 2020 | Generate to Adapt: Resolution Adaption Network for Surveillance Face Recognition
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu |
ECCV (15) | 4 |
| 2020 | FGAN: Fan-Shaped GAN for Racial TransformationabstractRacial bias in face recognition has recently been concerned by both general public and research community. Most face recognition systems have a strong bias in recognition accuracy for different races mainly because of the unbalanced ethnic distribution in their datasets. In this paper, we propose a novel generative adversarial network, which transfer the facial images of one race to corresponding images of other races, to facilitate the data augmentation to balance the ethnic distribution. Our approach can generate more realistic results and make the training process more stable than other image-to-image translation methods such as StarGAN and CycleGAN. Experiments results show the superiority of FGAN to the previous methods on the racial transformation task in terms of visual effects and quantitative results. Besides, we perform extensive experiments to show our data augmentation is beneficial to reduce the racial bias, improving the face recognition rate of non-Caucasian people. Finally, we show the possibility to generate the ethnic independent facial image by the average of various races. Jiancheng Ge, Weihong Deng, Jiani Hu |
IJCB | 4 |
| 2020 | H-AT: Hybrid Attention Transfer for Knowledge Distillation
Yan Qu, Weihong Deng, Jiani Hu |
PRCV (3) | 3 |
| 2020 | Identity-aware CycleGAN for face photo-sketch synthesis and recognition
Yuke Fang, Weihong Deng, Junping Du 0001, Jiani Hu |
Pattern Recognit. | 4 |
| 2019 | Unsupervised Face Normalization With Extreme Pose and Expression in the WildabstractFace recognition achieves great success thanks to the emergence of deep learning. However, many contemporary face recognition models still have limited invariance to strong intra-personal variations such as large pose changes. Face normalization provides an effective and cheap way to distil face identity and dispel face variances for recognition. We focus on face generation in the wild with unpaired data. To this end, we propose a Face Normalization Model (FNM) to generate a frontal, neutral expression, photorealistic face image for face recognition. FNM is a well-designed Generative Adversarial Network (GAN) with three distinct novelties. First, a face expert network is introduced to construct generator and provide the ability of retaining face identity. Second, with the reconstruction of normal face, pixel-wise loss is applied to stabilize optimization process. Third, we present a series of face attention discriminators to refine local textures. FNM could recover canonical-view, expression-free image and directly improve the performance of face recognition model. Extensive qualitative and quantitative experiments on both controlled and in-the-wild databases demonstrate the superiority of our face normalization method. Yichen Qian, Weihong Deng, Jiani Hu |
CVPR | 3 |
| 2019 | Unequal-Training for Deep Face Recognition With Long-Tailed Noisy DataabstractLarge-scale face datasets usually exhibit a massive number of classes, a long-tailed distribution, and severe label noise, which undoubtedly aggravate the difficulty of training. In this paper, we propose a training strategy that treats the head data and the tail data in an unequal way, accompanying with noise-robust loss functions, to take full advantage of their respective characteristics. Specifically, the unequal-training framework provides two training data streams: the first stream applies the head data to learn discriminative face representation supervised by Noise Resistance loss; the second stream applies the tail data to learn auxiliary information by gradually mining the stable discriminative information from confusing tail classes. Consequently, both training streams offer complementary information to deep feature learning. Extensive experiments have demonstrated the effectiveness of the new unequal-training framework and loss functions. Better yet, our method could save a significant amount of GPU memory. With our method, we achieve the best result on MegaFace Challenge 2 (MF2) given a large-scale noisy training data set. Yaoyao Zhong, Weihong Deng, Jiani Hu, Jianteng Peng, Xunqiang Tao, Yaohai Huang |
CVPR | 4 |
| 2019 | Mixed High-Order Attention Network for Person Re-IdentificationabstractAttention has become more attractive in person re-identification (ReID) as it is capable of biasing the allocation of available resources towards the most informative parts of an input signal. However, state-of-the-art works concentrate only on coarse or first-order attention design, e.g. spatial and channels attention, while rarely exploring higher-order attention mechanism. We take a step towards addressing this problem. In this paper, we first propose the High-Order Attention (HOA) module to model and utilize the complex and high-order statistics information in attention mechanism, so as to capture the subtle differences among pedestrians and to produce the discriminative attention proposals. Then, rethinking person ReID as a zero-shot learning problem, we propose the Mixed High-Order Attention Network (MHN) to further enhance the discrimination and richness of attention knowledge in an explicit manner. Extensive experiments have been conducted to validate the superiority of our MHN for person ReID over a wide variety of state-of-the-art methods on three large-scale datasets, including Market-1501, DukeMTMC-ReID and CUHK03-NP. Code is available at http://www.bhchen.cn. Binghui Chen, Weihong Deng, Jiani Hu |
ICCV | 3 |
| 2019 | Fair Loss: Margin-Aware Reinforcement Learning for Deep Face RecognitionabstractRecently, large-margin softmax loss methods, such as angular softmax loss (SphereFace), large margin cosine loss (CosFace), and additive angular margin loss (ArcFace), have demonstrated impressive performance on deep face recognition. These methods incorporate a fixed additive margin to all the classes, ignoring the class imbalance problem. However, imbalanced problem widely exists in various real-world face datasets, in which samples from some classes are in a higher number than others. We argue that the number of a class would influence its demand for the additive margin. In this paper, we introduce a new margin-aware reinforcement learning based loss function, namely fair loss, in which each class will learn an appropriate adaptive margin by Deep Q-learning. Specifically, we train an agent to learn a margin adaptive strategy for each class, and make the additive margins for different classes more reasonable. Our method has better performance than present large-margin loss functions on three benchmarks, Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace, which demonstrates that our method could learn better face representation on imbalanced face datasets. Weihong Deng, Yaoyao Zhong, Jiani Hu, Xunqiang Tao, Yaohai Huang |
ICCV | 5 |
| 2019 | Racial Faces in the Wild: Reducing Racial Bias by Information Maximization Adaptation NetworkabstractRacial bias is an important issue in biometric, but has not been thoroughly studied in deep face recognition. In this paper, we first contribute a dedicated dataset called Racial Faces in-the-Wild (RFW) database, on which we firmly validated the racial bias of four commercial APIs and four state-of-the-art (SOTA) algorithms. Then, we further present the solution using deep unsupervised domain adaptation and propose a deep information maximization adaptation network (IMAN) to alleviate this bias by using Caucasian as source domain and other races as target domains. This unsupervised method simultaneously aligns global distribution to decrease race gap at domain-level, and learns the discriminative target representations at cluster level. A novel mutual information loss is proposed to further enhance the discriminative ability of network output without label information. Extensive experiments on RFW, GBU, and IJB-A databases show that IMAN successfully learns features that generalize well across different races and across different databases. Weihong Deng, Jiani Hu, Xunqiang Tao, Yaohai Huang |
ICCV | 3 |
| 2019 | Unsupervised adaptive hashing based on feature clustering
Tongtong Yuan, Weihong Deng, Jiani Hu, Zhanfu An, Yinan Tang |
Neurocomputing | 3 |
| 2019 | Compressive Binary Patterns: Designing a Robust Binary Face Descriptor with Random-Field EigenfiltersabstractA binary descriptor typically consists of three stages: image filtering, binarization, and spatial histogram. This paper first demonstrates that the binary code of the maximum-variance filtering responses leads to the lowest bit error rate under Gaussian noise. Then, an optimal eigenfilter bank is derived from a universal assumption on the local stationary random field. Finally, compressive binary patterns (CBP) is designed by replacing the local derivative filters of local binary patterns (LBP) with these novel random-field eigenfilters, which leads to a compact and robust binary descriptor that characterizes the most stable local structures that are resistant to image noise and degradation. A scattering-like operator is subsequently applied to enhance the distinctiveness of the descriptor. Surprisingly, the results obtained from experiments on the FERET, LFW, and PaSC databases show that the scattering CBP (SCBP) descriptor, which is handcrafted by only 6 optimal eigenfilters under restrictive assumptions, outperforms the state-of-the-art learning-based face descriptors in terms of both matching accuracy and robustness. In particular, on probe images degraded with noise, blur, JPEG compression, and reduced resolution, SCBP outperforms other descriptors by a greater than 10 percent accuracy margin. Weihong Deng, Jiani Hu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Deep Transfer Network with 3D Morphable Models for Face RecognitionabstractData augmentation using 3D face models to synthesize faces has been demonstrated to be effective for face recognition. However, the model directly trained by using the synthesized faces together with the original real faces is not optimal. In this paper, we propose a novel approach that uses a deep transfer network (DTN) with 3D morphable models (3DMMs) for face recognition to overcome the shortage of labeled face images and the dataset bias between synthesized images and corresponding real images. We first utilize the 3DMM to synthesize faces with various poses to augment the training dataset. Then, we train a deep neural network using the synthesized face images and the original real face images. The results obtained on LFW show that the accuracy of the model utilizing synthesized data only is lower than that of the model using the original data, although the synthesized dataset contains much considerably images with more unconstrained poses. This result shows that a dataset bias exists between the synthesized faces and the real faces. We treat the synthesized faces as the source domain, and we treat the actual faces as the target domain. We use the DTN to alleviate the discrepancy between the source domain and the target domain. The DTN attempts to project source domain samples and target domain samples to a new space where they are fused together such that one cannot distinguish the domain from which a specific image is from. We optimize our DTN based on the maximum mean discrepancy (MMD) of the shared feature extraction layers and the discrimination layers. We choose AlexNet and Inception-ResNet-V1 as our benchmark models. The proposed method is also evaluated on the LFW and SLLFW databases. The experimental results show that our method can effectively address the domain discrepancy. Moreover, the dataset bias between the synthesized data and the real data is remarkably reduced, which can thus improve the performance of the convolutional neural network (CNN) model. Zhanfu An, Weihong Deng, Tongtong Yuan, Jiani Hu |
FG | 4 |
| 2018 | Deep Unsupervised Domain Adaptation for Face RecognitionabstractFace recognition is challenge task which involves determining the identity of facial images. With availability of a massive amount of labeled facial images gathered from Internet, deep convolution neural networks(DCNNs) have achieved great success in face recognition tasks. Those images are gathered from unconstrain environment, which contain people with different ethnicity, age, gender and so on. However, in the actual application scenario, the target face database may be gathered under different conditions compared with source training dataset, e.g. different ethnicity, different age distribution, disparate shooting environment. These factors increase domain discrepancy between source training database and target application database which makes the learnt model degenerate in target database. Meanwhile, for the target database where labeled data are lacking or unavailable, directly using target data to fine-tune pre-learnt model becomes intractable and impractical. In this paper, we adopt unsupervised transfer learning methods to address this issue. To alleviate the discrepancy between source and target face database and ensure the generalization ability of the model, we constrain the maximum mean discrepancy (MMD) between source database and target database and utilize the massive amount of labeled facial images of source database to training the deep neural network at the same time. We evaluate our method on two face recognition benchmarks and significantly enhance the performance without utilizing the target label. Zimeng Luo, Jiani Hu, Weihong Deng, Haifeng Shen |
FG | 2 |
| 2018 | Task Specific Networks for Identity and Face VariationabstractPose and illumination variations are considered as two main challenges that face recognition system encounters. Most existing methods perform face normalization, aiming at untangling identity representation from these variations to improve recognition accuracy. Taking into account face variation representations, this paper proposes Task Specific Networks for the two representations with two novelties. First, we rotate and normalize face image to multi-pose view for one subtask, and learn face variation representations for another. Second, we learn face variation representations in an unsupervised way, which is more robust and more universal. We couple these two representations in the part of reconstructing the original face, where the two representations effect and restrict each other. Extensive experiments demonstrate the superiority of our method in both learning representations and rotating non-frontal face image. Yichen Qian, Weihong Deng, Jiani Hu |
FG | 3 |
| 2018 | Unsupervised Domain Adaptation by regularizing Softmax ActivationabstractIn recent years, deep learning has achieved very good results with a large amount of labeled data but can't generalize well when there is a shift between train data distribution(source domain) and test data distribution(target domain). Deep domain adaptation is an effective way to solve this problem. Many previous deep domain adaptation methods are based on Maximum Mean Discrepancy(MMD). These methods use MMD to regularize the feature-layers to learn transferable features directly. However, in this paper, we propose to use MMD to regularize the softmax predictions to learn more transferable features by backpropagating. At the same time, in order to get discriminative classifiers, we propose to depart but bridge the domain-invariant feature, which is learned by matching the feature distribution, and the classifying feature, which is before the final softmax, by a Residual-block. Our method can be implemented in almost all deep networks with softmax classifiers. In order to compare with the recent deep domain adaptation methods, we implement our method on Alexnet and outperforms almost all state-of-the-art methods on standard domain adaptation benchmarks. Cunbin Gui, Jiani Hu |
ICPR | 2 |
| 2018 | Local Subclass Constraint for Facial Expression Recognition in the WildabstractThe Automated Facial Expression Recognition (FER) in the wild is still a challenge problem. Currently, most of Deep Convolutional Neural Networks (DCNNs) based FER methods adopt softmax cross-entropy loss to encourage the separability of inter-class features. Many deep embedding approaches (e.g. contrastive loss, triplet loss, center loss) have been extended to the field of FER to enhance the discriminative ability of deep expression features and obtain the predictive effect. In this work, we present a novel deep embedding approach explicitly designed to respect the huge intra-class variation of expression features while learning discriminative expression features. We aim at forming a locally compact representation space structure through minimizing the distance between samples and their nearest subclass center. We demonstrate the effectiveness of this idea on RAF (Real-world Affective Faces) database. The experiment results show that our approaches can not only improve the classification performance but also adaptively learn a locally compact and expression intensity-aware feature space structure. We further extend our models to Static Facial Expressions in the Wild (SFEW) dataset and the results show the generalized ability of our approaches. Zimeng Luo, Jiani Hu, Weihong Deng |
ICPR | 2 |
| 2018 | Facial landmark localization by enhanced convolutional neural network
Weihong Deng, Yuke Fang, Zhenqi Xu, Jiani Hu |
Neurocomputing | 4 |
| 2018 | Face Recognition via Collaborative Representation: Its Discriminant Nature and Superposed RepresentationabstractCollaborative representation methods, such as sparse subspace clustering (SSC) and sparse representation-based classification (SRC), have achieved great success in face clustering and classification by directly utilizing the training images as the dictionary bases. In this paper, we reveal that the superior performance of collaborative representation relies heavily on the sufficiently large class separability of the controlled face datasets such as Extended Yale B. On the uncontrolled or undersampled dataset, however, collaborative representation suffers from the misleading coefficients of the incorrect classes. To address this limitation, inspired by the success of linear discriminant analysis (LDA), we develop a superposed linear representation classifier (SLRC) to cast the recognition problem by representing the test image in term of a superposition of the class centroids and the shared intra-class differences. In spite of its simplicity and approximation, the SLRC largely improves the generalization ability of collaborative representation, and competes well with more sophisticated dictionary learning techniques, on the experiments of AR and FRGC databases. Enforced with the sparsity constraint, SLRC achieves the state-of-the-art performance on FERET database using single sample per person. Weihong Deng, Jiani Hu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | From one to many: Pose-Aware Metric Learning for single-sample face recognition
Weihong Deng, Jiani Hu, Zhongjun Wu, Jun Guo 0002 |
Pattern Recognit. | 2 |
| 2017 | Metric-Promoted Siamese Network for Gender ClassificationabstractGender classification is a fundamental and important application in computer vision, and it has become a research hotspot. Real-world applications require gender classification in unconstrained conditions where traditional methods are not appropriate. This paper proposes a Deep Convolutional Neural Network for feature extraction together with fully-connected layers for metric learning. A Siamese network is built for similarity measuring to promote the performance of classification. Extensive experiments on several databases demonstrate that a significant improvement can be obtained for gender classification tasks in both constrained and unconstrained conditions. Yipeng Huang 0002, Shuying Liu, Jiani Hu, Weihong Deng |
FG | 3 |
| 2017 | Learning Local Responses of Facial Landmarks with Conditional Variational Auto-Encoder for Face AlignmentabstractThis work proposes a novel convolutional neural network architecture which can locate landmarks accurately by learning local responses of facial landmarks. The network consists of a Conditional Variational Auto-Encoder(CVAE) and a Deep Convolutional Neural Network(DCNN). The CVAE is used to learn the response maps of facial landmarks from face images and the DCNN is used to learn accurate landmark locations from the response maps and facial textures. The CVAE consists of a face encoder, which extracts high-level information from raw pixels, and a decoder which outputs local response maps from high-level coding. We derive the CVAE used for catching local responses as an optimization problem, which can be solved through back-propagation. Extensive experiments show that the proposed CVAE can learn better local response maps than Fully Convolutional Network(FCN). Our method outperforms state-of-the-art methods on AFLW(5 points) and the challenging subset of 300-W(68 points), which means our method shows advantages in the condition of complex poses and expressions. Shuying Liu, Yipeng Huang 0002, Jiani Hu, Weihong Deng |
FG | 3 |
| 2017 | FolkPopularityRank: Tag Recommendation for Enhancing Social Popularity using Text Tags in Content Sharing ServicesabstractIn this study, we address two emerging yet challenging problems in social media: (1) scoring the text tags in terms of the influence to the numbers of views, comments, and favorite ratings of images and videos on content sharing services, and (2) recommending additional tags to increase such popularity-related numbers. For these purposes, we present the FolkPopularityRank algorithm, which can score text tags based on their ability to influence the popularity-related numbers. The FolkPopularityRank algorithm is inspired by the PageRank and FolkRank algorithms but the scores of the tags are calculated not only by the co-occurrence of the tags but also by considering the popularity-related numbers of the content. To the best of our knowledge, this is the first attempt to recommending tags that can enhance popularity attributes of social media. We conducted extensive experiments with about 1,000 images. We uploaded the photos with the recommended tags along with the original tags to Flickr as a real test, and obtained very promising results. Toshihiko Yamasaki, Jiani Hu, Shumpei Sano, Kiyoharu Aizawa |
IJCAI | 2 |
| 2017 | Become Popular in SNS: Tag Recommendation using FolkPopularityRank to Enhance Social PopularityabstractIn this demo, we address two emerging yet challenging problems in social media: (1) scoring the text tags in terms of the influence to the numbers of views, comments, and favorite ratings of images and videos on content sharing services, and (2) recommending additional tags to increase such popularity-related numbers. For these purposes, we present a demo using our FolkPopularityRank (FP-Rank) algorithm, which can score and recommend text tags based on their ability to influence the popularity-related numbers. Our experiments using 1,000 photos showed that we can achieve 1.6 times more views than the original tag sets in Flickr just by adding tags recommended by FP-Rank. Toshihiko Yamasaki, Yiwei Zhang 0014, Jiani Hu, Shumpei Sano, Kiyoharu Aizawa |
IJCAI | 3 |
| 2017 | A Tag Recommendation System for Popularity BoostingabstractIn order to support users in the tagging process and recommendation, we had proposed two tag ranking algorithms, Document Frequency-Weights from regression and Folk Popularity Rank, which can extract tags greatly influencing popularity. We have developed a tag recommendation system using the algorithm we proposed. The recommended tags are not only for appropriate annotations but also for popularity boosting. Yiwei Zhang 0014, Jiani Hu, Shumpei Sano, Toshihiko Yamasaki, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2017 | Deep transfer network for face recognition using 3D synthesized faceabstractFace recognition has experienced a flurry of advances with deep learning. However, training a model requires a lot of data. In order to meet this condition, some researchers use the 3D rendering technique to synthesize fake face images to expand the training data. Experimental results have demonstrated that this method is an effective way. There exist, however, dataset bias between the real 2D real face images and 3D synthesized face images. In this paper, we use Deep Transfer Network(DTN) to reduce dataset bias. First, we utilize the 3DMM face model to synthesize face images with various poses and natural expression. We choose the Inception-Resnet-V1 as our benchmark model. Then, we optimize our DTN based on maximum mean discrepancy(MMD) of the shared feature extraction layers and the discrimination layers. Our experiments demonstrate that the model jointly trained using synthesized images and real images is more robust than using either dataset (2D real faces or 3D synthesized faces). Furthermore, the performance obtained by our approach is comparable to the-state-of-the-art results to the systems trained on millions of real images. Zhanfu An, Weihong Deng, Jiani Hu |
VCIP | 3 |
| 2017 | Supervised hashing with extreme learning machineabstractSupervised hashing methods, which aim to generate semantic similarity-preserving binary codes, have been proposed to improve the performance of large-scale image retrieval. However, learning binary codes remains an NP-hard problem due to the binary constraints and complex computation. Existing hashing methods have never explored the potentiality of the label information, leading to a limited performance. To address these problems, we propose a simple supervised hashing method based on extreme learning machine (ELM). And we generate the supervised information in ELM by target code learning instead of using the traditional label code to fit the retrieval problem. With this modified label code, our method can produce high-quality binary codes and obtain high retrieval precision. Comprehensive experiments have shown our superiority to other state-of-the-art methods. Tongtong Yuan, Weihong Deng, Jiani Hu |
VCIP | 3 |
| 2017 | Deep probabilities for age estimationabstractHuman age can provide important demographic information. In this paper, we tackle the estimation of age in face images with probabilities. The design of the proposed method is based on the relative order of age labels in the database. The age estimation problem is transformed into a series of binary classifications achieved by convolution neural network. Each classifier is used to judge whether the age of input image is larger than a certain age and the estimated age is obtained by adding probability values of these classification problems. The proposed method: Deep Probabilities (DP) of facial age shows improvements over direct regression and multi-classification methods. Tianyue Zheng, Weihong Deng, Jiani Hu |
VCIP | 3 |
| 2017 | Lighting-aware face frontalization for unconstrained face recognition
Weihong Deng, Jiani Hu, Zhongjun Wu, Jun Guo 0002 |
Pattern Recognit. | 2 |
| 2017 | Fine-grained face verification: FGLFW database, baselines, and human-DCMN partnership
Weihong Deng, Jiani Hu, Nanhai Zhang, Binghui Chen, Jun Guo 0002 |
Pattern Recognit. | 2 |
| 2017 | Deep Correlation Feature Learning for Face Verification in the WildabstractConvolutional neural networks (CNNs) commonly uses the softmax loss function as the supervision signal. In order to enhance the discriminative power of the deeply learned features, this letter proposes a new supervision signal, called correlation loss, for face verification task. Specifically, the correlation loss encourages the large correlation between the deep feature vectors and their corresponding weight vectors in softmax loss. With the joint supervision of softmax loss and correlation loss, the deep correlation feature learning (DCFL) network can learn the deep features with both the interclass separability and the intraclass compactness, which are highly discriminative for face verification. More importantly, by applying the weight vector of softmax function as the class prototype, the proposed correlation loss function is easy to be optimized during the backpropatation of CNN. Finally, the DCFL method achieves 99.55% and 96.06% face verification accuracy using a 64-layer ResNet on the labeled face in-the-Wild (LFW) and you-tube face (YTF) benchmark, respectively. Weihong Deng, Binghui Chen, Yuke Fang, Jiani Hu |
IEEE Signal Process. Lett. | 4 |
| 2016 | Learning Facial Point Response for Alignment by Purely Convolutional Network
Zhenqi Xu, Weihong Deng, Jiani Hu |
ACCV (3) | 3 |
| 2016 | Recurrent convolutional neural network for video classificationabstractVideo classification is more difficult than image classification since additional motion feature between image frames and amount of redundancy in videos should be taken into account. In this work, we proposed a new deep learning architecture called recurrent convolutional neural network (RCNN) which combines convolution operation and recurrent links for video classification tasks. Our architecture can extract the local and dense features from image frames as well as learning the temporal features between consecutive frames. We also explore the effectiveness of sequential sampling and random sampling when training our models, and find out that random sampling is necessary for video classification. The feature maps from our learned model preserve motion from image frames, which is analogous to the persistence of vision in human visual system. We achieved 81.0% classification accuracy without optical flow and 86.3% with optical flow on the UCF-101 dataset, both are competitive to the state-of-the-art methods. Zhenqi Xu, Jiani Hu, Weihong Deng |
ICME | 2 |
| 2016 | Geometry-aware metric learning for similar face recognitionabstractNoticing that face images (from different persons) with high similarity computed by current state-of-the-art methods may be not visually similar, in this paper, we present a new verification problem on judging whether the given faces are similar or not. Similar to “view 2” of Labeled Faces in the Wild (LFW), we construct ten subsets' face pairs using images from LFW. Label of each pair comes from human annotation results. Since similar faces are not from the same person after all, pushing similar faces too close will easily contribute to wrong models. Therefore, we propose a new geometry-aware metric learning (GAML) method which can preserve the similarity of similar faces while enlarge the difference between dissimilar faces. Experimental results show that our method outperforms traditional face verification methods on our similar face dataset. Nanhai Zhang, Jiajie Han, Jiani Hu, Weihong Deng |
ICME | 3 |
| 2015 | DeepEmo: Real-world facial expression analysis via deep learningabstractRecent automatic facial expression recognition research has focused on optimizing performance on a few databases that were collected under controlled pose and lighting conditions, and has produced nearly perfect accuracy. This paper explores the necessary characteristics of the training dataset, feature representations and machine learning algorithms for a system that operates reliably in more realistic conditions. A new database, Real-world Affective Face Database (RAF-DB), is presented which contains about 30,000 greatly-diverse facial images from social networks. Crowdsourcing results suggest that real-world expression recognition problem is a typical imbalanced multi-label classification problem, and the balanced, single-label datasets currently used in the literature could potentially lead research into misleading algorithmic solutions. A deep learning architecture, DeepEmo, is proposed to address the real-world challenge of emotion recognition by learning the highlevel feature representations which are highly effective for discriminating realistic facial expressions. Extensive experimental results show that the deep learning method is significantly superior to handcrafted features, and with the near-frontal pose constraint, human-level recognition accuracy is achievable. Weihong Deng, Jiani Hu, Jun Guo 0002 |
VCIP | 2 |
| 2014 | Transformed Principal Gradient Orientation for Robust and Precise Batch Face Alignment
Weihong Deng, Jiani Hu, Jun Guo 0002 |
ACCV (4) | 2 |
| 2014 | Linear Ranking AnalysisabstractWe extend the classical linear discriminant analysis (LDA) technique to linear ranking analysis (LRA), by considering the ranking order of classes centroids on the projected subspace. Under the constrain on the ranking order of the classes, two criteria are proposed: 1) minimization of the classification error with the assumption that each class is homogenous Guassian distributed, 2) maximization of the sum (average) of the K minimum distances of all neighboring-class (centroid) pairs. Both criteria can be efficiently solved by the convex optimization for one-dimensional subspace. Greedy algorithm is applied to extend the results to the multi-dimensional subspace. Experimental results show that 1) LRA with both criteria achieve state-of-the-art performance on the tasks of ranking learning and zero-shot learning, and 2) the maximum margin criterion provides a discriminative subspace selection method, which can significantly remedy the class separation problem in comparing with several representative extensions of LDA. Weihong Deng, Jiani Hu, Jun Guo 0002 |
CVPR | 2 |
| 2014 | Precise eye localization by fast local linear SVMabstractRecently, discriminative methods such as SVM has been widely used in object location. But there has been no method to perform well enough both at accuracy and speed. For linear SVM, it is hard to separate the nonlinear samples exactly. For kernel SVM, it is hard to be applied to real-time application, because of the computational cost kernel function. Local linear SVM has been proved to be a good tradeoff between fast linear SVM and qualitative best kernel methods. However, it is still time-consuming for real-time application. To design a high efficiency and high precision eye locahzer, first, we deduce a fast variation for LL-S VM which can serve as a more fast and accurate substitute of the traditional nonlinear kernel SVM. Second, to further improve the speed, we also adopt a candidate selection strategy. Extensive experiments on the BioID, FERET, FRGC, and LFW database show that our proposed method achieves favorable localization accuracy against other state-of-the-art methods at a speed as fast as 5ms to localize two eyes. Xiang Sun 0003, Jiani Hu, Weihong Deng |
ICME | 3 |
| 2014 | Online Regression of Grandmother-Cell Responses with Visual Experience Learning for Face RecognitionabstractGrandmother cell is a term in neuroscience to imitate the simplistic notion that the brain has a separate neuron to represent every familiar face, with important properties of sparseness and invariance. This paper proposes a linear regression based classification model for face recognition, which learn a mapping from the training feature vectors to the grandmother-cell-like codes, with one unit corresponding to an individual. Two kinds of visual experiences are incorporated to enhance the generalization capability of the regression mapping. First, the regression model maps the intra-personal facial differences of the unknown faces to the zeros vectors, so that any similar variation on the familiar face would not affect the regression result. Second, to adapt to the evolution of facial appearance, the model feeds the selected testing images back to incrementally retrain the regression mapping, and decrement ally remove the influence of outdated training images, all in an unsupervised manner. Experiments results on Extended Yale B, FERET, and AR databases demonstrate the efficacy of the proposed regression based face recognition algorithms. Jiani Hu, Weihong Deng, Jun Guo 0002 |
ICPR | 1 |
| 2014 | Max-K-Min Distance Analysis for Dimension ReductionabstractWe propose a new criterion for discriminative dimension reduction, Max-K-Min Distance Analysis (MKMDA). Given a data set with C classes, MKMDA maximizes the sum of the K minimum pair wise distance of these C classes on the selected one-dimensional subspace. The set of the possible one-dimensional subspace, for which the order of the projected class centroids is identical, define a convex region with associated convex sum of K smallest margin functions. This allows for the maximization of the margin function using standard convex optimization algorithms. This result is further extended to obtain the d-dimensional subspace for any given d by iterative applying our algorithm to the null space of the (d -- 1)-dimensional subspace. The effectiveness of the proposed criterion and corresponding algorithm is shown by the visualization and classification experiments on both synthetic data and real data sets. Jiani Hu, Weihong Deng, Jun Guo 0002 |
ICPR | 1 |
| 2014 | Transform-Invariant PCA: A Unified Approach to Fully Automatic FaceAlignment, Representation, and RecognitionabstractWe develop a transform-invariant PCA (TIPCA) technique which aims to accurately characterize the intrinsic structures of the human face that are invariant to the in-plane transformations of the training images. Specially, TIPCA alternately aligns the image ensemble and creates the optimal eigenspace, with the objective to minimize the mean square error between the aligned images and their reconstructions. The learning from the FERET facial image ensemble of 1,196 subjects validates the mutual promotion between image alignment and eigenspace representation, which eventually leads to the optimized coding and recognition performance that surpasses the handcrafted alignment based on facial landmarks. Experimental results also suggest that state-of-the-art invariant descriptors, such as local binary pattern (LBP), histogram of oriented gradient (HOG), and Gabor energy filter (GEF), and classification methods, such as sparse representation based classification (SRC) and support vector machine (SVM), can benefit from using the TIPCA-aligned faces, instead of the manually eye-aligned faces that are widely regarded as the ground-truth alignment. Favorable accuracies against the state-of-the-art results on face coding and face recognition are reported. Weihong Deng, Jiani Hu, Jiwen Lu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Equidistant prototypes embedding for single sample based face recognition with generic learning and incremental learning
Weihong Deng, Jiani Hu, Xiuzhuang Zhou, Jun Guo 0002 |
Pattern Recognit. | 2 |
| 2013 | In Defense of Sparsity Based Face RecognitionabstractThe success of sparse representation based classification (SRC) has largely boosted the research of sparsity based face recognition in recent years. A prevailing view is that the sparsity based face recognition performs well only when the training images have been carefully controlled and the number of samples per class is sufficiently large. This paper challenges the prevailing view by proposing a ``prototype plus variation'' representation model for sparsity based face recognition. Based on the new model, a Superposed SRC (SSRC), in which the dictionary is assembled by the class centroids and the sample-to-centroid differences, leads to a substantial improvement on SRC. The experiments results on AR, FERET and FRGC databases validate that, if the proposed prototype plus variation representation model is applied, sparse coding plays a crucial role in face recognition, and performs well even when the dictionary bases are collected under uncontrolled conditions and only a single sample per classes is available. Weihong Deng, Jiani Hu, Jun Guo 0002 |
CVPR | 2 |
| 2012 | Extended SRC: Undersampled Face Recognition via Intraclass Variant DictionaryabstractSparse Representation-Based Classification (SRC) is a face recognition breakthrough in recent years which has successfully addressed the recognition problem with sufficient training images of each gallery subject. In this paper, we extend SRC to applications where there are very few, or even a single, training images per subject. Assuming that the intraclass variations of one subject can be approximated by a sparse linear combination of those of other subjects, Extended Sparse Representation-Based Classifier (ESRC) applies an auxiliary intraclass variant dictionary to represent the possible variation between the training and testing images. The dictionary atoms typically represent intraclass sample differences computed from either the gallery faces themselves or the generic faces that are outside the gallery. Experimental results on the AR and FERET databases show that ESRC has better generalization ability than SRC for undersampled face recognition under variable expressions, illuminations, disguises, and ages. The superior results of ESRC suggest that if the dictionary is properly constructed, SRC algorithms can generalize well to the large-scale face recognition problem, even with a single training image per class. Weihong Deng, Jiani Hu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | The small sample size problem of ICA: A comparative study and analysis
Weihong Deng, Yebin Liu, Jiani Hu, Jun Guo 0002 |
Pattern Recognit. | 3 |
| 2010 | Robust, accurate and efficient face recognition from a single training image: A uniform pursuit approach
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng |
Pattern Recognit. | 2 |
| 2010 | Emulating biological strategies for uncontrolled face recognition
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng |
Pattern Recognit. | 2 |
| 2009 | Semi-supervised Learning Based on Label Propagation through Submanifold
Jiani Hu, Weihong Deng, Jun Guo 0002 |
ISNN (1) | 1 |
| 2009 | Learning a locality discriminating projection for classification
Jiani Hu, Weihong Deng, Jun Guo 0002, Weiran Xu |
Knowl. Based Syst. | 1 |
| 2008 | Comments on "Globally Maximizing, Locally Minimizing: Unsupervised Discriminant Projection with Application to Face and Palm Biometrics"abstractIn [1], UDP is proposed to address the limitation of LPP for the clustering and classification tasks. In this communication, we show that the basic ideas of UDP and LPP are identical. In particular, UDP is just a simplified version of LPP on the assumption that the local density is uniform. Weihong Deng, Jiani Hu, Jun Guo 0002, Honggang Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Locality discriminating indexing for document classificationabstractThis paper introduces a locality discriminating indexing (LDI) algorithm for document classification. Based on the hypothesis that samples from different classes reside in class-specific manifold structures, LDI seeks for a projection which best preserves the within-class local structures while suppresses the between-class overlap. Comparative experiments show that the proposed method isable to derives compact discriminating document representations for classification. Jiani Hu, Weihong Deng, Jun Guo 0002, Weiran Xu |
SIGIR | 1 |