EDBT 2026 Demo / reviewers in the wild / expert
Weihong Deng
dblp:39/232
· DBLP profile ↗
160ranked-venue papers
18as first author
70since 2021 · last 2026
0000-0001-5952-6996ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 115 · 16 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 100 · 5 first-author · 39 since 2021Security and privacy · 7 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question AnsweringabstractMulti-hop question answering (MHQA) requires integrating knowledge scattered across multiple passages to derive the correct answer. Traditional retrieval-augmented generation (RAG) methods primarily focus on coarse-grained textual semantic similarity and ignore structural associations among dispersed knowledge, which limits their effectiveness in MHQA tasks. GraphRAG methods address this by leveraging knowledge graphs (KGs) to capture structural associations, but they tend to overly rely on structural information and fine-grained word- or phrase-level retrieval, resulting in an underutilization of textual semantics. In this paper, we propose a novel RAG approach called HGRAG for MHQA that achieves cross-granularity integration of structural and semantic information via hypergraphs. Structurally, we construct an entity hypergraph where fine-grained entities serve as nodes and coarse-grained passages as hyperedges, and establish knowledge association through shared entities. Semantically, we design a hypergraph retrieval method that integrates fine-grained entity similarity and coarse-grained passage similarity via hypergraph diffusion. Finally, we employ a retrieval enhancement module, which further refines the retrieved results both semantically and structurally, to obtain the most relevant passages as context for answer generation with the LLM. Experimental results on benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in QA performance, and achieves a 6× speedup in retrieval efficiency. Changjian Wang 0001, Weihong Deng, Weili Guan |
AAAI | 2 |
| 2025 | Device-aware Optical Adversarial Attack for a Portable Projector-camera SystemabstractDeep-learning-based face recognition (FR) systems are susceptible to adversarial examples in both digital and physical domains. Physical attacks present a greater threat to deployed systems as adversaries can easily access the input channel, allowing them to provide malicious inputs to impersonate a victim. This paper addresses the limitations of existing projector-camera-based adversarial light attacks in practical FR setups. By incorporating device-aware adaptations into the digital attack algorithm, such as resolution-aware and color-aware adjustments, we mitigate the degradation from digital to physical domains. Experimental validation showcases the efficacy of our proposed algorithm against real and spoof adversaries, achieving high physical similarity scores in FR models and state-of-the-art commercial systems. On average, there is only a 14% reduction in scores from digital to physical attacks, with high attack success rate in both white- and black-box scenarios. Dingheng Zeng, Weihong Deng, Ying Li 0012 |
ICASSP | 5 |
| 2025 | Seek and Solve Reasoning for Table Question AnsweringabstractThe complexities of table structures and question logic make table-based question answering (TQA) tasks challenging for Large Language Models (LLMs), often requiring task simplification before solving. This paper reveals that the reasoning process during task simplification may be more valuable than the simplified tasks themselves and aims to improve TQA performance by leveraging LLMs’ reasoning capabilities. We propose a Seek-and-Solve pipeline that instructs the LLM to first seek relevant information and then answer questions, integrating these two stages at the reasoning level into a coherent Seek-and-Solve Chain of Thought (SS-CoT). Additionally, we distill a single-step TQA-solving prompt from this pipeline, using demonstrations with SS-CoT paths to guide the LLM in solving complex TQA tasks under In-Context Learning settings. Our experiments show that our approaches result in improved performance and reliability while being efficient. Our findings emphasize the importance of eliciting LLMs’ reasoning capabilities to handle complex TQA tasks effectively. Ruya Jiang, Weihong Deng |
ICASSP | 3 |
| 2025 | MS-UFAD: A Large-Scale Dataset for Real-world Unified Face Attack Detection with Text DescriptionsabstractAs deepfake and adversarial attacks evolve, facial recognition systems are encountering increasingly diverse threats. Most existing face liveness detection algorithms focus on single tasks, like spoofing or deepfake attack detection. The corresponding datasets have limited coverage of attack methods, with original data mostly sourced from the internet or laboratory environments. Moreover, existing datasets lack textual annotations, particularly for attack clues, limiting algorithms’ ability to utilize semantic assistance from text. To address these issues, we propose a large-scale unified attack dataset, which includes newly collected facial videos from 5,000 individuals, along with generated videos corresponding to 52 face attack methods. The dataset contains 795k videos and 60k images across four different quality levels. Through semi-automated annotation, we provide detailed textual descriptions. This is the first face attack dataset with textual descriptions. Additionally, we propose a text-guided face attack detection method, demonstrating significant improvements in accuracy using fine-grained textual descriptions. Our dataset will be released at https://ms-ufad.github.io. Dingheng Zeng, Zhifei Kong, Tongtong Yuan, Weihong Deng, Ying Li 0012 |
ICASSP | 10 |
| 2025 | Towards Interactive Deepfake AnalysisabstractExisting deepfake analysis methods are primarily based on discriminative models, which significantly limit their application scenarios. This paper aims to explore interactive deepfake analysis by performing instruction tuning on multi-modal large language models (MLLMs). This will face challenges such as the lack of datasets and benchmarks, and low training efficiency. To address these issues, we introduce (1) a GPT-assisted data construction process resulting in an instruction-following dataset called DFA-Instruct, (2) a benchmark named DFA-Bench, designed to comprehensively evaluate the capabilities of MLLMs in deepfake detection, deepfake classification, and artifact description, and (3) construct an interactive deepfake analysis system called DFA-GPT, as a strong baseline for the community, with the Low-Rank Adaptation (LoRA) module. The dataset and code will be made available at https://github.com/lxq1000/DFA-Instruct to facilitate further research. Lixiong Qin, Yuhan Qiu, Dingheng Zeng, Jiani Hu, Weihong Deng |
ICASSP | 7 |
| 2025 | Transitive Inference in Large Language Models and Prompting InterventionabstractTransitive inference (TI) is a critical form of deductive reasoning, essential to both human and animal cognition. This study explores whether state-of-the-art large language models (LLMs) possess TI capabilities and examines the impact of two emerging prompting methods on model performance. Four LLMs—GPT-3.5-Turbo, GPT-4, Llama3-8B, and Qwen—are evaluated using a TI task involving a 10-item hierarchy. Results indicate that these models demonstrate solid performance, along with human-like behavioral effects such as the symbolic distance effect, terminal item effect, and context effect. The sequence of input premises significantly affects model accuracy, with all models showing a preference for the chain condition over the jump condition. Two prompting methods — Model Confidence prompts (likelihood tests) and Chain-of-Thought prompts—are applied in order to further enhance TI performance. While GPT-4 benefits the most from these prompts, other models experience a decline in performance. Additionally, behavioral biases like the terminal item effect persist and are even amplified following prompt adjustments. Therefore, specific prompting methods may not be universally effective across different models, and in some cases, may cause adverse effects. Further research is needed to better understand the nature of TI in LLMs. Wenya Wu, Weihong Deng |
ICASSP | 2 |
| 2025 | MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMsabstractCurrent multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods. Xuannan Liu, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Shuhan Xia, Xing Cui, Linzhi Huang, Weihong Deng, Zhaofeng He 0001 |
ICLR | 8 |
| 2025 | Dual Information Speech Language Models for Emotional ConversationsabstractConversational systems relying on text-based large language models (LLMs) often overlook paralinguistic cues, essential for understanding emotions and intentions. Speech-language models (SLMs), which use speech as input, are emerging as a promising solution. However, SLMs built by extending frozen LLMs struggle to capture paralinguistic information and exhibit reduced context understanding. We identify entangled information and improper training strategies as key issues. To address these issues, we propose two heterogeneous adapters and suggest a weakly supervised training strategy. Our approach disentangles paralinguistic and linguistic information, enabling SLMs to interpret speech through structured representations. It also preserves contextual understanding by avoiding the generation of task-specific vectors through controlled randomness. This approach trains only the adapters on common datasets, ensuring parameter and data efficiency. Experiments demonstrate competitive performance in emotional conversation tasks, showcasing the model’s ability to effectively integrate both paralinguistic and linguistic information within contextual settings. Wenze Xu, Weihong Deng |
ICME | 4 |
| 2025 | Federated Knowledge Distillation Based on Prompt for Matching Data Distribution
Yizhang Liu, Wenze Xu, Tongtong Yuan, Weihong Deng |
PRCV (18) | 5 |
| 2025 | AdvCloak: Customized adversarial cloak for privacy protection
Xuannan Liu, Yaoyao Zhong, Xing Cui, Yuhang Zhang 0016, Peipei Li 0002, Weihong Deng |
Pattern Recognit. | 6 |
| 2025 | Marginal debiased network for fair visual recognition
Mei Wang 0001, Weihong Deng, Jiani Hu, Sen Su |
Pattern Recognit. | 2 |
| 2025 | Implicit face model: Depth super-resolution for 3D face recognition
Mei Wang 0001, Ruizhuo Xu, Weihong Deng |
Pattern Recognit. | 3 |
| 2025 | Unsupervised evaluation for out-of-distribution detection
Jiani Hu, Dongchao Wen, Weihong Deng |
Pattern Recognit. | 4 |
| 2025 | A perturbed match filtering approach for face image quality assessment
Yuying Zhao, Mei Wang 0001, Jiani Hu, Weihong Deng, Chun-Guang Li |
Pattern Recognit. | 4 |
| 2025 | DDL: Dynamic Direction Learning for Semi-Supervised Facial Expression RecognitionabstractMost semi-supervised facial expression recognition (FER) algorithms leverage pseudo-labeling to mine additional information from unlabeled samples. Despite its good performance, two critical issues persist: class imbalance and domain shift. The former is a typical challenge due to the significant variation in sample numbers across different FER classes, resulting in highly imbalanced pseudo labels in existing semi-supervised methods. For the latter, given that labeled and unlabeled data usually come from different sources, a considerable domain gap might exist, leading the model to generate low-quality pseudo labels. To tackle these issues, we introduce a novel semi-supervised FER algorithm called Dynamic Direction Learning (DDL), which consists of adaptive balance learning (ABL) and adaptive alignment learning (AAL). ABL allows a balanced training process by dynamically adjusting the constraints of self-training based on the performance of a balanced validation dataset. Moreover, AAL adaptively aligns the feature distribution of labeled and unlabeled data by minimizing their distance in feature space. Additionally, a role rotation mechanism (RRM) is proposed to avoid confirmation bias, which further improves self-training. Extensive experiments demonstrate that DDL achieves state-of-the-art performance on different FER datasets. Yuhang Zhang 0016, Han Fang 0002, Jiani Hu, Weihong Deng |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Efficient Image Super-Resolution With Feature Interaction Weighted Hybrid NetworkabstractLightweight image super-resolution aims to reconstruct high-resolution images from low-resolution images using low computational costs. However, existing methods result in the loss of middle-layer features due to activation functions. To minimize the impact of intermediate feature loss on reconstruction quality, we propose a Feature Interaction Weighted Hybrid Network (FIWHN), which comprises a series of Wide-residual Distillation Interaction Block (WDIB) as the backbone. Every third WDIB forms a Feature Shuffle Weighted Group (FSWG) by applying mutual information shuffle and fusion. Moreover, to mitigate the negative effects of intermediate feature loss, we introduce Wide Residual Weighting units within WDIB. These units effectively fuse features of varying levels of detail through a Wide-residual Distillation Connection (WRDC) and a Self-Calibrating Fusion (SCF). To compensate for global feature deficiencies, we incorporate a Transformer and explore a novel architecture to combine CNN and Transformer. We show that our FIWHN achieves a favorable balance between performance and efficiency through extensive experiments on low-level and high-level tasks. Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jian Yang 0003, Guo-Jun Qi, Chia-Wen Lin |
IEEE Trans. Multim. | 4 |
| 2024 | Blind Face Restoration under Extreme Conditions: Leveraging 3D-2D Prior Fusion for Superior Structural and Texture RecoveryabstractBlind face restoration under extreme conditions involves reconstructing high-quality face images from severely degraded inputs. These input images are often in poor quality and have extreme facial poses, leading to errors in facial structure and unnatural artifacts within the restored images. In this paper, we show that utilizing 3D priors effectively compensates for structure knowledge deficiencies in 2D priors while preserving the texture details. Based on this, we introduce FREx (Face Restoration under Extreme conditions) that combines structure-accurate 3D priors and texture-rich 2D priors in pretrained generative networks for blind face restoration under extreme conditions. To fuse the different information in 3D and 2D priors, we introduce an adaptive weight module that adjusts the importance of features based on the input image's condition. With this approach, our model can restore structure-accurate and natural-looking faces even when the images have lost a lot of information due to degradation and extreme pose. Extensive experimental results on synthetic and real-world datasets validate the effectiveness of our methods. Zhengrui Chen, Liying Lu, Ziyang Yuan, Yu Li 0003, Chun Yuan 0003, Weihong Deng |
AAAI | 7 |
| 2024 | Open-Set Facial Expression RecognitionabstractFacial expression recognition (FER) models are typically trained on datasets with a fixed number of seven basic classes. However, recent research works (Cowen et al. 2021; Bryant et al. 2022; Kollias 2023) point out that there are far more expressions than the basic ones. Thus, when these models are deployed in the real world, they may encounter unknown classes, such as compound expressions that cannot be classified into existing basic classes. To address this issue, we propose the open-set FER task for the first time. Though there are many existing open-set recognition methods, we argue that they do not work well for open-set FER because FER data are all human faces with very small inter-class distances, which makes the open-set samples very similar to close-set samples. In this paper, we are the first to transform the disadvantage of small inter-class distance into an advantage by proposing a new way for open-set FER. Specifically, we find that small inter-class distance allows for sparsely distributed pseudo labels of open-set samples, which can be viewed as symmetric noisy labels. Based on this novel observation, we convert the open-set FER to a noisy label detection problem. We further propose a novel method that incorporates attention map consistency and cycle training to detect the open-set samples. Extensive experiments on various FER datasets demonstrate that our method clearly outperforms state-of-the-art open-set recognition methods by large margins. Code is available at https://github.com/zyh-uaiaaaa. Yue Yao 0001, Xuannan Liu, Lixiong Qin, Weihong Deng |
AAAI | 6 |
| 2024 | Faceptor: A Generalist Model for Face Perception
Lixiong Qin, Mei Wang 0001, Xuannan Liu, Yuhang Zhang 0016, Wei Deng 0004, Xiaoshuai Song, Weiran Xu, Weihong Deng |
ECCV (34) | 8 |
| 2024 | Generalizable Facial Expression Recognition
Yuhang Zhang 0016, Xiuqi Zheng, Chenyi Liang, Jiani Hu, Weihong Deng |
ECCV (14) | 5 |
| 2024 | Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient AccumulationabstractThe blooming of social media and face recognition (FR) systems has increased people’s concern about privacy and security. A new type of adversarial privacy cloak (class-universal) can be applied to all the images of regular users, to prevent malicious FR systems from acquiring their identity information. In this work, we discover the optimization dilemma in the existing methods – the local optima problem in large-batch optimization and the gradient information elimination problem in small-batch optimization. To solve these problems, we propose Gradient Accumulation (GA) to aggregate multiple small-batch gradients into a one-step iterative gradient to enhance the gradient stability and reduce the usage of quantization operations. Experiments show that our proposed method achieves high performance on the Privacy-Commons dataset against black-box face recognition models. Xuannan Liu, Yaoyao Zhong, Weihong Deng, Hongzhi Shi, Xingchen Cui, Yunfeng Yin, Dongchao Wen |
ICASSP | 3 |
| 2024 | FedSC: Federated Generalized Face Anti-Spoofing via Shuffled Codebook
Mei Wang 0001, Weihong Deng, Jiani Hu |
ICPR (3) | 3 |
| 2024 | SEQ-former: A context-enhanced and efficient automatic speech recognition framework
Kaixun Huang, Lei Xie 0001, Zongfeng Quan, Weihong Deng |
INTERSPEECH | 7 |
| 2024 | FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsabstractThe massive generation of multimodal fake news involving both text and images exhibits substantial distribution discrepancies, prompting the need for generalized detectors. However, the insulated nature of training restricts the capability of classical detectors to obtain open-world facts. While Large Vision-Language Models (LVLMs) have encoded rich world knowledge, they are not inherently tailored for combating fake news and struggle to comprehend local forgery details. In this paper, we propose FKA-Owl, a novel framework that leverages forgery-specific knowledge to augment LVLMs, enabling them to reason about manipulations effectively. The augmented forgery-specific knowledge includes semantic correlation between text and images, and artifact trace in image manipulation. To inject these two kinds of knowledge into the LVLM, we design two specialized modules to establish their representations, respectively. The encoded knowledge embeddings are then incorporated into LVLMs. Extensive experiments on the public benchmark demonstrate that FKA-Owl achieves superior cross-domain performance compared to previous methods. Code is publicly available at https://liuxuannan.github.io/FKA_Owl.github.io/. Xuannan Liu, Peipei Li 0002, Huaibo Huang, Zekun Li 0001, Xing Cui, Lixiong Qin, Weihong Deng, Zhaofeng He 0001 |
ACM Multimedia | 8 |
| 2024 | Joint recognition of basic and compound facial expressions by mining latent soft labels
Mei Wang 0001, Bo Xiao 0006, Jiani Hu, Weihong Deng |
Pattern Recognit. | 5 |
| 2024 | Oracle character recognition using unsupervised discriminative consistency networkabstractAncient history relies on the study of ancient characters. However, real-world scanned oracle characters are difficult to collect and annotate, posing a major obstacle for oracle character recognition (OrCR). Besides, serious abrasion and inter-class similarity also make OrCR more challenging. In this paper, we propose a novel unsupervised domain adaptation method for OrCR, which enables to transfer knowledge from labeled handprinted oracle characters to unlabeled scanned data. We leverage pseudo-labeling to incorporate the semantic information into adaptation and constrain augmentation consistency to make the predictions of scanned samples consistent under different perturbations, leading to the model robustness to abrasion, stain and distortion. Simultaneously, an unsupervised transition loss is proposed to learn more discriminative features on the scanned domain by optimizing both between-class and within-class transition probability . Extensive experiments show that our approach achieves state-of-the-art result on Oracle-241 dataset and substantially outperforms the recently proposed structure-texture separation network by 15.1%. Mei Wang 0001, Weihong Deng, Sen Su |
Pattern Recognit. | 2 |
| 2024 | Depth Map Denoising Network and Lightweight Fusion Network for Enhanced 3D Face Recognition
Ruizhuo Xu, Chao Deng 0002, Mei Wang 0001, Junlan Feng, Weihong Deng |
Pattern Recognit. | 8 |
| 2024 | Improving Multi-Label Facial Expression Recognition With Consistent and Distinct AttentionsabstractFacial expression recognition (FER) attracts much attention in computer vision. Previous works mostly study the single-label FER problem. The more complex multi-label facial expression recognition task is underexplored. Multi-label FER is more challenging than single-label task due to two primary causes. On one hand, there are less available multi-label facial expression data for analysis. On the other hand, the entanglement of expressions makes recognition more difficult. In this work, we leverage class activation map (CAM) to improve the performance of multi-label FER. Firstly, considering the shortage of training data, an attention flipping consistency (AFC) loss equipped with random rotation augmentation is proposed. It restrains the network to produce consistent CAMs under horizontally flipping transformation and thus improves the stability of network without any extra data. Secondly, based on the prior knowledge that different emotions have different predominant activated facial regions, we propose a label-guided spatial attention dispersing (SAD) loss to enable the model to learn from distinct expression-related regions. By combining the widely used multi-label classification loss (i.e., binary cross-entropy loss) and proposed AFC loss and SAD loss, our method achieves state-of-the-art performance on multi-label FER databases and the model's interpretability is improved. Weihong Deng |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | SwinFace: A Multi-Task Transformer for Face Recognition, Expression Recognition, Age Estimation and Attribute EstimationabstractIn recent years, vision transformers have been introduced into face recognition and analysis and have achieved performance breakthroughs. However, most previous methods generally train a single model or an ensemble of models to perform the desired task, which ignores the synergy among different tasks and fails to achieve improved prediction accuracy, increased data efficiency, and reduced training time. This paper presents a multi-purpose algorithm for simultaneous face recognition, facial expression recognition, age estimation, and face attribute estimation (40 attributes including gender) based on a single Swin Transformer. Our design, the SwinFace, consists of a single shared backbone together with a subnet for each set of related tasks. To address the conflicts among multiple tasks and meet the different demands of tasks, a Multi-Level Channel Attention (MLCA) module is integrated into each task-specific analysis subnet, which can adaptively select the features from optimal levels and channels to perform the desired tasks. Extensive experiments show that the proposed model has a better understanding of the face and achieves excellent performance for all tasks. Especially, it achieves 90.97% accuracy on RAF-DB and 0.22 ϵ-error on CLAP2015, which are state-of-the-art results on facial expression recognition and age estimation respectively. The code and models will be made publicly available at https://github.com/lxq1000/SwinFace. Lixiong Qin, Mei Wang 0001, Chao Deng 0002, Jiani Hu, Weihong Deng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Cross-Receptive Focused Inference Network for Lightweight Image Super-ResolutionabstractRecently, Transformer-based methods have shown impressive performance in single image super-resolution (SISR) tasks due to the ability of global feature extraction. However, the capabilities of Transformers that need to incorporate contextual information to extract features dynamically are neglected. To address this issue, we propose a lightweight Cross-receptive Focused Inference Network (CFIN) that consists of a cascade of CT Blocks mixed with CNN and Transformer. Specifically, in the CT block, we first propose a CNN-based Cross-Scale Information Aggregation Module (CIAM) to enable the model to better focus on potentially helpful information to improve the efficiency of the Transformer phase. Then, we design a novel Cross-receptive Field Guided Transformer (CFGT) to enable the selection of contextual information required for reconstruction by using a modulated convolutional kernel that understands the current semantic information and exploits the information interaction within different self-attention. Extensive experiments have shown that our proposed CFIN can effectively reconstruct images using contextual information, and it can strike a good balance between computational cost and model performance as an efficient model. Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jiantao Zhou 0001, Jian Yang 0003, Guo-Jun Qi |
IEEE Trans. Multim. | 4 |
| 2024 | Confusion-Based Metric Learning for Regularizing Zero-Shot Image Retrieval and ClusteringabstractDeep metric learning turns to be attractive in zero-shot image retrieval and clustering (ZSRC) task in which a good embedding/metric is requested such that the unseen classes can be distinguished well. Most existing works deem this "good" embedding just to be the discriminative one and race to devise the powerful metric objectives or the hard-sample mining strategies for learning discriminative deep metrics. However, in this article, we first emphasize that the generalization ability is also a core ingredient of this "good" metric and it largely affects the metric performance in zero-shot settings as a matter of fact. Then, we propose the confusion-based metric learning (CML) framework to explicitly optimize a robust metric. It is mainly achieved by introducing two interesting regularization terms, i.e., the energy confusion (EC) and diversity confusion (DC) terms. These terms daringly break away from the traditional deep metric learning idea of designing discriminative objectives and instead seek to "confuse" the learned model. These two confusion terms focus on local and global feature distribution confusions, respectively. We train these confusion terms together with the conventional deep metric objective in an adversarial manner. Although it seems weird to "confuse" the model learning, we show that our CML indeed serves as an efficient regularization framework for deep metric learning and it is applicable to various conventional metric methods. This article empirically and experimentally demonstrates the importance of learning an embedding/metric with good generalization, achieving the state-of-the-art performances on the popular CUB, CARS, Stanford Online Products, and In-Shop datasets for ZSRC tasks. Binghui Chen, Weihong Deng, Lei Zhang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Targeted Shilling Attacks on GNN-based Recommender SystemsabstractGNN-based recommender systems have shown their vulnerability to shilling attacks in recent studies. By conducting shilling attacks on recommender systems, the attackers aim to have homogeneous impacts on all users. However, such indiscriminate attacks suffer from a waste of resources because even if the target item is promoted to users who are not interested, they are unlikely to click on them. In this paper, we conduct targeted shilling attacks in GNN-based recommender systems. By automatically constructing the features and edges of the fake users, our proposed framework AutoAttack achieves accurate attacks on a specific group of users while minimizing the impact on non-target users. Specifically, the features of fake users are generated based on a similarity function, which is optimized according to the features of target users. The structure of fake users is learned by conducting spectral clustering on the target users based on their graph Laplacian matrix, which contains the degree and adjacency information that provides guidance to the edge generation of fake users. We conduct extensive experiments on four real-world datasets in different GNN-based RS and evaluate the performance of our method on the shilling attack and recommendation tasks comprehensively, showing the effectiveness and flexibility of our framework. Sihan Guo, Ting Bai 0004, Weihong Deng |
CIKM | 3 |
| 2023 | Semi-Supervised 2D Human Pose Estimation Driven by Position Inconsistency Pseudo Label Correction ModuleabstractIn this paper, we delve into semi-supervised 2D human pose estimation. The previous method ignored two problems: (i) When conducting interactive training between large model and lightweight model, the pseudo label of lightweight model will be used to guide large models. (ii) The negative impact of noise pseudo labels on training. Moreover, the labels used for 2D human pose estimation are relatively complex: keypoint category and keypoint position. To solve the problems mentioned above, we propose a semi-supervised 2D human pose estimation framework driven by a position inconsistency pseudo label correction module (SSPCM). We introduce an additional auxiliary teacher and use the pseudo labels generated by the two teacher model in different periods to calculate the inconsistency score and remove outliers. Then, the two teacher models are updated through interactive training, and the student model is updated using the pseudo labels generated by two teachers. To further improve the performance of the student model, we use the semi-supervised Cut-Occlude based on pseudo keypoint perception to generate more hard and effective samples. In addition, we also proposed a new indoor overhead fisheye human keypoint dataset WEPDTOF-Pose. Extensive experiments demonstrate that our method outperforms the previous best semi-supervised 2D human pose estimation method. We will release the code and dataset at https://github.com/hlz0606/SSPCM Linzhi Huang, Hongbo Tian, Xiangang Li, Weihong Deng, Jieping Ye |
CVPR | 6 |
| 2023 | Enhancing Generalization of Universal Adversarial Perturbation through Gradient AggregationabstractDeep neural networks are vulnerable to universal adversarial perturbation (UAP), an instance-agnostic perturbation capable of fooling the target model for most samples. Compared to instance-specific adversarial examples, UAP is more challenging as it needs to generalize across various samples and models. In this paper, we examine the serious dilemma of UAP generation methods from a generalization perspective – the gradient vanishing problem using small-batch stochastic gradient optimization and the local optima problem using large-batch optimization. To address these problems, we propose a simple and effective method called Stochastic Gradient Aggregation (SGA), which alleviates the gradient vanishing and escapes from poor local optima at the same time. Specifically, SGA employs the small-batch training to perform multiple iterations of inner pre-search. Then, all the inner gradients are aggregated as a one-step gradient estimation to enhance the gradient stability and reduce quantization errors. Extensive experiments on the standard ImageNet dataset demonstrate that our method significantly enhances the generalization ability of UAP and outperforms other state-of-the-art methods. The code is available at https://github.com/liuxuannan/Stochastic-Gradient-Aggregation. Xuannan Liu, Yaoyao Zhong, Yuhang Zhang 0016, Lixiong Qin, Weihong Deng |
ICCV | 5 |
| 2023 | Leave No Stone Unturned: Mine Extra Knowledge for Imbalanced Facial Expression RecognitionabstractFacial expression data is characterized by a significant imbalance, with most collected data showing happy or neutral expressions and fewer instances of fear or disgust. This imbalance poses challenges to facial expression recognition (FER) models, hindering their ability to fully understand various human emotional states. Existing FER methods typically report overall accuracy on highly imbalanced test sets but exhibit low performance in terms of the mean accuracy across all expression classes. In this paper, our aim is to address the imbalanced FER problem. Existing methods primarily focus on learning knowledge of minor classes solely from minor-class samples. However, we propose a novel approach to extract extra knowledge related to the minor classes from both major and minor class samples. Our motivation stems from the belief that FER resembles a distribution learning task, wherein a sample may contain information about multiple classes. For instance, a sample from the major class surprise might also contain useful features of the minor class fear. Inspired by that, we propose a novel method that leverages re-balanced attention maps to regularize the model, enabling it to extract transformation invariant information about the minor classes from all training samples. Additionally, we introduce re-balanced smooth labels to regulate the cross-entropy loss, guiding the model to pay more attention to the minor classes by utilizing the extra information regarding the label distribution of the imbalanced training data. Extensive experiments on different datasets and backbones show that the two proposed modules work together to regularize the model and achieve state-of-the-art performance under the imbalanced FER task. Code is available at https://github.com/zyh-uaiaaaa. Lixiong Qin, Xuannan Liu, Weihong Deng |
NeurIPS | 5 |
| 2023 | OPOM: Customized Invisible Cloak Towards Face Privacy ProtectionabstractWhile convenient in daily life, face recognition technologies also raise privacy concerns for regular users on the social media since they could be used to analyze face images and videos, efficiently and surreptitiously without any security restrictions. In this paper, we investigate the face privacy protection from a technology standpoint based on a new type of customized cloak, which can be applied to all the images of a regular user, to prevent malicious face recognition systems from uncovering their identity. Specifically, we propose a new method, named one person one mask (OPOM), to generate person-specific (class-wise) universal masks by optimizing each training sample in the direction away from the feature subspace of the source identity. To make full use of the limited training images, we investigate several modeling methods, including affine hulls, class centers and convex hulls, to obtain a better description of the feature subspace of source identities. The effectiveness of the proposed method is evaluated on both common and celebrity datasets against black-box face recognition models with different loss functions and network architectures. In addition, we discuss the advantages and potential problems of the proposed method. In particular, we conduct an application study on the privacy protection of a video dataset, Sherlock, to demonstrate the potential practical usage of the proposed method. Yaoyao Zhong, Weihong Deng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Boosting Facial Expression Recognition by A Semi-Supervised Progressive TeacherabstractIn this paper, we aim to improve the performance of in-the-wild Facial Expression Recognition (FER) by exploiting semi-supervised learning. Large-scale labeled data and deep learning methods have greatly improved the performance of image recognition. However, the performance of FER is still not ideal due to the lack of training data and incorrect annotations (e.g., label noises). Among existing in-the-wild FER datasets, reliable ones contain insufficient data to train robust deep models while large-scale ones are annotated in lower quality. To address this problem, we propose a semi-supervised learning algorithm named Progressive Teacher (PT) to utilize reliable FER datasets as well as large-scale unlabeled expression images for effective training. On the one hand, PT introduces semi-supervised learning method to relieve the shortage of data in FER. On the other hand, it selects useful labeled training samples automatically and progressively to alleviate label noise. PT uses selected clean labeled data for computing the supervised classification loss and unlabeled data for unsupervised consistency loss. Experiments on widely-used databases RAF-DB and FERPlus validate the effectiveness of our method, which achieves state-of-the-art performance with accuracy of 89.57% on RAF-DB. Additionally, when the synthetic noise rate reaches even 30%, the performance of our PT algorithm only degrades by 4.37%. Weihong Deng |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Federated Learning for Face Recognition with Gradient CorrectionabstractWith increasing appealing to privacy issues in face recognition, federated learning has emerged as one of the most prevalent approaches to study the unconstrained face recognition problem with private decentralized data. However, conventional decentralized federated algorithm sharing whole parameters of networks among clients suffers from privacy leakage in face recognition scene. In this work, we introduce a framework, FedGC, to tackle federated learning for face recognition and guarantees higher privacy. We explore a novel idea of correcting gradients from the perspective of backward propagation and propose a softmax-based regularizer to correct gradients of class embeddings by precisely injecting a cross-client gradient term. Theoretically, we show that FedGC constitutes a valid loss function similar to standard softmax. Extensive experiments have been conducted to validate the superiority of FedGC which can match the performance of conventional centralized methods utilizing full training dataset on several popular benchmark datasets. Yifan Niu, Weihong Deng |
AAAI | 2 |
| 2022 | Augmented Geometric Distillation for Data-Free Incremental Person ReIDabstractIncremental learning (IL) remains an open issue for Person Re-identification (ReID), where a ReID system is expected to preserve preceding knowledge while learning incrementally. However, due to the strict privacy licenses and the open-set retrieval setting, it is intractable to adapt existing class IL methods to ReID. In this work, we propose an Augmented Geometric Distillation (AGD) framework to tackle these issues. First, a general data-free incremental framework with dreaming memory is constructed to avoid privacy disclosure. On this basis, we reveal a “noisy distillation” problem stemming from the noise in dreaming memory, and further propose to augment distillation in a pairwise and cross-wise pattern over different views of memory to mitigate it. Second, for the open-set retrieval property, we propose to maintain feature space structure during evolving via a novel geometric way and preserve relationships between exemplars when representations drift. Extensive experiments demonstrate the superiority of our AGD to baseline with a margin of 6.0% mAP/7.9% R@1 and it could be generalized to class IL. Code is available here11†https://github.com/eddielyc/Augmented-Geometric-Distillation. Weihong Deng |
CVPR | 3 |
| 2022 | Domain Generalization via Shuffled Style Assembly for Face Anti-SpoofingabstractWith diverse presentation attacks emerging continually, generalizable face anti-spoofing (FAS) has drawn growing attention. Most existing methods implement domain generalization (DG) on the complete representations. However, different image statistics may have unique properties for the FAS tasks. In this work, we separate the complete representation into content and style ones. A novel Shuffled Style Assembly Network (SSAN) is proposed to extract and reassemble different content and style features for a stylized feature space. Then, to obtain a generalized representation, a contrastive learning strategy is developed to emphasize liveness-related style information while suppress the domain-specific one. Finally, the representations of the correct assemblies are used to distinguish between living and spoofing during the inferring. On the other hand, despite the decent performance, there still exists a gap between academia and industry, due to the difference in data quantity and distribution. Thus, a new large-scale benchmark for FAS is built up to further evaluate the performance of algorithms in reality. Both qualitative and quantitative results on existing and proposed benchmarks demonstrate the effectiveness of our methods. The codes will be available at https://github.com/wangzhuo2019/SSAN. Zezheng Wang 0002, Zitong Yu, Weihong Deng, Tingting Gao, Zhongyuan Wang 0006 |
CVPR | 4 |
| 2022 | DH-AUG: DH Forward Kinematics Model Driven Augmentation for 3D Human Pose Estimation
Linzhi Huang, Weihong Deng |
ECCV (6) | 3 |
| 2022 | Exploring Disentangled Content Information for Face Forgery Detection
Huafeng Shi, Weihong Deng |
ECCV (14) | 3 |
| 2022 | Learn from All: Erasing Attention Consistency for Noisy Label Facial Expression Recognition
Yuhang Zhang 0016, Xu Ling, Weihong Deng |
ECCV (26) | 4 |
| 2022 | Video Question Answering: Datasets, Algorithms and ChallengesabstractThis survey aims to organize the recent advances in video question answering (VideoQA) and point towards future directions.We firstly categorize the datasets into: 1) normal VideoQA, multi-modal VideoQA and knowledge-based VideoQA, according to the modalities invoked in the question-answer pairs, and 2) factoid VideoQA and inference VideoQA, according to the technical challenges in comprehending the questions and deriving the correct answers.We then summarize the VideoQA techniques, including those mainly designed for Factoid QA (such as the early spatio-temporal attention-based methods and the recent Transformer-based ones) and those targeted at explicit relation and logic inference (such as neural modular networks, neural symbolic methods, and graph-structured methods).Aside from the backbone techniques, we also delve into specific models and derive some common and useful insights either for video modeling, question answering, or for cross-modal correspondence learning.Finally, we present the research trends of studying beyond factoid VideoQA to inference VideoQA, as well as towards the robustness and interpretability.Additionally, we maintain a repository, https://github.com/VRU-NExT/ VideoQA, to keep trace of the latest VideoQA papers, datasets, and their open-source implementations if available.With these efforts, we strongly hope this survey could shed light on the follow-up VideoQA research. Yaoyao Zhong, Wei Ji 0008, Junbin Xiao, Yicong Li 0004, Weihong Deng, Tat-Seng Chua |
EMNLP | 5 |
| 2022 | Momentum Distillation Improves Multimodal Sentiment Analysis
Weihong Deng, Jiani Hu |
PRCV (1) | 2 |
| 2022 | Meta Balanced Network for Fair Face RecognitionabstractAlthough deep face recognition has achieved impressive progress in recent years, controversy has arisen regarding discrimination based on skin tone, questioning their deployment into real-world scenarios. In this paper, we aim to systematically and scientifically study this bias from both data and algorithm aspects. First, using the dermatologist approved Fitzpatrick Skin Type classification system and Individual Typology Angle, we contribute a benchmark called Identity Shades (IDS) database, which effectively quantifies the degree of the bias with respect to skin tone in existing face recognition algorithms and commercial APIs. Further, we provide two skin-tone aware training datasets, called BUPT-Globalface dataset and BUPT-Balancedface dataset, to remove bias in training data. Finally, to mitigate the algorithmic bias, we propose a novel meta-learning algorithm, called Meta Balanced Network (MBN), which learns adaptive margins in large margin loss such that the model optimized by this loss can perform fairly across people with different skin tones. To determine the margins, our method optimizes a meta skewness loss on a clean and unbiased meta set and utilizes backward-on-backward automatic differentiation to perform a second order gradient descent step on the current margins. Extensive experiments show that MBN successfully mitigates bias and learns more balanced performance for people with different skin tones in face recognition. The proposed datasets are available at http://www.whdeng.cn/RFW/index.html. Mei Wang 0001, Yaobin Zhang, Weihong Deng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Transferring discriminative knowledge via connective momentum clustering on person re-identification
Weihong Deng |
Pattern Recognit. | 2 |
| 2022 | Dual Gaussian Modeling for Deep Face Embeddings
Yuying Zhao, Weihong Deng |
Pattern Recognit. Lett. | 2 |
| 2022 | Disentangling Identity and Pose for Facial Expression RecognitionabstractFacial expression recognition (FER) is a challenging problem because the expression component is always entangled with other irrelevant factors, such as identity and head pose. In this work, we propose an identity and pose disentangled facial expression recognition (IPD-FER) model to learn more discriminative feature representation. We regard the holistic facial representation as the combination of identity, pose and expression. These three components are encoded with different encoders. For identity encoder, a well pre-trained face recognition model is utilized and fixed during training, which alleviates the restriction on specific expression training data in previous works and makes the disentanglement practicable on in-the-wild datasets. At the same time, the pose and expression encoder are optimized with corresponding labels. Combining identity and pose feature, a neutral face of input individual should be generated by the decoder. When expression feature is added, the input image should be reconstructed. By comparing the difference between synthesized neutral and expressional images of the same individual, the expression component is further disentangled from identity and pose. Experimental results verify the effectiveness of our method on both lab-controlled and in-the-wild databases and we achieve state-of-the-art recognition performance. Weihong Deng |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | A Deeper Look at Facial Expression Dataset BiasabstractDatasets play an important role in the progress of facial expression recognition algorithms, but they may suffer from obvious biases caused by different cultures and collection conditions. To look deeper into this bias, we first conduct comprehensive experiments on dataset recognition and cross-dataset generalization tasks, and for the first time, explore the intrinsic causes of the dataset discrepancy. The results quantitatively verify that current datasets have a strong build-in bias, and corresponding analyses indicate that the conditional probability distributions between source and target datasets are different. However, previous researches are mainly based on shallow features with limited discriminative ability under the assumption that the conditional distribution remains unchanged across domains. To address these issues, we further propose a novel deep Emotion-Conditional Adaption Network (ECAN) to learn domain-invariant and discriminative feature representations, which can match not only the marginal distribution but also the class-conditional distribution across domains by exploring the underlying label information of the target dataset. Moreover, the largely ignored expression class distribution bias is also addressed so that the training and testing domains can share similar class distribution. Extensive cross-database experiments on both lab-controlled datasets (CK+, JAFFE, MMI, and Oulu-CASIA) and real-world databases (AffectNet, FER2013, RAF-DB 2.0, and SFEW 2.0) demonstrate that our ECAN can yield competitive performances across various cross-dataset facial expression recognition tasks and outperform the state-of-the-art methods. Shan Li 0001, Weihong Deng |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Deep Facial Expression Recognition: A SurveyabstractWith the transition of facial expression recognition (FER) from laboratory-controlled to challenging in-the-wild conditions and the recent success of deep learning techniques in various fields, deep neural networks have increasingly been leveraged to learn discriminative representations for automatic FER. Recent deep FER systems generally focus on two important issues: overfitting caused by a lack of sufficient training data and expression-unrelated variations, such as illumination, head pose, and identity bias. In this survey, we provide a comprehensive review of deep FER, including datasets and algorithms that provide insights into these intrinsic problems. First, we introduce the available datasets that are widely used in the literature and provide accepted data selection and evaluation principles for these datasets. We then describe the standard pipeline of a deep FER system with the related background knowledge and suggestions for applicable implementations for each stage. For the state-of-the-art in deep FER, we introduce existing novel deep neural networks and related training strategies that are designed for FER based on both static images and dynamic image sequences and discuss their advantages and limitations. Competitive performances and experimental comparisons on widely used benchmarks are also summarized. We then extend our survey to additional related issues and application scenarios. Finally, we review the remaining challenges and corresponding opportunities in this field as well as future directions for the design of robust deep FER systems. Shan Li 0001, Weihong Deng |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Point Adversarial Self-Mining: A Simple Method for Facial Expression RecognitionabstractIn this article, we propose a simple yet effective approach, called point adversarial self mining (PASM), to improve the recognition accuracy in facial expression recognition (FER). Unlike previous works focusing on designing specific architectures or loss functions to solve this problem, PASM boosts the network capability by simulating human learning processes: providing updated learning materials and guidance from more capable teachers. Specifically, to generate new learning materials, PASM leverages a point adversarial attack method and a trained teacher network to locate the most informative position related to the target task, generating harder learning samples to refine the network. The searched position is highly adaptive since it considers both the statistical information of each sample and the teacher network capability. Other than being provided new learning materials, the student network also receives guidance from the teacher network. After the student network finishes training, the student network changes its role and acts as a teacher, generating new learning materials and providing stronger guidance to train a better student network. The adaptive learning materials generation and teacher/student update can be conducted more than one time, improving the network capability iteratively. Extensive experimental results validate the efficacy of our method over the existing state of the arts for FER. Ping Liu 0004, Yuewei Lin, Zibo Meng, Weihong Deng, Joey Tianyi Zhou, Yi Yang 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | Learning Multi-Granularity Temporal Characteristics for Face Anti-SpoofingabstractFace anti-spoofing (FAS) is essential for securing face recognition systems. Despite the decent performance, few existing works fully leverage temporal information. This would inevitably lead to inferior performance because real and fake faces tend to share highly similar spatial appearances, while important temporal features between consecutive frames are neglected. In this work, we propose a temporal transformer network (TTN) to learn multi-granularity temporal characteristics for FAS. It mainly consists of temporal difference attentions (TDA), a pyramid temporal aggregation (PTA), and a temporal depth difference loss (TDL). Firstly, the vision transformer (ViT) is used as the backbone where comprehensive local patches are utilized to provide subtle differences between live and spoof faces. Then, instead of learning temporal features on global faces which may miss some important local cues, the TDA is developed to extract motion-sensitive cues on each of the comprehensive local patches. Moreover, the TDA is inserted into different layers of the ViT, learning multi-scale motion-sensitive local cues to improve the FAS performance. Secondly, it is observed that different subjects may have different visual tempos in some actions, making it necessary to model different temporal speeds. Our PTA aggregates temporal features at various tempos, which could build short-range and long-range relations among multiple frames. Thirdly, depth maps for real parts may change continuously, while they remain zeros for spoof regions. In order to locate motion features on facial parts, the TDL is proposed to guide the network to locate spoof facial parts where motion patterns between neighboring frames are set as the ground truth. To the best of our knowledge, this work is the first attempt to learn temporal characteristics via transformers. Both qualitative and quantitative results on several challenging tasks demonstrate the usefulness and effectiveness of our proposed methods. Qiangchang Wang, Weihong Deng, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Detecting Overlapped Objects in X-Ray Security Imagery by a Label-Aware MechanismabstractOne of the key challenges to the X-ray security check is to detect the overlapped items in backpacks or suitcases in the X-ray images. Most existing methods improve the robustness of models to the object overlapping problem by enhancing the underlying visual information such as colors and edges. However, this strategy ignores the situations that the objects have similar visual clues as to the background, and objects overlapping each other. Since the two cases rarely appear in existing datasets, we contribute a novel dataset – Cutters and Liquid Containers X-ray Dataset (CLCXray) to complete the related research. Furthermore, we propose a novel Label-aware Mechanism (LA) to tackle the object overlapping problem. Particularly, LA establishes the associations between feature channels and different labels and adjusts the features according to the assigned labels (or pseudo labels) to help improve the prediction results. Extensive experiments demonstrate that the LA is accurate and robust to detect overlapped objects, and also validate the effectiveness and the good generalization of the LA for arbitrary state-of-the-art (SOTA) methods. Furthermore, experimental results show that the network constructed by the LA is superior to the SOTA models on OPIXray and CLCXray, especially solving the challenges of the subset of the highly overlapped objects. Cairong Zhao, Shuguang Dou, Weihong Deng, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Adaptive Face Recognition Using Adversarial Information NetworkabstractIn many real-world applications, face recognition models often degenerate when training data (referred to as source domain) are different from testing data (referred to as target domain). To alleviate this mismatch caused by some factors like pose and skin tone, the utilization of pseudo-labels generated by clustering algorithms is an effective way in unsupervised domain adaptation. However, they always miss some hard positive samples. Supervision on pseudo-labeled samples attracts them towards their prototypes and would cause an intra-domain gap between pseudo-labeled samples and the remaining unlabeled samples within target domain, which results in the lack of discrimination in face recognition. In this paper, considering the particularity of face recognition, we propose a novel adversarial information network (AIN) to address it. First, a novel adversarial mutual information (MI) loss is proposed to alternately minimize MI with respect to the target classifier and maximize MI with respect to the feature extractor. By this min-max manner, the positions of target prototypes are adaptively modified which makes unlabeled images clustered more easily such that intra-domain gap can be mitigated. Second, to assist adversarial MI loss, we utilize a graph convolution network to predict linkage likelihoods between target data and generate pseudo-labels. It leverages valuable information in the context of nodes and can achieve more reliable results. The proposed method is evaluated under two scenarios, i.e., domain adaptation across poses and image conditions, and domain adaptation across faces with different skin tones. Extensive experiments show that AIN successfully improves cross-domain generalization and offers a new state-of-the-art on RFW dataset. Mei Wang 0001, Weihong Deng |
IEEE Trans. Image Process. | 2 |
| 2022 | Unsupervised Structure-Texture Separation Network for Oracle Character RecognitionabstractOracle bone script is the earliest-known Chinese writing system of the Shang dynasty and is precious to archeology and philology. However, real-world scanned oracle data are rare and few experts are available for annotation which make the automatic recognition of scanned oracle characters become a challenging task. Therefore, we aim to explore unsupervised domain adaptation to transfer knowledge from handprinted oracle data, which are easy to acquire, to scanned domain. We propose a structure-texture separation network (STSN), which is an end-to-end learning framework for joint disentanglement, transformation, adaptation and recognition. First, STSN disentangles features into structure (glyph) and texture (noise) components by generative models, and then aligns handprinted and scanned data in structure feature space such that the negative influence caused by serious noises can be avoided when adapting. Second, transformation is achieved via swapping the learned textures across domains and a classifier for final classification is trained to predict the labels of the transformed scanned characters. This not only guarantees the absolute separation, but also enhances the discriminative ability of the learned features. Extensive experiments on Oracle-241 dataset show that STSN outperforms other adaptation methods and successfully improves recognition performance on scanned data even when they are contaminated by long burial and careless excavation. Mei Wang 0001, Weihong Deng, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Dynamic Training Data Dropout for Robust Deep Face RecognitionabstractLearning with noise is a practically challenging problem in deep face recognition. Despite the success of large margin softmax loss functions, these methods are designed for clean face databases. Considering the inevitable noise in the large scale databases, we first analyze the performance of noise in the training databases. For noise-robust deep face recognition, we propose a dynamic training data dropout (DTDD) method to dynamically filter the noise in the training database and gradually form a stable refined database for model learning. Specifically, we leverage the information provided by the model predictions of accumulated training epochs, which can distinguish regular samples and noise effectively and accurately. The proposed DTDD method is easy and stable for implementation, and can be combined with existing state-of-the-art loss functions and network architectures. Extensive experiments on CASIA-WebFace, VGGFace2, and MS-Celeb-1 M databases empirically demonstrate that our proposed method can robustly train deep face recognition models in the presence of label noise and low quality images. Yaoyao Zhong, Weihong Deng, Han Fang 0002, Jiani Hu, Dongyue Zhao, Dongchao Wen |
IEEE Trans. Multim. | 2 |
| 2021 | Selective Pseudo-Labeling with Reinforcement Learning for Semi-Supervised Domain Adaptation
Yuhong Guo, Jieping Ye, Weihong Deng |
BMVC | 4 |
| 2021 | Representative Forgery Mining for Fake Face DetectionabstractAlthough vanilla Convolutional Neural Network (CNN) based detectors can achieve satisfactory performance on fake face detection, we observe that the detectors tend to seek forgeries on a limited region of face, which reveals that the detectors is short of understanding of forgery. Therefore, we propose an attention-based data augmentation framework to guide detector refine and enlarge its attention. Specifically, our method tracks and occludes the Top-N sensitive facial regions, encouraging the detector to mine deeper into the regions ignored before for more representative forgery. Especially, our method is simple-to-use and can be easily integrated with various CNN models. Extensive experiments show that the detector trained with our method is capable to separately point out the representative forgery of fake faces generated by different manipulation techniques, and our method enables a vanilla CNN-based detector to achieve state-of-the-art performance without structure modification. Our code is available at https://github.com/crywang/RFM. Weihong Deng |
CVPR | 2 |
| 2021 | Augmented Face Representation Learning via Transitive DistillationabstractThe wild face of large variations is hard to recognize in unconstrained scenarios. To tackle this issue, existing works synthesize and augment the variation-specific faces for recognition. However, directly feeding generated samples results in negative transfer, because the feature spaces are shifted compared with normal samples. Instead, we propose a transitive distillation network (TDNet) that introduces a transitive domain to transfer cross-variation representations, which alleviates the negative influence of synthesized data. Specifically, data of diverse variations are firstly synthesized. Then we construct distributions from different variations as teachers to distill student. The negative transfer is mitigated by adopting adaptor as a bridge to break large domain distance. To handle faces of different quality, we propose a novel strategy to define easy and hard samples, which are utilized to select specific transitive status. Meanwhile, bilateral classification with curriculum learning is proposed to improve confidence of synthesized data gradually, enhancing the robustness of representation learning. Experiments show that our method achieves superiority on unconstrained face benchmarks such as IJB-C and SCface, while maintaining competence on general test sets. Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen |
FG | 2 |
| 2021 | Identifying Rhythmic Patterns for Face Forgery Detection and CategorizationabstractWith the emergence of GAN, face forgery technologies have been heavily abused. Achieving accurate face forgery detection is imminent. Inspired by remote photoplethysmography (rPPG) that PPG signal corresponds to the periodic change of skin color caused by heartbeat in face videos, we observe that despite the inevitable loss of PPG signal during the forgery process, there is still a mixture of PPG signals in the forgery video with a unique rhythmic pattern depending on its generation method. Motivated by this key observation, we propose a two-stage network for face forgery detection and categorization consisting of: 1) a Spatial-Temporal Filter Module (STFM) for PPG signals filtering, and 2) an Adjacency Interaction Module (AIM) for constraint and interaction of PPG signals. Moreover, with insight into the generation of forgery methods, we further propose Spatial-Temporal Mixup (ST-Mixup) to boost the performance of the network. Overall, extensive experiments have proved the superiority of our method. Weihong Deng |
IJCB | 2 |
| 2021 | Adaptive Label Noise Cleaning with Meta-Supervision for Deep Face RecognitionabstractThe training of a deep face recognition system usually faces the interference of label noise in the training data. However, it is difficult to obtain a high-precision cleaning model to remove these noises. In this paper, we propose an adaptive label noise cleaning algorithm based on meta-learning for face recognition datasets, which can learn the distribution of the data to be cleaned and make automatic adjustments based on class differences. It first learns re-liable cleaning knowledge from well-labeled noisy data, then gradually transfers it to the target data with meta-supervision to improve performance. A threshold adapter module is also proposed to address the drift problem in transfer learning methods. Extensive experiments clean two noisy in-the-wild face recognition datasets and show the effectiveness of the proposed method to reach state-of-the-art performance on the IJB-C face recognition benchmark. Yaobin Zhang, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen |
ICCV | 2 |
| 2021 | Multi-view Correlation based Black-box Adversarial Attack for 3D Object DetectionabstractDeep neural networks have made tremendous progress in 3D object detection, which is an important task especially in autonomous driving scenarios. Benefited from the breakthroughs in deep learning and sensor technologies, 3D object detection methods based on different sensors, such as camera and LiDAR, have developed rapidly. Meanwhile, more and more researches notice that the abundant information contained in the multi-view data can be used to obtain more accurate understanding of the 3D surrounding environment. Therefore, many sensor-fusion 3D object detection methods have been proposed. As safety is critical in autonomous driving and the deep neural networks are known to be vulnerable to adversarial examples with visually imperceptible perturbations, it is significant to investigate adversarial attacks for 3D object detection. Recent works have shown that both image-based and LiDAR-based networks can be attacked by the adversarial examples while the attacks to the sensor-fusion models, which tend to be more robust, haven't been studied. To this end, we propose a simple multi-view correlation based adversarial attack method for the camera-LiDAR fusion 3D object detection models and focus on the black-box attack setting which is more practical in real-world systems. Specifically, we first design a generative network to generate image adversarial examples based on an auxiliary image semantic segmentation network. Then, we develop a cross-view perturbation projection method by exploiting the camera-LiDAR correlations to map each image adversarial example to the space of the point cloud data to form the point cloud adversarial examples in the LiDAR view. Extensive experiments on the KITTI dataset demonstrate the effectiveness of the proposed method. Yuhong Guo, Jianan Jiang, Jian Tang 0008, Weihong Deng |
KDD | 5 |
| 2021 | Relative Uncertainty Learning for Facial Expression RecognitionabstractIn facial expression recognition (FER), the uncertainties introduced by inherent noises like ambiguous facial expressions and inconsistent labels raise concerns about the credibility of recognition results. To quantify these uncertainties and achieve good performance under noisy data, we regard uncertainty as a relative concept and propose an innovative uncertainty learning method called Relative Uncertainty Learning (RUL). Rather than assuming Gaussian uncertainty distributions for all datasets, RUL builds an extra branch to learn uncertainty from the relative difficulty of samples by feature mixup. Specifically, we use uncertainties as weights to mix facial features and design an add-up loss to encourage uncertainty learning. It is easy to implement and adds little or no extra computation overhead. Extensive experiments show that RUL outperforms state-of-the-art FER uncertainty learning methods in both real-world and synthetic noisy FER datasets. Besides, RUL also works well on other datasets such as CIFAR and Tiny ImageNet. The code is available at https://github.com/zyh-uaiaaaa/Relative-Uncertainty-Learning. Weihong Deng |
NeurIPS | 3 |
| 2021 | FIE-GAN: Illumination Enhancement Network for Face Recognition
Weihong Deng, Jiancheng Ge |
PRCV (3) | 2 |
| 2021 | Cycle label-consistent networks for unsupervised domain adaptation
Mei Wang 0001, Weihong Deng |
Neurocomputing | 2 |
| 2021 | Deep face recognition: A survey
Mei Wang 0001, Weihong Deng |
Neurocomputing | 2 |
| 2021 | Orthogonality Loss: Learning Discriminative Representations for Face RecognitionabstractConvolutional neural networks have achieved excellent performance on face recognition (FR) by learning the high discriminative features with advanced loss functions. These improved loss functions share the similar idea for maximizing inter-class variance or minimizing intra-class variance. In this article, from a different perspective, we consider enlarging the inter-class variance by directly penalizing weight vectors of last fully connected layer, which represent the center of classes. To the end, we propose Orthogonality loss as an elegant penalty item appends to common classification loss to learn the discriminative representations. The main idea is that in order for weight vectors to be discriminative, it should be as close as possible to be orthogonal to each other in the vector space. More specifically, the optimization objective of Orthogonality loss is the first moment and second moment of cosine similarity of weight vectors. We performed the empirical studies through simulating the long-tail datasets to show the generalization ability of the proposed approach on long-tail distribution datasets. Further, extensive experiments on large-scale face recognition benchmarks including the Labeled Face in the Wild (LFW), the IARPA Janus Benchmark A (IJB-A), IJB-B, IJB-C, MegaFace Challenge 1 (MF1) and MS-Celeb-1M Low-shot Learning demonstrated that Orthogonality loss outperforms strong baselines, which showcases the extensive suitability and effectiveness of Orthogonality loss. Shan-Ming Yang, Weihong Deng, Mei Wang 0001, Junping Du 0001, Jiani Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Towards Transferable Adversarial Attack Against Deep Face RecognitionabstractFace recognition has achieved great success in the last five years due to the development of deep learning methods. However, deep convolutional neural networks (DCNNs) have been found to be vulnerable to adversarial examples. In particular, the existence of transferable adversarial examples can severely hinder the robustness of DCNNs since this type of attacks can be applied in a fully black-box manner without queries on the target system. In this work, we first investigate the characteristics of transferable adversarial attacks in face recognition by showing the superiority of feature-level methods over label-level methods. Then, to further improve transferability of feature-level adversarial examples, we propose DFANet, a dropout-based method used in convolutional layers, which can increase the diversity of surrogate models and obtain ensemble-like effects. Extensive experiments on state-of-the-art face models with various training databases, loss functions and network architectures show that the proposed method can significantly enhance the transferability of existing attack methods. Finally, by applying DFANet to the LFW database, we generate a new set of adversarial face pairs that can successfully attack four commercial APIs without any queries. This TALFW database is available to facilitate research on the robustness and defense of deep face recognition. Yaoyao Zhong, Weihong Deng |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face RecognitionabstractDeep face recognition has achieved great success due to large-scale training databases and rapidly developing loss functions. The existing algorithms devote to realizing an ideal idea: minimizing the intra-class distance and maximizing the inter-class distance. However, they may neglect that there are also low quality training images which should not be optimized in this strict way. Considering the imperfection of training databases, we propose that intra-class and inter-class objectives can be optimized in a moderate way to mitigate overfitting problem, and further propose a novel loss function, named sigmoid-constrained hypersphere loss (SFace). Specifically, SFace imposes intra-class and inter-class constraints on a hypersphere manifold, which are controlled by two sigmoid gradient re-scale functions respectively. The sigmoid curves precisely re-scale the intra-class and inter-class gradients so that training samples can be optimized to some degree. Therefore, SFace can make a better balance between decreasing the intra-class distances for clean examples and preventing overfitting to the label noise, and contributes more robust deep face recognition models. Extensive experiments of models trained on CASIA-WebFace, VGGFace2, and MS-Celeb-1M databases, and evaluated on several face recognition benchmarks, such as LFW, MegaFace and IJB-C databases, have demonstrated the superiority of SFace. Yaoyao Zhong, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen |
IEEE Trans. Image Process. | 2 |
| 2020 | RAF-AU Database: In-the-Wild Facial Expressions with Subjective Emotion Judgement and Objective AU Annotations
Shan Li 0001, Chengtao Que, JiQuan Pei, Weihong Deng |
ACCV (6) | 5 |
| 2020 | PropagationNet: Propagate Points to Curve to Learn Structure InformationabstractDeep learning technique has dramatically boosted the performance of face alignment algorithms. However, due to large variability and lack of samples, the alignment problem in unconstrained situations, e.g. large head poses, exaggerated expression, and uneven illumination, is still largely unsolved. In this paper, we explore the instincts and reasons behind our two proposals, i.e. Propagation Module and Focal Wing Loss, to tackle the problem. Concretely, we present a novel structure-infused face alignment algorithm based on heatmap regression via propagating landmark heatmaps to boundary heatmaps, which provide structure information for further attention map generation. Moreover, we propose a Focal Wing Loss for mining and emphasizing the difficult samples under in-the-wild condition. In addition, we adopt methods like CoordConv and Anti-aliased CNN from other fields that address the shift variance problem of CNN for face alignment. When implementing extensive experiments on different benchmarks, i.e. WFLW, 300W, and COFW, our method outperforms the state-of-the-arts by a significant margin. Our proposed approach achieves 4.05% mean error on WFLW, 2.93% mean error on 300W full-set, and 3.71% mean error on COFW. Xiehe Huang, Weihong Deng, Haifeng Shen, Xiubao Zhang, Jieping Ye |
CVPR | 2 |
| 2020 | Mitigating Bias in Face Recognition Using Skewness-Aware Reinforcement LearningabstractRacial equality is an important theme of international human rights law, but it has been largely obscured when the overall face recognition accuracy is pursued blindly. More facts indicate racial bias indeed degrades the fairness of recognition system and the error rates on non-Caucasians are usually much higher than Caucasians. To encourage fairness, we introduce the idea of adaptive margin to learn balanced performance for different races based on large margin losses. A reinforcement learning based race balance network (RL-RBN) is proposed. We formulate the process of finding the optimal margins for non-Caucasians as a Markov decision process and employ deep Q-learning to learn policies for an agent to select appropriate margin by approximating the Q-value function. Guided by the agent, the skewness of feature scatter between races can be reduced. Besides, we provide two ethnicity aware training datasets, called BUPT-Globalface and BUPT-Balancedface dataset, which can be utilized to study racial bias from both data and algorithm aspects. Extensive experiments on RFW database show that RL-RBN successfully mitigates racial bias and learns more balanced performance. Weihong Deng |
CVPR | 2 |
| 2020 | Global-Local GCN: Large-Scale Label Noise Cleansing for Face RecognitionabstractIn the field of face recognition, large-scale web-collected datasets are essential for learning discriminative representations, but they suffer from noisy identity labels, such as outliers and label flips. It is beneficial to automatically cleanse their label noise for improving recognition accuracy. Unfortunately, existing cleansing methods cannot accurately identify noise in the wild. To solve this problem, we propose an effective automatic label noise cleansing framework for face recognition datasets, FaceGraph. Using two cascaded graph convolutional networks, FaceGraph performs global-to-local discrimination to select useful data in a noisy environment. Extensive experiments show that cleansing widely used datasets, such as CASIA-WebFace, VGGFace2, MegaFace2, and MS-Celeb-1M, using the proposed method can improve the recognition performance of state-of-the-art representation learning methods like Arcface. Further, we cleanse massive self-collected celebrity data, namely MillionCelebs, to provide 18.8M images of 636K identities. Training with the new data, Arcface surpasses state-of-the-art performance by a notable margin to reach 95.62% TPR at 1e-5 FPR on the IJB-C benchmark. Yaobin Zhang, Weihong Deng, Jiani Hu, Dongyue Zhao, Dongchao Wen |
CVPR | 2 |
| 2020 | Generate to Adapt: Resolution Adaption Network for Surveillance Face Recognition
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu |
ECCV (15) | 2 |
| 2020 | FGAN: Fan-Shaped GAN for Racial TransformationabstractRacial bias in face recognition has recently been concerned by both general public and research community. Most face recognition systems have a strong bias in recognition accuracy for different races mainly because of the unbalanced ethnic distribution in their datasets. In this paper, we propose a novel generative adversarial network, which transfer the facial images of one race to corresponding images of other races, to facilitate the data augmentation to balance the ethnic distribution. Our approach can generate more realistic results and make the training process more stable than other image-to-image translation methods such as StarGAN and CycleGAN. Experiments results show the superiority of FGAN to the previous methods on the racial transformation task in terms of visual effects and quantitative results. Besides, we perform extensive experiments to show our data augmentation is beneficial to reduce the racial bias, improving the face recognition rate of non-Caucasian people. Finally, we show the possibility to generate the ethnic independent facial image by the average of various races. Jiancheng Ge, Weihong Deng, Jiani Hu |
IJCB | 2 |
| 2020 | A Multi-Modal Approach for Driver Gaze Prediction to Remove Identity BiasabstractDriver gaze prediction is an important task in Advanced Driver Assistance System (ADAS). Although the Convolutional Neural Network (CNN) can greatly improve the recognition ability, there are still several unsolved problems due to the challenge of illumination, pose and camera placement. To solve these difficulties, we propose an effective multi-model fusion method for driver gaze estimation. Rich appearance representations, i.e. holistic and eyes regions, and geometric representations, i.e. landmarks and Delaunay angles, are separately learned to predict the gaze, followed by a score-level fusion system. Moreover, pseudo-3D appearance supervision and identity-adaptive geometric normalization are proposed to further enhance the prediction accuracy. Finally, the proposed method achieves state-of-the-art accuracy of 82.5288% on the test data, which ranks 1st at the EmotiW2020 driver gaze prediction sub-challenge. Zehui Yu, Xiehe Huang, Xiubao Zhang, Haifeng Shen, Qun (Tracy) Li, Weihong Deng, Jian Tang 0008, Jieping Ye |
ICMI | 6 |
| 2020 | H-AT: Hybrid Attention Transfer for Knowledge Distillation
Yan Qu, Weihong Deng, Jiani Hu |
PRCV (3) | 2 |
| 2020 | Deep face recognition with clustering based domain adaptation
Mei Wang 0001, Weihong Deng |
Neurocomputing | 2 |
| 2020 | Improved community structure discovery algorithm based on combined clique percolation method and K-means algorithm
Zhou Zhou 0001, Zhuopeng Xiao, Weihong Deng |
Peer-to-Peer Netw. Appl. | 3 |
| 2020 | Identity-aware CycleGAN for face photo-sketch synthesis and recognition
Yuke Fang, Weihong Deng, Junping Du 0001, Jiani Hu |
Pattern Recognit. | 2 |
| 2019 | Energy Confused Adversarial Metric Learning for Zero-Shot Image Retrieval and ClusteringabstractDeep metric learning has been widely applied in many computer vision tasks, and recently, it is more attractive in zeroshot image retrieval and clustering (ZSRC) where a good embedding is requested such that the unseen classes can be distinguished well. Most existing works deem this ’good’ embedding just to be the discriminative one and thus race to devise powerful metric objectives or hard-sample mining strategies for leaning discriminative embedding. However, in this paper, we first emphasize that the generalization ability is a core ingredient of this ’good’ embedding as well and largely affects the metric performance in zero-shot settings as a matter of fact. Then, we propose the Energy Confused Adversarial Metric Learning (ECAML) framework to explicitly optimize a robust metric. It is mainly achieved by introducing an interesting Energy Confusion regularization term, which daringly breaks away from the traditional metric learning idea of discriminative objective devising, and seeks to ’confuse’ the learned model so as to encourage its generalization ability by reducing overfitting on the seen classes. We train this confusion term together with the conventional metric objective in an adversarial manner. Although it seems weird to ’confuse’ the network, we show that our ECAML indeed serves as an efficient regularization technique for metric learning and is applicable to various conventional metric methods. This paper empirically and experimentally demonstrates the importance of learning embedding with good generalization, achieving state-of-theart performances on the popular CUB, CARS, Stanford Online Products and In-Shop datasets for ZSRC tasks. Code available at http://www.bhchen.cn/. Binghui Chen, Weihong Deng |
AAAI | 2 |
| 2019 | Hybrid-Attention Based Decoupled Metric Learning for Zero-Shot Image RetrievalabstractIn zero-shot image retrieval (ZSIR) task, embedding learning becomes more attractive, however, many methods follow the traditional metric learning idea and omit the problems behind zero-shot settings. In this paper, we first emphasize the importance of learning visual discriminative metric and preventing the partial/selective learning behavior of learner in ZSIR, and then propose the Decoupled Metric Learning (DeML) framework to achieve these individually. Instead of coarsely optimizing an unified metric, we decouple it into multiple attention-specific parts so as to recurrently induce the discrimination and explicitly enhance the generalization. And they are mainly achieved by our object-attention module based on random walk graph propagation and the channel-attention module based on the adversary constraint, respectively. We demonstrate the necessity of addressing the vital problems in ZSIR on the popular benchmarks, outperforming the state-of-the-art methods by a significant margin. Code is available at http://www.bhchen.cn Binghui Chen, Weihong Deng |
CVPR | 2 |
| 2019 | Unsupervised Face Normalization With Extreme Pose and Expression in the WildabstractFace recognition achieves great success thanks to the emergence of deep learning. However, many contemporary face recognition models still have limited invariance to strong intra-personal variations such as large pose changes. Face normalization provides an effective and cheap way to distil face identity and dispel face variances for recognition. We focus on face generation in the wild with unpaired data. To this end, we propose a Face Normalization Model (FNM) to generate a frontal, neutral expression, photorealistic face image for face recognition. FNM is a well-designed Generative Adversarial Network (GAN) with three distinct novelties. First, a face expert network is introduced to construct generator and provide the ability of retaining face identity. Second, with the reconstruction of normal face, pixel-wise loss is applied to stabilize optimization process. Third, we present a series of face attention discriminators to refine local textures. FNM could recover canonical-view, expression-free image and directly improve the performance of face recognition model. Extensive qualitative and quantitative experiments on both controlled and in-the-wild databases demonstrate the superiority of our face normalization method. Yichen Qian, Weihong Deng, Jiani Hu |
CVPR | 2 |
| 2019 | Signal-To-Noise Ratio: A Robust Distance Metric for Deep Metric LearningabstractDeep metric learning, which learns discriminative features to process image clustering and retrieval tasks, has attracted extensive attention in recent years. A number of deep metric learning methods, which ensure that similar examples are mapped close to each other and dissimilar examples are mapped farther apart, have been proposed to construct effective structures for loss functions and have shown promising results. In this paper, different from the approaches on learning the loss structures, we propose a robust SNR distance metric based on Signal-to-Noise Ratio (SNR) for measuring the similarity of image pairs for deep metric learning. By exploring the properties of our SNR distance metric from the view of geometry space and statistical theory, we analyze the properties of our metric and show that it can preserve the semantic similarity between image pairs, which well justify its suitability for deep metric learning. Compared with Euclidean distance metric, our SNR distance metric can further jointly reduce the intra-class distances and enlarge the inter-class distances for learned features. Leveraging our SNR distance metric, we propose Deep SNR-based Metric Learning (DSML) to generate discriminative feature embeddings. By extensive experiments on three widely adopted benchmarks, including CARS196, CUB200-2011 and CIFAR10, our DSML has shown its superiority over other state-of-the-art methods. Additionally, we extend our SNR distance metric to deep hashing learning, and conduct experiments on two benchmarks, including CIFAR10 and NUS-WIDE, to demonstrate the effectiveness and generality of our SNR distance metric. Tongtong Yuan, Weihong Deng, Jian Tang 0008, Yinan Tang, Binghui Chen |
CVPR | 2 |
| 2019 | Unequal-Training for Deep Face Recognition With Long-Tailed Noisy DataabstractLarge-scale face datasets usually exhibit a massive number of classes, a long-tailed distribution, and severe label noise, which undoubtedly aggravate the difficulty of training. In this paper, we propose a training strategy that treats the head data and the tail data in an unequal way, accompanying with noise-robust loss functions, to take full advantage of their respective characteristics. Specifically, the unequal-training framework provides two training data streams: the first stream applies the head data to learn discriminative face representation supervised by Noise Resistance loss; the second stream applies the tail data to learn auxiliary information by gradually mining the stable discriminative information from confusing tail classes. Consequently, both training streams offer complementary information to deep feature learning. Extensive experiments have demonstrated the effectiveness of the new unequal-training framework and loss functions. Better yet, our method could save a significant amount of GPU memory. With our method, we achieve the best result on MegaFace Challenge 2 (MF2) given a large-scale noisy training data set. Yaoyao Zhong, Weihong Deng, Jiani Hu, Jianteng Peng, Xunqiang Tao, Yaohai Huang |
CVPR | 2 |
| 2019 | Exploring Features and Attributes in Deep Face Recognition Using Visualization TechniquesabstractDeep convolutional neural networks (CNNs) currently have achieved state-of-the-art results on face recognition; yet, the understanding behind the success of the deep face model is still lacking. In particular, it is still unclear the inner workings of deep face model. What effective features does a deep face model learn? What do these features represent and what is the sematic meaning of them? This work explores this problem by analyzing the classic network VGGFace using deep visualization techniques. We first explore features computed by neurons, investigating characters of features like diversity, invariance, discrimination. It's worth noting that the middle layer is the least robust to transform, which contradicts the conventional view that robustness to transform increases as the network going deeper. The most significant phenomenon we find is that high level features are correspond with complex face attributes which human could not describe using a few words. We present a quantitative analysis on these face attributes perceived by deep CNNs, understanding them and the complex relationships between them. Additionally, we also focus on the significant point, the pose invariance in face recognition. Our research is the first work to understand the inner works of deep face models, elucidating some particular phenomena in deep face recognition. Yaoyao Zhong, Weihong Deng |
FG | 2 |
| 2019 | Mixed High-Order Attention Network for Person Re-IdentificationabstractAttention has become more attractive in person re-identification (ReID) as it is capable of biasing the allocation of available resources towards the most informative parts of an input signal. However, state-of-the-art works concentrate only on coarse or first-order attention design, e.g. spatial and channels attention, while rarely exploring higher-order attention mechanism. We take a step towards addressing this problem. In this paper, we first propose the High-Order Attention (HOA) module to model and utilize the complex and high-order statistics information in attention mechanism, so as to capture the subtle differences among pedestrians and to produce the discriminative attention proposals. Then, rethinking person ReID as a zero-shot learning problem, we propose the Mixed High-Order Attention Network (MHN) to further enhance the discrimination and richness of attention knowledge in an explicit manner. Extensive experiments have been conducted to validate the superiority of our MHN for person ReID over a wide variety of state-of-the-art methods on three large-scale datasets, including Market-1501, DukeMTMC-ReID and CUHK03-NP. Code is available at http://www.bhchen.cn. Binghui Chen, Weihong Deng, Jiani Hu |
ICCV | 2 |
| 2019 | Fair Loss: Margin-Aware Reinforcement Learning for Deep Face RecognitionabstractRecently, large-margin softmax loss methods, such as angular softmax loss (SphereFace), large margin cosine loss (CosFace), and additive angular margin loss (ArcFace), have demonstrated impressive performance on deep face recognition. These methods incorporate a fixed additive margin to all the classes, ignoring the class imbalance problem. However, imbalanced problem widely exists in various real-world face datasets, in which samples from some classes are in a higher number than others. We argue that the number of a class would influence its demand for the additive margin. In this paper, we introduce a new margin-aware reinforcement learning based loss function, namely fair loss, in which each class will learn an appropriate adaptive margin by Deep Q-learning. Specifically, we train an agent to learn a margin adaptive strategy for each class, and make the additive margins for different classes more reasonable. Our method has better performance than present large-margin loss functions on three benchmarks, Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace, which demonstrates that our method could learn better face representation on imbalanced face datasets. Weihong Deng, Yaoyao Zhong, Jiani Hu, Xunqiang Tao, Yaohai Huang |
ICCV | 2 |
| 2019 | Racial Faces in the Wild: Reducing Racial Bias by Information Maximization Adaptation NetworkabstractRacial bias is an important issue in biometric, but has not been thoroughly studied in deep face recognition. In this paper, we first contribute a dedicated dataset called Racial Faces in-the-Wild (RFW) database, on which we firmly validated the racial bias of four commercial APIs and four state-of-the-art (SOTA) algorithms. Then, we further present the solution using deep unsupervised domain adaptation and propose a deep information maximization adaptation network (IMAN) to alleviate this bias by using Caucasian as source domain and other races as target domains. This unsupervised method simultaneously aligns global distribution to decrease race gap at domain-level, and learns the discriminative target representations at cluster level. A novel mutual information loss is proposed to further enhance the discriminative ability of network output without label information. Extensive experiments on RFW, GBU, and IJB-A databases show that IMAN successfully learns features that generalize well across different races and across different databases. Weihong Deng, Jiani Hu, Xunqiang Tao, Yaohai Huang |
ICCV | 2 |
| 2019 | Adversarial Learning With Margin-Based Triplet Embedding RegularizationabstractThe Deep neural networks (DNNs) have achieved great success on a variety of computer vision tasks, however, they are highly vulnerable to adversarial attacks. To address this problem, we propose to improve the local smoothness of the representation space, by integrating a margin-based triplet embedding regularization term into the classification objective, so that the obtained models learn to resist adversarial examples. The regularization term consists of two steps optimizations which find potential perturbations and punish them by a large margin in an iterative way. Experimental results on MNIST, CASIA-WebFace, VGGFace2 and MS-Celeb-1M reveal that our approach increases the robustness of the network against both feature and label adversarial attacks in simple object classification and deep face recognition. Yaoyao Zhong, Weihong Deng |
ICCV | 2 |
| 2019 | Blended Emotion in-the-Wild: Multi-label Facial Expression Recognition Using Crowdsourced Annotations and Deep Locality Feature Learning
Weihong Deng |
Int. J. Comput. Vis. | 2 |
| 2019 | Unsupervised adaptive hashing based on feature clustering
Tongtong Yuan, Weihong Deng, Jiani Hu, Zhanfu An, Yinan Tang |
Neurocomputing | 2 |
| 2019 | Compressive Binary Patterns: Designing a Robust Binary Face Descriptor with Random-Field EigenfiltersabstractA binary descriptor typically consists of three stages: image filtering, binarization, and spatial histogram. This paper first demonstrates that the binary code of the maximum-variance filtering responses leads to the lowest bit error rate under Gaussian noise. Then, an optimal eigenfilter bank is derived from a universal assumption on the local stationary random field. Finally, compressive binary patterns (CBP) is designed by replacing the local derivative filters of local binary patterns (LBP) with these novel random-field eigenfilters, which leads to a compact and robust binary descriptor that characterizes the most stable local structures that are resistant to image noise and degradation. A scattering-like operator is subsequently applied to enhance the distinctiveness of the descriptor. Surprisingly, the results obtained from experiments on the FERET, LFW, and PaSC databases show that the scattering CBP (SCBP) descriptor, which is handcrafted by only 6 optimal eigenfilters under restrictive assumptions, outperforms the state-of-the-art learning-based face descriptors in terms of both matching accuracy and robustness. In particular, on probe images degraded with noise, blur, JPEG compression, and reduced resolution, SCBP outperforms other descriptors by a greater than 10 percent accuracy margin. Weihong Deng, Jiani Hu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Deep embedding learning with adaptive large margin N-pair loss for image retrieval and clustering
Binghui Chen, Weihong Deng |
Pattern Recognit. | 2 |
| 2019 | Reliable Crowdsourcing and Deep Locality-Preserving Learning for Unconstrained Facial Expression RecognitionabstractFacial expression is central to human experience, but most previous databases and studies are limited to posed facial behavior under controlled conditions. In this paper, we present a novel facial expression database, Real-world Affective Face Database (RAF-DB), which contains approximately 30 000 facial images with uncontrolled poses and illumination from thousands of individuals of diverse ages and races. During the crowdsourcing annotation, each image is independently labeled by approximately 40 annotators. An expectation-maximization algorithm is developed to reliably estimate the emotion labels, which reveals that real-world faces often express compound or even mixture emotions. A cross-database study between RAF-DB and CK+ database further indicates that the action units of real-world emotions are much more diverse than, or even deviate from, those of laboratory-controlled emotions. To address the recognition of multi-modal expressions in the wild, we propose a new deep locality-preserving convolutional neural network (DLP-CNN) method that aims to enhance the discriminative power of deep features by preserving the locality closeness while maximizing the inter-class scatter. Benchmark experiments on 7-class basic expressions and 11-class compound expressions, as well as additional experiments on CK+, MMI, and SFEW 2.0 databases, show that the proposed DLP-CNN outperforms the state-of-the-art handcrafted features and deep learning-based methods for expression recognition in the wild. To promote further study, we have made the RAF database, benchmarks, and descriptor encodings publicly available to the research community. Shan Li 0001, Weihong Deng |
IEEE Trans. Image Process. | 2 |
| 2018 | Deep Transfer Network with 3D Morphable Models for Face RecognitionabstractData augmentation using 3D face models to synthesize faces has been demonstrated to be effective for face recognition. However, the model directly trained by using the synthesized faces together with the original real faces is not optimal. In this paper, we propose a novel approach that uses a deep transfer network (DTN) with 3D morphable models (3DMMs) for face recognition to overcome the shortage of labeled face images and the dataset bias between synthesized images and corresponding real images. We first utilize the 3DMM to synthesize faces with various poses to augment the training dataset. Then, we train a deep neural network using the synthesized face images and the original real face images. The results obtained on LFW show that the accuracy of the model utilizing synthesized data only is lower than that of the model using the original data, although the synthesized dataset contains much considerably images with more unconstrained poses. This result shows that a dataset bias exists between the synthesized faces and the real faces. We treat the synthesized faces as the source domain, and we treat the actual faces as the target domain. We use the DTN to alleviate the discrepancy between the source domain and the target domain. The DTN attempts to project source domain samples and target domain samples to a new space where they are fused together such that one cannot distinguish the domain from which a specific image is from. We optimize our DTN based on the maximum mean discrepancy (MMD) of the shared feature extraction layers and the discrimination layers. We choose AlexNet and Inception-ResNet-V1 as our benchmark models. The proposed method is also evaluated on the LFW and SLLFW databases. The experimental results show that our method can effectively address the domain discrepancy. Moreover, the dataset bias between the synthesized data and the real data is remarkably reduced, which can thus improve the performance of the convolutional neural network (CNN) model. Zhanfu An, Weihong Deng, Tongtong Yuan, Jiani Hu |
FG | 2 |
| 2018 | Cross-Generating GAN for Facial Identity PreservingabstractThe large variations of pose and illumination have been the great challenges to face recognition for many years. Because of these variations, many classical recognition methods fail to work. The key to solve this problem is to extract identity feature from face images. In recent years, people have been concentrating on synthesizing rotated faces, however, neglected the form of facial identity representation. In this paper, we propose Cross-generating Generative Adversarial Network (CG-GAN) to generate rotated faces while extracting discriminative identity. CG-GAN is allowed to learn a network to exchange poses and illuminations of two different subjects' picture. Within the network, each input image is resolved into a variation code and a identity code at the representation layer; then these codes are randomly combined for generating corresponding pictures. Not only does CG-GAN synthesis vivid face under desired pose from one picture, but also the represention layer is very suitable for face recognition task. We train and test CG-GAN on the Multi-PIE dataset and achieve state-of-the-art results. Weilong Chai, Weihong Deng, Haifeng Shen |
FG | 2 |
| 2018 | Deep Unsupervised Domain Adaptation for Face RecognitionabstractFace recognition is challenge task which involves determining the identity of facial images. With availability of a massive amount of labeled facial images gathered from Internet, deep convolution neural networks(DCNNs) have achieved great success in face recognition tasks. Those images are gathered from unconstrain environment, which contain people with different ethnicity, age, gender and so on. However, in the actual application scenario, the target face database may be gathered under different conditions compared with source training dataset, e.g. different ethnicity, different age distribution, disparate shooting environment. These factors increase domain discrepancy between source training database and target application database which makes the learnt model degenerate in target database. Meanwhile, for the target database where labeled data are lacking or unavailable, directly using target data to fine-tune pre-learnt model becomes intractable and impractical. In this paper, we adopt unsupervised transfer learning methods to address this issue. To alleviate the discrepancy between source and target face database and ensure the generalization ability of the model, we constrain the maximum mean discrepancy (MMD) between source database and target database and utilize the massive amount of labeled facial images of source database to training the deep neural network at the same time. We evaluate our method on two face recognition benchmarks and significantly enhance the performance without utilizing the target label. Zimeng Luo, Jiani Hu, Weihong Deng, Haifeng Shen |
FG | 3 |
| 2018 | Task Specific Networks for Identity and Face VariationabstractPose and illumination variations are considered as two main challenges that face recognition system encounters. Most existing methods perform face normalization, aiming at untangling identity representation from these variations to improve recognition accuracy. Taking into account face variation representations, this paper proposes Task Specific Networks for the two representations with two novelties. First, we rotate and normalize face image to multi-pose view for one subtask, and learn face variation representations for another. Second, we learn face variation representations in an unsupervised way, which is more robust and more universal. We couple these two representations in the part of reconstructing the original face, where the two representations effect and restrict each other. Extensive experiments demonstrate the superiority of our method in both learning representations and rotating non-frontal face image. Yichen Qian, Weihong Deng, Jiani Hu |
FG | 2 |
| 2018 | Deep Emotion Transfer Network for Cross-database Facial Expression RecognitionabstractDue to the large domain discrepancy between the training and testing data and the inaccessibility of annotating sufficient training samples, cross-database facial expression recognition which has more application value remains to be challenging in the literature. Previous researches on this problem are based on shallow features with limited discrimination ability. In this paper, we propose to address this problem with a Deep Emo-transfer Network (DETN). Specifically, maximum mean discrepancy was embedded in the deep architecture to reduce dataset bias. Furthermore, a very common but widely ignored bottleneck in facial expression, imbalanced class distribution, has been taken into account. A learnable class-wise weighting parameter was introduced to our network by exploring class prior distribution on unlabeled data so that the training and testing domains can share similar class distribution. Extensive empirical evidences involving both lab-controlled vs. real-world and small-scale vs. large-scale facial expression databases show that our DETN can yield competitive performances across various facial expression transfer tasks. Shan Li 0001, Weihong Deng |
ICPR | 2 |
| 2018 | Local Subclass Constraint for Facial Expression Recognition in the WildabstractThe Automated Facial Expression Recognition (FER) in the wild is still a challenge problem. Currently, most of Deep Convolutional Neural Networks (DCNNs) based FER methods adopt softmax cross-entropy loss to encourage the separability of inter-class features. Many deep embedding approaches (e.g. contrastive loss, triplet loss, center loss) have been extended to the field of FER to enhance the discriminative ability of deep expression features and obtain the predictive effect. In this work, we present a novel deep embedding approach explicitly designed to respect the huge intra-class variation of expression features while learning discriminative expression features. We aim at forming a locally compact representation space structure through minimizing the distance between samples and their nearest subclass center. We demonstrate the effectiveness of this idea on RAF (Real-world Affective Faces) database. The experiment results show that our approaches can not only improve the classification performance but also adaptively learn a locally compact and expression intensity-aware feature space structure. We further extend our models to Static Facial Expressions in the Wild (SFEW) dataset and the results show the generalized ability of our approaches. Zimeng Luo, Jiani Hu, Weihong Deng |
ICPR | 3 |
| 2018 | Deep Difference Analysis in Similar-looking Face recognitionabstractDeep convolutional neural networks (DCNNs) have recently demonstrated impressive performance in face recognition. However, there is no clear understanding of what difference they find between two similar-looking faces. In this paper, we propose a visualization method that gives insight into difference of similar-looking faces found by DCNNs. This method, used as an assistant role, could help human to identify people who try to invade the biometric system using a similar-looking face. We design a crowdsourcing task to evaluate our method. With assistance of our method, accuracy of participants is greatly increased by 8%, which is also better than the accuracy of network, while participants get little improvement with assistance of Deconvolutional network or Gradient Back-propagation. The experiment result suggests that our method makes a difference in human-machine cooperation. Yaoyao Zhong, Weihong Deng |
ICPR | 2 |
| 2018 | Virtual Class Enhanced Discriminative Embedding LearningabstractRecently, learning discriminative features to improve the recognition performances gradually becomes the primary goal of deep learning, and numerous remarkable works have emerged. In this paper, we propose a novel yet extremely simple method Virtual Softmax to enhance the discriminative property of learned features by injecting a dynamic virtual negative class into the original softmax. Injecting virtual class aims to enlarge inter-class margin and compress intra-class distribution by strengthening the decision boundary constraint. Although it seems weird to optimize with this additional virtual class, we show that our method derives from an intuitive and clear motivation, and it indeed encourages the features to be more compact and separable. This paper empirically and experimentally demonstrates the superiority of Virtual Softmax, improving the performances on a variety of object classification and face verification tasks. Binghui Chen, Weihong Deng, Haifeng Shen |
NeurIPS | 2 |
| 2018 | Deep Local Descriptors with Domain Adaptation
Shuwen Qiu, Weihong Deng |
PRCV (2) | 2 |
| 2018 | Facial landmark localization by enhanced convolutional neural network
Weihong Deng, Yuke Fang, Zhenqi Xu, Jiani Hu |
Neurocomputing | 1 |
| 2018 | Deep visual domain adaptation: A survey
Mei Wang 0001, Weihong Deng |
Neurocomputing | 2 |
| 2018 | Face Recognition via Collaborative Representation: Its Discriminant Nature and Superposed RepresentationabstractCollaborative representation methods, such as sparse subspace clustering (SSC) and sparse representation-based classification (SRC), have achieved great success in face clustering and classification by directly utilizing the training images as the dictionary bases. In this paper, we reveal that the superior performance of collaborative representation relies heavily on the sufficiently large class separability of the controlled face datasets such as Extended Yale B. On the uncontrolled or undersampled dataset, however, collaborative representation suffers from the misleading coefficients of the incorrect classes. To address this limitation, inspired by the success of linear discriminant analysis (LDA), we develop a superposed linear representation classifier (SLRC) to cast the recognition problem by representing the test image in term of a superposition of the class centroids and the shared intra-class differences. In spite of its simplicity and approximation, the SLRC largely improves the generalization ability of collaborative representation, and competes well with more sophisticated dictionary learning techniques, on the experiments of AR and FRGC databases. Enforced with the sparsity constraint, SLRC achieves the state-of-the-art performance on FERET database using single sample per person. Weihong Deng, Jiani Hu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | From one to many: Pose-Aware Metric Learning for single-sample face recognition
Weihong Deng, Jiani Hu, Zhongjun Wu, Jun Guo 0002 |
Pattern Recognit. | 1 |
| 2018 | Generative Model With Coordinate Metric Learning for Object Recognition Based on 3D ModelsabstractOne of the bottlenecks in acquiring a perfect database for deep learning is the tedious process of collecting and labeling data. In this paper, we propose a generative model trained with synthetic images rendered from 3D models which can reduce the burden on collecting real training data and make the background conditions more realistic. Our architecture is composed of two sub-networks: a semantic foreground object reconstruction network based on Bayesian inference and a classification network based on multi-triplet cost training for avoiding overfitting on the monotone synthetic object surface and utilizing accurate information of synthetic images like object poses and lighting conditions which are helpful for recognizing regular photos. First, our generative model with metric learning utilizes additional foreground object channels generated from semantic foreground object reconstruction sub-network for recognizing the original input images. Multi-triplet cost function based on poses is used for metric learning which makes it possible to train an effective categorical classifier purely based on synthetic data. Second, we design a coordinate training strategy with the help of adaptive noise applied on the inputs of both of the concatenated sub-networks to make them benefit from each other and avoid inharmonious parameter tuning due to different convergence speeds of two sub-networks. Our architecture achieves the state-of-the-art accuracy of 50.5% on the ShapeNet database with data migration obstacle from synthetic images to real images. This pipeline makes it applicable to do recognition on real images only based on 3D models. Our codes are available at https://github.com/wangyida/gm-cml. Yida Wang 0001, Weihong Deng |
IEEE Trans. Image Process. | 2 |
| 2017 | Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the WildabstractPast research on facial expressions have used relatively limited datasets, which makes it unclear whether current methods can be employed in real world. In this paper, we present a novel database, RAF-DB, which contains about 30000 facial images from thousands of individuals. Each image has been individually labeled about 40 times, then EM algorithm was used to filter out unreliable labels. Crowdsourcing reveals that real-world faces often express compound emotions, or even mixture ones. For all we know, RAF-DB is the first database that contains compound expressions in the wild. Our cross-database study shows that the action units of basic emotions in RAF-DB are much more diverse than, or even deviate from, those of lab-controlled ones. To address this problem, we propose a new DLP-CNN (Deep Locality-Preserving CNN) method, which aims to enhance the discriminative power of deep features by preserving the locality closeness while maximizing the inter-class scatters. The benchmark experiments on the 7-class basic expressions and 11-class compound expressions, as well as the additional experiments on SFEW and CK+ databases, show that the proposed DLP-CNN outperforms the state-of-the-art handcrafted features and deep learning based methods for the expression recognition in the wild. Shan Li 0001, Weihong Deng, Junping Du 0001 |
CVPR | 2 |
| 2017 | Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation
Binghui Chen, Weihong Deng, Junping Du 0001 |
CVPR | 2 |
| 2017 | Attention-Based Template Adaptation for Face VerificationabstractIn this paper, we propose an Attention-Based Template Adaptation (termed as ABTA) algorithm for face recognition in the unconstrained environment. This ABTA algorithm can be divided into two modules, which consist of an attention-based neural network (feature extractor module) to integrate the template features of various lengths to a single fixed length feature representation according to the attention mechanism, and a template adaptation module (transfer module) which is used to transfer the knowledge of a hold-out dataset to the test templates to improve the performance via transfer learning. The feature extractor module is invariant to the order of the images and videos and can save both memory and computation resources due to its compactness. As for the transfer module, we apply the one-shot similarity to get the scores between the test template pairs, which demonstrates its power in recent research. Our method produces results comparable to the state-of-the-art in the challenging face dataset, IJB-A. Zhanfu An, Weihong Deng |
FG | 4 |
| 2017 | Metric-Promoted Siamese Network for Gender ClassificationabstractGender classification is a fundamental and important application in computer vision, and it has become a research hotspot. Real-world applications require gender classification in unconstrained conditions where traditional methods are not appropriate. This paper proposes a Deep Convolutional Neural Network for feature extraction together with fully-connected layers for metric learning. A Siamese network is built for similarity measuring to promote the performance of classification. Extensive experiments on several databases demonstrate that a significant improvement can be obtained for gender classification tasks in both constrained and unconstrained conditions. Yipeng Huang 0002, Shuying Liu, Jiani Hu, Weihong Deng |
FG | 4 |
| 2017 | Learning Local Responses of Facial Landmarks with Conditional Variational Auto-Encoder for Face AlignmentabstractThis work proposes a novel convolutional neural network architecture which can locate landmarks accurately by learning local responses of facial landmarks. The network consists of a Conditional Variational Auto-Encoder(CVAE) and a Deep Convolutional Neural Network(DCNN). The CVAE is used to learn the response maps of facial landmarks from face images and the DCNN is used to learn accurate landmark locations from the response maps and facial textures. The CVAE consists of a face encoder, which extracts high-level information from raw pixels, and a decoder which outputs local response maps from high-level coding. We derive the CVAE used for catching local responses as an optimization problem, which can be solved through back-propagation. Extensive experiments show that the proposed CVAE can learn better local response maps than Fully Convolutional Network(FCN). Our method outperforms state-of-the-art methods on AFLW(5 points) and the challenging subset of 300-W(68 points), which means our method shows advantages in the condition of complex poses and expressions. Shuying Liu, Yipeng Huang 0002, Jiani Hu, Weihong Deng |
FG | 4 |
| 2017 | Boosting-POOF: Boosting Part Based One vs One Feature for Facial Expression Recognition in the WildabstractRecent years, facial expression recognition has remained a challenging and interesting problem, especially for faces in the real world. Most of the traditional approaches are based on Action Units (AUs) detection or low-level features (e.g. LBP, HOG, SIFT and Gabor). Thus, when recognizing real-world facial expressions, these methods might result in poor performance. In this paper, we propose an automatic framework called `Boosting-POOF' to extract discriminative Mid-Level features using low-level features extracted from local face regions. Rather than cascade local features altogether, we adopt class-pairwise Mid-Level descriptors for each local region to extract Mid-Level features and Adaboost feature selection to choose more discriminative features. In experiments, four facial expression benchmarks (CK+, SFEW, RAF-BASIC, RAF-COMPOUND) are evaluated. The `Boosting-POOF' achieves state-of-the-art performance compared with recent approaches. What' more, the `Boosting-POOF' can automatically provide the most significant difference between two expression categories, which is more useful than AUs detection for real world images. Shan Li 0001, Weihong Deng |
FG | 3 |
| 2017 | Deep transfer network for face recognition using 3D synthesized faceabstractFace recognition has experienced a flurry of advances with deep learning. However, training a model requires a lot of data. In order to meet this condition, some researchers use the 3D rendering technique to synthesize fake face images to expand the training data. Experimental results have demonstrated that this method is an effective way. There exist, however, dataset bias between the real 2D real face images and 3D synthesized face images. In this paper, we use Deep Transfer Network(DTN) to reduce dataset bias. First, we utilize the 3DMM face model to synthesize face images with various poses and natural expression. We choose the Inception-Resnet-V1 as our benchmark model. Then, we optimize our DTN based on maximum mean discrepancy(MMD) of the shared feature extraction layers and the discrimination layers. Our experiments demonstrate that the model jointly trained using synthesized images and real images is more robust than using either dataset (2D real faces or 3D synthesized faces). Furthermore, the performance obtained by our approach is comparable to the-state-of-the-art results to the systems trained on millions of real images. Zhanfu An, Weihong Deng, Jiani Hu |
VCIP | 2 |
| 2017 | Supervised hashing with extreme learning machineabstractSupervised hashing methods, which aim to generate semantic similarity-preserving binary codes, have been proposed to improve the performance of large-scale image retrieval. However, learning binary codes remains an NP-hard problem due to the binary constraints and complex computation. Existing hashing methods have never explored the potentiality of the label information, leading to a limited performance. To address these problems, we propose a simple supervised hashing method based on extreme learning machine (ELM). And we generate the supervised information in ELM by target code learning instead of using the traditional label code to fit the retrieval problem. With this modified label code, our method can produce high-quality binary codes and obtain high retrieval precision. Comprehensive experiments have shown our superiority to other state-of-the-art methods. Tongtong Yuan, Weihong Deng, Jiani Hu |
VCIP | 2 |
| 2017 | Deep probabilities for age estimationabstractHuman age can provide important demographic information. In this paper, we tackle the estimation of age in face images with probabilities. The design of the proposed method is based on the relative order of age labels in the database. The age estimation problem is transformed into a series of binary classifications achieved by convolution neural network. Each classifier is used to judge whether the age of input image is larger than a certain age and the estimated age is obtained by adding probability values of these classification problems. The proposed method: Deep Probabilities (DP) of facial age shows improvements over direct regression and multi-classification methods. Tianyue Zheng, Weihong Deng, Jiani Hu |
VCIP | 2 |
| 2017 | Regularization techniques for high-dimensional data analysis
Jiwen Lu, Xi Peng 0001, Weihong Deng, Ajmal Mian |
Image Vis. Comput. | 3 |
| 2017 | Lighting-aware face frontalization for unconstrained face recognition
Weihong Deng, Jiani Hu, Zhongjun Wu, Jun Guo 0002 |
Pattern Recognit. | 1 |
| 2017 | Fine-grained face verification: FGLFW database, baselines, and human-DCMN partnership
Weihong Deng, Jiani Hu, Nanhai Zhang, Binghui Chen, Jun Guo 0002 |
Pattern Recognit. | 1 |
| 2017 | Deep Correlation Feature Learning for Face Verification in the WildabstractConvolutional neural networks (CNNs) commonly uses the softmax loss function as the supervision signal. In order to enhance the discriminative power of the deeply learned features, this letter proposes a new supervision signal, called correlation loss, for face verification task. Specifically, the correlation loss encourages the large correlation between the deep feature vectors and their corresponding weight vectors in softmax loss. With the joint supervision of softmax loss and correlation loss, the deep correlation feature learning (DCFL) network can learn the deep features with both the interclass separability and the intraclass compactness, which are highly discriminative for face verification. More importantly, by applying the weight vector of softmax function as the class prototype, the proposed correlation loss function is easy to be optimized during the backpropatation of CNN. Finally, the DCFL method achieves 99.55% and 96.06% face verification accuracy using a 64-layer ResNet on the labeled face in-the-Wild (LFW) and you-tube face (YTF) benchmark, respectively. Weihong Deng, Binghui Chen, Yuke Fang, Jiani Hu |
IEEE Signal Process. Lett. | 1 |
| 2016 | ZigzagNet: Efficient Deep Learning for Real Object Recognition Based on 3D Models
Yida Wang 0001, Xiuzhuang Zhou, Weihong Deng |
ACCV (4) | 4 |
| 2016 | Illumination-Recovered Pose Normalization for Unconstrained Face Recognition
Zhongjun Wu, Weihong Deng, Zhanfu An |
ACCV (3) | 2 |
| 2016 | Learning Facial Point Response for Alignment by Purely Convolutional Network
Zhenqi Xu, Weihong Deng, Jiani Hu |
ACCV (3) | 2 |
| 2016 | Face Recognition Using a Unified 3D Morphable Model
Guosheng Hu, Fei Yan 0001, Chi-Ho Chan, Weihong Deng, William J. Christmas, Josef Kittler, Neil Robertson 0002 |
ECCV (8) | 4 |
| 2016 | Self-restraint object recognition by model based CNN learningabstractCNN has shown excellent performance on object recognition based on huge amount of real images. For training with synthetic data rendered from 3D models alone to reduce the workload of collecting real images, we propose a concatenated self-restraint learning structure lead by a triplet and softmax jointed loss function for object recognition. Locally connected auto encoder trained from rendered images with and without background used for object reconstruction against environment variables produces an additional channel automatically concatenated to RGB channels as input of classification network. This structure makes it possible training a softmax classifier directly from CNN based on synthetic data with our rendering strategy. Our structure halves the gap between training based on real photos and 3D model in both PASCAL and ImageNet database compared to GoogleNet. Yida Wang 0001, Weihong Deng |
ICIP | 2 |
| 2016 | Weakly-supervised deep self-learning for face recognitionabstractFor recent years, state-of-the-art deep learning systems for face recognition task completely use supervised training. Their performances depend critically on the amount of manually-labeled examples and the correctness of label data. In real life, however, it is very costly and time-consuming to collect and label such database. Therefore, we intend to build a feasible self-learning system, handling the face images which are unlabeled. In this paper, we first build a challenging unlabeled database and propose an efficient Self-Learning DCNN structure (SL-DCNN) to handle weakly-supervised training for face recognition using complicated and unlabeled training data. Our main contribution is that we introduce a novel modification signal as an ingenious supervision to distinguish the misclassifications, and to correctly reduce intra-class variations and enlarge inter-class distances in combination with identification-verification. Then, we investigate the method of feature merging and whether rich identity improves feature learning under unlabeled data. Finally, 97.47% face verification accuracy on LFW [1] is impressively achieved by our method, which is much higher than state-of-the-art methods under noise. Binghui Chen, Weihong Deng |
ICME | 2 |
| 2016 | One-shot deep neural network for pose and illumination normalization face recognitionabstractPose and illumination are considered as two main challenges that face recognition system encounters. In this paper, we consider face recognition problem across pose and illumination variations, given small amount of training samples and single sample per gallery (a.k.a., one shot classification). We combine the strength of 3D models in generating multiviews and various illumination samples and the ability of deep learning in learning non-linear transformation, which is very suitable for pose and illumination normalization, by using a multi-task deep neural network. By the pose and illumination augmentation strategy, we train a pose and illumination normalization neural network with much less training data compared to other methods. Experiments on MultiPIE database achieve competitive recognition results, demonstrating the effectiveness of proposed method. Zhongjun Wu, Weihong Deng |
ICME | 2 |
| 2016 | Recurrent convolutional neural network for video classificationabstractVideo classification is more difficult than image classification since additional motion feature between image frames and amount of redundancy in videos should be taken into account. In this work, we proposed a new deep learning architecture called recurrent convolutional neural network (RCNN) which combines convolution operation and recurrent links for video classification tasks. Our architecture can extract the local and dense features from image frames as well as learning the temporal features between consecutive frames. We also explore the effectiveness of sequential sampling and random sampling when training our models, and find out that random sampling is necessary for video classification. The feature maps from our learned model preserve motion from image frames, which is analogous to the persistence of vision in human visual system. We achieved 81.0% classification accuracy without optical flow and 86.3% with optical flow on the UCF-101 dataset, both are competitive to the state-of-the-art methods. Zhenqi Xu, Jiani Hu, Weihong Deng |
ICME | 3 |
| 2016 | Geometry-aware metric learning for similar face recognitionabstractNoticing that face images (from different persons) with high similarity computed by current state-of-the-art methods may be not visually similar, in this paper, we present a new verification problem on judging whether the given faces are similar or not. Similar to “view 2” of Labeled Faces in the Wild (LFW), we construct ten subsets' face pairs using images from LFW. Label of each pair comes from human annotation results. Since similar faces are not from the same person after all, pushing similar faces too close will easily contribute to wrong models. Therefore, we propose a new geometry-aware metric learning (GAML) method which can preserve the similarity of similar faces while enlarge the difference between dissimilar faces. Experimental results show that our method outperforms traditional face verification methods on our similar face dataset. Nanhai Zhang, Jiajie Han, Jiani Hu, Weihong Deng |
ICME | 4 |
| 2015 | Multi-manifold deep metric learning for image set classificationabstractIn this paper, we propose a multi-manifold deep metric learning (MMDML) method for image set classification, which aims to recognize an object of interest from a set of image instances captured from varying viewpoints or under varying illuminations. Motivated by the fact that manifold can be effectively used to model the nonlinearity of samples in each image set and deep learning has demonstrated superb capability to model the nonlinearity of samples, we propose a MMDML method to learn multiple sets of nonlinear transformations, one set for each object class, to nonlinearly map multiple sets of image instances into a shared feature subspace, under which the manifold margin of different class is maximized, so that both discriminative and class-specific information can be exploited, simultaneously. Our method achieves the state-of-the-art performance on five widely used datasets. Jiwen Lu, Gang Wang 0012, Weihong Deng, Pierre Moulin, Jie Zhou 0001 |
CVPR | 3 |
| 2015 | DeepEmo: Real-world facial expression analysis via deep learningabstractRecent automatic facial expression recognition research has focused on optimizing performance on a few databases that were collected under controlled pose and lighting conditions, and has produced nearly perfect accuracy. This paper explores the necessary characteristics of the training dataset, feature representations and machine learning algorithms for a system that operates reliably in more realistic conditions. A new database, Real-world Affective Face Database (RAF-DB), is presented which contains about 30,000 greatly-diverse facial images from social networks. Crowdsourcing results suggest that real-world expression recognition problem is a typical imbalanced multi-label classification problem, and the balanced, single-label datasets currently used in the literature could potentially lead research into misleading algorithmic solutions. A deep learning architecture, DeepEmo, is proposed to address the real-world challenge of emotion recognition by learning the highlevel feature representations which are highly effective for discriminating realistic facial expressions. Extensive experimental results show that the deep learning method is significantly superior to handcrafted features, and with the near-frontal pose constraint, human-level recognition accuracy is achievable. Weihong Deng, Jiani Hu, Jun Guo 0002 |
VCIP | 1 |
| 2015 | Face recognition based on random featureabstractThis paper presents a simple, yet very efficient facial image representation based on random feature. We describe the face on three different levels of locality. Firstly, random features are extracted from local image patches with random projection and then we use BoW model to get labels on a pixel-level by coding the random features to the closest textons. Secondly, the labels are statistically calculated over a region to get histogram containing information on a regional level. Thirdly, the regional histograms are concatenated to build a global description of the face. Experiments conducted on the FERET database show that our approach has outstanding robustness to variations, especially to noise. Shasha Li 0001, Weihong Deng |
VCIP | 2 |
| 2015 | Reconstruction-Based Metric Learning for Unconstrained Face VerificationabstractIn this paper, we propose a reconstruction-based metric learning method to learn a discriminative distance metric for unconstrained face verification. Unlike conventional metric learning methods, which only consider the label information of training samples and ignore the reconstruction residual information in the learning procedure, we apply a reconstruction criterion to learn a discriminative distance metric. For each training example, the distance metric is learned by enforcing a margin between the interclass sparse reconstruction residual and interclass sparse reconstruction residual, so that the reconstruction residual of training samples can be effectively exploited to compute the between-class and within-class variations. To better use multiple features for distance metric learning, we propose a reconstruction-based multimetric learning method to collaboratively learn multiple distance metrics, one for each feature descriptor, to remove uncorrelated information for recognition. We evaluate our proposed methods on the Labelled Faces in the Wild (LFW) and YouTube face data sets and our experimental results clearly show the superiority of our methods over both previous metric learning methods and several state-of-the-art unconstrained face verification methods. Jiwen Lu, Gang Wang 0012, Weihong Deng, Kui Jia |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Transformed Principal Gradient Orientation for Robust and Precise Batch Face Alignment
Weihong Deng, Jiani Hu, Jun Guo 0002 |
ACCV (4) | 1 |
| 2014 | Linear Ranking AnalysisabstractWe extend the classical linear discriminant analysis (LDA) technique to linear ranking analysis (LRA), by considering the ranking order of classes centroids on the projected subspace. Under the constrain on the ranking order of the classes, two criteria are proposed: 1) minimization of the classification error with the assumption that each class is homogenous Guassian distributed, 2) maximization of the sum (average) of the K minimum distances of all neighboring-class (centroid) pairs. Both criteria can be efficiently solved by the convex optimization for one-dimensional subspace. Greedy algorithm is applied to extend the results to the multi-dimensional subspace. Experimental results show that 1) LRA with both criteria achieve state-of-the-art performance on the tasks of ranking learning and zero-shot learning, and 2) the maximum margin criterion provides a discriminative subspace selection method, which can significantly remedy the class separation problem in comparing with several representative extensions of LDA. Weihong Deng, Jiani Hu, Jun Guo 0002 |
CVPR | 1 |
| 2014 | Simultaneous Feature and Dictionary Learning for Image Set Based Face Recognition
Jiwen Lu, Gang Wang 0012, Weihong Deng, Pierre Moulin |
ECCV (1) | 3 |
| 2014 | Precise eye localization by fast local linear SVMabstractRecently, discriminative methods such as SVM has been widely used in object location. But there has been no method to perform well enough both at accuracy and speed. For linear SVM, it is hard to separate the nonlinear samples exactly. For kernel SVM, it is hard to be applied to real-time application, because of the computational cost kernel function. Local linear SVM has been proved to be a good tradeoff between fast linear SVM and qualitative best kernel methods. However, it is still time-consuming for real-time application. To design a high efficiency and high precision eye locahzer, first, we deduce a fast variation for LL-S VM which can serve as a more fast and accurate substitute of the traditional nonlinear kernel SVM. Second, to further improve the speed, we also adopt a candidate selection strategy. Extensive experiments on the BioID, FERET, FRGC, and LFW database show that our proposed method achieves favorable localization accuracy against other state-of-the-art methods at a speed as fast as 5ms to localize two eyes. Xiang Sun 0003, Jiani Hu, Weihong Deng |
ICME | 4 |
| 2014 | Online Regression of Grandmother-Cell Responses with Visual Experience Learning for Face RecognitionabstractGrandmother cell is a term in neuroscience to imitate the simplistic notion that the brain has a separate neuron to represent every familiar face, with important properties of sparseness and invariance. This paper proposes a linear regression based classification model for face recognition, which learn a mapping from the training feature vectors to the grandmother-cell-like codes, with one unit corresponding to an individual. Two kinds of visual experiences are incorporated to enhance the generalization capability of the regression mapping. First, the regression model maps the intra-personal facial differences of the unknown faces to the zeros vectors, so that any similar variation on the familiar face would not affect the regression result. Second, to adapt to the evolution of facial appearance, the model feeds the selected testing images back to incrementally retrain the regression mapping, and decrement ally remove the influence of outdated training images, all in an unsupervised manner. Experiments results on Extended Yale B, FERET, and AR databases demonstrate the efficacy of the proposed regression based face recognition algorithms. Jiani Hu, Weihong Deng, Jun Guo 0002 |
ICPR | 2 |
| 2014 | Max-K-Min Distance Analysis for Dimension ReductionabstractWe propose a new criterion for discriminative dimension reduction, Max-K-Min Distance Analysis (MKMDA). Given a data set with C classes, MKMDA maximizes the sum of the K minimum pair wise distance of these C classes on the selected one-dimensional subspace. The set of the possible one-dimensional subspace, for which the order of the projected class centroids is identical, define a convex region with associated convex sum of K smallest margin functions. This allows for the maximization of the margin function using standard convex optimization algorithms. This result is further extended to obtain the d-dimensional subspace for any given d by iterative applying our algorithm to the null space of the (d -- 1)-dimensional subspace. The effectiveness of the proposed criterion and corresponding algorithm is shown by the visualization and classification experiments on both synthetic data and real data sets. Jiani Hu, Weihong Deng, Jun Guo 0002 |
ICPR | 2 |
| 2014 | A high-performance training-free approach for hand gesture recognition with accelerometer
Mingzhi Dong, Ying Duan, Weihong Deng, Kaili Zhao, Jun Guo 0002 |
Multim. Tools Appl. | 4 |
| 2014 | Transform-Invariant PCA: A Unified Approach to Fully Automatic FaceAlignment, Representation, and RecognitionabstractWe develop a transform-invariant PCA (TIPCA) technique which aims to accurately characterize the intrinsic structures of the human face that are invariant to the in-plane transformations of the training images. Specially, TIPCA alternately aligns the image ensemble and creates the optimal eigenspace, with the objective to minimize the mean square error between the aligned images and their reconstructions. The learning from the FERET facial image ensemble of 1,196 subjects validates the mutual promotion between image alignment and eigenspace representation, which eventually leads to the optimized coding and recognition performance that surpasses the handcrafted alignment based on facial landmarks. Experimental results also suggest that state-of-the-art invariant descriptors, such as local binary pattern (LBP), histogram of oriented gradient (HOG), and Gabor energy filter (GEF), and classification methods, such as sparse representation based classification (SRC) and support vector machine (SVM), can benefit from using the TIPCA-aligned faces, instead of the manually eye-aligned faces that are widely regarded as the ground-truth alignment. Favorable accuracies against the state-of-the-art results on face coding and face recognition are reported. Weihong Deng, Jiani Hu, Jiwen Lu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Equidistant prototypes embedding for single sample based face recognition with generic learning and incremental learning
Weihong Deng, Jiani Hu, Xiuzhuang Zhou, Jun Guo 0002 |
Pattern Recognit. | 1 |
| 2014 | Discriminative Multimetric Learning for Kinship VerificationabstractIn this paper, we propose a new discriminative multimetric learning method for kinship verification via facial image analysis. Given each face image, we first extract multiple features using different face descriptors to characterize face images from different aspects because different feature descriptors can provide complementary information. Then, we jointly learn multiple distance metrics with these extracted multiple features under which the probability of a pair of face image with a kinship relation having a smaller distance than that of the pair without a kinship relation is maximized, and the correlation of different features of the same face sample is maximized, simultaneously, so that complementary and discriminative information is exploited for verification. Experimental results on four face kinship data sets show the effectiveness of our proposed method over the existing single-metric and multimetric learning methods. Haibin Yan, Jiwen Lu, Weihong Deng, Xiuzhuang Zhou |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2013 | A Maximum K-Min Approach for Classification
Mingzhi Dong, Weihong Deng, Jun Guo 0002, Honggang Zhang 0002 |
AAAI | 3 |
| 2013 | In Defense of Sparsity Based Face RecognitionabstractThe success of sparse representation based classification (SRC) has largely boosted the research of sparsity based face recognition in recent years. A prevailing view is that the sparsity based face recognition performs well only when the training images have been carefully controlled and the number of samples per class is sufficiently large. This paper challenges the prevailing view by proposing a ``prototype plus variation'' representation model for sparsity based face recognition. Based on the new model, a Superposed SRC (SSRC), in which the dictionary is assembled by the class centroids and the sample-to-centroid differences, leads to a substantial improvement on SRC. The experiments results on AR, FERET and FRGC databases validate that, if the proposed prototype plus variation representation model is applied, sparse coding plays a crucial role in face recognition, and performs well even when the dictionary bases are collected under uncontrolled conditions and only a single sample per classes is available. Weihong Deng, Jiani Hu, Jun Guo 0002 |
CVPR | 1 |
| 2012 | Integrative labeling based statistical color models with application to skin detectionabstractTo alleviate the workload of labeling before estimating certain color distributions, integrative labeling is introduced, which merely needs to figure out whether a picture contains positive-class regions or not and then all pixels of the picture are treated as positive or negative class training samples. Integrative labeling, however, results in heavy mixture of training samples. Thus traditional generative density estimation methods can't be used directly in that they perform poorly with heavily polluted training samples. In this paper, by utilizing the prior knowledge of high separability between positive and negative class color distributions, a discriminative learning based GMM(DiscGMM) is proposed for integrative labeling. Besides generating the polluted positive-class samples with comparatively high probability, optimal parameters found by DiscGMM also enjoy a comparatively low probability of generating negative-class samples. The parameter learning problem is solved by a modified Expectation Maximization (EM) algorithm. In an integrative labeling experiment of skin detection, DiscGMM is testified to enjoy much better performance than generative density estimation methods and shows qualified results. Mingzhi Dong, Jun Guo 0002, Weihong Deng, Weiran Xu |
ICIP | 4 |
| 2012 | A Linear Max K-min classifier
Mingzhi Dong, Weihong Deng, Qiang Wang 0048, Caixia Yuan, Jun Guo 0002, Liwei Ma |
ICPR | 3 |
| 2012 | Extended SRC: Undersampled Face Recognition via Intraclass Variant DictionaryabstractSparse Representation-Based Classification (SRC) is a face recognition breakthrough in recent years which has successfully addressed the recognition problem with sufficient training images of each gallery subject. In this paper, we extend SRC to applications where there are very few, or even a single, training images per subject. Assuming that the intraclass variations of one subject can be approximated by a sparse linear combination of those of other subjects, Extended Sparse Representation-Based Classifier (ESRC) applies an auxiliary intraclass variant dictionary to represent the possible variation between the training and testing images. The dictionary atoms typically represent intraclass sample differences computed from either the gallery faces themselves or the generic faces that are outside the gallery. Experimental results on the AR and FERET databases show that ESRC has better generalization ability than SRC for undersampled face recognition under variable expressions, illuminations, disguises, and ages. The superior results of ESRC suggest that if the dictionary is properly constructed, SRC algorithms can generalize well to the large-scale face recognition problem, even with a single training image per class. Weihong Deng, Jiani Hu, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | The small sample size problem of ICA: A comparative study and analysis
Weihong Deng, Yebin Liu, Jiani Hu, Jun Guo 0002 |
Pattern Recognit. | 1 |
| 2010 | Locality preserving and global discriminant projection with prior information
Honggang Zhang 0002, Weihong Deng, Jun Guo 0002, Jie Yang 0001 |
Mach. Vis. Appl. | 2 |
| 2010 | Robust, accurate and efficient face recognition from a single training image: A uniform pursuit approach
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng |
Pattern Recognit. | 1 |
| 2010 | Emulating biological strategies for uncontrolled face recognition
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng |
Pattern Recognit. | 1 |
| 2009 | Semi-supervised Learning Based on Label Propagation through Submanifold
Jiani Hu, Weihong Deng, Jun Guo 0002 |
ISNN (1) | 2 |
| 2009 | Learning a locality discriminating projection for classification
Jiani Hu, Weihong Deng, Jun Guo 0002, Weiran Xu |
Knowl. Based Syst. | 2 |
| 2008 | Handwritten Chinese character recognition using Local Discriminant Projection with Prior InformationabstractIn this paper, we propose a new method to model the manifold of handwritten Chinese characters using the local discriminant projection. We utilize a cascade framework that combines global similarity with local discriminative cues to recognize Chinese characters. We find the similarity of different characters using a nearest-neighbor (NN) classifier, and followed by the Local Discriminant Projection with Prior Information (LDPPI) to map similar characters within a cluster to a low-dimensional space. We evaluate the proposed method on two large public datasets, ETL9B which contains 607,200 handwritten characters from 200 people, and HCL2000 which contains 3,755,000 characters written by 1,000 people. The experimental results demonstrate that the proposed method achieves 0.74% error rate on ETL9B database and 1.88% on HCL2000 database. Honggang Zhang 0002, Jie Yang 0001, Weihong Deng, Jun Guo 0002 |
ICPR | 3 |
| 2008 | Comments on "Globally Maximizing, Locally Minimizing: Unsupervised Discriminant Projection with Application to Face and Palm Biometrics"abstractIn [1], UDP is proposed to address the limitation of LPP for the clustering and classification tasks. In this communication, we show that the basic ideas of UDP and LPP are identical. In particular, UDP is just a simplified version of LPP on the assumption that the local density is uniform. Weihong Deng, Jiani Hu, Jun Guo 0002, Honggang Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Locality discriminating indexing for document classificationabstractThis paper introduces a locality discriminating indexing (LDI) algorithm for document classification. Based on the hypothesis that samples from different classes reside in class-specific manifold structures, LDI seeks for a projection which best preserves the within-class local structures while suppresses the between-class overlap. Comparative experiments show that the proposed method isable to derives compact discriminating document representations for classification. Jiani Hu, Weihong Deng, Jun Guo 0002, Weiran Xu |
SIGIR | 2 |