EDBT 2026 Demo / reviewers in the wild / expert
Dihong Gong
dblp:136/1107
· DBLP profile ↗
26ranked-venue papers
6as first author
7since 2021 · last 2023
0000-0002-6426-6380ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
21 papers |
Face, body and person analysis · 68% Trustworthy machine learning · 6% Deep learning architectures and training · 6% | |
| Network and information security
4 papers |
Security and privacy of machine learning · 37% Hardware security and side channels · 37% Biometric security · 26% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 90% Multimedia analysis and retrieval · 10% |
Topics — the 30 heaviest of 42, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
face recognition |
5.2 | 16 | 2023 | LARNeXt: End-to-End Lie Algebra Residual Network for Face Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 End2End Occluded Face Recognition by Masking Corrupted Features · IEEE Trans. Pattern Anal. Mach. Intell. 2022 LARNet: Lie Algebra Residual Network for Face Recognition · ICML 2021 |
Computer vision › Face, body and person analysis › face recognition › robust face recognition
age-invariant face recognition |
1.3 | 5 | 2019 | Decorrelated Adversarial Learning for Age-Invariant Face Recognition · CVPR 2019 Orthogonal Deep Features Decomposition for Age-Invariant Face Recognition · ECCV (15) 2018 Aging Face Recognition: A Hierarchical Learning Model Based on Local Patterns Selection · IEEE Trans. Image Process. 2016 |
Computer vision › Face, body and person analysis › face recognition › robust face recognition
pose-invariant face recognition |
1.2 | 2 | 2023 | LARNeXt: End-to-End Lie Algebra Residual Network for Face Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023 LARNet: Lie Algebra Residual Network for Face Recognition · ICML 2021 |
Computer vision › Face, body and person analysis › face recognition
occluded face recognition |
1.0 | 2 | 2022 | End2End Occluded Face Recognition by Masking Corrupted Features · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Occlusion Robust Face Recognition Based on Mask Learning With Pairwise Differential Siamese Network · ICCV 2019 |
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
feature disentanglement |
0.7 | 2 | 2019 | Decorrelated Adversarial Learning for Age-Invariant Face Recognition · CVPR 2019 Orthogonal Deep Features Decomposition for Age-Invariant Face Recognition · ECCV (15) 2018 |
Machine learning › Trustworthy machine learning › robustness › adversarial attack
hard-label black-box attack |
0.6 | 1 | 2022 | Triangle Attack: A Query-Efficient Decision-Based Adversarial Attack · ECCV (5) 2022 |
Security and privacy of machine learning
adversarial attack |
0.6 | 1 | 2022 | Triangle Attack: A Query-Efficient Decision-Based Adversarial Attack · ECCV (5) 2022 |
Security and privacy of machine learning › adversarial attack
backdoor attack |
0.6 | 1 | 2022 | Hardly Perceptible Trojan Attack Against Neural Networks with Bit Flips · ECCV (5) 2022 |
Hardware security and side channels › fault attacks › fault injection attack
bit-flip attack |
0.6 | 1 | 2022 | Hardly Perceptible Trojan Attack Against Neural Networks with Bit Flips · ECCV (5) 2022 |
Hardware security and side channels
fault attacks |
0.6 | 1 | 2022 | Hardly Perceptible Trojan Attack Against Neural Networks with Bit Flips · ECCV (5) 2022 |
Machine learning › Transfer learning and domain adaptation
domain gap |
0.5 | 1 | 2021 | SynFace: Face Recognition with Synthetic Data · ICCV 2021 |
Machine learning › Deep learning architectures and training › convolutional neural network
residual network |
0.5 | 1 | 2021 | LARNet: Lie Algebra Residual Network for Face Recognition · ICML 2021 |
Machine learning › Generative modeling
synthetic training data |
0.5 | 1 | 2021 | SynFace: Face Recognition with Synthetic Data · ICCV 2021 |
Image and video processing › super-resolution › image super-resolution
face super-resolution |
0.5 | 1 | 2021 | Learning Spatial Attention for Face Super-Resolution · IEEE Trans. Image Process. 2021 |
Image and video processing › super-resolution
image super-resolution |
0.5 | 1 | 2021 | Learning Spatial Attention for Face Super-Resolution · IEEE Trans. Image Process. 2021 |
Image and video processing › perceptual modeling › visual perception modeling
spatial attention |
0.5 | 1 | 2021 | Learning Spatial Attention for Face Super-Resolution · IEEE Trans. Image Process. 2021 |
Computer vision › Face, body and person analysis › face recognition › cross-spectral face recognition
NIR-VIS face recognition |
0.4 | 2 | 2015 | Learning Compact Feature Descriptor and Adaptive Matching Framework for Face Recognition · IEEE Trans. Image Process. 2015 Common Feature Discriminant Analysis for Matching Infrared Face Images to Optical Face Images · IEEE Trans. Image Process. 2014 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2019 | Occlusion Robust Face Recognition Based on Mask Learning With Pairwise Differential Siamese Network · ICCV 2019 |
Biometric security
anti-spoofing |
0.4 | 1 | 2019 | Face Anti-Spoofing: Model Matters, so Does Data · CVPR 2019 |
Biometric security
face anti-spoofing |
0.4 | 1 | 2019 | Face Anti-Spoofing: Model Matters, so Does Data · CVPR 2019 |
Machine learning › Deep learning architectures and training
loss function design |
0.3 | 1 | 2018 | CosFace: Large Margin Cosine Loss for Deep Face Recognition · CVPR 2018 |
Computer vision › Face, body and person analysis › face recognition › heterogeneous face recognition
cross-modal face matching |
0.3 | 1 | 2017 | Heterogeneous Face Recognition: A Common Encoding Feature Discriminant Approach · IEEE Trans. Image Process. 2017 |
Computer vision › Face, body and person analysis › face recognition
heterogeneous face recognition |
0.3 | 1 | 2017 | Heterogeneous Face Recognition: A Common Encoding Feature Discriminant Approach · IEEE Trans. Image Process. 2017 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.3 | 1 | 2017 | Multimodal Learning for Web Information Extraction · ACM Multimedia 2017 |
Computer vision › 3D vision › local feature descriptor
feature descriptor |
0.2 | 1 | 2015 | A maximum entropy feature descriptor for age invariant face recognition · CVPR 2015 |
Computer vision › Face, body and person analysis › facial age estimation
age estimation |
0.2 | 1 | 2014 | Orthogonal Gaussian Process for Automatic Age Estimation · ACM Multimedia 2014 |
Computer vision › Face, body and person analysis
facial age estimation |
0.2 | 1 | 2014 | Orthogonal Gaussian Process for Automatic Age Estimation · ACM Multimedia 2014 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.2 | 1 | 2014 | Orthogonal Gaussian Process for Automatic Age Estimation · ACM Multimedia 2014 |
Computer vision › Face, body and person analysis
face alignment |
0.1 | 1 | 2021 | Learning Spatial Attention for Face Super-Resolution · IEEE Trans. Image Process. 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2019 | Occlusion Robust Face Recognition Based on Mask Learning With Pairwise Differential Siamese Network · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 1.2canonical correlation analysis · 0.7pose estimation · 0.7attention · 0.7trojan attack · 0.6end-to-end learning · 0.6bit flip · 0.6spatial attention · 0.5residual network · 0.5multi-scale discriminator · 0.5lie algebra · 0.5identity mixup · 0.5gating mechanism · 0.5face synthesis · 0.5domain mixup · 0.5spatio-temporal network · 0.4data synthesis · 0.4semi-supervised learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | LARNeXt: End-to-End Lie Algebra Residual Network for Face RecognitionabstractFace recognition has always been courted in computer vision and is especially amenable to situations with significant variations between frontal and profile faces. Traditional techniques make great strides either by synthesizing frontal faces from sizable datasets or by empirical pose invariant learning. In this paper, we propose a completely integrated embedded end-to-end Lie algebra residual architecture (LARNeXt) to achieve pose robust face recognition. First, we explore how the face rotation in the 3D space affects the deep feature generation process of convolutional neural networks (CNNs), and prove that face rotation in the image space is equivalent to an additive residual component in the feature space of CNNs, which is determined solely by the rotation. Second, on the basis of this theoretical finding, we further design three critical subnets to leverage a soft regression subnet with novel multi-fusion attention feature aggregation for efficient pose estimation, a residual subnet for decoding rotation information from input face images, and a gating subnet to learn rotation magnitude for controlling the strength of the residual component that contributes to the feature learning process. Finally, we conduct a large number of ablation experiments, and our quantitative and visualization results both corroborate the credibility of our theory and corresponding network designs. Our comprehensive experimental evaluations on frontal-profile face datasets, general unconstrained face recognition datasets, and industrial-grade tasks demonstrate that our method consistently outperforms the state-of-the-art ones. Xiaohong Jia 0001, Dihong Gong, Dong-Ming Yan 0001, Zhifeng Li 0001, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hardly Perceptible Trojan Attack Against Neural Networks with Bit Flips
Jiawang Bai, Kuofeng Gao, Dihong Gong, Shutao Xia, Zhifeng Li 0001, Wei Liu 0005 |
ECCV (5) | 3 |
| 2022 | Triangle Attack: A Query-Efficient Decision-Based Adversarial Attack
Xiaosen Wang, Zeliang Zhang 0001, Kangheng Tong, Dihong Gong, Kun He 0001, Zhifeng Li 0001, Wei Liu 0005 |
ECCV (5) | 4 |
| 2022 | End2End Occluded Face Recognition by Masking Corrupted FeaturesabstractWith the recent advancement of deep convolutional neural networks, significant progress has been made in general face recognition. However, the state-of-the-art general face recognition models do not generalize well to occluded face images, which are exactly the common cases in real-world scenarios. The potential reasons are the absences of large-scale occluded face data for training and specific designs for tackling corrupted features brought by occlusions. This article presents a novel face recognition method that is robust to occlusions based on a single end-to-end deep neural network. Our approach, named FROM (Face Recognition with Occlusion Masks), learns to discover the corrupted features from the deep convolutional neural networks, and clean them by the dynamically learned masks. In addition, we construct massive occluded face images to train FROM effectively and efficiently. FROM is simple yet powerful compared to the existing methods that either rely on external detectors to discover the occlusions or employ shallow models which are less discriminative. Experimental results on the LFW, Megaface challenge 1, RMF2, AR dataset and other simulated occluded/masked datasets confirm that FROM dramatically improves the accuracy under occlusions, and generalizes well on general face recognition. Haibo Qiu, Dihong Gong, Zhifeng Li 0001, Wei Liu 0005, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | SynFace: Face Recognition with Synthetic DataabstractWith the recent success of deep neural networks, remarkable progress has been achieved on face recognition. However, collecting large-scale real-world training data for face recognition has turned out to be challenging, especially due to the label noise and privacy issues. Meanwhile, existing face recognition datasets are usually collected from web images, lacking detailed annotations on attributes (e.g., pose and expression), so the influences of different attributes on face recognition have been poorly investigated. In this paper, we address the above-mentioned issues in face recognition using synthetic face images, i.e., SynFace. Specifically, we first explore the performance gap between recent state-of-the-art face recognition models trained with synthetic and real face images. We then analyze the underlying causes behind the performance gap, e.g., the poor intraclass variations and the domain gap between synthetic and real face images. Inspired by this, we devise the SynFace with identity mixup (IM) and domain mixup (DM) to mitigate the above performance gap, demonstrating the great potentials of synthetic data for face recognition. Furthermore, with the controllable face synthesis model, we can easily manage different factors of synthetic face generation, including pose, expression, illumination, the number of identities, and samples per identity. Therefore, we also perform a systematically empirical analysis on synthetic face images to provide some insights on how to effectively utilize synthetic data for face recognition. Haibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li 0001, Wei Liu 0005, Dacheng Tao |
ICCV | 3 |
| 2021 | LARNet: Lie Algebra Residual Network for Face RecognitionabstractFace recognition is an important yet challenging problem in computer vision. A major challenge in practical face recognition applications lies in significant variations between profile and frontal faces. Traditional techniques address this challenge either by synthesizing frontal faces or by pose invariant learning. In this paper, we propose a novel method with Lie algebra theory to explore how face rotation in the 3D space affects the deep feature generation process of convolutional neural networks (CNNs). We prove that face rotation in the image space is equivalent to an additive residual component in the feature space of CNNs, which is determined solely by the rotation. Based on this theoretical finding, we further design a Lie Algebraic Residual Network (LARNet) for tackling pose robust face recognition. Our LARNet consists of a residual subnet for decoding rotation information from input face images, and a gating subnet to learn rotation magnitude for controlling the strength of the residual component contributing to the feature learning process. Comprehensive experimental evaluations on both frontal-profile face datasets and general face recognition datasets convincingly demonstrate that our method consistently outperforms the state-of-the-art ones. Xiaohong Jia 0001, Dihong Gong, Dong-Ming Yan 0001, Zhifeng Li 0001, Wei Liu 0005 |
ICML | 3 |
| 2021 | Learning Spatial Attention for Face Super-ResolutionabstractGeneral image super-resolution techniques have difficulties in recovering detailed face structures when applying to low resolution face images. Recent deep learning based methods tailored for face images have achieved improved performance by jointly trained with additional task such as face parsing and landmark prediction. However, multi-task learning requires extra manually labeled data. Besides, most of the existing works can only generate relatively low resolution face images (e.g., 128×128 ), and their applications are therefore limited. In this paper, we introduce a novel SPatial Attention Residual Network (SPARNet) built on our newly proposed Face Attention Units (FAUs) for face super-resolution. Specifically, we introduce a spatial attention mechanism to the vanilla residual blocks. This enables the convolutional layers to adaptively bootstrap features related to the key face structures and pay less attention to those less feature-rich regions. This makes the training more effective and efficient as the key face structures only account for a very small portion of the face image. Visualization of the attention maps shows that our spatial attention network can capture the key face structures well even for very low resolution faces (e.g., 16×16 ). Quantitative comparisons on various kinds of metrics (including PSNR, SSIM, identity similarity, and landmark detection) demonstrate the superiority of our method over current state-of-the-arts. We further extend SPARNet with multi-scale discriminators, named as SPARNetHD, to produce high resolution results (i.e., 512×512 ). We show that SPARNetHD trained with synthetic data can not only produce high quality and high resolution outputs for synthetically degraded face images, but also show good generalization ability to real world low quality face images. Codes are available at https://github.com/chaofengc/Face-SPARNet. Chaofeng Chen, Dihong Gong, Hao Wang 0050, Zhifeng Li 0001, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 2 |
| 2019 | Decorrelated Adversarial Learning for Age-Invariant Face RecognitionabstractThere has been an increasing research interest in age-invariant face recognition. However, matching faces with big age gaps remains a challenging problem, primarily due to the significant discrepancy of face appearance caused by aging. To reduce such discrepancy, in this paper we present a novel algorithm to remove age-related components from features mixed with both identity and age information. Specifically, we factorize a mixed face feature into two uncorrelated components: identity-dependent component and age-dependent component, where the identity-dependent component contains information that is useful for face recognition. To implement this idea, we propose the Decorrelated Adversarial Learning (DAL) algorithm, where a Canonical Mapping Module (CMM) is introduced to find maximum correlation of the paired features generated by the backbone network, while the backbone network and the factorization module are trained to generate features reducing the correlation. Thus, the proposed model learns the decomposed features of age and identity whose correlation is significantly reduced. Simultaneously, the identity-dependent feature and the age-dependent feature are supervised by ID and age preserving signals respectively to ensure they contain the correct information. Extensive experiments have been conducted on the popular public-domain face aging datasets (FG-NET, MORPH Album 2, and CACD-VS) to demonstrate the effectiveness of the proposed approach. Hao Wang 0050, Dihong Gong, Zhifeng Li 0001, Wei Liu 0005 |
CVPR | 2 |
| 2019 | Face Anti-Spoofing: Model Matters, so Does DataabstractFace anti-spoofing is an important task in full-stack face applications including face detection, verification, and recognition. Previous approaches build models on datasets which do not simulate the real-world data well (e.g., small scale, insignificant variance, etc.). Existing models may rely on auxiliary information, which prevents these anti-spoofing solutions from generalizing well in practice. In this paper, we present a data collection solution along with a data synthesis technique to simulate digital medium-based face spoofing attacks, which can easily help us obtain a large amount of training data well reflecting the real-world scenarios. Through exploiting a novel Spatio-Temporal Anti-Spoof Network (STASN), we are able to push the performance on public face anti-spoofing datasets over state-of-the-art methods by a large margin. Since the proposed model can automatically attend to discriminative regions, it makes analyzing the behaviors of the network possible.We conduct extensive experiments and show that the proposed model can distinguish spoof faces by extracting features from a variety of regions to seek out subtle evidences such as borders, moire patterns, reflection artifacts, etc. Wenhan Luo, Linchao Bao, Yuan Gao 0015, Dihong Gong, Shibao Zheng, Zhifeng Li 0001, Wei Liu 0005 |
CVPR | 5 |
| 2019 | Occlusion Robust Face Recognition Based on Mask Learning With Pairwise Differential Siamese NetworkabstractDeep Convolutional Neural Networks (CNNs) have been pushing the frontier of face recognition over past years. However, existing CNN models are far less accurate when handling partially occluded faces. These general face models generalize poorly for occlusions on variable facial areas. Inspired by the fact that human visual system explicitly ignores the occlusion and only focuses on the non-occluded facial areas, we propose a mask learning strategy to find and discard corrupted feature elements from recognition. A mask dictionary is firstly established by exploiting the differences between the top conv features of occluded and occlusion-free face pairs using innovatively designed pairwise differential siamese network (PDSN). Each item of this dictionary captures the correspondence between occluded facial areas and corrupted feature elements, which is named Feature Discarding Mask (FDM). When dealing with a face image with random partial occlusions, we generate its FDM by combining relevant dictionary items and then multiply it with the original features to eliminate those corrupted feature elements from recognition. Comprehensive experiments on both synthesized and realistic occluded face datasets show that the proposed algorithm significantly outperforms the state-of-the-art systems. Lingxue Song, Dihong Gong, Zhifeng Li 0001, Changsong Liu, Wei Liu 0005 |
ICCV | 2 |
| 2018 | CosFace: Large Margin Cosine Loss for Deep Face RecognitionabstractFace recognition has made extraordinary progress owing to the advancement of deep convolutional neural networks (CNNs). The central task of face recognition, including face verification and identification, involves face feature discrimination. However, the traditional softmax loss of deep CNNs usually lacks the power of discrimination. To address this problem, recently several loss functions such as center loss, large margin softmax loss, and angular softmax loss have been proposed. All these improved losses share the same idea: maximizing inter-class variance and minimizing intra-class variance. In this paper, we propose a novel loss function, namely large margin cosine loss (LMCL), to realize this idea from a different perspective. More specifically, we reformulate the softmax loss as a cosine loss by L2 normalizing both features and weight vectors to remove radial variations, based on which a cosine margin term is introduced to further maximize the decision margin in the angular space. As a result, minimum intra-class variance and maximum inter-class variance are achieved by virtue of normalization and cosine decision margin maximization. We refer to our model trained with LMCL as CosFace. Extensive experimental evaluations are conducted on the most popular public-domain face recognition datasets such as MegaFace Challenge, Youtube Faces (YTF) and Labeled Face in the Wild (LFW). We achieve the state-of-the-art performance on these benchmarks, which confirms the effectiveness of our proposed approach. Hao Wang 0050, Dihong Gong, Jingchao Zhou, Zhifeng Li 0001, Wei Liu 0005 |
CVPR | 5 |
| 2018 | Orthogonal Deep Features Decomposition for Age-Invariant Face Recognition
Dihong Gong, Hao Wang 0050, Zhifeng Li 0001, Wei Liu 0005, Tong Zhang 0001 |
ECCV (15) | 2 |
| 2017 | Extracting Visual Knowledge from the Web with Multimodal LearningabstractWe consider the problem of automatically extracting visual objects from web images. Despite the extraordinary advancement in deep learning, visual object detection remains a challenging task. To overcome the deficiency of pure visual techniques, we propose to make use of meta text surrounding images on the Web for enhanced detection accuracy. In this paper we present a multimodal learning algorithm to integrate text information into visual knowledge extraction. To demonstrate the effectiveness of our approach, we developed a system that takes raw webpages as input, and automatically extracts visual knowledge (e.g. object bounding boxes) from tens of millions of images crawled from the Web. Experimental results based on 46 object categories show that the extraction precision is improved significantly from 73% (with state-of-the-art deep learning programs) to 81%, which is equivalent to a 31% reduction in error rates. Dihong Gong, Daisy Zhe Wang |
IJCAI | 1 |
| 2017 | Multimodal Learning for Web Information ExtractionabstractWe consider the problem of extracting text instances of predefined categories from the Web. Instances of a category may be scattered across thousands of independent sources in many different formats with potential noises, which makes open-domain information extraction a challenging problem. Learning syntactic rules like "cities such as _" or "_ is a city" in a semi-supervised manner using a few labeled examples is usually unreliable because 1) high quality syntactic rules are rare and 2) the learning task is usually underconstrained. To address these problems, in this paper we propose to learn multimodal rules to combat the difficulty of syntactic rules. The multimodal rules are learned from information sources of different modalities, which is motivated by an intuition that information that is difficult to disambiguate correctly in one modality may be easily recognized in another. To demonstrate the effectiveness of this method, we have built a sophisticated end-to-end multimodal information extraction system that takes unannotated raw web pages as input, and generates a set of extracted instances as outputs. More specifically, our system learns reliable relationship between multimodal information by multimodal relation analysis on big unstructured data. Based on the learned relationship, we further train a set of multimodal rules for information extraction. Experimental evaluation shows that a greater accuracy for information extraction can be achieved by multimodal learning. Dihong Gong, Daisy Zhe Wang |
ACM Multimedia | 1 |
| 2017 | Heterogeneous Face Recognition: A Common Encoding Feature Discriminant ApproachabstractHeterogeneous face recognition is an important, yet challenging problem in face recognition community. It refers to matching a probe face image to a gallery of face images taken from alternate imaging modality. The major challenge of heterogeneous face recognition lies in the great discrepancies between different image modalities. Conventional face feature descriptors, e.g., local binary patterns, histogram of oriented gradients, and scale-invariant feature transform, are mostly designed in a handcrafted way and thus generally fail to extract the common discriminant information from the heterogeneous face images. In this paper, we propose a new feature descriptor called common encoding model for heterogeneous face recognition, which is able to capture common discriminant information, such that the large modality gap can be significantly reduced at the feature extraction stage. Specifically, we turn a face image into an encoded one with the encoding model learned from the training data, where the difference of the encoded heterogeneous face images of the same person can be minimized. Based on the encoded face images, we further develop a discriminant matching method to infer the hidden identity information of the cross-modality face images for enhanced recognition performance. The effectiveness of the proposed approach is demonstrated (on several public-domain face datasets) in two typical heterogeneous face recognition scenarios: matching NIR faces to VIS faces and matching sketches to photographs. Dihong Gong, Zhifeng Li 0001, Xuelong Li 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2017 | Multifeature Anisotropic Orthogonal Gaussian Process for Automatic Age EstimationabstractAutomatic age estimation is an important yet challenging problem. It has many promising applications in social media. Of the existing age estimation algorithms, the personalized approaches are among the most popular ones. However, most person-specific approaches rely heavily on the availability of training images across different ages for a single subject, which is usually difficult to satisfy in practical application of age estimation. To address this limitation, we first propose a new model called Orthogonal Gaussian Process (OGP), which is not restricted by the number of training samples per person. In addition, without sacrifice of discriminative power, OGP is much more computationally efficient than the standard Gaussian Process. Based on OGP, we then develop an effective age estimation approach, namely anisotropic OGP (A-OGP), to further reduce the estimation error. A-OGP is based on an anisotropic noise level learning scheme that contributes to better age estimation performance. To finally optimize the performance of age estimation, we propose a multifeature A-OGP fusion framework that uses multiple features combined with a random sampling method in the feature space. Extensive experiments on several public domain face aging datasets (FG-NET, MORPH Album1, and MORPH Album 2) are conducted to demonstrate the state-of-the-art estimation accuracy of our new algorithms. Zhifeng Li 0001, Dihong Gong, Dacheng Tao, Xuelong Li 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | Aging Face Recognition: A Hierarchical Learning Model Based on Local Patterns SelectionabstractAging face recognition refers to matching the same person's faces across different ages, e.g., matching a person's older face to his (or her) younger one, which has many important practical applications, such as finding missing children. The major challenge of this task is that facial appearance is subject to significant change during the aging process. In this paper, we propose to solve the problem with a hierarchical model based on two-level learning. At the first level, effective features are learned from low-level microstructures, based on our new feature descriptor called local pattern selection (LPS). The proposed LPS descriptor greedily selects low-level discriminant patterns in a way, such that intra-user dissimilarity is minimized. At the second level, higher level visual information is further refined based on the output from the first level. To evaluate the performance of our new method, we conduct extensive experiments on the MORPH data set (the largest face aging data set available in the public domain), which show a significant improvement in accuracy over the state-of-the-art methods. Zhifeng Li 0001, Dihong Gong, Xuelong Li 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2016 | Mutual Component Analysis for Heterogeneous Face RecognitionabstractHeterogeneous face recognition, also known as cross-modality face recognition or intermodality face recognition, refers to matching two face images from alternative image modalities. Since face images from different image modalities of the same person are associated with the same face object, there should be mutual components that reflect those intrinsic face characteristics that are invariant to the image modalities. Motivated by this rationality, we propose a novel approach called Mutual Component Analysis (MCA) to infer the mutual components for robust heterogeneous face recognition. In the MCA approach, a generative model is first proposed to model the process of generating face images in different modalities, and then an Expectation Maximization (EM) algorithm is designed to iteratively learn the model parameters. The learned generative model is able to infer the mutual components (which we call the hidden factor , where hidden means the factor is unreachable and invisible, and can only be inferred from observations) that are associated with the person’s identity, thus enabling fast and effective matching for cross-modality face recognition. To enhance recognition performance, we propose an MCA-based multiclassifier framework using multiple local features. Experimental results show that our new approach significantly outperforms the state-of-the-art results on two typical application scenarios: sketch-to-photo and infrared-to-visible face recognition. Zhifeng Li 0001, Dihong Gong, Qiang Li 0024, Dacheng Tao, Xuelong Li 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2015 | A maximum entropy feature descriptor for age invariant face recognitionabstractIn this paper, we propose a new approach to overcome the representation and matching problems in age invariant face recognition. First, a new maximum entropy feature descriptor (MEFD) is developed that encodes the microstructure of facial images into a set of discrete codes in terms of maximum entropy. By densely sampling the encoded face image, sufficient discriminatory and expressive information can be extracted for further analysis. A new matching method is also developed, called identity factor analysis (IFA), to estimate the probability that two faces have the same underlying identity. The effectiveness of the framework is confirmed by extensive experimentation on two face aging datasets, MORPH (the largest public-domain face aging dataset) and FGNET. We also conduct experiments on the famous LFW dataset to demonstrate the excellent generalizability of our new approach. Dihong Gong, Zhifeng Li 0001, Dacheng Tao, Jianzhuang Liu, Xuelong Li 0001 |
CVPR | 1 |
| 2015 | Probabilistic Ensemble Fusion for Multimodal Word Sense DisambiguationabstractWith the advent of abundant multimedia data on the Internet, there have been research efforts on multimodal machine learning to utilize data from different modalities. Current approaches mostly focus on developing models to fuse low-level features from multiple modalities and learn unified representation from different modalities. But most related work failed to justify why we should use multimodal data and multimodal fusion, and few of them leveraged the complementary relation among different modalities. In this paper, we first identify the correlative and complementary relations among multiple modalities. Then we propose a probabilistic ensemble fusion model to capture the complementary relation between two modalities (images and text). Experimental results on the UIUC-ISD dataset show our ensemble approach outperforms approaches using only single modality. Word sense disambiguation (WSD) is the use case we studied to demonstrate the effectiveness of our probabilistic ensemble fusion model. Daisy Zhe Wang, Ishan Patwa, Dihong Gong, Chunsheng Victor Fang |
ISM | 4 |
| 2015 | Hierarchical decomposition of dynamically evolving regulatory networksabstractBACKGROUND: Gene regulatory networks describe the interplay between genes and their products. These networks control almost every biological activity in the cell through interactions. The hierarchy of genes in these networks as defined by their interactions gives important insights into how these functions are governed. Accurately determining the hierarchy of genes is however a computationally difficult problem. This problem is further complicated by the fact that an intrinsic characteristic of regulatory networks is that the wiring of interactions can change over time. Determining how the hierarchy in the gene regulatory networks changes with dynamically evolving network topology remains to be an unsolved challenge. RESULTS: In this study, we develop a new method, named D-HIDEN (Dynamic-HIerarchical DEcomposition of Networks) to find the hierarchy of the genes in dynamically evolving gene regulatory network topologies. Unlike earlier methods, which recompute the hierarchy from scratch when the network topology changes, our method adapts the hierarchy based on the wiring of the interactions only for the nodes which have the potential to move in the hierarchy. CONCLUSIONS: We compare D-HIDEN to five currently available hierarchical decomposition methods on synthetic and real gene regulatory networks. Our experiments demonstrate that D-HIDEN significantly outperforms existing methods in running time, accuracy, or both. Furthermore, our method is robust against dynamic changes in hierarchy. Our experiments on human gene regulatory networks suggest that our method may be used to reconstruct hierarchy in gene regulatory networks. Ahmet Ay, Dihong Gong, Tamer Kahveci |
BMC Bioinform. | 2 |
| 2015 | Learning Compact Feature Descriptor and Adaptive Matching Framework for Face RecognitionabstractDense feature extraction is becoming increasingly popular in face recognition tasks. Systems based on this approach have demonstrated impressive performance in a range of challenging scenarios. However, improvements in discriminative power come at a computational cost and with a risk of over-fitting. In this paper, we propose a new approach to dense feature extraction for face recognition, which consists of two steps. First, an encoding scheme is devised that compresses high-dimensional dense features into a compact representation by maximizing the intrauser correlation. Second, we develop an adaptive feature matching algorithm for effective classification. This matching method, in contrast to the previous methods, constructs and chooses a small subset of training samples for adaptive matching, resulting in further performance gains. Experiments using several challenging face databases, including labeled Faces in the Wild data set, Morph Album 2, CUHK optical-infrared, and FERET, demonstrate that the proposed approach consistently outperforms the current state of the art. Zhifeng Li 0001, Dihong Gong, Xuelong Li 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2014 | Orthogonal Gaussian Process for Automatic Age EstimationabstractAge Estimation from facial images has been receiving increasing interest due to its important applications. Among the existing age estimation algorithms, the personalized approaches have been shown to be the most effective ones. However, most of the person-specific approaches (e.g. MTWGP [1], AGES [2], WAS [3]) rely heavily on the availability of training images across different ages for a single subject, which is very difficult to satisfy in practical applications. In order to overcome this problem, we propose a new approach to age estimation, called Orthogonal Gaussian Process (OGP). Compared to standard Gaussian Process, OGP is much more efficient while maintaining the discriminatory power of the standard Gaussian Process. Based on OGP, we further propose an improvement of OGP called anisotropic OGP (A-OGP) to enhance the age estimation performance. Extensive experiments are conducted to demonstrate the state-of-the-art estimation accuracy of our new algorithm on several public-domain face aging datasets: FG-NET face dataset with 82 different subjects, Morph Album 1 dataset with more than 600 subjects, and Morph Album 2 with about 20,000 different subjects. Dihong Gong, Zhifeng Li 0001, Xiaoou Tang |
ACM Multimedia | 2 |
| 2014 | Common Feature Discriminant Analysis for Matching Infrared Face Images to Optical Face ImagesabstractIn biometrics research and industry, it is critical yet a challenge to match infrared face images to optical face images. The major difficulty lies in the fact that a great discrepancy exists between the infrared face image and corresponding optical face image because they are captured by different devices (optical imaging device and infrared imaging device). This paper presents a new approach called common feature discriminant analysis to reduce this great discrepancy and improve optical-infrared face recognition performance. In this approach, a new learning-based face descriptor is first proposed to extract the common features from heterogeneous face images (infrared face images and optical face images), and an effective matching method is then applied to the resulting features to obtain the final decision. Extensive experiments are conducted on two large and challenging optical-infrared face data sets to show the superiority of our approach over the state-of-the-art. Zhifeng Li 0001, Dihong Gong, Yu Qiao 0001, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2013 | Hidden Factor Analysis for Age Invariant Face RecognitionabstractAge invariant face recognition has received increasing attention due to its great potential in real world applications. In spite of the great progress in face recognition techniques, reliably recognizing faces across ages remains a difficult task. The facial appearance of a person changes substantially over time, resulting in significant intra-class variations. Hence, the key to tackle this problem is to separate the variation caused by aging from the person-specific features that are stable. Specifically, we propose a new method, called Hidden Factor Analysis (HFA). This method captures the intuition above through a probabilistic model with two latent factors: an identity factor that is age-invariant and an age factor affected by the aging process. Then, the observed appearance can be modeled as a combination of the components generated based on these factors. We also develop a learning algorithm that jointly estimates the latent factors and the model parameters using an EM procedure. Extensive experiments on two well-known public domain face aging datasets: MORPH (the largest public face aging database) and FGNET, clearly show that the proposed method achieves notable improvement over state-of-the-art algorithms. Dihong Gong, Zhifeng Li 0001, Dahua Lin, Jianzhuang Liu, Xiaoou Tang |
ICCV | 1 |
| 2013 | Multi-feature canonical correlation analysis for face photo-sketch image retrievalabstractAutomatic face photo-sketch image retrieval has attracted great attention in recent years due to its important applications in real life. The major difficulty in automatic face photo-sketch image retrieval lies in the fact that there exists great discrepancy between the different image modalities (photo and sketch). In order to reduce such discrepancy and improve the performance of automatic face photo-sketch image retrieval, we propose a new framework called multi-feature canonical correlation analysis (MCCA) to effectively address this problem. The MCCA is an extension and improvement of the canonical correlation analysis (CCA) algorithmusing multiple features combined with two different random sampling methods in feature space and sample space. In this framework, we first represent each photo or sketch using a patch-based local feature representation scheme, in which histograms of oriented gradients (HOG) and multi-scale local binary pattern (MLBP) serve as the local descriptors. Canonical correlation analysis (CCA) is then performed on a collection of random subspaces to construct an ensemble of classifiers for photo-sketch image retrieval. Extensive experiments on two public-domain face photo-sketch datasets (CUFS and CUFSF) clearly show that the proposed approach obtains a substantial improvement over the state-of-the-art. Dihong Gong, Zhifeng Li 0001, Jianzhuang Liu, Yu Qiao 0001 |
ACM Multimedia | 1 |