VLDB 2026 Research / reviewers in the wild / expert
Jia Guo 0003
dblp:57/4142-3
· DBLP profile ↗
13ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-0709-261XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unifying and Conquering Adversarial Attacks against Deep Face RecognitionabstractAdversarial attack severely degrades the performance of deep face recognition models with minimal image perturbations. Facial attacks can be broadly categorized into dodging and impersonation. Although existing approaches have made satisfactory progress by separately addressing each type of attack, the unification of them still remains unexplored and the potential benefit might be ignored. In this work, we aim to construct a unified facial attack framework that can simultaneously perform both attacks. We first design Unified Attack Modeling, which aligns optimization directions and spaces of two attacks to harmonize optimizing forces of dodging and impersonation. Based on the unified optimization goals, we introduce an Attacking by Decomposing scheme. The proposed scheme leverages the unified optimization goal by decomposing adversarial perturbation synthesis into two steps, orientation exploration and amplified forwarding. Extensive experiments on several facial datasets LFW, CelebA-HQ, and FFHQ show that our proposed unified framework achieves significant performance improvement on dodging and impersonation attacks, compared with the state-of-the-art methods. Jia Guo 0003, Zhipeng Du, Jiankang Deng |
FG | 1 |
| 2024 | Monocular Identity-Conditioned Facial Reflectance ReconstructionabstractRecent 3D face reconstruction methods have made re-markable advancements, yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing methods rely on a large amount of light-stage captured data to learn facial reflectance models. However, the lack of subject diversity poses challenges in achieving good generalization and widespread applicability. In this paper, we learn the reflectance prior in image space rather than UV space and present a framework named ID2Reflectance. Our framework can directly estimate the reflectance maps of a single image while using limited reflectance data for training. Our key insight is that reflectance data shares facial structures with RGB faces, which enables obtaining expressive facial prior from inexpensive RGB data thus re-ducing the dependency on reflectance data. We first learn a high-quality prior for facial reflectance. Specifically, we pretrain multi-domain facial feature code books and design a codebook fusion method to align the reflectance and RGB domains. Then, we propose an identity-conditioned swapping module that injects facial identity from the target image into the pre-trained autoencoder to modify the identity of the source reflectance image. Finally, we stitch multi-view swapped reflectance images to obtain renderable assets. Extensive experiments demonstrate that our method exhibits excellent generalization capability and achieves state-of-the-art facial reflectance reconstruction results for in-the-wild faces. Our project page is https://xingyuren.github.io/id2reflectance. Xingyu Ren, Jiankang Deng, Yuhao Cheng, Jia Guo 0003, Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001 |
CVPR | 4 |
| 2024 | 3DGazeNet: Generalizing 3D Gaze Estimation with Weak-Supervision from Synthetic Views
Evangelos Ververas, Polydefkis Gkagkos, Jiankang Deng, Michail C. Doukas, Jia Guo 0003, Stefanos Zafeiriou |
ECCV (21) | 5 |
| 2023 | ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionabstractContrastive Language-Image Pre-training (CLIP) has significantly boosted the performance of various vision-language tasks by scaling up the dataset with image-text pairs collected from the web. However, the presence of intrinsic noise and unmatched image-text pairs in web data can potentially affect the performance of representation learning. To address this issue, we first utilize the OFA model to generate synthetic captions that focus on the image content. The generated captions contain complementary information that is beneficial for pre-training. Then, we propose an Adaptive Language-Image Pre-training (ALIP), a bi-path model that integrates supervision from both raw text and synthetic caption. As the core components of ALIP, the Language Consistency Gate (LCG) and Description Consistency Gate (DCG) dynamically adjust the weights of samples and image-text/caption pairs during the training process. Meanwhile, the adaptive contrastive loss can effectively reduce the impact of noise data and enhances the efficiency of pre-training data. We validate ALIP with experiments on different scales of models and pre-training datasets. Experiments results show that ALIP achieves state-of-the-art performance on multiple downstream tasks including zero-shot image-text retrieval and linear probe. To facilitate future research, the code and pre-trained models are released at https://github.com/deepglint/ALIP. Kaicheng Yang 0002, Jiankang Deng, Xiang An, Ziyong Feng, Jia Guo 0003, Jing Yang 0038, Tongliang Liu |
ICCV | 6 |
| 2023 | Unicom: Universal and Compact Representation Learning for Image Retrieval
Xiang An, Jiankang Deng, Kaicheng Yang 0002, Jaiwei Li, Ziyong Feng, Jia Guo 0003, Jing Yang 0038, Tongliang Liu |
ICLR | 6 |
| 2022 | Killing Two Birds with One Stone: Efficient and Robust Training of Face Recognition CNNs by Partial FCabstractLearning discriminative deep feature embeddings by using million-scale in-the-wild datasets and margin-based softmax loss is the current state-of-the-art approach for face recognition. However, the memory and computing cost of the Fully Connected (FC) layer linearly scales up to the number of identities in the training set. Besides, the largescale training data inevitably suffers from inter-class conflict and long-tailed distribution. In this paper, we propose a sparsely updating variant of the FC layer, named Partial FC (PFC). In each iteration, positive class centers and a random subset of negative class centers are selected to compute the margin-based softmax loss. All class centers are still maintained throughout the whole training process, but only a subset is selected and updated in each iteration. Therefore, the computing requirement, the probability of inter-class conflict, and the frequency of passive update on tail class centers, are dramatically reduced. Extensive experiments across different training data and backbones (e.g. CNN and ViT) confirm the effectiveness, robustness and efficiency of the proposed PFC. The source code is available at https://github.com/deepinsight/insightface/tree/master/recognition. Xiang An, Jiankang Deng, Jia Guo 0003, Ziyong Feng, Xuhan Zhu, Jing Yang 0038, Tongliang Liu |
CVPR | 3 |
| 2022 | Sample and Computation Redistribution for Efficient Face Detection
Jia Guo 0003, Jiankang Deng, Alexander Lattas, Stefanos Zafeiriou |
ICLR | 1 |
| 2022 | ArcFace: Additive Angular Margin Loss for Deep Face Recognition
Jiankang Deng, Jia Guo 0003, Jing Yang 0038, Niannan Xue, Irene Kotsia, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Variational Prototype Learning for Deep Face RecognitionabstractDeep face recognition has achieved remarkable improvements due to the introduction of margin-based softmax loss, in which the prototype stored in the last linear layer represents the center of each class. In these methods, training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin. However, we argue that prototype learning only employs sample-to-prototype comparisons without considering sample-to-sample comparisons during training and the low loss value gives us an illusion of perfect feature embedding, impeding the further exploration of SGD. To this end, we propose Variational Prototype Learning (VPL), which represents every class as a distribution instead of a point in the latent space. By identifying the slow feature drift phenomenon, we directly inject memorized features into prototypes to approximate variational prototype sampling. The proposed VPL can simulate sample-to-sample comparisons within the classification framework, encouraging the SGD solver to be more exploratory, while boosting performance. Moreover, VPL is conceptually simple, easy to implement, computationally efficient and memory saving. We present extensive experimental results on popular benchmarks, which demonstrate the superiority of the proposed VPL method over the state-of-the-art competitors. Jiankang Deng, Jia Guo 0003, Jing Yang 0038, Alexander Lattas, Stefanos Zafeiriou |
CVPR | 2 |
| 2020 | RetinaFace: Single-Shot Multi-Level Face Localisation in the WildabstractThough tremendous strides have been made in uncontrolled face detection, accurate and efficient 2D face alignment and 3D face reconstruction in-the-wild remain an open challenge. In this paper, we present a novel single-shot, multi-level face localisation method, named RetinaFace, which unifies face box prediction, 2D facial landmark localisation and 3D vertices regression under one common target: point regression on the image plane. To fill the data gap, we manually annotated five facial landmarks on the WIDER FACE dataset and employed a semi-automatic annotation pipeline to generate 3D vertices for face images from the WIDER FACE, AFLW and FDDB datasets. Based on extra annotations, we propose a mutually beneficial regression target for 3D face reconstruction, that is predicting 3D vertices projected on the image plane constrained by a common 3D topology. The proposed 3D face reconstruction branch can be easily incorporated, without any optimisation difficulty, in parallel with the existing box and 2D landmark regression branches during joint training. Extensive experimental results show that RetinaFace can simultaneously achieve stable face detection, accurate 2D face alignment and robust 3D face reconstruction while being efficient through single-shot inference. Jiankang Deng, Jia Guo 0003, Evangelos Ververas, Irene Kotsia, Stefanos Zafeiriou |
CVPR | 2 |
| 2020 | Sub-center ArcFace: Boosting Face Recognition by Large-Scale Noisy Web Faces
Jiankang Deng, Jia Guo 0003, Tongliang Liu, Mingming Gong, Stefanos Zafeiriou |
ECCV (11) | 2 |
| 2019 | ArcFace: Additive Angular Margin Loss for Deep Face RecognitionabstractRecently, a popular line of research in face recognition is adopting margins in the well-established softmax loss function to maximize class separability. In this paper, we first introduce an Additive Angular Margin Loss (ArcFace), which not only has a clear geometric interpretation but also significantly enhances the discriminative power. Since ArcFace is susceptible to the massive label noise, we further propose sub-center ArcFace, in which each class contains K sub-centers and training samples only need to be close to any of the K positive sub-centers. Sub-center ArcFace encourages one dominant sub-class that contains the majority of clean faces and non-dominant sub-classes that include hard or noisy faces. Based on this self-propelled isolation, we boost the performance through automatically purifying raw web faces under massive real-world noise. Besides discriminative feature embedding, we also explore the inverse problem, mapping feature vectors to face images. Without training any additional generator or discriminator, the pre-trained ArcFace model can generate identity-preserved face images for both subjects inside and outside the training data only by using the network gradient and Batch Normalization (BN) priors. Extensive experiments demonstrate that ArcFace can enhance the discriminative feature embedding as well as strengthen the generative face synthesis. Jiankang Deng, Jia Guo 0003, Niannan Xue, Stefanos Zafeiriou |
CVPR | 2 |
| 2018 | Stacked Dense U-Nets with Dual Transformers for Robust Face Alignment
Jia Guo 0003, Jiankang Deng, Niannan Xue, Stefanos Zafeiriou |
BMVC | 1 |