Jing Yang 0038

dblp:62/5839-38 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-8794-4842ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Graph in Graph Neural Network
abstract
Abstract Existing Graph Neural Networks (GNNs) are limited to process graphs each of whose vertices is represented by a vector or a single value, limited their representing capability to describe complex objects. In this paper, we propose a novel GNN (called Graph in Graph Neural (GIG) Network) which can process graph-style data (called GIG sample) whose vertices are further represented by graphs. Given a set of graphs or a data sample whose components can be represented by a set of graphs (called multi-graph data sample), our GIG network starts with a GIG sample generation (GSG) module which encodes the input as a GIG sample , where each GIG vertex includes a graph. Then, a set of GIG hidden layers are stacked, with each consisting of: (1) a GIG vertex-level updating (GVU) module that individually updates the graph in every GIG vertex based on its internal information; and (2) a global-level GIG sample updating (GGU) module that updates graphs in all GIG vertices based on their relationships, making the updated GIG vertices become global context-aware. This way, both internal cues within the graph contained in each GIG vertex and the relationships among GIG vertices could be utilized for down-stream tasks. Experimental results demonstrate that our GIG network generalizes well for not only various generic graph analysis tasks but also real-world multi-graph data analysis (e.g., human skeleton video-based action recognition), which achieved the new state-of-the-art results on 15 out of 16 evaluated datasets. Our code is publicly available at https://github.com/wangjs96/Graph-in-Graph-Neural-Network .
Jiongshu Wang, Jing Yang 0038, Jiankang Deng, Hatice Gunes, Siyang Song
Int. J. Comput. Vis.2
2025 HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization
Zimin Ran, Xingyu Ren, Xiang An, Kaicheng Yang 0002, Ziyong Feng, Jing Yang 0038, Rolandos Alexandros Potamias, Linchao Zhu, Jiankang Deng
ICCV6
2025 Knowledge Distillation Meets Open-Set Semi-supervised Learning
abstract
Abstract Existing knowledge distillation methods mostly focus on distillation of teacher’s prediction and intermediate activation. However, the structured representation, which arguably is one of the most critical ingredients of deep models, is largely overlooked. In this work, we propose a novel semantic representational distillation (SRD) method dedicated for distilling representational knowledge semantically from a pretrained teacher to a target student. The key idea is that we leverage the teacher’s classifier as a semantic critic for evaluating the representations of both teacher and student and distilling the semantic knowledge with high-order structured information over all feature dimensions. This is accomplished by introducing a notion of cross-network logit computed through passing student’s representation into teacher’s classifier. Further, considering the set of seen classes as a basis for the semantic space in a combinatorial perspective, we scale SRD to unseen classes for enabling effective exploitation of largely available, arbitrary unlabeled training data. At the problem level, this establishes an interesting connection between knowledge distillation with open-set semi-supervised learning (SSL). Extensive experiments show that our SRD outperforms significantly previous state-of-the-art knowledge distillation methods on both coarse object classification and fine face recognition tasks, as well as less studied yet practically crucial binary network distillation. Under more realistic open-set SSL settings we introduce, we reveal that knowledge distillation is generally more effective than existing out-of-distribution sample detection, and our proposed SRD is superior over both previous distillation and SSL competitors. The source code is available at https://github.com/jingyang2017/SRD_ossl .
Jing Yang 0038, Xiatian Zhu, Adrian Bulat, Brais Martínez, Georgios Tzimiropoulos
Int. J. Comput. Vis.1
2023 ALIP: Adaptive Language-Image Pre-training with Synthetic Caption
abstract
Contrastive Language-Image Pre-training (CLIP) has significantly boosted the performance of various vision-language tasks by scaling up the dataset with image-text pairs collected from the web. However, the presence of intrinsic noise and unmatched image-text pairs in web data can potentially affect the performance of representation learning. To address this issue, we first utilize the OFA model to generate synthetic captions that focus on the image content. The generated captions contain complementary information that is beneficial for pre-training. Then, we propose an Adaptive Language-Image Pre-training (ALIP), a bi-path model that integrates supervision from both raw text and synthetic caption. As the core components of ALIP, the Language Consistency Gate (LCG) and Description Consistency Gate (DCG) dynamically adjust the weights of samples and image-text/caption pairs during the training process. Meanwhile, the adaptive contrastive loss can effectively reduce the impact of noise data and enhances the efficiency of pre-training data. We validate ALIP with experiments on different scales of models and pre-training datasets. Experiments results show that ALIP achieves state-of-the-art performance on multiple downstream tasks including zero-shot image-text retrieval and linear probe. To facilitate future research, the code and pre-trained models are released at https://github.com/deepglint/ALIP.
Kaicheng Yang 0002, Jiankang Deng, Xiang An, Ziyong Feng, Jia Guo 0003, Jing Yang 0038, Tongliang Liu
ICCV7
2023 Unicom: Universal and Compact Representation Learning for Image Retrieval
Xiang An, Jiankang Deng, Kaicheng Yang 0002, Jaiwei Li, Ziyong Feng, Jia Guo 0003, Jing Yang 0038, Tongliang Liu
ICLR7
2023 FAN-Trans: Online Knowledge Distillation for Facial Action Unit Detection
abstract
Due to its importance in facial behaviour analysis, facial action unit (AU) detection has attracted increasing attention from the research community. Leveraging the online knowledge distillation framework, we propose the "FAN-Trans" method for AU detection. Our model consists of a hybrid network of convolution and transformer blocks to learn per-AU features and to model AU co-occurrences. The model uses a pre-trained face alignment network as the feature extractor. After further transformation by a small learnable add-on convolutional subnet, the per-AU features are fed into transformer blocks to enhance their representation. As multiple AUs often appear together, we propose a learnable attention drop mechanism in the transformer block to learn the correlation between the features for different AUs. We also design a classifier that predicts AU presence by considering all AUs’ features, to explicitly capture label dependencies. Finally, we make the attempt of adapting online knowledge distillation in the training stage for this task, further improving the model’s performance. Experiments on the BP4D and DISFA datasets demonstrating the effectiveness of proposed method.
Jing Yang 0038, Jie Shen 0008, Yiming Lin 0001, Yordan Hristov, Maja Pantic
WACV1
2023 Toward Robust Facial Action Units' Detection
abstract
Facial action unit (AU) detection plays an important role in performing facial behavioral analysis of raw video inputs. Overall, there are three key factors that contribute toward the optimal performance of AU detectors: 1) being able to capture local AU-centered features; 2) exploiting the fact that some AUs co-occur with others; and 3) utilizing appearance changes across frames. We briefly review current techniques addressing each factor and discuss the challenges they meet. Given that very few works consider how to effectively and efficiently merge them all into a single framework that can be trained in an end-to-end manner, we propose facial AU detection with face alignment(AUNet), a simple yet strong baseline for landmark-based AU detection. AUNet implements the abovementioned key factors by: 1) using the intermediate layers of a pretrained face alignment model to act as our AU features’ space; 2) optimized to satisfy a correlation constraint, derived from the AU labels, and 3) temporal constraint, derived from variations in the contents of consecutive frames in the input videos. The proposed model, with its three key components, remains simple in nature and aligns with the primary AU detection task. Experiments on several benchmarks show that it substantially improves the AU detector’s accuracy and achieves new state-of-the-art AU detection results on popular benchmarks: BP4D and DISFA. Code is available athttps://github.com/jingyang2017/AU-Net.
Jing Yang 0038, Yordan Hristov, Jie Shen 0008, Yiming Lin 0001, Maja Pantic
Proc. IEEE1
2022 Killing Two Birds with One Stone: Efficient and Robust Training of Face Recognition CNNs by Partial FC
abstract
Learning discriminative deep feature embeddings by using million-scale in-the-wild datasets and margin-based softmax loss is the current state-of-the-art approach for face recognition. However, the memory and computing cost of the Fully Connected (FC) layer linearly scales up to the number of identities in the training set. Besides, the largescale training data inevitably suffers from inter-class conflict and long-tailed distribution. In this paper, we propose a sparsely updating variant of the FC layer, named Partial FC (PFC). In each iteration, positive class centers and a random subset of negative class centers are selected to compute the margin-based softmax loss. All class centers are still maintained throughout the whole training process, but only a subset is selected and updated in each iteration. Therefore, the computing requirement, the probability of inter-class conflict, and the frequency of passive update on tail class centers, are dramatically reduced. Extensive experiments across different training data and backbones (e.g. CNN and ViT) confirm the effectiveness, robustness and efficiency of the proposed PFC. The source code is available at https://github.com/deepinsight/insightface/tree/master/recognition.
Xiang An, Jiankang Deng, Jia Guo 0003, Ziyong Feng, Xuhan Zhu, Jing Yang 0038, Tongliang Liu
CVPR6
2022 Pre-training Strategies and Datasets for Facial Representation Learning
Adrian Bulat, Shiyang Cheng 0001, Jing Yang 0038, Andrew Garbett, Enrique Sánchez-Lozano, Georgios Tzimiropoulos
ECCV (13)3
2022 ArcFace: Additive Angular Margin Loss for Deep Face Recognition
Jiankang Deng, Jia Guo 0003, Jing Yang 0038, Niannan Xue, Irene Kotsia, Stefanos Zafeiriou
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Variational Prototype Learning for Deep Face Recognition
abstract
Deep face recognition has achieved remarkable improvements due to the introduction of margin-based softmax loss, in which the prototype stored in the last linear layer represents the center of each class. In these methods, training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin. However, we argue that prototype learning only employs sample-to-prototype comparisons without considering sample-to-sample comparisons during training and the low loss value gives us an illusion of perfect feature embedding, impeding the further exploration of SGD. To this end, we propose Variational Prototype Learning (VPL), which represents every class as a distribution instead of a point in the latent space. By identifying the slow feature drift phenomenon, we directly inject memorized features into prototypes to approximate variational prototype sampling. The proposed VPL can simulate sample-to-sample comparisons within the classification framework, encouraging the SGD solver to be more exploratory, while boosting performance. Moreover, VPL is conceptually simple, easy to implement, computationally efficient and memory saving. We present extensive experimental results on popular benchmarks, which demonstrate the superiority of the proposed VPL method over the state-of-the-art competitors.
Jiankang Deng, Jia Guo 0003, Jing Yang 0038, Alexander Lattas, Stefanos Zafeiriou
CVPR3
2021 Knowledge distillation via softmax regression representation learning
Jing Yang 0038, Brais Martínez, Adrian Bulat, Georgios Tzimiropoulos
ICLR1
2020 FAN-Face: a Simple Orthogonal Improvement to Deep Face Recognition
abstract
It is known that facial landmarks provide pose, expression and shape information. In addition, when matching, for example, a profile and/or expressive face to a frontal one, knowledge of these landmarks is useful for establishing correspondence which can help improve recognition. However, in prior work on face recognition, facial landmarks are only used for face cropping in order to remove scale, rotation and translation variations. This paper proposes a simple approach to face recognition which gradually integrates features from different layers of a facial landmark localization network into different layers of the recognition network. To this end, we propose an appropriate feature integration layer which makes the features compatible before integration. We show that such a simple approach systematically improves recognition on the most difficult face recognition datasets, setting a new state-of-the-art on IJB-B, IJB-C and MegaFace datasets.
Jing Yang 0038, Adrian Bulat, Georgios Tzimiropoulos
AAAI1
2020 Training binary neural networks with real-to-binary convolutions
Brais Martínez, Jing Yang 0038, Adrian Bulat, Georgios Tzimiropoulos
ICLR2
2018 To Learn Image Super-Resolution, Use a GAN to Learn How to Do Image Degradation First
Adrian Bulat, Jing Yang 0038, Georgios Tzimiropoulos
ECCV (6)2
2017 Face image retrieval based on shape and texture feature fusion
abstract
Humongous amounts of data bring various challenges to face image retrieval. This paper proposes an efficient method to solve those problems. Firstly, we use accurate facial landmark locations as shape features. Secondly, we utilise shape priors to provide discriminative texture features for convolutional neural networks. These shape and texture features are fused to make the learned representation more robust. Finally, in order to increase efficiency, a coarse-tofine search mechanism is exploited to efficiently find similar objects. Extensive experiments on the CASIAWebFace, MSRA-CFW, and LFW datasets illustrate the superiority of our method.
Zongguang Lu, Jing Yang 0038, Qingshan Liu 0001
Comput. Vis. Media2
2017 Robust facial landmark tracking via cascade regression
Qingshan Liu 0001, Jing Yang 0038, Jiankang Deng, Kaihua Zhang 0001
Pattern Recognit.2
2017 Adaptive Compressive Tracking via Online Vector Boosting Feature Selection
abstract
Recently, the compressive tracking (CT) method has attracted much attention due to its high efficiency, but it cannot well deal with the large scale target appearance variations due to its data-independent random projection matrix that results in less discriminative features. To address this issue, in this paper, we propose an adaptive CT approach, which selects the most discriminative features to design an effective appearance model. Our method significantly improves CT in three aspects. First, the most discriminative features are selected via an online vector boosting method. Second, the object representation is updated in an effective online manner, which preserves the stable features while filtering out the noisy ones. Furthermore, a simple and effective trajectory rectification approach is adopted that can make the estimated location more accurate. Finally, a multiple scale adaptation mechanism is explored to estimate object size, which helps to relieve interference from background information. Extensive experiments on the CVPR2013 tracking benchmark and the VOT2014 challenges demonstrate the superior performance of our method.
Qingshan Liu 0001, Jing Yang 0038, Kaihua Zhang 0001, Yi Wu 0001
IEEE Trans. Cybern.2
2017 Adaptive Cascade Regression Model For Robust Face Alignment
abstract
Cascade regression is a popular face alignment approach, and it has achieved good performances on the wild databases. However, it depends heavily on local features in estimating reliable landmark locations and therefore suffers from corrupted images, such as images with occlusion, which often exists in real-world face images. In this paper, we present a new adaptive cascade regression model for robust face alignment. In each iteration, the shape-indexed appearance is introduced to estimate the occlusion level of each landmark, and each landmark is then weighted according to its estimated occlusion level. Also, the occlusion levels of the landmarks act as adaptive weights on the shape-indexed features to decrease the noise on the shape-indexed features. At the same time, an exemplar-based shape prior is designed to suppress the influence of local image corruption. Extensive experiments are conducted on the challenging benchmarks, and the experimental results demonstrate that the proposed method achieves better results than the state-of-the-art methods for facial landmark localization and occlusion detection.
Qingshan Liu 0001, Jiankang Deng, Jing Yang 0038, Guangcan Liu, Dacheng Tao
IEEE Trans. Image Process.3
2016 Robust object tracking by online Fisher discrimination boosting feature selection
Jing Yang 0038, Kaihua Zhang 0001, Qingshan Liu 0001
Comput. Vis. Image Underst.1
2016 M3 CSR: Multi-view, multi-scale and multi-component cascade shape regression
Jiankang Deng, Qingshan Liu 0001, Jing Yang 0038, Dacheng Tao
Image Vis. Comput.3
2016 FaceHunter: A multi-task convolutional neural network based face detector
Jing Yang 0038, Jiankang Deng, Qingshan Liu 0001
Signal Process. Image Commun.2
2015 Hierarchical Convolutional Neural Network for Face Detection
Jing Yang 0038, Jiankang Deng, Qingshan Liu 0001
ICIG (2)2