VLDB 2026 Research / reviewers in the wild / expert
Hangyu Li 0001
dblp:75/10203-1
· DBLP profile ↗
9ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-7153-6676ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Enhanced Adaptive Confidence Margin for Semi-Supervised Facial Expression RecognitionabstractSemi-supervised learning (SSL) provides a practical framework for leveraging massive unlabeled samples, especially when labels are expensive for facial expression recognition (FER). Typical SSL methods like FixMatch select unlabeled samples with confidence scores above a fixed threshold for training. However, these methods face two primary limitations: failing to consider the varying confidence across facial expression categories and failing to utilize unlabeled facial expression samples efficiently. To address these challenges, we propose an Enhanced Adaptive Confidence Margin (EACM), consisting of dynamic thresholds for different categories, to fully learn unlabeled samples. Specifically, we employ the predictions on labeled samples at each training iteration to learn an EACM. It then partitions unlabeled samples into two subsets: (1) subset I, including samples whose confidence scores are no less than the margin; (2) subset II, including samples whose confidence scores are less than the margin. For samples in subset I, we constrain their predictions on strongly-augmented versions to match the pseudo-labels derived from the predictions on weakly-augmented versions. Meanwhile, we introduce a feature-level contrastive objective to enhance the similarity between two weakly-augmented features of a sample in subset II. We extensively evaluate EACM on image-based and video-based facial expression datasets, showing that our method achieves superior performance, significantly surpassing fully-supervised baselines in a semi-supervised manner. Additionally, our EACM is promising to leverage cross-dataset unlabeled samples for practical training to boost fully-supervised performance. Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Towards Regularized Mixture of Predictions for Class-Imbalanced Semi-Supervised Facial Expression RecognitionabstractSemi-supervised facial expression recognition (SSFER) effectively assigns pseudo-labels to confident unlabeled samples when only limited emotional annotations are available. Existing SSFER methods are typically built upon an assumption of the class-balanced distribution. However, they are far from real-world applications due to biased pseudo-labels caused by class imbalance. To alleviate this issue, we propose Regularized Mixture of Predictions (ReMoP), a simple yet effective method to generate high-quality pseudo-labels for imbalanced samples. Specifically, we first integrate feature similarity into the linear prediction to learn a mixture of predictions. Furthermore, we introduce a class regularization term that constrains the feature geometry to mitigate imbalance bias. Being practically simple, our method can be integrated with existing semi-supervised learning and SSFER methods to tackle the challenge associated with class-imbalanced SSFER effectively. Extensive experiments on four facial expression datasets demonstrate the effectiveness of the proposed method across various imbalanced conditions. The source code is made publicly available at https://github.com/hangyu94/ReMoP. Hangyu Li 0001, Jiangchao Yao, Nannan Wang 0001, Bo Han 0003 |
IJCAI | 1 |
| 2025 | Boosting Semi-Supervised Facial Attribute Recognition With Dynamic Threshold PairsabstractSemi-supervised learning (SSL) has proven effective in assigning a pseudo-label to a confident sample whose largest class probability is above a fixed threshold. However, in the context of semi-supervised facial attribute recognition (SSFAR), where a sample is associated with multiple presence and absence pseudo-labels, directly applying existing SSL methods is challenging due to two issues: 1) the lack of a clear boundary between presence and absence predictions for an attribute makes it difficult to distinguish them using a single threshold; 2) the learning difficulty varies across attributes, so the fixed strategy fails to adaptively learn different attributes. To address these challenges, we propose Dynamic thrEShold Pairs (DESP), a simple yet effective method to handle the SSFAR problem. Specifically, during each training stage, we derive two sets for each attribute from labeled samples, which contain the predicted probabilities of presence and absence, respectively. We then compute the mid-ranges of the two sets as paired presence and absence thresholds. Finally, we assign a presence or absence pseudo-label for the attribute to an unlabeled sample when its prediction exceeds the presence threshold or falls below the absence threshold. Extensive experiments on the CelebA and LFWA datasets demonstrate that DESP achieves superior performance compared to state-of-the-art methods, especially in the case of scarce labeled samples. Also, DESP performs well on multi-label datasets such as Pascal VOC and MS-COCO. The code will be publicly available athttps://github.com/yihanxxu/DESP. Hangyu Li 0001, Mingrui Zhu, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Knowledge-Enhanced Facial Expression Recognition With Emotional-to-Neutral TransformationabstractExisting facial expression recognition (FER) methods typically fine-tune a pre-trained visual encoder using discrete labels. However, this form of supervision limits to specify the emotional concept of different facial expressions. In this paper, we observe that the rich knowledge in text embeddings, generated by vision-language models, is a promising alternative for learning discriminative facial expression representations. Inspired by this, we propose a novel knowledge-enhanced FER method with an emotional-to-neutral transformation. Specifically, we formulate the FER problem as a process to match the similarity between a facial expression representation and text embeddings. Then, we transform the facial expression representation to a neutral representation by simulating the difference in text embeddings from textual facial expression to textual neutral. Finally, a self-contrast objective is introduced to pull the facial expression representation closer to the textual facial expression, while pushing it farther from the neutral representation. We conduct evaluation with diverse pre-trained visual encoders including ResNet-18 and Swin-T on four challenging facial expression datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art FER methods. The code is made publicly available athttps://github.com/hangyu94/KE2NT. Hangyu Li 0001, Jiangchao Yao, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003 |
IEEE Trans. Multim. | 1 |
| 2024 | Unconstrained Facial Expression Recognition With No-Reference De-Elements LearningabstractMost unconstrained facial expression recognition (FER) methods take original facial images as inputs to learn discriminative features by well-designed loss functions, which cannot reflect important visual information in faces. Although existing methods have explored the visual information of constrained facial expressions, there is no explicit modeling of what visual information is important for unconstrained FER. To find out valuable information of unconstrained facial expressions, we pose a new problem of no-reference de-elements learning: we decompose any unconstrained facial image into the facial expression element and a neutral face without the reference of corresponding neutral faces. Importantly, the element provides visualization results to understand important facial expression information and improves the discriminative power of features. Moreover, we propose a simple yet effectiveDe-ElementsNetwork (DENet) to learn the element and introduce appropriate constraints to overcome no ground truth of corresponding neutral faces during the de-elements learning. We extensively evaluate the proposed method on in-the-wild FER datasets including RAF-DB, AffectNet, SFEW and FERPlus. The comparable results show that our method is promising to improve classification performance and achieves equivalent performance compared with state-of-the-art methods. Also, we demonstrate the strong generalization performance on realistic occlusion and pose variation datasets and the cross-dataset evaluation. Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Towards Semi-Supervised Deep Facial Expression Recognition with An Adaptive Confidence MarginabstractOnly parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all unlabeled data. In this paper, we learn an Adaptive Confidence Margin (Ada-CM) to fully leverage all unlabeled data for semi-supervised deep facial expression recognition. All unlabeled samples are partitioned into two subsets by comparing their confidence scores with the adaptively learned confidence margin at each training epoch: (1) subset I including samples whose confidence scores are no lower than the margin; (2) subset II including samples whose confidence scores are lower than the margin. For samples in subset I, we constrain their predictions to match pseudo labels. Meanwhile, samples in subset II participate in the feature-level contrastive objective to learn effective facial expression features. We extensively evaluate Ada-CM on four challenging datasets, showing that our method achieves state-of-the-art performance, especially surpassing fully-supervised baselines in a semi-supervised manner. Ablation study further proves the effectiveness of our method. The source code is available at https://github.com/hangyu94/Ada-CM. Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001 |
CVPR | 1 |
| 2022 | CRS-CONT: A Well-Trained General Encoder for Facial Expression AnalysisabstractExisting facial expression recognition (FER) methods train encoders with different large-scale training data for specific FER applications. In this paper, we propose a new task in this field. This task aims to pre-train a general encoder to extract any facial expression representations without fine-tuning. To tackle this task, we extend the self-supervised contrastive learning to pre-train a general encoder for facial expression analysis. To be specific, given a batch of facial expressions, some positive and negative pairs are firstly constructed based on coarse-grained labels and a FER-specified data augmentation strategy. Secondly, we propose the coarse-contrastive (CRS-CONT) learning, where the features of positive pairs are pulled together, while pushed away from the features of negative pairs. Moreover, one key event is that the excessive constraint on the coarse-grained feature distribution will affect fine-grained FER applications. To address this, a weight vector is designed to control the optimization of the CRS-CONT learning. As a result, a well-trained general encoder with frozen weights could preferably adapt to different facial expressions and realize the linear evaluation on any target datasets. Extensive experiments on both in- the-wild and in- the-lab FER datasets show that our method provides superior or comparable performance against state-of-the-art FER methods, especially on unseen facial expressions and cross-dataset evaluation. We hope that this work will help to reduce the training burden and develop a new solution against the fully-supervised feature learning with fine-grained labels. Code and the general encoder will be publicly available at https://github.com/hangyu94/CRS-CONT. Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | LBAN-IL: A novel method of high discriminative representation for facial expression recognition
Hangyu Li 0001, Nannan Wang 0001, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2021 | Adaptively Learning Facial Expression Representation via C-F Labels and DistillationabstractFacial expression recognition is of significant importance in criminal investigation and digital entertainment. Under unconstrained conditions, existing expression datasets are highly class-imbalanced, and the similarity between expressions is high. Previous methods tend to improve the performance of facial expression recognition through deeper or wider network structures, resulting in increased storage and computing costs. In this paper, we propose a new adaptive supervised objective named AdaReg loss, re-weighting category importance coefficients to address this class imbalance and increasing the discrimination power of expression representations. Inspired by human beings' cognitive mode, an innovative coarse-fine (C-F) labels strategy is designed to guide the model from easy to difficult to classify highly similar representations. On this basis, we propose a novel training framework named the emotional education mechanism (EEM) to transfer knowledge, composed of a knowledgeable teacher network (KTN) and a self-taught student network (STSN). Specifically, KTN integrates the outputs of coarse and fine streams, learning expression representations from easy to difficult. Under the supervision of the pre-trained KTN and existing learning experience, STSN can maximize the potential performance and compress the original KTN. Extensive experiments on public benchmarks demonstrate that the proposed method achieves superior performance compared to current state-of-the-art frameworks with 88.07% on RAF-DB, 63.97% on AffectNet and 90.49% on FERPlus. Hangyu Li 0001, Nannan Wang 0001, Xinpeng Ding, Xi Yang 0011, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |