VLDB 2026 Research / reviewers in the wild / expert
Fang Liu 0030
dblp:67/5807-30
· DBLP profile ↗
14ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-7611-7876ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A dual uncertainty-aware fusion framework for face expression recognition in the wildabstractFacial Expression Recognition(FER) is a key task in the broader landscape of affective computing and human-computer interaction, enabling machines to interpret human emotions. To better learn discriminative features under complex facial variations, recent FER research has increasingly adopted multi-branch fusion architectures that aim to capture complementary features from diverse perspectives. However, existing multi-branch fusion strategies, including static weighting, simple concatenation, or uncertainty-aware modeling, lack the capacity to comprehensively capture and reconcile the reliability variations across both individual instances and structural branches. To overcome these limitations, we propose a novel multi-branch fusion strategy, named Dual Uncertainty-Aware Fusion Framework(DUAFF), which improves the discriminability of integrated features by simultaneously modeling instance-wise uncertainty and inter-branch correlations. Specifically, the proposed method comprises two complementary modules: Instance-Discrepant Uncertainty-Aware Fusion Module (ID-UAFM) and Branch-Discrepant Uncertainty-Aware Fusion Module (BD-UAFM). ID-UAFM is introduced to perform channel-wise entropy analysis between semantically distinct samples to estimate instance-level uncertainty, enabling selective channel-wise fusion that emphasizes reliable representations while suppressing uncertain responses. BD-UAFM is further proposed to capture structural uncertainty by evaluating the relative reliability of features across multiple branches and adaptively weighting their contributions based on inter-branch discrepancies. Experimental results demonstrate that the proposed DUAFF consistently outperforms POSTER across three benchmark datasets, achieving accuracy improvements of 0.23 % on RAF-DB, 0.69 % on FER2013, and 0.29 % on AffectNet (7-class), thereby confirming its effectiveness in enhancing the reliability and discriminability of facial representations. Wenfeng Jiang, Lin Wang 0004, Fang Liu 0030, Chunmei Qing, Xiaofen Xing, Xiangmin Xu 0001, Weiquan Fan, Zhanpeng Jin |
Expert Syst. Appl. | 4 |
| 2026 | PEGCL: Pseudo-Entropy Guided Complementary Learning for Robust Facial Expression Recognition Under Label NoiseabstractFacial Expression Recognition (FER) has recently plays a crucial role in advancing human-computer interaction systems, aiming to understand users' inner states and underlying intentions. However, FER in real-world scenarios remains challenging due to significant label noise, caused by ambiguous facial expressions in low-quality images and annotation bias. To tackle this issue, this paper proposes a novel framework, Pseudo-Entropy Guided Complementary Learning (PEGCL), designed to robustly handle noisy labels by leveraging complementary information, which trains networks using all complementary labels defined as “facial expression images that do not belong to complementary emotion labels.” This approach effectively utilizes non-target emotion labels to mitigate the impact of label noise, rather than relying solely on annotated emotion labels. Specifically, the proposed PEGCL framework consists of three components: logit normalization to stabilize predicted probabilities and prevent gradient explosions, transformed complementary learning to redistribute the optimization focus across complementary categories by leveraging pseudo-entropy guided, and random complementary label dropping to dynamically exclude subsets of complementary labels, enhancing generalization and preventing overfitting. These components collectively ensure robust and efficient optimization under noisy label conditions. Importantly, the proposed PEGCL does not require explicit noise estimation or complex label correction mechanisms, making it a simple and effective solution for real-world FER tasks. Extensive experiments on benchmark FER datasets demonstrate that PEGCL consistently outperforms existing methods, achieving the state-of-the-art robustness against label noise while maintaining high classification accuracy. Lin Wang 0004, Dan Liao, Fang Liu 0030, Xiangmin Xu 0001, Kailing Guo, Zhanpeng Jin |
IEEE Trans. Multim. | 3 |
| 2025 | Multimodal speech emotion recognition via dynamic multilevel contrastive loss under local enhancement network
Weiquan Fan, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing |
Expert Syst. Appl. | 3 |
| 2025 | rPPG-TFCL: Time-frequency consistency learning for robust remote physiological measurement
Kailing Guo, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004, Xiangmin Xu 0001, Zhanpeng Jin |
Knowl. Based Syst. | 3 |
| 2025 | CSE-GResNet: A Simple and Highly Efficient Network for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted extensive attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources. Hence, achieving competitive performance while maintaining the model efficiency is still a huge challenge. To tackle these issues, we propose a highly lightweight yet effective Channel Shift-Enhancement Gabor-ResNet (CSE-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the robust GResNet as our backbone with limited memory cost. Furthermore, we propose extremely efficient Channel-Shift Module and Channel-Enhancement Module to insert in the GResNet in cascade. They are adopted to obtain and aggregate the facial informative representation from adjacent channels for extracting the subtle facial expression representation. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that the proposed CSE-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost. Shaoping Jiang, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001, Lin Wang 0004, Kailing Guo |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Compact Model Training by Low-Rank Projection With Energy TransferabstractLow-rankness plays an important role in traditional machine learning but is not so popular in deep learning. Most previous low-rank network compression methods compress networks by approximating pretrained models and retraining. However, the optimal solution in the Euclidean space may be quite different from the one with low-rank constraint. A well-pretrained model is not a good initialization for the model with low-rank constraints. Thus, the performance of a low-rank compressed network degrades significantly. Compared with other network compression methods such as pruning, low-rank methods attract less attention in recent years. In this article, we devise a new training method, low-rank projection with energy transfer (LRPET), that trains low-rank compressed networks from scratch and achieves competitive performance. We propose to alternately perform stochastic gradient descent training and projection of each weight matrix onto the corresponding low-rank manifold. Compared to retraining on the compact model, this enables full utilization of model capacity since solution space is relaxed back to Euclidean space after projection. The matrix energy (the sum of squares of singular values) reduction caused by projection is compensated by energy transfer. We uniformly transfer the energy of the pruned singular values to the remaining ones. We theoretically show that energy transfer eases the trend of gradient vanishing caused by projection. In modern networks, a batch normalization (BN) layer can be merged into the previous convolution layer for inference, thereby influencing the optimal low-rank approximation (LRA) of the previous layer. We propose BN rectification to cut off its effect on the optimal LRA, which further improves the performance. Comprehensive experiments on CIFAR-10 and ImageNet have justified that our method is superior to other low-rank compression methods and also outperforms recent state-of-the-art pruning methods. For object detection and semantic segmentation, our method still achieves good compression results. In addition, we combine LRPET with quantization and hashing methods and achieve even better compression than the original single method. We further apply it in Transformer-based models to demonstrate its transferability. Our code is available at https://github.com/BZQLin/LRPET. Kailing Guo, Zhenquan Lin, Canyang Chen, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Cross Range Quantization for Network CompressionabstractQuantization is effective in reducing model memory and accelerating inference, and is an important way to deploy deep neural networks on mobile smart devices. However, current popular learnable quantization functions often take simple truncation operations for values beyond the quantization range. We find that the truncation operation cause information loss and and restricts the update of values out of quantization range. To address this problem, we propose a universal cross range quantization (CRQ) method to reduce the information loss caused by the conventional truncation operation. CRQ splits the values exceeding the quantization range into two parts for separate quantization, and thus retain the information efficiently for performance improvement. In addition, we define a new metric named performance improvement efficiency (PIE) to measure the relationship between increased computation and performance improvement. Experiments on public benchmark image classification datasets show that CRQ achieves a significant accuracy gain with only a small increase in computation compared to the original learnable quantization method, and also outperforms many sophisticated designed state-of-the-art quantization methods in terms of accuracy and PIE. Yicai Yang, Xiaofen Xing, Kailing Guo, Xiangmin Xu 0001, Fang Liu 0030 |
IJCNN | 6 |
| 2023 | Reflective Learning With Label NoiseabstractLearning with noisy labels is one of the most challenging tasks in semi-supervised learning, and it poses significant problems in various practical applications. In the network learning process, the noisy labels concealed in the training dataset are easy to remember, resulting in poor generalization performance. To overcome this problem, inspired by the correction ability of humans – “think and learn from the past,” an end-to-end dynamic correction framework against label noise called Reflective Learning (RL) is proposed. This solution incorporates valuable knowledge from the past network training process to assist in correcting noisy labels. Specifically, during network training, a dynamic iterative function is implemented to adaptively correct noisy labels by employing the network’s predictive distribution information of all training epochs. This dynamic iterative function takes the form of a Standard Normal Distribution function to effectively match the changes of noisy label correction information contained in the network’s predictive probabilities. The proposed method is general and applicable to any backbone network and different types of noise without auxiliary information. Experiments are conducted on datasets with synthetic and real-world label noise datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and Clothing1M. They demonstrate that the proposed method is superior to the state-of-the-art results. Lin Wang 0004, Xiangmin Xu 0001, Kailing Guo, Bolun Cai, Fang Liu 0030 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | CS-GResNet: A Simple and Highly Efficient Network for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources and memory consumption. Hence, achieving promising performance while maintaining the efficiency of models is still a huge challenge. In this work, we propose a highly efficient Channel-Shift Gabor-ResNet (CS-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the significant GResNet as our backbone with limited memory cost. Furthermore, we adopt an extremely simple yet effective Channel-Shift Module inserted into the GResNet to obtain the facial informative representation via facilitating information exchanged among neighboring channels. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that our proposed CS-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost. Codes are available at https://github.com/jsesr/CS-GResNet-PyTorch. Shaoping Jiang, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004 |
ICASSP | 3 |
| 2022 | Simple-action-guided dictionary learning for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Xiaofen Xing, Kailing Guo, Lin Wang 0004 |
Neurocomputing | 1 |
| 2021 | Two-stream Gabor-AGraph Convolutional Networks for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted much attention in computer vision. However, existing methods mostly focus on the texture information of faces and overlook their inherent topological features. Hence more informative and significant contents are ignored for expression recognition. In this work, we propose a Two-stream Gabor-AGraph Convolutional Network (2s-GAGCN) to exploit the facial texture and topological features simultaneously. The Gabor Stream and the Attention-Graph (AGraph) Stream are respectively introduced to capture the salient visual properties and discriminative landmark features of faces. In particular, we adopt a flexible node attention mechanism in AGraph Stream through utilizing global and local information to enhance the potential relationships among landmarks. Furthermore, a novel landmark feature descriptor is proposed to alleviate the redundant topological features, which shows promising improvement for the recognition accuracy. We conduct extensive experiments on two wild datasets: RAF-DB and SFEW. The results show that the proposed 2s-GAGCN achieves superior performance against the state-of-the-art methods. Shaoping Jiang, Xiangmin Xu 0001, Xiaofen Xing, Lin Wang 0004, Fang Liu 0030 |
FG | 5 |
| 2021 | Two-stream Global-Guided Attention Network for Facial Expression RecognitionabstractFacial expression recognition (FER) in the wild is an important yet challenging problem because of uncontrolled conditions, such as occlusions, pose, and illumination. Most existing methods utilize the global and local information, but ignore the potential correlation between the global and local faces. In this paper, we propose a two-stream global-guided attention network (TGGAN) for FER in the wild. To further exploit the complementary relationship between local and global, we design a global-guided attention module (GGAM). Especially, a global guidance mechanism is proposed in GGAM to utilize the global information to guide the capture of key local features. Furthermore, inspired by the transformer, a self-attention mechanism is introduced in GGAM to emphasize salient face regions and fully integrate features extracted from global and local streams. We validate the proposed TGGAN on two wild datasets (FERPlus, RAF -DB) and further conduct experiments on their test subsets of occlusion and multi-poses. Extensive experiments show that the proposed TGGAN achieves superior performance against the state-of-the-art methods. Yaoli Wen, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004 |
FG | 3 |
| 2020 | Exploring privileged information from simple actions for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Tong Zhang 0015, Kailing Guo, Lin Wang 0004 |
Neurocomputing | 1 |
| 2016 | Simple to Complex Transfer Learning for Action RecognitionabstractRecognizing complex human actions is very challenging, since training a robust learning model requires a large amount of labeled data, which is difficult to acquire. Considering that each complex action is composed of a sequence of simple actions which can be easily obtained from existing data sets, this paper presents a simple to complex action transfer learning model (SCA-TLM) for complex human action recognition. SCA-TLM improves the performance of complex action recognition by leveraging the abundant labeled simple actions. In particular, it optimizes the weight parameters, enabling the complex actions to be learned to be reconstructed by simple actions. The optimal reconstruct coefficients are acquired by minimizing the objective function, and the target weight parameters are then represented as a combination of source weight parameters. The main advantage of the proposed SCA-TLM compared with existing approaches is that we exploit simple actions to recognize complex actions instead of only using complex actions as training samples. To validate the proposed SCA-TLM, we conduct extensive experiments on two well-known complex action data sets: 1) Olympic Sports data set and 2) UCF50 data set. The results show the effectiveness of the proposed SCA-TLM for complex action recognition. Fang Liu 0030, Xiangmin Xu 0001, Shuoyang Qiu, Chunmei Qing, Dacheng Tao |
IEEE Trans. Image Process. | 1 |