Rui Wang 0156

dblp:06/2293-156 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-1614-9884ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Facial Expression Recognition With Vision Transformer Using Fused Shifted Windows
abstract
Facial expressions contain massive affective information. Previous methods have focused on using diverse models based on CNN or Transformer to handle the facial expression recognition(FER) task. However, most of them treat the FER task as a general image classification task and neglect the impact of regions of interest (ROIs) for the performance of FER. To verify the influence of different ROIs, in this paper, we propose a vision Transformer based onFusedShiftedwindows (FSwin), called FSwin Transformer. The semantic information of the face, obtained by ROIs, guides the FSwin Transformer to focus more on the key regions. The fused shifted windows enable the model to perform global semantic interactions, allowing it to concentrate on key regions without losing the topological structural information of the entire face. Additionally, learnable parameters are introduced to learn the feature weights expressed by each ROI, helping the model dynamically adjust the attention distribution. We have conducted controlled experiments to quantitatively verify the impact of ROIs. And the experimental results show that with the increase of the number of ROIs, the accuracy of FER is significantly improved, demonstrating that the key ROIs play an important role in feature extraction. The results on Jaffe, CK+, FER2013, AffectNet have reached 99.8%, 99.0%, 74.8%, and 68.9%, respectively, which all set new state-of-the-art. Extensive cross-dataset experiments also show that the FSwin Transformer has good generalization ability, proving that our proposed model has a beneficial effect on FER tasks.
Xiao Sun 0003, Rui Wang 0156, Shaokai Chen, Meng Wang 0001
IEEE Trans. Affect. Comput.2
2024 MCAN: An Efficient Multi-Task Network for Facial Expression Analysis
abstract
With the development of artificial intelligence, artificial intelligence technology is widely used in robots. For example, emotional computing robots need to be able to complete the function of facial expression analysis. The multi-task deep learning model MCAN proposed in this paper is designed to complete facial expression analysis for robots which can predict discrete expressions and dimensional measures. The class center loss function proposed in this paper can increase the inter-class distance while reducing the intra-class distance. In addition, multi-head attention network was improved to focus on different areas of the input image. Finally, feature pyramid network allows information exchange and fusion between feature maps at different levels. Experiments have proven that MCAN achieves excellent results on both the AffectNet dataset and the AFEW-VA dataset, and also speeds up inference.
Rui Wang 0156, Qingjian Ni, Xiao Sun 0003
CSCWD2
2024 WR-Former: Vision Transformer with Weight Reallocation Module for Robust Facial Expression Recognition
abstract
Facial expression recognition(FER) is an exceedingly challenging task in the field of computer vision, primarily due to variant head poses, occlusions and illumination conditions. Previous methods focus on extracting more facial features to improve the performance on the FER benchmarks. However, most of them ignore the weights of the key regions. To address the above problem, this paper proposes a vision Transformer with weight reallocation module termed as WR-Former. The weight reallocation module utilizes a hybrid multi-head self-attention and can dynamically iterate during the training process. It guides the model to reallocate tokens based on the importance of different facial regions. Besides, a simple yet effective data augmentation method is proposed to expand training samples and improve the robustness. Experimental results and visualizations demonstrate that our WR-Former outperforms the previous state-of-the-art methods and has the ability to better extract the features conveyed by the key facial regions.
Rui Wang 0156, Xiao Sun 0003
CSCWD1
2023 Dynamic Facial Expression Recognition Based on Vision Transformer with Deformable Module
abstract
Facial expressions convey a great deal of information during human emotional interaction. However, due to the potential for various types of in-the-wild interferences, such as occlusions and variant head poses, dynamic facial expression recognition (DFER) has been a desperately complicated task. Previous methods focus on applying more robust models to extract the spatial-temporal features but ignore the impact of the key features of the regions of interest (ROIs). This inhibits further improvement of recognition accuracy. In this paper, we propose a 3D vision Transformer with a deformable module termed 3D-DSwin Transformer to guide our model to capture more discriminative features. The deformable module can gradually shift the deformable points to guide our model to pay more attention to the ROIs. A simple yet effective video augmentation method is proposed to expand the number of training samples and avoid overfitting. Visualizations and extensive experimental results demonstrate that our proposed 3D-DSwin Transformer has the ability to obtain the key feature maps, and outperforms the previous state-of-the-art methods on both the FERV39k and DFEW benchmarks.
Rui Wang 0156, Xiao Sun 0003
SMC1