Weicheng Xie 0001

dblp:28/6098-1 · also Wei-Cheng Xie 0001 · DBLP profile ↗
← Back
64ranked-venue papers
18as first author
46since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 9 first-author · 36 since 2021Artificial intelligence and machine learning · 40 · 8 first-author · 29 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explainable Depression Assessment from Face Videos by Weakly Supervised Learning
abstract
Existing video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanced deep learning models, they typically fail to explicitly consider the varied importance of depression-related behavioural cues across different video segments, i.e., segments within one video may contain behaviours reflecting varying levels of depression. Underestimating segment-level variations can obscure the detection of facial behaviour cues associated with depression, thereby undermining the accuracy and interpretability of video-based depression detection systems. In this paper, we propose a novel video-based ADA approach that specifically identifies and differentiates video segments that exhibit depression-related facial behaviours across varying temporal durations, providing clear insights into how each segment contributes to the video-level depression prediction. To achieve this, a novel weakly supervised strategy is proposed to compare segment-level behaviours with video-level depression label, enabling the model to assign depression-relevant scores to multiple temporal scale video segments and attend selectively to those most indicative of depressive states. Extensive experiments on the AVEC 2013 and AVEC 2014 face video depression datasets demonstrate the effectiveness of our approach.
Rongfan Liao, Xiangyu Kong 0001, Shiqing Tang, Changzeng Fu, Weicheng Xie 0001, Lu Liu 0001, Siyang Song
AAAI6
2026 PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
abstract
In recent years, face anti-spoofing (FAS) has made notable progress in multimodal fusion, cross-domain generalization, and interpretability. With the development of large language models and reinforcement learning (RL), strategy-based training paradigms offer new opportunities for jointly modeling multimodality, generalization, and interpretability. However, compared to unimodal reasoning, multimodal reasoning introduces more complex logic, such as accurate feature representation and cross-modal verification, which significantly increases reasoning complexity and labeling difficulty. Due to the lack of high-quality annotations in existing multimodal FAS datasets, directly applying RL strategies is sub-optimal, hindering robust multimodal reasoning. In this paper, we find two key issues of supervised fine-tuning combined with reinforcement learning (SFT+RL) paradigms in multimodal FAS reasoning: 1) limited multimodal reasoning paths not only hinder the full utilization of multimodal information but also constrain the model’s exploration space after SFT, thereby affecting the effectiveness of subsequent RL; and 2) the mismatch between single-task supervision and the diversity of multimodal reasoning paths leads to reasoning confusion, where models may exploit shortcuts by directly mapping input images to answers, bypassing the intended reasoning process. These issues further increase the complexity of multimodal reasoning and hinder the effective application of RL strategies. To address these challenges, we propose the PA-FAS framework with a reasoning path enhancement strategy for high-quality extended reasoning sequences construction based on limited annotated data to enrich the reasoning paths and alleviate exploration constraints. Additionally, we introduce an answer shuffling mechanism during SFT for comprehensive multimodal analysis rather than mining superficial cues, thus encouraging deeper reasoning and avoiding shortcut learning. Our method significantly improves multimodal reasoning accuracy and generalization, and successfully unifies multimodal fusion, cross-domain generalization, and interpretability towards trustworthy multimodal FAS.
Xun Lin, Yong Xu 0001, Weicheng Xie 0001, Zitong Yu
AAAI4
2026 SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
abstract
Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand the skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visual-motion knowledge for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods.
Qilang Ye, Yu Zhou 0015, Jie Zhang 0081, Xuanming Guo, Mingkui Tan, Weicheng Xie 0001, Yue Sun 0001, Tao Tan 0002, Xiaochen Yuan, Ghada Khoriba, Zitong Yu
AAAI8
2026 EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
abstract
Abstract Facial expression recognition (FER) has emerged as an important research topic in recent years. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. Nonetheless, directly applying pre-trained MLLMs to FER remains challenging due to insufficient instruction datasets and the inability of vision encoders to extract fine-grained facial information. Our zero-shot evaluations of existing open-source MLLMs on FER reveal a significant performance gap compared to GPT-4V and state-of-the-art supervised methods. In this paper, we aim to enhance MLLMs’ capabilities in understanding facial expressions. We first introduce a facial expression recognition instruction dataset ( FERID ), which has 376k category instructions and 339k conversational instructions. We then propose a novel MLLM, named EMO-LLaMA , which incorporates facial priors from a pretrained facial analysis network to enhance its understanding of human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Furthermore, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across diverse human groups. Extensive experiments show that EMO-LLaMA achieves results comparable to or competitive with SOTA on both static and dynamic FER datasets. The instruction dataset and code will be available at https://github.com/xxtars/EMO-LLaMA .
Bohao Xing, Zitong Yu, Xin Liu 0012, Kaishen Yuan, Qilang Ye, Weicheng Xie 0001, Huanjing Yue, Heikki Kälviäinen
Int. J. Comput. Vis.6
2026 Compressed video-driven multimodal modeling and interaction for dynamic expression recognition
Weicheng Xie 0001, Junliang Zhang, Haijian Liang, LinLin Shen, Zhihui Lai 0001, Siyang Song, Zitong Yu
Knowl. Based Syst.1
2026 Multi-Granularity Facial Emotional Representation With Unlabeled Data and Textual Supervision
abstract
Facial expressions (FEs) and action units (AUs) are facial emotional representations at different levels of granularity. In the past, recognizing them has often been treated as two separate tasks. There are also some methods that use the knowledge of one to aid in recognizing the other, but currently, unified models capable of recognizing both FEs and AUs simultaneously remain rare. In this paper, we construct a unified model with strong generalization capability to jointly perform facial expression recognition (FER) and action unit detection (AUD). Considering the extremely limited training samples annotated with both FEs and AUs, we introduce a large amount of unlabeled facial data from the wild. We carefully design category-specific confidence margins and leverage the correspondences between FEs and AUs to assign credible pseudo-labels to the unlabeled facial data. Furthermore, we incorporate semantically richer textual descriptions as supervision and refine them through visual perception, leveraging the inherent correlations between AUs and between FEs and AUs to enhance their precision. Extensive experiments demonstrate the superiority of the proposed method from various perspectives, including a unified zero-shot benchmark for exploring the model's comprehensive generalization capability to recognize facial emotional representations across multiple datasets, as well as within-domain and cross-domain evaluations after fine-tuning. The code for the proposed method is available at https://github.com/yuankaishen2001/MGFER.
Kaishen Yuan, Zitong Yu, Xin Liu 0012, Bohao Xing, Yuting Zhang 0008, Weicheng Xie 0001, LinLin Shen, Björn W. Schuller
IEEE Trans. Image Process.6
2025 MSAmba: Exploring Multimodal Sentiment Analysis with State Space Models
abstract
Multimodal sentiment analysis, which learns a model to process multiple modalities simultaneously and predict a sentiment value, is an important area of affective computing. Modeling sequential intra-modal information and enhancing cross-modal interactions are crucial to multimodal sentiment analysis. In this paper, we propose MSAmba, a novel hybrid Mamba-based architecture for multimodal sentiment analysis, consisting of two core blocks: Intra-Modal Sequential Mamba (ISM) block and Cross-Modal Hybrid Mamba (CHM) block, to comprehensively address the above-mentioned challenges with hybrid state space models. Firstly, the ISM block models the sequential information within each modality in a bi-directional manner with the assistance of global information. Subsequently, the CHM blocks explicitly model centralized cross-modal interaction with a hybrid combination of Mamba and attention mechanism to facilitate information fusion across modalities. Finally, joint learning of the intra-modal tokens and cross-modal tokens is utilized to predict the sentiment values. This paper serves as one of the pioneering works to unravel the outstanding performances and great research potential of Mamba-based methods in the task of multimodal sentiment analysis. Experiments on CMU-MOSI, CMU-MOSEI and CH-SIMS demonstrate the superior performance of the proposed MSAmba over prior Transformer-based and CNN-based methods.
Xilin He, Haijian Liang, Boyi Peng, Weicheng Xie 0001, Muhammad Haris Khan, Siyang Song, Zitong Yu
AAAI4
2025 CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing
abstract
For efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternatively, inpainting methods can edit the target image region while preserving external areas. However, current inpainting methods still suffer from the generation misalignment with facial attributes description and the loss of facial skin details. To address these challenges, (i) a novel data utilization strategy is introduced to construct datasets consisting of attribute-text-image triples from a data-driven perspective, (ii) a Causality-Aware Condition Adapter is proposed to enhance the contextual causality modeling of specific details, which encodes the skin details from the original image while preventing conflicts between these cues and textual conditions. In addition, a Skin Transition Frequency Guidance technique is introduced for the local modeling of contextual causality via sampling guidance driven by low-frequency alignment. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method in boosting both fidelity and editability for localized attribute editing. Our codes will be made publicly available.
Xiaole Xian, Xilin He, Zenghao Niu, Junliang Zhang, Weicheng Xie 0001, Siyang Song, Zitong Yu, LinLin Shen
AAAI5
2025 PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation
abstract
In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strategies lead the training of previous automatic facial reaction generation (FRG) models to not only suffer from the 'one-to-many mapping' problem, but also fail to comprehensively consider the quality of the generated facial reactions. Besides, none of them considered such personalised behaviour patterns in generating facial reactions. In this paper, we propose the first adversarial FRG model training strategy which jointly learns appropriateness and realism discriminators to provide comprehensive task-specific supervision for training the target facial reaction generators, and reformulates the 'one-to-many (facial reactions) mapping' training problem as a 'one-to-one (distribution) mapping' training task, i.e., the FRG model is trained to output a distribution representing multiple appropriate/plausible facial reaction from each input human behaviour. In addition, our approach also serves as the first offline FRG approach that considers personalised behaviour patterns in generating of target individuals' facial reactions. Experiments show that our PerReactor not only largely outperformed all existing offline solutions for generating more appropriate, diverse and realistic facial reactions, but also is the first approach that can effectively generate personalised appropriate facial reactions.
Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, Xilin He, Lu Liu 0001, LinLin Shen, Wei Zhang 0243, Hatice Gunes, Siyang Song
AAAI3
2025 DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face Synthesis
abstract
Accurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code is available at https://github.com/CVI-SZU/DEGSTalk.
Kaijun Deng, Dezhi Zheng, Jindong Xie, Jinbao Wang 0001, Weicheng Xie 0001, LinLin Shen, Siyang Song
ICASSP5
2025 Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-Spoofing
abstract
In the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain shifts. Our research introduces a Mixture of Experts (MoE) model to address these issues effectively. We identified three limitations in traditional MoE approaches to multimodal FAS: (1) Coarse-grained experts’ inability to capture nuanced spoofing indicators; (2) Gated networks’ susceptibility to input noise affecting decision-making; (3) MoE’s sensitivity to prompt tokens leading to overfitting with conventional learning methods. To mitigate these, we propose the Bypass Isolated Gating MoE (BIG-MoE) framework, featuring: (1) Fine-grained experts for enhanced detection of subtle spoofing cues; (2) An isolation gating mechanism to counteract input noise; (3) A novel differential convolutional prompt bypass enriching the gating network with critical local features, thereby improving perceptual capabilities. Extensive experiments on four benchmark datasets demonstrate significant generalization performance improvement in multimodal FAS task. The code is released at https://github.com/murInJ/BIG-MoE.
Zitong Yu, Xun Lin, Weicheng Xie 0001, LinLin Shen
ICASSP4
2025 SynFER: Towards Boosting Facial Expression Recognition With Synthetic Data
abstract
Facial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foundation models, that rely on large-scale data for optimal performance. To tackle the overarching and complex challenge, instead of introducing a new large-scale dataset, we introduce SynFER (Synthesis of Facial Expressions with Refined Control), a novel synthetic framework for synthesizing facial expression image data based on high-level textual descriptions as well as more fine-grained and precise control through facial action units. To ensure the quality and reliability of the synthetic data, we propose a semantic guidance technique to steer the generation process and a pseudo-label generator to help rectify the facial expression labels for the synthetic images. To demonstrate the generation fidelity and the effectiveness of the synthetic data from SynFER, we conduct extensive experiments on representation learning using both synthetic data and real-world data. Results validate the efficacy of our approach and the synthetic data. Notably, our approach achieves a 67.23% classification accuracy on AffectNet when training solely with synthetic data equivalent to the AffectNet training set size, which increases to 69.84% when scaling up to five times the original size. Code is available here.
Xilin He, Xiaole Xian, Bing Li 0024, Muhammad Haris Khan, ZongYuan Ge, Weicheng Xie 0001, Siyang Song, LinLin Shen, Bernard Ghanem, Xiangyu Yue 0001
ICCV7
2025 Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
abstract
Adversarial attacks present a critical challenge to deep neural networks' robustness, particularly in transfer scenarios across different model architectures. However, the transferability of adversarial attacks faces a fundamental dilemma between Exploitation (maximizing attack potency) and Exploration (enhancing cross-model generalization). Traditional momentum-based methods over-prioritize Exploitation, i.e., higher loss maxima for attack potency but weakened generalization (narrow loss surface). Conversely, recent methods with inner-iteration sampling over-prioritize Exploration, i.e., flatter loss surfaces for cross-model generalization but weakened attack potency (suboptimal local maxima). To resolve this dilemma, we propose a simple yet effective Gradient-Guided Sampling (GGS), which harmonizes both objectives through guiding sampling along the gradient ascent direction to improve both sampling efficiency and stability. Specifically, based on MI-FGSM, GGS introduces inner-iteration random sampling and guides the sampling direction using the gradient from the previous inner-iteration (the sampling's magnitude is determined by a random distribution). This mechanism encourages adversarial examples to reside in balanced regions with both flatness for cross-model generalization and higher local maxima for strong attack potency. Comprehensive experiments across multiple DNN architectures and multimodal large language models (MLLMs) demonstrate the superiority of our method over state-of-the-art transfer attacks. Code is made available at https://github.com/anuin-cat/GGS.
Zenghao Niu, Weicheng Xie 0001, Siyang Song, Zitong Yu, Feng Liu 0013, LinLin Shen
ICCV2
2025 DeeperForward: Enhanced Forward-Forward Training for Deeper and Better Performance
abstract
While backpropagation effectively trains models, it presents challenges related to bio-plausibility, resulting in high memory demands and limited parallelism. Recently, Hinton (2022) proposed the Forward-Forward (FF) algorithm for high-parallel local updates. FF leverages squared sums as the local update target, termed goodness, and decouples goodness by normalizing the vector length to extract new features. However, this design encounters issues with feature scaling and deactivated neurons, limiting its application mainly to shallow networks. This paper proposes a novel goodness design utilizing **layer normalization** and **mean goodness** to overcome these challenges, demonstrating performance improvements even in 17-layer CNNs. Experiments on CIFAR-10, MNIST, and Fashion-MNIST show significant advantages over existing FF-based algorithms, highlighting the potential of FF in deep models. Furthermore, the model parallel strategy is proposed to achieve highly efficient training based on the property of local updates.
Yang Zhang 0012, Weizhao He, Jiajun Wen 0001, LinLin Shen, Weicheng Xie 0001
ICLR6
2025 TRRG: Towards Truthful Radiology Report Generation With Cross-Modal Disease Clue Enhanced Large Language Models
Yue Sun 0001, Tao Tan 0002, Chao Hao, Yawen Cui, Xinqi Su, Weicheng Xie 0001, LinLin Shen, Zitong Yu
MICCAI (7)7
2025 Smooth Online Multiple Appropriate Facial Reaction Generation
abstract
In dyadic interactions, facial reactions are crucial for conveying an individuals' responses to their conversational partners. Individuals may exhibit varied but appropriate facial reactions (AFRs) when perceiving the same behavioral expression. Although some recent methods can already respond multiple appropriate facial reactions to the given human speaker behaviors, the AFRs generated by these methods often fail to adequately preserve crucial head motions, leading to visual jitter and unnatural transitions between generated AFR segments. In this paper, we propose a novel and generic PFLPosNet framework which addresses the aforementioned problems at both pre-processing and post-processing stages, where a new pose-aware face behavior localization method PFL is introduced to retain the head pose displacement information from the source data. In addition, the framework proposes a real-time head pose adjustment method, PosNet, to ensure continuity and smoothness in the visual output of the model when using data with correct head pose displacement. Experimental results demonstrate that our approach not only generates more coherent and natural facial reaction sequences but also significantly outperforms existing online MAFRG methods in terms of continuity and smoothness. Our code is made available at https://github.com/rainforcetime/PFLPosNet.
Weicheng Xie 0001, Chunlin Yan, Siyang Song, Zitong Yu, LinLin Shen, Laizhong Cui
ACM Multimedia1
2025 Towards Robust Training via Gradient-Diversified Backpropagation
abstract
Neural networks are prone to be vulnerable to adversarial attacks and domain shifts. Adversarial-driven methods including adversarial training and adversarial augmentation, have been frequently proposed to improve the model's robustness against adversarial attacks and distribution-shifted samples. Nonetheless, recent research on adversarial attacks has cast a spotlight on the robustness lacuna against attacks targeted at deep semantic layers. Our analysis reveals that previous adversarial-driven methods tend to generate overpowering perturbations in deep semantic layers, leading to distortion of the training for these layers. This can be primarily attributed to the exclusive utilization of loss functions on the output layer for adversarial gradient generation. This inherent practice projects an excessive adversarial impact on the deep semantic layers, elevating the difficulty of training such layers. Therefore, from the standing point of relaxing the excessive perturbations in the deep semantic layer and diversifying the adversarial gradients to ensure robust training for deep semantic layers, this paper proposes a novel Stochastic Loss Integration Method (SLIM), which can be instantiated into the existing adversarial-driven methods in a plug-and-play manner. Experimental results across diverse tasks, including classification and segmentation, as well as various areas such as adversarial robustness and domain generalization, validate the effectiveness of our proposed method. Furthermore, we provide an in-depth analysis to offer a comprehensive understanding of layer-wise training involving various loss terms.
Xilin He, Qinliang Lin, Weicheng Xie 0001, Muhammad Haris Khan, Siyang Song, LinLin Shen
WACV4
2025 Self-inferring incomplete multi-view clustering
abstract
Abstract With the advantage of exploiting complementary and consensus information across multiple views, techniques for Multi‐view Clustering have attracted increasing attention in recent years. However, it is common that data on some views is not completed in real‐world applications, which brings the challenge of partial mapping between the views. To explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances, a self‐inferring incomplete multi‐view clustering algorithm is proposed. Firstly, the incomplete multi‐view data is replenished directly and exploited as variables for inferring the missing instances. And then, a feature graph constraint is united in consensus learning. Besides, a similarity graph learning method is imposed to preserve the local manifold structure. At last, the inferred instances are filled in the missing instances for learning better consensus representation in the iterative process. Extensive experiment results show that this method can improve the clustering performance compared with the state‐of‐the‐art methods.
Junjun Fan, Zeqi Ma, Jiajun Wen 0001, Zhihui Lai 0001, Weicheng Xie 0001, Wai Keung Wong
IET Comput. Vis.5
2025 SymGraphAU: Prior knowledge based symbolic graph for action unit recognition
Weicheng Xie 0001, Junliang Zhang, Siyang Song, LinLin Shen, Zitong Yu
Pattern Recognit.1
2025 Frequency Restoration and Modality Enforcement towards Resisting-corruption Multimodal Sentiment Analysis
abstract
For Multimodal Sentiment Analysis (MSA), previous methods concentrate on designing sophisticated fusion strategies and performing representation learning across heterogeneous modalities, aiming to leverage multimodal signals to detect human sentiment. However, these approaches fail to address the long-standing issue of corrupted modal details in videos, which may be caused by the challenge of the excessive loss of emotionally relevant semantics resulted from the degradation of detailed information. In this work, we aim to improve the robustness capacity of resisting corruption in MSA, by introducing a Hierarchical Frequency Restoration and Adaptive Modality Enforcement (HFR-AME) approach. The HFR-AME progressively recovers blurred detailed cues in each modality while enhancing the discriminative power of modal representations. Specifically, to reconstruct distinct frequency band features, we propose to equip the HFR module with a key component called the Frequency Multimodal UNet (FM-UNet), so as to utilize complementary modal features as conditions. This meticulous restoration process, performed from low to high frequency, facilitates the comprehensive recovery of intricate details. Meanwhile, to adaptively integrate these diverse frequency features, we introduce the AME module to enhance the beneficial modal frequencies while suppressing irrelevant ones, with the goal of strengthening the restored modal representations. Extensive experiments show our HFR-AME outperforms state-of-the-art methods on the CMU-MOSI and CMU-MOSEI datasets, improving 7-class accuracy by 0.5% and 0.6%, respectively. Further analysis also confirms its cross-lingual generalization and competitive computational efficiency. Our code is made available at https://github.com/nianhua20/HFR-AME .
Weicheng Xie 0001, Haijian Liang, Zenghao Niu, Xianxu Hou, Siyang Song, Zitong Yu, LinLin Shen
ACM Trans. Multim. Comput. Commun. Appl.1
2025 ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic Interactions
abstract
In dyadic interaction, predicting the listener's facial reactions is challenging as different reactions could be appropriate in response to the same speaker's behaviour. Previous approaches predominantly treated this task as an interpolation or fitting problem, emphasizing deterministic outcomes but ignoring the diversity and uncertainty of human facial reactions. Furthermore, these methods often failed to model short-range and long-range dependencies within the interaction context, leading to issues in the synchrony and appropriateness of the generated facial reactions. To address these limitations, this paper reformulates the task as an extrapolation or prediction problem, and proposes an novel framework (called ReactFace) to generate multiple different but appropriate facial reactions from a speaker behaviour rather than merely replicating the corresponding listener facial behaviours. Our ReactFace generates multiple different but appropriate photo-realistic human facial reactions by: (i) learning an appropriate facial reaction distribution representing multiple different but appropriate facial reactions; and (ii) synchronizing the generated facial reactions with the speaker verbal and non-verbal behaviours at each time stamp, resulting in realistic 2D facial reaction sequences. Experimental results demonstrate the effectiveness of our approach in generating multiple diverse, synchronized, and appropriate facial reactions from each speaker's behaviour. The quality of the generated facial reactions is intimately tied to the speaker's speech and facial expressions, achieved through our novel speaker-listener interaction modules.
Siyang Song, Weicheng Xie 0001, Micol Spitale, ZongYuan Ge, LinLin Shen, Hatice Gunes
IEEE Trans. Vis. Comput. Graph.3
2024 Boosting Adversarial Transferability across Model Genus by Deformation-Constrained Warping
abstract
Adversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances in attacking systems having different model genera from the surrogate model. In this paper, we propose a novel and generic attacking strategy, called Deformation-Constrained Warping Attack (DeCoWA), that can be effectively applied to cross model genus attack. Specifically, DeCoWA firstly augments input examples via an elastic deformation, namely Deformation-Constrained Warping (DeCoW), to obtain rich local details of the augmented input. To avoid severe distortion of global semantics led by random deformation, DeCoW further constrains the strength and direction of the warping transformation by a novel adaptive control strategy. Extensive experiments demonstrate that the transferable examples crafted by our DeCoWA on CNN surrogates can significantly hinder the performance of Transformers (and vice versa) on various tasks, including image classification, video action recognition, and audio recognition. Code is made available at https://github.com/LinQinLiang/DeCoWA.
Qinliang Lin, Zenghao Niu, Xilin He, Weicheng Xie 0001, Yuanbo Hou, LinLin Shen, Siyang Song
AAAI5
2024 Multi-Scale Dynamic and Hierarchical Relationship Modeling for Facial Action Units Recognition
abstract
Human facial action units (AUs) are mutually related in a hierarchical manner, as not only they are associated with each other in both spatial and temporal domains but also AUs located in the same/close facial regions show stronger relationships than those of different facial regions. While none of existing approach thoroughly model such hi-erarchical inter-dependencies among AUs, this paper proposes to comprehensively model multi-scale AU-related dynamic and hierarchical spatiotemporal relationship among AUs for their occurrences recognition. Specifically, we first propose a novel multi-scale temporal differencing network with an adaptive weighting block to explicitly capture facial dynamics across frames at different spatial scales, which specifically considers the heterogeneity of range and mag-nitude in different AUs' activation. Then, a two-stage strategy is introduced to hierarchically model the relationship among AUs based on their spatial distribution (i.e., local and cross-region AU relationship modelling). Experimental results achieved on BP4D and DISFA show that our approach is the new state-of-the-art in the field of AU occurrence recognition. Our code is publicly available at https://github.com/CVI-SZU/MDHR.
Zihan Wang 0005, Siyang Song, Songhe Deng, Weicheng Xie 0001, LinLin Shen
CVPR5
2024 MTaDCS: Moving Trace and Feature Density-Based Confidence Sample Selection Under Label Noise
Qingzheng Huang, Xilin He, Xiaole Xian, Qinliang Lin, Weicheng Xie 0001, Siyang Song, LinLin Shen, Zitong Yu
ECCV (71)5
2024 AUFormer: Vision Transformers Are Parameter-Efficient Facial Action Unit Detectors
Kaishen Yuan, Zitong Yu, Xin Liu 0012, Weicheng Xie 0001, Huanjing Yue, Jing-Yu Yang 0002
ECCV (50)4
2024 Expression-Aware Masking and Progressive Decoupling for Cross-Database Facial Expression Recognition
abstract
Cross-database facial expression recognition (CD-FER) has been widely studied due to its promising applicability in real-life situations, while the generalization performance is the main concern in this task. For improving cross-database generalization, current works frequently resort to masked auto encoder (MAE) to learn the expression representation in an unsupervised manner, and disentanglement of expression and domain features. (i) For MAE, current algorithms mainly employ random masking, and leverage the reconstruction of these masked regions to enable networks to learn the expression representation. However, these masked regions are expression-irrelevant, can not well reflect the characteristics of expression, thus are not efficient enough in representation learning. To this end, we propose an expression-aware masking in MAE to improve the learning efficiency of expression representation, by guiding MAE to mask out expression-aware regions during training. (ii) For disentanglement of expression and domain features, current algorithms realize it mainly in the deep layers. However, the coupling of these features in the shallow layers are rarely concerned, which may largely affect the disentanglement performance in deep layers. Thus, we propose a progressive decoupler to disentangle these features block by block, to use the feature disentanglement in shallow layers to facilitate that in deep layers. Extensive quantitative and qualitative results on multiple expression datasets show that our method can largely outperform the state of the arts in terms of cross-database generalization performance.
Xiaole Xian, Zihan Wang 0005, Weicheng Xie 0001, LinLin Shen
FG4
2024 Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting Generator
abstract
Recent CNN generator-based attack approaches can synthe-size unrestricted and semantically meaningful entities to the image, which are able to improve the transferability and robustness. However, such methods attack images by either synthesizing local adversarial entities, which are only suitable for attacking specific contents, or performing global attacks, which are only applicable to a specific image scale. In this paper, we propose a novel Patch Quilting Generative Adversarial Networks (PQ-GAN) to learn the first scale-free CNN generator that can be applied to attack images with arbitrary scales for various computer vision tasks. The principal investigation on transferability of the generated adversarial examples, robustness to defense frameworks, and visual quality assessment show that the proposed PQG-based attack framework outperforms the other nine state-of-the-art adversarial attack approaches when attacking the neural networks trained on two standard evaluation datasets (i.e., ImageNet and CityScapes). Our code is made available at https://github.com/XiangboGaoBarry/PQAttack.
Xiangbo Gao, Qinliang Lin, Weicheng Xie 0001, LinLin Shen, Keerthy Kusumam, Siyang Song
ICASSP4
2024 Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis is a burgeoning research area, leveraging various modalities to predict the sentiment score. Nevertheless, previous studies have disregarded the impact of noise interference on specific modal sentiments during video recording, thereby compromising the accuracy of sentiment prediction. In this paper, we propose the Guided Circular Decomposition and Cross-Modal Recombination (GCD-CMR) model, which aims to eliminate contaminated sentiment features in a fine-grained way. To achieve this, we utilize tailored global information specific to each modality to guide the circular decomposing process in the GCD module, to produce a set of sentiment prototypes. Subsequently, in the CMR module, we align cross-modal sentiment prototypes and remove the contaminated prototypes for recombination. Experimental results on two publicly available datasets demonstrate that our model surpasses state-of-the-art models, confirming the effectiveness of our proposed method. We release the code at: https://github.com/nianhua20/GCD-CMR.
Haijian Liang, Weicheng Xie 0001, Xilin He, Siyang Song, LinLin Shen
ICASSP2
2024 MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural Networks
abstract
Edges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. Although some recent Graph Neural Networks (GNNs) can process graphs containing multi-dimensional edge features, they cannot convert single-value edge graphs to multi-dimensional edge graphs during propagation. This paper proposes a generic Multi-dimensional Edge Representation Generation (MERG) layer that can be inserted into any GNNs for heterogeneous graph analysis. It assigns multi-dimensional edge features for the input single-value edge graph, describing multiple task-specific and global context-aware relationship cues between each connected node pair. Results on eight graph benchmark datasets demonstrate that inserting the MERG layer into widely-used GNNs (e.g., GatedGCN and GAT) leads to major performance improvements, resulting in state-of-the-art (SOTA) results on seven out of eight evaluated datasets. Our code is publicly available at1.
YuXin Song 0001, Aaron S. Jackson, Xi Jia, Weicheng Xie 0001, LinLin Shen, Hatice Gunes, Siyang Song
ICASSP5
2024 CLIP-Guided Bidirectional Prompt and Semantic Supervision for Dynamic Facial Expression Recognition
abstract
Due to the insufficient semantic information supervision in existing works for dynamic facial expression recognition (DFER), videos with similar facial changes but different expressions may be easily confused. Thanks to the potential textual information for semantic supervision, contrastive language-image pretraining (CLIP) model provides a new direction for DFER. However, pre-trained CLIP based on image-text pairs has difficulty in capturing temporal features in the video domain. Therefore, we propose a novel visual language model that captures and aggregates dynamic features of expressions in semantic supervision via Inter-Frame Interaction Transformer (Inter-FIT) and Multi-Scale Temporal Aggregation (MSTA). Furthermore, though prompt learning is often used in CLIP to enhance semantic supervision, previous studies have only focused on the role of textual prompts, ignoring the importance of visual prompts in facilitating the relationality between the two. Therefore, we designed a Bidirectional Enhanced Prompt (BiEhPro) to facilitate the learning of this relationality between text and visual cues in enhancing semantic supervision. Extensive experiments and ablation studies on three benchmark datasets, i.e., DFEW, FERV39K, and MAFW, validate the effectiveness of our modules and algorithm. Code is publicly available at https://github.com/JunLiangZ/CLIP-Guided-DFER.
Junliang Zhang, Xiaole Xian, Weicheng Xie 0001, LinLin Shen, Siyang Song
IJCB5
2024 PerFRDiff: Personalised Weight Editing for Multiple Appropriate Facial Reaction Generation
abstract
Human facial reactions play crucial roles in dyadic human-human interactions, where individuals (i.e., listeners) with varying cognitive process styles may display different but appropriate facial reactions in response to an identical behaviour expressed by their conversational partners. While several existing facial reaction generation approaches are capable of generating multiple appropriate facial reactions (AFRs) in response to each given human behaviour, they fail to take human's personalised cognitive process in AFRs generation. In this paper, we propose the first online personalised multiple appropriate facial reaction generation (MAFRG) approach which learns a unique personalised cognitive style from the target human listener's previous facial behaviours and represents it as a set of network weight shifts. These personalised weight shifts are then applied to edit the weights of a pre-trained generic MAFRG model, allowing the obtained personalised model to naturally mimic the target human listener's cognitive process in its reasoning for multiple AFRs generations. Experimental results show that our approach not only largely outperformed all existing approaches in generating more appropriate and diverse generic AFRs, but also serves as the first reliable personalised MAFRG solution. Our code is made available at https://github.com/xk0720/PerFRDiff.
Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, LinLin Shen, Lu Liu 0001, Hatice Gunes, Siyang Song
ACM Multimedia3
2024 Towards Combating Frequency Simplicity-biased Learning for Domain Generalization
abstract
Domain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains. Recent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, namely as frequency shortcuts, instead of semantic information, resulting in poor generalization performance. Despite previous data augmentation techniques successfully enhancing generalization performances, they intend to apply more frequency shortcuts, thereby causing hallucinations of generalization improvement. In this paper, we aim to prevent such learning behavior of applying frequency shortcuts from a data-driven perspective. Given the theoretical justification of models' biased learning behavior on different spatial frequency components, which is based on the dataset frequency properties, we argue that the learning behavior on various frequency components could be manipulated by changing the dataset statistical structure in the Fourier domain. Intuitively, as frequency shortcuts are hidden in the dominant and highly dependent frequencies of dataset structure, dynamically perturbating the over-reliance frequency components could prevent the application of frequency shortcuts. To this end, we propose two effective data augmentation modules designed to collaboratively and adaptively adjust the frequency characteristic of the dataset, aiming to dynamically influence the learning behavior of the model and ultimately serving as a strategy to mitigate shortcut learning. Our code will be made publicly available.
Xilin He, Qinliang Lin, Weicheng Xie 0001, Siyang Song, Muhammad Haris Khan, LinLin Shen
NeurIPS5
2024 Network characteristics adaption and hierarchical feature exploration for robust object recognition
Weicheng Xie 0001, Gui Wang, LinLin Shen, Zhihui Lai 0001, Siyang Song
Pattern Recognit.1
2024 Generative Imperceptible Attack With Feature Learning Bias Reduction and Multi-Scale Variance Regularization
abstract
Existing studies have shown that malicious and imperceptible adversarial samples may significantly weaken the reliability and validity of deep learning systems. Since gradient-based attack algorithms may result in higher generation latency or demand large computation overhead, generative attack methods are frequently considered. However, the effectiveness and imperceptibility are still the main concerns for these generative attacks, 1) biased feature learning may occur, i.e., these algorithms may generate undesirable feature perturbations for samples that are less likely to be successfully attacked; 2) the produced perturbation noises may be easily perceived by human eyes. To this end, we propose a novel generative attack by manipulating the feature update. The proposed algorithm has two main merits, 1) our Bias-reduced Feature Manipulation (BrFM) that differentiates the hard-to-attack (Hard2Attack) and easy-to-attack (Easy2Attack) features, can avoid the possible learning shortcut for different difficulties of features in attack process, by customizing perturbations for Hard2Attack features to make them behave oppositely to those of benign features; 2) our Multi-scale Variance Regularization (MsVR) can reduce the unnatural transitions of perturbations in mask edges and flat areas with low contrast, while simultaneously trading off a reasonable attack capacity. Extensive experiments on the datasets of Caltech-101 and Imagenette in terms of the attack success rate and four imperceptibility metrics, show the effectiveness of our attack paradigm over the related state-of-the-art generative attack methods. Our codes will be made publicly available.
Weicheng Xie 0001, Zenghao Niu, Qinliang Lin, Siyang Song, LinLin Shen
IEEE Trans. Inf. Forensics Secur.1
2024 Cross-Layer Contrastive Learning of Latent Semantics for Facial Expression Recognition
abstract
Convolutional neural networks (CNNs) have achieved significant improvement for the task of facial expression recognition. However, current training still suffers from the inconsistent learning intensities among different layers, i.e., the feature representations in the shallow layers are not sufficiently learned compared with those in deep layers. To this end, this work proposes a contrastive learning framework to align the feature semantics of shallow and deep layers, followed by an attention module for representing the multi-scale features in the weight-adaptive manner. The proposed algorithm has three main merits. First, the learning intensity, defined as the magnitude of the backpropagation gradient, of the features on the shallow layer is enhanced by cross-layer contrastive learning. Second, the latent semantics in the shallow-layer and deep-layer features are explored and aligned in the contrastive learning, and thus the fine-grained characteristics of expressions can be taken into account for the feature representation learning. Third, by integrating the multi-scale features from multiple layers with an attention module, our algorithm achieved the state-of-the-art performances, i.e. 92.21%, 89.50%, 62.82%, on three in-the-wild expression databases, i.e. RAF-DB, FERPlus, SFEW, and the second best performance, i.e. 65.29% on AffectNet dataset. Our codes will be made publicly available.
Weicheng Xie 0001, Zhibin Peng, LinLin Shen, Wenya Lu, Yang Zhang 0012, Siyang Song
IEEE Trans. Image Process.1
2023 Shift from Texture-bias to Shape-bias: Edge Deformation-based Augmentation for Robust Object Recognition
abstract
Recent studies have shown the vulnerability of CNNs under perturbation noises, which is partially caused by the reason that the well-trained CNNs are too biased toward the object texture, i.e., they make predictions mainly based on texture cues. To reduce this texture-bias, current studies resort to learning augmented samples with heavily perturbed texture to make networks be more biased toward relatively stable shape cues. However, such methods usually fail to achieve real shape-biased networks due to the insufficient diversity of the shape cues. In this paper, we propose to augment the training dataset by generating semantically meaningful shapes and samples, via a shape deformation-based online augmentation, namely as SDbOA. The samples generated by our SDbOA have two main merits. First, the augmented samples with more diverse shape variations enable networks to learn the shape cues more elaborately, which encourages the network to be shape-biased. Second, semantic-meaningful shape-augmentation samples could be produced by jointly regularizing the generator with object texture and edge-guidance soft constraint, where the edges are represented more robustly with a self information guided map to better against the noises on them. Extensive experiments under various perturbation noises demonstrate the obvious superiority of our shape-bias-motivated model over the state of the arts in terms of robustness performance. Code is available at https://github.com/C0notSilly/-ICCV-23-Edge-Deformation-based-Online-Augmentation.
Xilin He, Qinliang Lin, Weicheng Xie 0001, Siyang Song, Feng Liu 0013, LinLin Shen
ICCV4
2023 Multi Task-Based Facial Expression Synthesis with Supervision Learning and Feature Disentanglement of Image Style
abstract
Image-to-Image synthesis paradigms have been widely used for facial expression synthesis. However, current generators are apt to either produce artifacts for largely posed and non-aligned faces or unduly change the identity information like AdaIN-based generator. In this work, we suggest to use image style feature to surrogate the expression cues in the generator, and propose a multi-task learning paradigm to explore this style information via the supervision learning and feature disentanglement. While the supervision learning can make the encoded style specifically represent the expression cues and enable the generator to produce correct expression, the feature disentanglement of content and style cues enables the generator to better preserve the identity information in expression synthesis. Experimental results show that the proposed algorithm can well reduce the artifacts for the synthesis of posed and non-aligned expressions, and achieves competitive performances in terms of FID, PNSR and classification accuracy, compared with four publicly available GANs. The code and pre-trained models are available at https://github.com/lumanxi236/MTSS.
Wenya Lu, Zhibin Peng, Weicheng Xie 0001, Jiajun Wen 0001, Zhihui Lai 0001, LinLin Shen
ICIP4
2023 Joint Prediction of Audio Event and Annoyance Rating in an Urban Soundscape by Hierarchical Graph Representation Learning
abstract
Sound events in daily life carry rich information about the objective world. The composition of these sounds affects the mood of people in a soundscape. Most previous approaches only focus on classifying and detecting audio events and scenes, but may ignore their perceptual quality that may impact humans' listening mood for the environment, e.g. annoyance. To this end, this paper proposes a novel hierarchical graph representation learning (HGRL) approach which links objective audio events (AE) with subjective annoyance ratings (AR) of the soundscape perceived by humans. The hierarchical graph consists of fine-grained event (fAE) embeddings with single-class event semantics, coarse-grained event (cAE) embeddings with multi-class event semantics, and AR embeddings. Experiments show the proposed HGRL successfully integrates AE with AR for AEC and ARP tasks, while coordinating the relations between cAE and fAE and further aligning the two different grains of AE information with the AR.
Yuanbo Hou, Siyang Song, Qiaoqiao Ren, Weicheng Xie 0001, Jian Kang 0002, Wenwu Wang 0001, Dick Botteldooren
INTERSPEECH6
2023 Consistency Preservation and Feature Entropy Regularization for GAN Based Face Editing
abstract
Generative Adversarial Network (GAN) has been widely used for image-to-image translation-based facial attribute editing. Existing GAN networks are likely to generate samples with anomalies, which may be caused by the lack of consistency preservation and feature entanglement. For preserving image consistency, many studies resorted to the design of the network framework and loss functions, e.g. cycle-consistency loss. However, the generator with the cycle-consistency loss could not well preserve the attribute-irrelevant features, and its feature-level noises may possibly cause synthesis abnormalities. For feature disentanglement, previous works were devoted to mining the implicit semantics of feature spaces, while these semantics are not stable and intuitive enough. For consistency preservation, we propose a target consistency loss to complement the cycle-consistency loss, and enable the network to learn to preserve features of the image more directly. Meanwhile, we filter out outlier feature maps to reduce the synthesis abnormalities and propose a dynamic dropout to better preserve the attribute-irrelevant features. For feature disentanglement, we encode the image semantics more stably and intuitively and propose an entropy regularization to decouple these semantics to allow independent editing of different attributes. The proposed modules are general and can be easily integrated with available image-to-image-based GAN models like StarGAN, AttGAN, and STGAN. Extensive experiments on CelebA dataset show that the our strategy can largely reduce the artifacts and better preserve the subtle facial features, and thus significantly improve the facial editing performance of these mainstream GAN models, in terms of FID, PSNR and SSIM. Additional experiments on realistic expression editing show that our method outperforms StarGAN on RaFD, and achieves much better generalization performances than the three baselines on datasets of FFHQ, RaFD and LFW.
Weicheng Xie 0001, Wenya Lu, Zhibin Peng, LinLin Shen
IEEE Trans. Multim.1
2022 Frequency-driven Imperceptible Adversarial Attack on Semantic Similarity
abstract
Current adversarial attack research reveals the vulnerability of learning-based classifiers against carefully crafted perturbations. However, most existing attack methods have inherent limitations in cross-dataset generalization as they rely on a classification layer with a closed set of categories. Furthermore, the perturbations generated by these methods may appear in regions easily perceptible to the human visual system (HVS). To circumvent the former problem, we propose a novel algorithm that attacks semantic similarity on feature representations. In this way, we are able to fool classifiers without limiting attacks to a specific dataset. For imperceptibility, we introduce the low-frequency constraint to limit perturbations within high-frequency components, ensuring perceptual similarity between adversarial examples and originals. Extensive experiments on three datasets (CIFAR-10, CIFAR-100, and ImageNet-1K) and three public online platforms indicate that our attack can yield misleading and transferable adversarial examples across architectures and datasets. Additionally, visualization results and quantitative performance (in terms of four different metrics) show that the proposed algorithm generates more imperceptible perturbations than the state-of-the-art methods. Code is made available at https://github.com/LinQinLiang/SSAH-adversarial-attack.
Qinliang Lin, Weicheng Xie 0001, Bizhu Wu, Jinheng Xie, LinLin Shen
CVPR3
2022 Scene Consistency Representation Learning for Video Scene Segmentation
abstract
A long-term video, such as a movie or TV show, is composed of various scenes, each of which represents a series of shots sharing the same semantic story. Spotting the correct scene boundary from the long-term video is a challenging task, since a model must understand the storyline of the video to figure out where a scene starts and ends. To this end, we propose an effective Self-Supervised Learning (SSL) framework to learn better shot representations from unlabeled long-term videos. More specifically, we present an SSL scheme to achieve scene consistency, while exploring considerable data augmentation and shuffling methods to boost the model generalizability. Instead of explicitly learning the scene boundary features as in the previous methods, we introduce a vanilla temporal model with less inductive bias to verify the quality of the shot features. Our method achieves the state-of-the-art performance on the task of Video Scene Segmentation. Additionally, we suggest a more fair and reasonable benchmark to evaluate the performance of Video Scene Segmentation methods. The code is made available.11https://github.com/TencentYoutuResearch/SceneSegmentation-SCRL.
Haoqian Wu, Yanan Luo, Ruizhi Qiao, Bo Ren 0002, Weicheng Xie 0001, LinLin Shen
CVPR7
2022 Learning Multi-dimensional Edge Feature-based AU Relation Graph for Facial Action Unit Recognition
abstract
The activations of Facial Action Units (AUs) mutually influence one another. While the relationship between a pair of AUs can be complex and unique, existing approaches fail to specifically and explicitly represent such cues for each pair of AUs in each facial display. This paper proposes an AU relationship modelling approach that deep learns a unique graph to explicitly describe the relationship between each pair of AUs of the target facial display. Our approach first encodes each AU's activation status and its association with other AUs into a node feature. Then, it learns a pair of multi-dimensional edge features to describe multiple task-specific relationship cues between each pair of AUs. During both node and edge feature learning, our approach also considers the influence of the unique facial display on AUs' relationship by taking the full face representation as an input. Experimental results on BP4D and DISFA datasets show that both node and edge feature learning modules provide large performance improvements for CNN and transformer-based backbones, with our best systems achieving the state-of-the-art AU recognition results. Our approach not only has a strong capability in modelling relationship cues for AU recognition but also can be easily incorporated into various backbones. Our PyTorch code is made available at https://github.com/CVI-SZU/ME-GraphAU.
Siyang Song, Weicheng Xie 0001, LinLin Shen, Hatice Gunes
IJCAI3
2022 Triplet Loss With Multistage Outlier Suppression and Class-Pair Margins for Facial Expression Recognition
abstract
Deep metric based triplet loss has been widely used to enhance inter-class separability and intra-class compactness of network features. However, the margin parameters in the triplet loss for current approaches are usually fixed and not adaptive to the variations among different expression pairs. Meanwhile, outlier samples like faces with confusing expressions, occlusion and large head poses may be introduced during the selection of the hard triplets, which may deteriorate the generalization performance of the learned features for normal testing samples. In this work, a new triplet loss based on class-pair margins and multistage outlier suppression is proposed for facial expression recognition (FER). In this approach, each expression pair is assigned with an order-insensitive or two order-aware adaptive margin parameters. While expression samples with large head poses or occlusion are firstly detected and excluded, abnormal hard triplets are discarded if their feature distances do not fit the model of normal feature distance distribution. Extensive experiments on seven public benchmark expression databases show that the network using the proposed loss achieves much better accuracy than that using the original triplet loss and the network without using the proposed strategies, and the most balanced performances among state-of-the-art algorithms in the literature.
Weicheng Xie 0001, Haoqian Wu, Mengchao Bai, LinLin Shen
IEEE Trans. Circuits Syst. Video Technol.1
2021 Group-wise Inhibition based Feature Regularization for Robust Classification
abstract
The convolutional neural network (CNN) is vulnerable to degraded images with even very small variations (e.g. corrupted and adversarial samples). One of the possible reasons is that CNN pays more attention to the most discriminative regions, but ignores the auxiliary features when learning, leading to the lack of feature diversity for final judgment. In our method, we propose to dynamically suppress significant activation values of CNN by group-wise inhibition, but not fixedly or randomly handle them when training. The feature maps with different activation distribution are then processed separately to take the feature independence into account. CNN is finally guided to learn richer discriminative features hierarchically for robust classification according to the proposed regularization. Our method is comprehensively evaluated under multiple settings, including classification against corruptions, adversarial attacks and low data regime. Extensive experimental results show that the proposed method can achieve significant improvements in terms of both robustness and generalization performances, when compared with the state-of-the-art methods. Code is available at https://github.com/LinusWu/TENET_Training.
Haoqian Wu, Weicheng Xie 0001, Feng Liu 0013, LinLin Shen
ICCV3
2021 Surrogate network-based sparseness hyper-parameter optimization for deep expression recognition
Weicheng Xie 0001, Wenting Chen, LinLin Shen, Jinming Duan 0001, Meng Yang 0001
Pattern Recognit.1
2021 Adaptive Weighting of Handcrafted Feature Losses for Facial Expression Recognition
abstract
Due to the importance of facial expressions in human-machine interaction, a number of handcrafted features and deep neural networks have been developed for facial expression recognition. While a few studies have shown the similarity between the handcrafted features and the features learned by deep network, a new feature loss is proposed to use feature bias constraint of handcrafted and deep features to guide the deep feature learning during the early training of network. The feature maps learned with and without the proposed feature loss for a toy network suggest that our approach can fully explore the complementarity between handcrafted features and deep features. Based on the feature loss, a general framework for embedding the traditional feature information into deep network training was developed and tested using the FER2013, CK+, Oulu-CASIA, and MMI datasets. Moreover, adaptive loss weighting strategies are proposed to balance the influence of different losses for different expression databases. The experimental results show that the proposed feature loss with adaptive weighting achieves much better accuracy than the original handcrafted feature and the network trained without using our feature loss. Meanwhile, the feature loss with adaptive weighting can provide complementary information to compensate for the deficiency of a single feature.
Weicheng Xie 0001, LinLin Shen, Jinming Duan 0001
IEEE Trans. Cybern.1
2020 Group-Wise Dynamic Dropout Based on Latent Semantic Variations
abstract
Dropout regularization has been widely used in various deep neural networks to combat overfitting. It works by training a network to be more robust on information-degraded data points for better generalization. Conventional dropout and variants are often applied to individual hidden units in a layer to break up co-adaptations of feature detectors. In this paper, we propose an adaptive dropout to reduce the co-adaptations in a group-wise manner by coarse semantic information to improve feature discriminability. In particular, we showed that adjusting the dropout probability based on local feature densities can not only improve the classification performance significantly but also enhance the network robustness against adversarial examples in some cases. The proposed approach was evaluated in comparison with the baseline and several state-of-the-art adaptive dropouts over four public datasets of Fashion-MNIST, CIFAR-10, CIFAR-100 and SVHN.
Zhiwei Ke, Zhiwei Wen, Weicheng Xie 0001, Yi Wang 0017, LinLin Shen
AAAI3
2020 Geometry Constrained Weakly Supervised Object Localization
Weizeng Lu, Xi Jia, Weicheng Xie 0001, LinLin Shen, Yicong Zhou, Jinming Duan 0001
ECCV (26)3
2020 Feature map masking based single-stage face detection
abstract
Although great progress has been made in face detection, a trade-off between speed and accuracy is still a great challenge. We propose in this paper a feature map masking based approach for single-stage face detection. As feature maps extracted from feature pyramid network might contain face unrelated features, we propose a mask generation branch to predict those significant units for face detection. The masked feature maps, where only important features are left, are then passed through the following detection process. Ground truth masks, directly generated from the training images, based on the face bounding boxes, are used to train the feature mask generation module. A mask constrained dropout module has also been proposed to drop out significant units of the shared feature maps, such that the detection performance can be further improved. The proposed approach is extensively tested using the WIDER FACE dataset. The results suggest that our detector with ResNet-152 backbone, achieves the best precision-recall performance among competing methods. As high as 95.4%, 94.0% and 86.9% accuracies have been achieved on the easy, medium and hard subsets, respectively.
Junliang Chen 0002, Weicheng Xie 0001, LinLin Shen
IJCB3
2020 Group-wise Feature Orthogonalization and Suppression for GAN based Facial Attribute Translation
abstract
Generative Adversarial Network (GAN) has been widely used for object attribute editing. However, the semantic correlation, resulted from the feature map interaction in the generative network of GAN, may impair the generalization ability of the generative network. In this work, semantic disentanglement is introduced in GAN to reduce the attribute correlation. The feature maps of the generative network are first grouped with an efficient clustering algorithm based on hash encoding, which are used to excavate hidden semantic attributes and calculate the group-wise orthogonality loss for the reduction of attribute entanglement. Meanwhile, the feature maps falling in the intersection regions of different groups are further suppressed to reduce the attribute-wise interaction. Extensive experiments reveal that the proposed GAN generated more genuine objects than the state of the arts. Quantitative results of classification accuracy, inception score and FID score further justify the effectiveness of the proposed GAN.
Zhiwei Wen, Haoqian Wu, Weicheng Xie 0001, LinLin Shen
ICPR3
2019 Disentangled Feature Based Adversarial Learning for Facial Expression Recognition
abstract
A facial expression image can be considered as an addition of expressive component to a neutral expression face. With this in mind, in this paper, we propose a novel end-to-end adversarial disentangled feature learning (ADFL) framework for facial expression recognition. The ADFL framework is mainly composed of three branches: expression disentangling branch ADFL-d, neutral expression branch ADFL-n and residual expression branch ADFL-r. The ADFL-d and ADFL-n aim to extract the expressive component and neutral component, respectively. The ADFL-r extracts the residual expression by calculating the difference between feature maps of ADFL-d and ADFL-n, and uses the residual expression feature for expression classification. Experimental results on several benchmark databases (CK+, MMI and Oulu-CASIA) show that the proposed method has remarkable performance compared to state-of-the-art methods.
Mengchao Bai, Weicheng Xie 0001, LinLin Shen
ICIP2
2019 Outlier-Suppressed Triplet Loss with Adaptive Class-Aware Margins for Facial Expression Recognition
abstract
Triplet loss has been proposed to increase the inter-class distance and decrease the intra-class distance for various tasks of image recognition. However, for facial expression recognition (FER) problem, the fixed margin parameter does not fit the diversity of scales between different expressions. Meanwhile, the strategy of selecting the hardest triplets can introduce noisy guidance information since various persons may present significantly different expressions. In this work, we propose a new triplet loss based on class-aware margins and outlier-suppressed triplet for FER, where each pair of expressions, e.g. 'happy' and 'fear', is assigned with an adaptive margin parameter and the abnormal hard triplets are discarded according to the feature distance distribution. Experimental results of the proposed triplet loss on the FER2013 and CK+ expression databases show that the proposed network achieves much better accuracy than the original triplet loss and the network without using the proposed strategies, and competitive performance compared with the state-of-the-art algorithms.
Zhiwei Wen, Weicheng Xie 0001, LinLin Shen, Jinming Duan 0001
ICIP3
2019 Adversarial Feature Distillation for Facial Expression Recognition
Mengchao Bai, Xi Jia, Weicheng Xie 0001, LinLin Shen
PRICAI (3)3
2019 Sparse deep feature learning for facial expression recognition
Weicheng Xie 0001, Xi Jia, LinLin Shen, Meng Yang 0001
Pattern Recognit.1
2018 Hand-Crafted Feature Guided Deep Learning for Facial Expression Recognition
abstract
A number of facial expression recognition algorithms based on hand-crafted features and deep neutral networks have been developed. Motivated by the similarity between the hand-crafted features and features learned by deep network, a new feature loss is proposed to embed the information of hand-crafted features into the training process of network, which tries to reduce the difference between the two features. Based on the feature loss, a general framework for embedding the traditional feature information was developed and tested using CK+, JAFFE and FER2013 datasets. Experimental results show that the proposed network achieves much better accuracy than the original hand-crafted feature and the network without using our feature loss. When compared with other algorithms in literature, our network also achieved the best performance on CK+ dataset, i.e. 97.35% accuracy has been achieved.
Guohang Zeng, Jiancan Zhou, Xi Jia, Weicheng Xie 0001, LinLin Shen
FG4
2018 Open snake model based on global guidance field for embryo vessel location
abstract
The development of vessels can provide important information about the growth status of animal embryos. It is, therefore, important to automatically locate the deformed vessel branches from the embryo images. However, very few vessel detectors can accurately locate all vessel branches when the captured images are low quality and the implied vessel shapes are complex. In this study, a new framework consisting of vessel region extraction and snake shape optimisation is proposed. The main contribution in this detector is a novel open snake model based on the global guidance field and deformation template initialisation. Experimental results on a specific application of an embryo vessel database [Database and source codes: https://github.com/wcxie/Egg‐embryro‐vessel‐location/ .] demonstrate that the proposed algorithm not only locates the vessel shape properly but also obtains the orientations of embryo vessel branches accurately. Comparison to traditional guidance fields and the active appearance model illustrates the effectiveness and competitiveness of the proposed model.
Weicheng Xie 0001, Jinming Duan 0001, LinLin Shen, Yuexiang Li, Meng Yang 0001, Guojun Lin
IET Comput. Vis.1
2018 Facial expression synthesis with direction field preservation based mesh deformation and lighting fitting based wrinkle mapping
Weicheng Xie 0001, LinLin Shen, Meng Yang 0001, Jianmin Jiang
Multim. Tools Appl.1
2018 Robust, discriminative and comprehensive dictionary learning for face recognition
Guojun Lin, Meng Yang 0001, Jian Yang 0003, LinLin Shen, Weicheng Xie 0001
Pattern Recognit.5
2017 A Novel Transient Wrinkle Detection Algorithm and Its Application for Expression Synthesis
abstract
Because facial wrinkle is a representative feature of facial expression, automatic wrinkle detection has been an important and challenging topic for expression simulation, recognition, and animation. Recently, most works about wrinkle detection have focused on permanent wrinkles (e.g., age wrinkles), which are usually linear shapes, whereas the detection of transient wrinkles (e.g., expression wrinkles) has not been sufficiently studied because of their shape diversity and complexity. In this work, a novel algorithm for automatic detection of transient wrinkles with linear, fixed, and chaotic shapes is proposed, which largely consists of edge pair matching, active-appearance-model-based wrinkle structure location, and support-vector-machine-based wrinkle classification. The proposed wrinkle detector is applied for expression synthesis and an improved Poisson wrinkle mapping approach is proposed. Experimental results illustrate the competitiveness of the proposed wrinkle detector in detecting different transient wrinkles. Compared with state-of-the-art algorithms, the proposed approach yields complete and accurate wrinkle centers. The expression synthesized by the improved wrinkle mapping is also much more realistic.
Weicheng Xie 0001, LinLin Shen, Jianmin Jiang
IEEE Trans. Multim.1
2015 A binary differential evolution algorithm learning from explored solutions
Yu Chen 0001, Weicheng Xie 0001, Xiu-Fen Zou
Neurocomputing2
2013 A triangulation-based hole patching method using differential evolution
Weicheng Xie 0001, Xiu-Fen Zou
Comput. Aided Des.1
2013 Diversity-maintained differential evolution embedded with gradient-based local search
Weicheng Xie 0001, Wei Yu 0009, Xiu-Fen Zou
Soft Comput.1
2012 Iteration and optimization scheme for the reconstruction of 3D surfaces based on non-uniform rational B-splines
Weicheng Xie 0001, Xiu-Fen Zou, Jian-Dong Yang, Jie-Bin Yang
Comput. Aided Des.1
2011 Convergence of multi-objective evolutionary algorithms to a uniformly distributed representation of the Pareto front
Yu Chen 0001, Xiu-Fen Zou, Weicheng Xie 0001
Inf. Sci.3