Xingxun Jiang

dblp:251/0975 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-2139-8623ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Towards consistent and controllable image synthesis for identity-preserving face editing
abstract
Face editing involves modifying facial attributes like expression, head pose, or lighting, with the goal of preserving the subject’s unique identity features. Diffusion models have recently emerged as the dominant approach in visual generation, driven by their strong generative power. However, challenges persist in the realm of face editing, where independently and correctly editing target attributes while preserving high-fidelity identity information remains a formidable problem. In this paper, we present RigFace, a novel framework that combines controllable signals derived from a 3D Morphable Model (3DMM) with a fine-tuned Stable Diffusion (SD) model. Our basic idea to achieve by leveraging disentangled facial attributes provided by 3DMM and harnessing the strong generative capacity of Stable Diffusion. Specifically, our method contains: 1) A Spatial Attribute Encoder that provides robust and decoupled conditions of background, pose, expression and lighting; 2) A FaceFusion module that transfers identity information at different resolutions from the Identity Encoder to the Denoising UNet of a pre-trained SD model through self-attention, which facilitates detailed identity preservation. Our model achieves superior performance in both identity preservation and photorealism compared to existing face editing models.
Mengting Wei, Tuomas Varanka, Yante Li, Xingxun Jiang, Huai-Qian Khor, Guoying Zhao 0001
Pattern Recognit.4
2026 LatentMag: Self-Supervised 3D Magnification for Micro Expressions via Latent Extrapolation
abstract
Micro-expressions (MEs) are subtle and brief facial movements that reveal genuine emotional states but are often imperceptible due to their low intensity. While motion magnification has proven effective for enhancing ME visibility in 2D settings, its extension to 3D remains largely unexplored. In this work, we presentLatentMag, the first controllable 3D micro-expression magnification framework. Unlike traditional editing methods that rely on fixed labels or expression targets, our approach models expression intensity as a relative, input-dependent signal. We adopt registered 3D meshes as our representation, enabling vertex-level correspondence and interpretable displacement analysis. To guide magnification, we introduce a geometric prior that models amplification as a spatially adaptive transformation, where the change in pairwise distance between points on the output mesh scales with that observed between the input shapes, ensuring natural, localized deformation. We operationalize this prior in a generative framework by disentangling a latent intensity code, whose extrapolation drives controllable shape amplification. Trained in a self-supervised manner using unlabeled mesh sequences, LatentMag generalizes well to unseen identities and expressions, offering a novel solution that bridges geometric interpretability with realistic 3D expression modeling.
Mengting Wei, Xingxun Jiang, Haoyu Chen 0001, Yante Li, Guoying Zhao 0001
IEEE Trans. Affect. Comput.2
2024 A Novel Decoupled Prototype Completion Network for Incomplete Multimodal Emotion Recognition
abstract
Reconstructing missing modality based on available modalities is widely used to address inevitable modality-missing for Multimodal Emotion Recognition (MER). However, due to explicit distribution gap across heterogeneous modalities, they fail to guarantee the consistency between the reconstructed data and the ground truth. To mitigate this problem, we propose a novel method to restore the missing modality using its weighted prototypes rather than other modalities. Specifically, prototypes of different classes of missing modality are used to encapsulate its representative knowledge. Then sample-to-prototype affinity measuring class similarity is used as weights to combine these prototypes for reconstruction, thereby effectively restoring the distribution-consistent modality. Furthermore, to improve the efficacy of prototype-based completion under seriously missing, we devise an adaptive knowledge distillation from the strong modality to the weaker ones. This reinforces the representation ability of weak modality features. Extensive experiments on CMU-MOSI and IEMOCAP datasets demonstrate the superiority of our method.
Zhangfeng Hu, Wenming Zheng, Yuan Zong, Mengting Wei, Xingxun Jiang, Mengxin Shi
ICME5
2023 CMNet: Contrastive Magnification Network for Micro-Expression Recognition
abstract
Micro-Expression Recognition (MER) is challenging because the Micro-Expressions' (ME) motion is too weak to distinguish. This hurdle can be tackled by enhancing intensity for a more accurate acquisition of movements. However, existing magnification strategies tend to use the features of facial images that include not only intensity clues as intensity features, leading to the intensity representation deficient of credibility. In addition, the intensity variation over time, which is crucial for encoding movements, is also neglected. To this end, we provide a reliable scheme to extract intensity clues while considering their variation on the time scale. First, we devise an Intensity Distillation (ID) loss to acquire the intensity clues by contrasting the difference between frames, given that the difference in the same video lies only in the intensity. Then, the intensity clues are calibrated to follow the trend of the original video. Specifically, due to the lack of truth intensity annotation of the original video, we build the intensity tendency by setting each intensity vacancy an uncertain value, which guides the extracted intensity clues to converge towards this trend rather some fixed values. A Wilcoxon rank sum test (Wrst) method is enforced to implement the calibration. Experimental results on three public ME databases i.e. CASME II, SAMM, and SMIC-HS validate the superiority against state-of-the-art methods.
Mengting Wei, Xingxun Jiang, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Jiateng Liu
AAAI2
2023 Geometric Magnification-based Attention Graph Convolutional Network for Skeleton-based Micro-Gesture Recognition
abstract
Micro-Gesture (MG) recognition is an emerging and challenging task due to the short duration and small amplitude of joints. MGs indicate subtle movements of the body in response to stress, which are more difficult to recognize than regular gestures. To solve the above problems, for the modeling of micro-gesture skeleton data, we propose a Geometric Magnification-Based Attention Graph Convolutional Network (MA-GCN) to magnify and select features. The network mainly consists of two modules: the geometric magnification module (GM module) controls the magnification of different joints, and the spatial temporal attention graph convolution module (STA module) selects valid information by weighting different joints and frames to focus on subtle movements. Extensive experiments on two MG datasets prove that our method achieves remarkable performance.
Haolin Jiang, Wenming Zheng, Yuan Zong, Xingxun Jiang, Yunlong Xue
ICIP5
2022 A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive Networks
abstract
Micro-Expression recognition (MER) is a challenging task due to the short duration and low intensity of Micro-Expressions. A popular method to tackle this is magnifying MEs so as to enlarge the expression intensity to make recognition easier. However, the single fixed magnification strategy, widely used in existing works of MER, is not appropriate for different subjects, because each subject has specific expression intensity corresponding to different MEs. To cope with this issue, we propose a novel Attention-based Magnification-Adaptive Network (AMAN) to learn adaptive magnification levels for the ME representation. The network consists of two modules: magnification attention (MA module) to adaptively focus on appropriate magnification levels of different MEs, and frame attention (FA module) to focus on discriminative aggregated frames in a ME video. Extensive experiments on three widely used databases manifest that our method yields state-of-art results compared with other methods.
Mengting Wei, Wenming Zheng, Yuan Zong, Xingxun Jiang, Cheng Lu 0005, Jiateng Liu
ICASSP4
2022 Seeking Salient Facial Regions for Cross-Database Micro-Expression Recognition
abstract
Cross-Database Micro-Expression Recognition (CD-MER) aims to develop the Micro-Expression Recognition (MER) methods with strong domain adaptability, i.e., the ability to recognize the Micro-Expressions (MEs) of different subjects captured by different imaging devices in different scenes. The development of CDMER is faced with two key problems: 1) the severe feature distribution gap between the source and target databases; 2) the feature representation bottleneck of ME such local and subtle facial expressions. To solve these problems, this paper proposes a novel Transfer Group Sparse Regression method, namely TGSR, which aims to 1) optimize the measurement and better alleviate the difference between the source and target databases, and 2) highlight the valid facial regions to enhance extracted features, by the operation of selecting the group features from the raw face feature, where each region is associated with a group of raw face feature, i.e., the salient facial region selection. Compared with previous transfer group sparse methods, our proposed TGSR has the ability to select the salient facial regions, which is effective in alleviating aforementioned problems for better performance and reducing the computational cost at the same time. We use two public ME databases, i.e., CASME II and SMIC, to evaluate our proposed TGSR method. Experimental results show that our proposed TGSR learns the discriminative and explicable regions, and outperforms most state-of-the-art subspace-learning-based domain-adaptive methods for CDMER.
Xingxun Jiang, Yuan Zong, Wenming Zheng, Jiateng Liu, Mengting Wei
ICPR1
2022 A Novel Magnification-Robust Network with Sparse Self-Attention for Micro-expression Recognition
abstract
Existing works for spontaneous Micro-Expression Recognition (MER) tend to encode Micro-Expression (ME) movements to get more discriminative features. However, MEs’ low intensity makes the capture for motion extremely difficult, and the widely adopted unified-magnification strategy is prone to noise and lacks flexibility. To this end, this paper provides a new insight to encode ME motion and tackle magnification noise. Specifically, we reconstruct a new sequence via magnification techniques to make subtle ME movements more distinguishable. Afterward, Sparse Self-Attention (SSA) rectifies self-attention with Locality Sensitive Hashing (LSH), cutting the space into several hush buckets of related features. Only keys in the same bucket are operated in the attention term for every query feature. The resulting sparsity in the attention matrix prevents the network from attending features stemming from less-informative magnification degrees which could be regarded as noise, while retains the sequence modelling capability of standard self-attention. Extensive experiments on three public MER databases demonstrate our superiority against the state-of-the-art methods.
Mengting Wei, Wenming Zheng, Xingxun Jiang, Yuan Zong, Cheng Lu 0005, Jiateng Liu
ICPR3
2022 Sample Self-Revised Network for Cross-Dataset Facial Expression Recognition
abstract
Facial images with low quality, subjective annotation, severe occlusion, and rare subject identity can lead to the existence of outlier samples in facial expression datasets. These outlier samples are usually far from the center of the dataset in the feature space, resulting in huge differences in feature distribution, which severely restricts the performance of cross-dataset facial expression recognition (FER). To eliminate the influence of outlier samples on cross-dataset FER, we propose an unsupervised domain adaptation (UDA) method called Sample Self-Revised Network (SSRN), which 1) dynamically detects the outlier level of each sample in the source domain to reduce the disturbance of outlier samples to the model training, as well as 2) adaptively revises outlier samples in the source domain to improve transferability of the learned features. Experimental results show that our SSRN outperforms both classic deep UDA methods and state-of-the-art cross-dataset FER results.
Wenming Zheng, Yuan Zong, Cheng Lu 0005, Xingxun Jiang
IJCNN5
2022 Adaptive Hierarchical Graph Convolutional Network for EEG Emotion Recognition
abstract
Human emotion is closely related to multiple distributed brain regions, and functional connections exist between the regions. However, how to abstract the region-level information to improve electroencephalograph (EEG) emotion recognition performance has not been well considered. To address this problem, we proposed a novel Adaptive Hierarchical Graph Convolutional Network (AHGCN), which includes the basic channel-level graph of EEG channels and the region-level graph of brain regions. Different from previous methods, we propose an adaptive pooling operation to automatically partition brain regions rather than manually define them. To capture the intrinsic functional connections between the brain regions or EEG channels, we design a gated adaptive graph convolution operation. Besides, we develop a graph unpooling operation to integrate the region-level graph and channel-level graph to extract more discrimination features for classification. Experiments on two widely-used datasets show that our proposed method is superior to many state-of-the-art methods on EEG emotion recognition and could find some interesting combinations of EEG channels.
Yunlong Xue, Wenming Zheng, Yuan Zong, Hongli Chang, Xingxun Jiang
IJCNN5
2020 DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the Wild
abstract
Recently, facial expression recognition (FER) in the wild has gained a lot of researchers' attention because it is a valuable topic to enable the FER techniques to move from the laboratory to the real applications. In this paper, we focus on this challenging but interesting topic and make contributions from three aspects. First, we present a new large-scale 'in-the-wild' dynamic facial expression database, DFEW (Dynamic Facial Expression in the Wild), consisting of over 16,000 video clips from thousands of movies. These video clips contain various challenging interferences in practical scenarios such as extreme illumination, occlusions, and capricious pose changes. Second, we propose a novel method called Expression-Clustered Spatiotemporal Feature Learning (EC-STFL) framework to deal with dynamic FER in the wild. Third, we conduct extensive benchmark experiments on DFEW using a lot of spatiotemporal deep feature learning methods as well as our proposed EC-STFL. Experimental results show that DFEW is a well-designed and challenging database, and the proposed EC-STFL can promisingly improve the performance of existing spatiotemporal deep neural networks in coping with the problem of dynamic FER in the wild. Our DFEW database is publicly available and can be freely downloaded from https://dfew-dataset.github.io/.
Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu 0005, Jiateng Liu
ACM Multimedia1
2019 Bi-modality Fusion for Emotion Recognition in the Wild
abstract
The emotion recognition in the wild has been a hot research topic in the field of affective computing. Though some progresses have been achieved, the emotion recognition in the wild is still an unsolved problem due to the challenge of head movement, face deformation, illumination variation etc. To deal with these unconstrained challenges, we propose a bi-modality fusion method for video based emotion recognition in the wild. The proposed framework takes advantages of the visual information from facial expression sequences and the speech information from audio. The state-of-the-art CNN based object recognition models are employed to facilitate the facial expression recognition performance. A bi-direction long short term Memory (Bi-LSTM) is employed to capture dynamic information of the learned features. Additionally, to take full advantages of the facial expression information, the VGG16 network is trained on AffectNet dataset to learn a specialized facial expression recognition model. On the other hand, the audio based features, like low level descriptor (LLD) and deep features obtained by spectrogram image, are also developed to improve the emotion recognition performance. The best experimental result shows that the overall accuracy of our algorithm on the Test dataset of the EmotiW challenge is 62.78, which outperforms the best result of EmotiW2018 and ranks 2nd at the EmotiW2019 challenge.
Sunan Li, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Chuangao Tang, Xingxun Jiang, Jiateng Liu, Wanchuang Xia
ICMI6