Yong Li 0032

dblp:93/2334-32 · DBLP profile ↗
← Back
30ranked-venue papers
13as first author
27since 2021 · last 2026
0000-0002-6521-5921ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 16 · 8 first-author · 14 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training
abstract
Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre-training by selectively removing noisy and redundant samples from large EEG datasets. EEG-DLite begins by encoding EEG segments into compact latent representations using a self-supervised autoencoder, allowing sample selection to be performed efficiently and with reduced sensitivity to noise. Based on these representations, EEG-DLite filters out outliers and minimizes redundancy, resulting in a smaller yet informative subset that retains the diversity essential for effective foundation model training. Through extensive experiments, we demonstrate that training on only 5 percent of a 2,500-hour dataset curated with EEG-DLite yields performance comparable to, and in some cases better than, training on the full dataset across multiple downstream tasks. To our knowledge, this is the first systematic study of pre-training data distillation in the context of EEG foundation models. EEG-DLite provides a scalable and practical path toward more effective and efficient physiological foundation modeling.
Yuting Tang, Wei-Bang Jiang, Shanglin Li, Yong Li 0032, Xinliang Zhou, Yi Ding 0012, Cuntai Guan
AAAI4
2026 Two-stream attentive spatial-temporal graph convolutional network for P300 detection in brain-computer interface
Jincen Wang, Yan Zhao 0037, Cunhang Fan, Yong Li 0032, Fan Liu 0003, Hailun Lian, Cheng Lu 0005
Expert Syst. Appl.4
2026 Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
abstract
Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1.3%/2.4% (ACC$_{7}$7), 1.3%/1.9% (ACC$_{2}$2) and 1.9%/1.8% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces.
Yong Li 0032, Yuanzhi Wang, Yi Ding 0012, Shiqing Zhang, Ke Lu 0002, Cuntai Guan
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Hierarchical Vision-Language Interaction for Facial Action Unit Detection
abstract
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Inter action for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal atten tion architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language based AU details, culminating in robust and semantically en riched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis.
Yong Li 0032, Yizhe Zhang 0001, Tianyi Zhang 0013, Muyun Jiang, Guosen Xie, Cuntai Guan
IEEE Trans. Affect. Comput.1
2026 Training-Free Controllable Text-Guided Video Editing
abstract
Decomposition-based text-guided video editing paradigm aims to utilize the layered neural atlas model to decompose the input video into foreground and background parts and edit the video in a divide-and-conquer manner, which is meaningful and improves the controllability of editing. However, they may suffer from some limitations: 1) High computational cost of per-video training (i.e, 7∼8 hours for training a single atlas model). 2) Foreground object deformation is restricted by the foreground opacity value. 3) Restricted flexibility in manipulating multiple objects. In this paper, we propose TraFrCo, aTraining-Free Controllable Text-guided Video Editingframework to mitigate these challenges. Instead of training complex atlas models, our method leverages pre-trained segmentation to rapidly decompose videos into foreground and background parts. This allows users to perform independent edits on foreground objects using existing video diffusion editing models without affecting the environment. To ensure visual consistency, we introduce a training-free mechanism that effectively propagates information across frames to fill missing background regions caused by the segmentation-derived foreground masks and reconstructs the scene behind moving objects. Finally, the edited components are seamlessly composited by re-predicting the new foreground masks. In contrast to prior works, TraFrCo enables efficient, fine-grained manipulation of video content without the burden of training. Experimental results verify that our TraFrCo consistently reduces the costs of decomposing video and achieves superior text-guided video editing performance. Codes and video demos will be released at https://github.com/mdswyz/TraFrCo.
Yuanzhi Wang, Yong Li 0032, Zhen Cui 0001, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2026 Location Matters: Frequency-Spatial Dual-Space Adaptation for Cross-Domain Few-Shot Segmentation
abstract
Current cross-domain few-shot semantic segmentation (CD-FSS) methods tend to overlook a fundamental yet domain-agnostic prior: the spatial correspondence between support and query images driven by the task itself. Unlike semantic similarity, this spatial correlation arises from the consistent structural layout of foreground objects across domains. To exploit this structural prior, we propose a novel frequency-spatial dual space adaptation (FDSA) framework, to learn domain-invariant structures and task-specific priors by jointly suppressing domain-specific redundancy in frequency domain and reinforcing geometric priors in spatial domain. Specifically, FDSA consists of two sequential modules, i.e., the frequency structural adapter (FSA) and the spatial geometry adapter (SGA). FSA performs image modulation in the frequency domain by emphasizing low-frequency foreground semantics and attenuating high-frequency noise, thus maintaining structural integrity of these input images. By contrast, SGA leverages handcrafted local descriptors to extract keypoints from both support and query images, generating Gaussian-based geometric priors that highlight desirable aligned regions. Additionally, we introduce spatial-guided SAM refinement (SSR) to extend our spatial geometric prior into the Segment Anything Model (SAM). SSR generates a soft Gaussian point prompt centered on the coarse mask, enabling SAM to refine segmentation masks without manual intervention. This integration effectively bridges task-specific localization with high-quality segmentation. Extensive experiments on four standard CD-FSS benchmarks demonstrate that our method achieves new state-of-the-art performance. Code is available at https://github.com/CVL-hub/FDSA.git.
Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Xiangbo Shu, Guosen Xie
IEEE Trans. Image Process.3
2025 SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot Segmentation
abstract
The primary challenge of cross-domain few-shot segmentation (CD-FSS) is the domain disparity between the training and inference phases, which can exist in either the input data or the target classes. Previous models struggle to learn feature representations that generalize to various unknown domains from limited training domain samples. In contrast, the large-scale visual model SAM, pre-trained on tens of millions of images from various domains and classes, possesses excellent generalizability. In this work, we propose a SAM-aware graph prompt reasoning network (GPRN) that fully leverages SAM to guide CD-FSS feature representation learning and improve prediction accuracy. Specifically, we propose a SAM-aware prompt initialization module (SPI) to transform the masks generated by SAM into visual prompts enriched with high-level semantic information. Since SAM tends to divide an object into many sub-regions, this may lead to visual prompts representing the same semantic object having inconsistent or fragmented features. We further propose a graph prompt reasoning (GPR) module that constructs a graph among visual prompts to reason about their interrelationships and enable each visual prompt to aggregate information from similar prompts, thus achieving global semantic consistency. Subsequently, each visual prompt embeds its semantic information into the corresponding mask region to assist in feature representation learning. To refine the segmentation mask during testing, we also design a non-parameter adaptive point selection module (APS) to select representative point prompts from query predictions and feed them back to SAM to refine inaccurate segmentation results. Experiments on four standard CD-FSS datasets demonstrate that our method establishes new state-of-the-art results.
Shi-Feng Peng, Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Guosen Xie
AAAI3
2025 Re-Attentional Controllable Video Diffusion Editing
abstract
Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resulting in remarkable video editing capabilities. However, they may still suffer from some limitations such as mislocated objects, incorrect number of objects. Therefore, the controllability of video editing remains a formidable challenge. In this paper, we aim to challenge the above limitations by proposing a Re-Attentional Controllable Video Diffusion Editing (ReAtCo) method. Specially, to align the spatial placement of the target objects with the edited text prompt in a training-free manner, we propose a Re-Attentional Diffusion (RAD) to refocus the cross-attention activation responses between the edited text prompt and the target video during the denoising stage, resulting in a spatially location-aligned and semantically high-fidelity manipulated video. In particular, to faithfully preserve the invariant region content with less border artifacts, we propose an Invariant Region-guided Joint Sampling (IRJS) strategy to mitigate the intrinsic sampling errors w.r.t the invariant regions at each denoising timestep and constrain the generated content to be harmonized with the invariant region content. Experimental results verify that ReAtCo consistently improves the controllability of video diffusion editing and achieves superior video editing performance.
Yuanzhi Wang, Yong Li 0032, Xin Liu 0011, Zhen Cui 0001, Antoni B. Chan
AAAI2
2025 Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
abstract
Fine-grained hashing has become a powerful solution for rapid and efficient image retrieval, particularly in scenarios requiring high discrimination between visually similar categories. To enable each hash bit to correspond to specific visual attributes, we propose a novel method that harnesses learnable queries for attribute-aware hash code learning. This method deploys a tailored set of queries to capture and represent nuanced attribute-level information within the hashing process, thereby enhancing both the interpretability and relevance of each hash bit. Building on this query-based optimization framework, we incorporate an auxiliary branch to help alleviate the challenges of complex landscape optimization often encountered with low-bit hash codes. This auxiliary branch models high-order attribute interactions, reinforcing the robustness and specificity of the generated hash codes. Experimental results on benchmark datasets demonstrate that our method generates attribute-aware hash codes and consistently outperforms state-of-the-art techniques in retrieval accuracy and robustness, especially for low-bit hash codes, underscoring its potential in fine-grained image hashing tasks.
Peng Wang 0107, Yong Li 0032, Lin Zhao 0003, Xiu-Shen Wei
ICML2
2025 MER 2025: When Affective Computing Meets Large Language Models
abstract
MER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality).
Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia8
2025 Sera: Separated Coarse-to-fine Representation Alignment for Cross-subject EEG-based Emotion Recognition
abstract
Neuropsychology-inspired models have been utilized in recent advances in EEG emotion recognition, such as convolutional networks for spatial features and Transformers for temporal dependencies. While these methods benefit from domain knowledge like frequency-band features and spatial correlations, most overlook the fundamental fact that EEG signals are complex mixtures of neural source activities recorded at the scalp. EEG signals presenting challenges for emotion recognition, particularly in cross-subject scenarios due to significant inter-subject variance. Inspired by neurophysiological principles, we propose a novel framework, named Sera, for EEG-based emotion recognition that explicitly separates source activities and aligns representations across subjects. Sera introduces two key components: (1) a variational autoencoder (VAE) with multiple multi-stage decoders (M2VAE) designed to disentangle EEG signals into independent sources, mimicking the neural generation process, and (2) a coarse-to-fine representation alignment block (CFRA) to mitigate subject-to-subject variability. The coarse alignment employs adversarial training with a domain discriminator, while the fine-grained alignment matches covariance matrices to capture temporal correlations within EEG segments. Extensive experiments demonstrate that Sera outperforms the state-of-the-art methods with improvements ranging from 1% to 5%, averaging 3.14% and 3.05% on the DEAP and DREAMER datasets, respectively, confirming its effectiveness and neurophysiological grounding. The code is available at: https://github.com/JZH98/Sera-code.
Meiyan Xu, Ziyu Jia, Yong Li 0032, Xinliang Zhou, Junfeng Yao, Yi Ding 0012
ACM Multimedia5
2025 Instance-Consistent Fair Face Recognition
abstract
The fairness of face recognition (FR) is a challenging issue to numerous FR algorithms in the modern pluralistic and egalitarian society. In this work, we propose an instance-consistent fair face recognition (IC-FFR) method by fulfilling complete instance fairness on false positive rate (FPR) and true positive rate (TPR). In view of the misalignment of testing and training metrics, not yet considered by the current fair FR algorithms, in theory, we inspect the correlation between the testing metrics (FPR and TPR) and the label classification loss, and we derive a high-probability consistency of unfairness penalties from FPR and TPR to the softmax loss. According to the theoretical analysis, we further develop an instance-consistent fairness solution by introducing customized instance margins, which well preserve consistent FPR and TPR of all instances during the label classification in training. To encourage more fine-grained fairness evaluation, we contribute a dataset called national faces in the world (NFW) to measure the fairness of individuals and countries. Extensive experiments on our NFW as well as the RFW and BFW benchmarks demonstrate the effectiveness and superiority of our method compared to those state-of-the-art fair FR methods.
Yong Li 0032, Zhen Cui 0001, Pengcheng Shen, Shiguang Shan
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Collaborative contrastive learning for cross-domain gaze estimation
Lifan Xia, Yong Li 0032, Zhen Cui 0001, Chunyan Xu, Antoni B. Chan
Pattern Recognit.2
2025 Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection
abstract
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed, they often overfit to specific datasets and individual features, limiting their cross-domain applicability. To overcome these limitations, we propose a doubly adaptive dropout approach for cross-domain AU detection, which enhances the robustness of convolutional feature maps and spatial tokens against domain shifts. This approach includes a Channel Drop Unit (CD-Unit) and a Token Drop Unit (TD-Unit), which work together to reduce domain-specific noise at both the channel and token levels. The CD-Unit preserves domain-agnostic local patterns in feature maps, while the TD-Unit helps the model identify AU relationships generalizable across domains. An auxiliary domain classifier, integrated at each layer, guides the selective omission of domain-sensitive features. To prevent excessive feature dropout, a progressive training strategy is used, allowing for selective exclusion of sensitive features at any model layer. Our method consistently outperforms existing techniques in cross-domain AU detection, as demonstrated by extensive experimental evaluations. Visualizations of attention maps also highlight clear and meaningful patterns related to both individual and combined AUs, further validating the approach's effectiveness.
Yong Li 0032, Xuesong Niu, Yi Ding 0012, Xiu-Shen Wei, Cuntai Guan
IEEE Trans. Affect. Comput.1
2025 Decoupled Doubly Contrastive Learning for Cross-Domain Facial Action Unit Detection
abstract
Despite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D2CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D2CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D2CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D2CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D2CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6%-14% across various cross-domain scenarios.
Yong Li 0032, Menglin Liu, Zhen Cui 0001, Yi Ding 0012, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan
IEEE Trans. Image Process.1
2025 EEG-Deformer: A Dense Convolutional Transformer for Brain-Computer Interfaces
abstract
Effectively learning the temporal dynamics in electroencephalogram (EEG) signals is challenging yet essential for decoding brain activities using brain-computer interfaces (BCIs). Although Transformers are popular for their long-term sequential learning ability in the BCI field, most methods combining Transformers with convolutional neural networks (CNNs) fail to capture the coarse-to-fine temporal dynamics of EEG signals. To overcome this limitation, we introduce EEG-Deformer, which incorporates two main novel components into a CNN-Transformer: (1) a Hierarchical Coarse-to-Fine Transformer (HCT) block that integrates a Fine-grained Temporal Learning (FTL) branch into Transformers, effectively discerning coarse-to-fine temporal patterns; and (2) a Dense Information Purification (DIP) module, which utilizes multi-level, purified temporal information to enhance decoding accuracy. Comprehensive experiments on three representative cognitive tasksâcognitive attention, driving fatigue, and mental workload detectionâconsistently confirm the generalizability of our proposed EEG-Deformer, demonstrating that it either outperforms or performs comparably to existing state-of-the-art methods. Visualization results show that EEG-Deformer learns from neurophysiologically meaningful brain regions for the corresponding cognitive tasks.
Yi Ding 0012, Yong Li 0032, Rui Liu 0034, Chengxuan Tong, Xinliang Zhou, Cuntai Guan
IEEE J. Biomed. Health Informatics2
2025 EmT: A Novel Transformer for Generalized Cross-Subject EEG Emotion Recognition
abstract
Integrating prior knowledge of neurophysiology into neural network architecture enhances the performance of emotion decoding. While numerous techniques emphasize learning spatial and short-term temporal patterns, there has been a limited emphasis on capturing the vital long-term contextual information associated with emotional cognitive processes. In order to address this discrepancy, we introduce a novel transformer model called emotion transformer (EmT). EmT is designed to excel in both generalized cross-subject electroencephalography (EEG) emotion classification and regression tasks. In EmT, EEG signals are transformed into a temporal graph format, creating a sequence of EEG feature graphs using a temporal graph construction (TGC) module. A novel residual multiview pyramid graph convolutional neural network (RMPG) module is then proposed to learn dynamic graph representations for each EEG feature graph within the series, and the learned representations of each graph are fused into one token. Furthermore, we design a temporal contextual transformer (TCT) module with two types of token mixers to learn the temporal contextual information. Finally, the task-specific output (TSO) module generates the desired outputs. Experiments on four publicly available datasets show that EmT achieves higher results than the baseline methods for both EEG emotion classification and regression tasks. The code is available at https://github.com/yi-ding-cs/EmT.
Yi Ding 0012, Chengxuan Tong, Shuailei Zhang, Muyun Jiang, Yong Li 0032, Kevin Junliang Lim, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.5
2024 Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual Reconstruction
abstract
Pre-training plays a vital role in various vision tasks, such as object recognition and detection. Commonly used pre-training methods, which typically rely on randomized approaches like uniform or Gaussian distributions to initialize model parameters, often fall short when confronted with long-tailed distributions, especially in detection tasks. This is largely due to extreme data imbalance and the issue of simplicity bias. In this paper, we introduce a novel pre-training framework for object detection, called Dynamic Rebalancing Contrastive Learning with Dual Reconstruction (2DRCL). Our method builds on a Holistic-Local Contrastive Learning mechanism, which aligns pre-training with object detection by capturing both global contextual semantics and detailed local patterns. To tackle the imbalance inherent in long-tailed data, we design a dynamic rebalancing strategy that adjusts the sampling of underrepresented instances throughout the pre-training process, ensuring better representation of tail classes. Moreover, Dual Reconstruction addresses simplicity bias by enforcing a reconstruction task aligned with the self-consistency principle, specifically benefiting underrepresented tail classes. Experiments on COCO and LVIS v1.0 datasets demonstrate the effectiveness of our method, particularly in improving the mAP/AP scores for tail classes.
Chen-Long Duan, Yong Li 0032, Xiu-Shen Wei, Lin Zhao 0003
NeurIPS2
2024 Multi-Level Information Aggregation Based Graph Attention Networks Towards Fake Speech Detection
abstract
It is widely acknowledged that distinguishing genuine speech from spoofed speech encompasses various subbands and temporal segments within speech signals. However, prevailing spoofing detection methods tend to oversimplify the relationships between these cues by employing linear models. In this paper, we introduce a multi-level information aggregation Graph Attention Networks (MiaGATs) to generate highly discriminative features for fake speech detection (FSD). In MiaGATs, each subband and temporal segment of a speech signal is represented as distinct nodes. MiaGATs incorporates channel information aggregation within each node to effectively harness the unique spectral and temporal characteristics during the feature encoding stage. In particular, MiaGATs address the interactions between nodes through indirect node aggregation and integrates both indirect and direct node aggregation by max-pooling operation. Experimental results on ASVspoof2019 and ASVspoof2021 LA databases show significant relative improvement compared to the current state-of-the-art. In comparison to the leading integrated spectro-temporal graph attention networks, MiaGATs gains an impressive performance improvement in various conditions, underscoring MiaGATs's position as a new benchmark in spoofing detection performance.
Jian Zhou 0006, Yong Li 0032, Cunhang Fan, Hon Keung Kwan
IEEE Signal Process. Lett.2
2024 Edit Temporal-Consistent Videos with Image Diffusion Model
abstract
Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal inconsistencies as the temporal characteristics of videos have not been faithfully modeled. In this article, we propose an elegant yet effective Temporal-Consistent Video Editing (TCVE) method to mitigate the temporal inconsistency challenge for robust text-guided video editing. In addition to the utilization of a pretrained T2I 2D Unet for spatial content manipulation, we establish a dedicated temporal Unet architecture to faithfully capture the temporal coherence of the input video sequences. Furthermore, to establish coherence and interrelation between the spatial-focused and temporal-focused components, a cohesive spatial-temporal modeling unit is formulated. This unit effectively interconnects the temporal Unet with the pretrained 2D Unet, thereby enhancing the temporal consistency of the generated videos while preserving the capacity for video content manipulation. Quantitative experimental results and visualization results demonstrate that TCVE achieves state-of-the-art performance in both video temporal consistency and video editing capability, surpassing existing benchmarks in the field. Codes are released at https://github.com/mdswyz/TCVE .
Yuanzhi Wang, Yong Li 0032, Xin Liu 0011, Anbo Dai, Antoni B. Chan, Zhen Cui 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Meta Auxiliary Learning for Facial Action Unit Detection
abstract
Despite the success of deep neural networks on facial action unit (AU) detection, better performance depends on a large number of training images with accurate AU annotations. However, labeling AU is time-consuming, expensive, and error-prone. Considering AU detection and facial expression recognition (FER) are two highly correlated tasks, and facial expression (FE) is relatively easy to annotate, we consider learning AU detection and FER in a multi-task manner. However, the performance of the AU detection task cannot be always enhanced due to the negative transfer in the multi-task scenario. To alleviate this issue, we propose a Meta Auxiliary Learning method (MAL) that automatically selects highly related FE samples by learning adaptative weights for the training FE samples in a meta learning manner. The learned sample weights alleviate the negative transfer from two aspects: 1) balance the loss of each task automatically, and 2) suppress the weights of FE samples that have large uncertainties. Experimental results on several popular AU datasets demonstrate MAL consistently improves the AU detection performance compared with the state-of-the-art multi-task and auxiliary learning methods. MAL automatically estimates adaptive weights for the auxiliary FE samples according to their semantic relevance with the primary AU detection task.
Yong Li 0032, Shiguang Shan
IEEE Trans. Affect. Comput.1
2023 Contrastive Learning of Person-Independent Representations for Facial Action Unit Detection
abstract
Facial action unit (AU) detection, aiming to classify AU present in the facial image, has long suffered from insufficient AU annotations. In this paper, we aim to mitigate this data scarcity issue by learning AU representations from a large number of unlabelled facial videos in a contrastive learning paradigm. We formulate the self-supervised AU representation learning signals in two-fold: 1) AU representation should be frame-wisely discriminative within a short video clip; 2) Facial frames sampled from different identities but show analogous facial AUs should have consistent AU representations. As to achieve these goals, we propose to contrastively learn the AU representation within a video clip and devise a cross-identity reconstruction mechanism to learn the person-independent representations. Specially, we adopt a margin-based temporal contrastive learning paradigm to perceive the temporal AU coherence and evolution characteristics within a clip that consists of consecutive input facial frames. Moreover, the cross-identity reconstruction mechanism facilitates pushing the faces from different identities but show analogous AUs close in the latent embedding space. Experimental results on three public AU datasets demonstrate that the learned AU representation is discriminative for AU detection. Our method outperforms other contrastive learning methods and significantly closes the performance gap between the self-supervised and supervised AU detection approaches.
Yong Li 0032, Shiguang Shan
IEEE Trans. Image Process.1
2023 3D3M: 3D Modulated Morphable Model for Monocular Face Reconstruction
abstract
3D face reconstruction from a single image is a vital task in various multimedia applications. A key challenge for 3D face shape reconstruction is to build the correct dense face correspondence between the monocular input face and the deformable mesh. Most existing methods rely on shape labels fitted by traditional methods or strong priors such as multi-view geometry consistency. In contrast, we propose an innovative 3D Modulated Morphable Model (3D3M) to learn the dense shape correspondence from monocular images in a self-supervised manner. Specifically, given a batch of input faces, 3D3M encodes their 3DMM attributes (shape, texture, lighting, etc.) and then randomly shuffles the 3DMM attributes to generate the attribute-changed faces. The attribute-changed faces can be encoded and rendered back in a cycle-consistent manner, which enables us to utilize the self-supervised consistencies in dense mesh vertices and reconstructed pixels. The dense shape and pixel correspondence enable us to adopt a series of self-supervised constraints to fit the 3D face model accurately and learn the per-vertex correctives end-to-end. 3D3M builds excellent high-quality 3D face reconstruction results from monocular images. Both quantitative and qualitative experimental results have verified the superiority of 3D3M over prior arts on 3D face reconstruction and face alignment.
Yong Li 0032, Jianguo Hu, Xinmiao Pan, Zechao Li, Zhen Cui 0001
IEEE Trans. Multim.1
2022 Learning Representations for Facial Actions From Unlabeled Videos
abstract
Facial actions are usually encoded as anatomy-based action units (AUs), the labelling of which demands expertise and thus is time-consuming and expensive. To alleviate the labelling demand, we propose to leverage the large number of unlabelled videos by proposing a twin-cycle autoencoder (TAE) to learn discriminative representations for facial actions. TAE is inspired by the fact that facial actions are embedded in the pixel-wise displacements between two sequential face images (hereinafter, source and target) in the video. Therefore, learning the representations of facial actions can be achieved by learning the representations of the displacements. However, the displacements induced by facial actions are entangled with those induced by head motions. TAE is thus trained to disentangle the two kinds of movements by evaluating the quality of the synthesized images when either the facial actions or head pose is changed, aiming to reconstruct the target image. Experiments on AU detection show that TAE can achieve accuracy comparable to other existing AU detection methods including some supervised methods, thus validating the discriminant capacity of the representations learned by TAE. TAE's ability in decoupling the action-induced and pose-induced movements is also validated by visualizing the generated images and analyzing the facial image retrieval results qualitatively and quantitatively.
Yong Li 0032, Jiabei Zeng, Shiguang Shan
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Graph Jigsaw Learning for Cartoon Face Recognition
abstract
Cartoon face recognition is challenging as they typically have smooth color regions and emphasized edges, the key to recognizing cartoon faces is to precisely perceive their sparse and critical shape patterns. However, it is quite difficult to learn a shape-oriented representation for cartoon face recognition with convolutional neural networks (CNNs). To mitigate this issue, we propose the GraphJigsaw that constructs jigsaw puzzles at various stages in the classification network and solves the puzzles with the graph convolutional network (GCN) in a progressive manner. Solving the puzzles requires the model to spot the shape patterns of the cartoon faces as the texture information is quite limited. The key idea of GraphJigsaw is constructing a jigsaw puzzle by randomly shuffling the intermediate convolutional feature maps in the spatial dimension and exploiting the GCN to reason and recover the correct layout of the jigsaw fragments in a self-supervised manner. The proposed GraphJigsaw avoids training the classification model with the deconstructed images that would introduce noisy patterns and are harmful for the final classification. Specially, GraphJigsaw can be incorporated at various stages in a top-down manner within the classification model, which facilitates propagating the learned shape patterns gradually. GraphJigsaw does not rely on any extra manual annotation during the training process and incorporates no extra computation burden at inference time. Both quantitative and qualitative experimental results have verified the feasibility of our proposed GraphJigsaw, which consistently outperforms other face recognition or jigsaw-based methods on two popular cartoon face datasets with considerable improvements.
Yong Li 0032, Lingjie Lao, Zhen Cui 0001, Shiguang Shan, Jian Yang 0003
IEEE Trans. Image Process.1
2021 Transfer Vision Patterns for Multi-Task Pixel Learning
abstract
Multi-task pixel perception is one of the most important topics in the field of machine intelligence. Inspired by the observation of cross-task interdependencies of visual patterns, we propose a multi-task vision pattern transformation (VPT) method to adaptively correlate and transfer cross-task visual patterns by leveraging the powerful transformer mechanism. To better transfer visual patterns, specifically, we build two types of pattern transformation based on the statistic prior that the affinity relations across tasks are correlated. One aims to transfer feature patterns for the integration of different task features; the other aims to exchange structure patterns for mining and leveraging the latent interaction cues. These two types of transformations are encapsulated into two VPT units, which provide universal matching interfaces for multi-task learning, complement each other to guide the transmission of feature/structure patterns, and finally realize an adaptive selection of important patterns across tasks. Extensive experiments on the joint learning of semantic segmentation, depth prediction and surface normal estimation demonstrate that our proposed method is more effective than those baselines and achieve the state-of-that-art performance in three pixel-level visual tasks.
Yong Li 0032, Zhen Cui 0001, Jin Xie 0001, Jian Yang 0003
ACM Multimedia3
2021 Localizing Anomalies From Weakly-Labeled Videos
abstract
Video anomaly detection under video-level labels is currently a challenging task. Previous works have made progresses on discriminating whether a video sequence contains anomalies. However, most of them fail to accurately localize the anomalous events within videos in the temporal domain. In this paper, we propose a Weakly Supervised Anomaly Localization (WSAL) method focusing on temporally localizing anomalous segments within anomalous videos. Inspired by the appearance difference in anomalous videos, the evolution of adjacent temporal segments is evaluated for the localization of anomalous segments. To this end, a high-order context encoding model is proposed to not only extract semantic representations but also measure the dynamic variations so that the temporal context could be effectively utilized. In addition, in order to fully utilize the spatial context information, the immediate semantics are directly derived from the segment representations. The dynamic variations as well as the immediate semantics, are efficiently aggregated to obtain the final anomaly scores. An enhancement strategy is further proposed to deal with noise interference and the absence of localization guidance in anomaly detection. Moreover, to facilitate the diversity requirement for anomaly detection benchmarks, we also collect a new traffic anomaly (TAD) dataset which specifies in the traffic conditions, differing greatly from the current popular anomaly detection evaluation benchmarks. Thedataset and the benchmark test codes, as well as experimental results, are made public on http://vgg-ai.cn/pages/Resource/ and https://github.com/ktr-hubrt/WSAL. Extensive experiments are conducted to verify the effectiveness of different components, and our proposed method achieves new state-of-the-art performance on the UCF-Crime and TAD datasets.
Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Yong Li 0032, Jian Yang 0003
IEEE Trans. Image Process.5
2019 Self-Supervised Representation Learning From Videos for Facial Action Unit Detection
abstract
In this paper, we aim to learn discriminative representation for facial action unit (AU) detection from large amount of videos without manual annotations. Inspired by the fact that facial actions are the movements of facial muscles, we depict the movements as the transformation between two face images in different frames and use it as the self-supervisory signal to learn the representations. However, under the uncontrolled condition, the transformation is caused by both facial actions and head motions. To remove the influence by head motions, we propose a Twin-Cycle Autoencoder (TCAE) that can disentangle the facial action related movements and the head motion related ones. Specifically, TCAE is trained to respectively change the facial actions and head poses of the source face to those of the target face. Our experiments validate TCAE's capability of decoupling the movements. Experimental results also demonstrate that the learned representation is discriminative for AU detection, where TCAE outperforms or is comparable with the state-of-the-art self-supervised learning methods and supervised AU detection methods.
Yong Li 0032, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001
CVPR1
2019 Occlusion Aware Facial Expression Recognition Using CNN With Attention Mechanism
abstract
Facial expression recognition in the wild is challenging due to various un-constrained conditions. Although existing facial expression classifiers have been almost perfect on analyzing constrained frontal faces, they fail to perform well on partially occluded faces that are common in the wild. In this paper, we propose a Convolution Neutral Network with attention mechanism (ACNN) that can perceive the occlusion regions of the face and focus on the most discriminative unoccluded regions. ACNN is an end to end learning framework. It combines the multiple representations from facial regions of interest (ROIs). Each representation is weighed via a proposed Gate Unit that computes an adaptive weight from the region itself according to the unobstructed-ness and importance. Considering different RoIs, we introduce two versions of ACNN: patch based ACNN (pACNN) and global-local based ACNN (gACNN). pACNN only pays attention to local facial patches. gACNN integrates local representations at patch-level with global representation at image-level. The proposed ACNNs are evaluated on both real and synthetic occlusions, including a self-collected facial expression dataset with real-world occlusions (FED-RO), two largest in-the-wild facial expression datasets (RAF-DB and AffectNet) and their modifications with synthesized facial occlusions. Experimental results show that ACNNs improve the recognition accuracy on both the non-occluded faces and occluded faces. Visualization results demonstrate that, compared with the CNN without Gate Unit, ACNNs are capable of shifting the attention from the occluded patches to other related but unobstructed ones. ACNNs also outperform other state-of-the-art methods on several widely used in-the-lab facial expression datasets under the cross-dataset evaluation protocol.
Yong Li 0032, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Image Process.1
2018 Patch-Gated CNN for Occlusion-aware Facial Expression Recognition
abstract
Facial expression recognition in the wild is challenging due to various un-constrained conditions. Although existing facial expression classifiers have been almost perfect on analyzing constrained frontal faces, they fail to perform well on partially occluded faces that are common in the wild. In this paper, we propose an end-to-end trainable Patch-Gated Convolution Neutral Network (PG-CNN) that can automatically percept the occluded region of the face and focus on the most discriminative un-occluded regions. To determine the possible regions of interest on the face, PG-CNN decomposes an intermediate feature map into several patches according to the positions of related facial landmarks. Then, via a proposed Patch-Gated Unit, PG-CNN reweighs each patch by the unobstructed-ness or importance that is computed from the patch itself. The proposed PG-CNN is evaluated on two largest in-the-wild facial expression datasets (RAF-DB and AffectNet) and their modifications with synthesized facial occlusions. Experimental results show that PG-CNN improves the recognition accuracy on both the original faces and faces with synthesized occlusions. Visualization results demonstrate that, compared with the CNN without Patch-Gated Unit, PG-CNN is capable of shifting the attention from the occluded patch to other related but unobstructed ones. Experiments also show that PG-CNN outperforms other state-of-the-art methods on several widely used in-the-lab facial expression datasets under the cross-dataset evaluation protocol.
Yong Li 0032, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001
ICPR1