Dongmei Jiang

dblp:47/1926 · DBLP profile ↗
← Back
89ranked-venue papers
7as first author
55since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 4 first-author · 33 since 2021Artificial intelligence and machine learning · 34 · 30 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-authorComputer networks · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bolster Hallucination Detection via Prompt-Guided Data Augmentation
abstract
Large language models (LLMs) have garnered significant interest in AI community. Despite their impressive generation capabilities, they have been found to produce misleading or fabricated information, a phenomenon known as hallucinations. Consequently, hallucination detection has become critical to ensure the reliability of LLM-generated content. One primary challenge in hallucination detection is the scarcity of well-labeled datasets containing both truthful and hallucinated outputs. To address this issue, we introduce Prompt-guided data Augmented haLlucination dEtection (PALE), a novel framework that leverages prompt-guided responses from LLMs as data augmentation for hallucination detection. This strategy can generate both truthful and hallucinated data under prompt guidance at a relatively low cost. To more effectively evaluate the truthfulness of the sparse intermediate embeddings produced by LLMs, we introduce an estimation metric called the Contrastive Mahalanobis Score (CM Score). This score is based on modeling the distributions of truthful and hallucinated data in the activation space. CM Score employs a matrix decomposition approach to more accurately capture the underlying structure of these distributions. Importantly, our framework does not require additional human annotations, offering strong generalizability and practicality for real-world applications. Extensive experiments demonstrate that PALE achieves superior hallucination detection performance, outperforming the competitive baseline by a significant margin of 6.55%.
Wenyun Li 0001, Zheng Zhang 0006, Dongmei Jiang, Xiangyuan Lan
AAAI3
2026 Reusable Experiences: Latent Routing and Modular Composition in LLMs
abstract
Large language models (LLMs) have remarkable capabilities, but adapting them to specialized domains poses a fundamental question: how should accumulated experience be represented and leveraged?Existing approaches represent experience either as explicit textual artifacts in prompts (e.g., retrieved documents or dialogues) or implicitly within model weights via fine-tuning (e.g., LoRA adapters).However, textual methods are limited by context windows and cannot internalize knowledge, while parametric fine-tuning yields one adapter per task with minimal cross-task skill reuse.We propose ReX (Reusable eXperience), an experience-centric adaptation framework that treats latent experiences -recurring reasoning patterns and skills -as fundamental units for LLM specialization.Our method learns a shared Experience Bank of foundational skill vectors and uses a VAE-based encoder to map each input to a low-dimensional experience code.An Experience Router then dynamically composes the relevant skill vectors from this bank into a lightweight adapter for that input.By reusing skills across inputs, ReX enables implicit knowledge sharing across tasks without any explicit task identifiers.Experiments on multi-task NLP benchmarks show that this approach outperforms standard task-specific fine-tuning, yielding improved generalization through flexible skill reuse.Code is available at https://github. com/iLearn-Lab/ACL26-ReX.
Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan
ACL (1)3
2026 RA3-FDA: Resource-adaptive federated domain adaptation with dual heterogeneity awareness for EEG-based depression detection
Siyang Song, Huaning Wang, Jiewei Jiang, Dongmei Jiang, Jie Zhang 0028, Prayag Tiwari, Jiaqing Liu
Expert Syst. Appl.8
2026 DyToS: Budget-aware dynamic token scheduling for efficient multi-modal large language models
Yifei Xing 0001, Ruiping Wang 0001, Dongmei Jiang, Xiangyuan Lan
Neurocomputing6
2026 CoSI-Gaze: Context-Spatial Integration for gaze target detection and social gaze prediction
abstract
Understanding gaze behavior is a fundamental aspect of human social perception and a challenging problem in computer vision. Social gaze encompass not only where a person is looking but also how their gaze functions within a social context to establish connections, regulate conversations, and convey meaning. Existing social gaze analysis methods either primarily focus on detecting the spatial location of gaze targets while overlooking contextual cues, or rely exclusively on semantic representations, neglecting the spatial consistency between gaze targets and social gaze patterns. To bridge this gap, we propose CoSI-Gaze(CoSI), a Context-Spatial Integration framework for joint gaze target detection and social gaze prediction. To balance the contributions of contextual information and spatial consistency, CoSI estimates the reliability of gaze target predictions and adaptively adjusts this reliability in the integration process. To evaluate the framework’s ability to understand social gaze behaviors, we introduce DyGaze, the first dataset of dyadic interactions annotated with both gaze targets and five social gaze patterns (mutual, shared, single, miss, and void). Extensive experiments demonstrate that CoSI achieves state-of-the-art performance across DyGaze and other gaze pattern prediction benchmarks.
Fei Chang, Jiabei Zeng, Dongmei Jiang, Shiguang Shan
Pattern Recognit.3
2026 Dual-Attention based prompt generation and catalyzing for instance-wise continual learning
Xiaopeng Hong, Yabin Wang 0001, Zhiheng Ma, Jinfeng Yang, Dongmei Jiang, Yaowei Wang 0001
Pattern Recognit.6
2026 AlignMamba-2: Enhancing multimodal fusion and sentiment analysis with modality-aware Mamba
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
Pattern Recognit.6
2026 Facial Action Units Generation via Cross-Modality Attention Fusion and Calibrated Denoising
Chenyue Liang, Zhenliang He, Jiabei Zeng, Dongmei Jiang, Shiguang Shan
IEEE Signal Process. Lett.4
2026 Appearance- and Relation-Aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
abstract
The key to facial expression recognition is to learn discriminative spatial-temporal representations that embed facial expression dynamics. Previous studies predominantly rely on pre-trained Convolutional Neural Networks (CNNs) to learn facial appearance representations, overlooking the relationships between facial regions. To address this issue, this paper presents an Appearance- and Relation-aware Parallel Graph attention fusion Network (ARPGNet) to learn mutually enhanced spatial-temporal representations of appearance and relation information. Specifically, we construct a facial region relation graph and leverage the graph attention mechanism to model the relationships between facial regions. The resulting relational representation sequences, along with CNN-based appearance representation sequences, are then fed into a parallel graph attention fusion module for mutual interaction and enhancement. This module simultaneously explores the complementarity between different representation sequences and the temporal dynamics within each sequence. Experimental results on three facial expression recognition datasets demonstrate that the proposed ARPGNet outperforms or is comparable to state-of-the-art methods.
Yan Li 0121, Xiaohan Xia, Dongmei Jiang
IEEE Trans. Affect. Comput.4
2026 FedDAAM: Federated Domain Adversarial Learning With Attention Mechanism for Privacy Preserving Multimodal Depression Assessment
abstract
Major depressive disorder (MDD) is projected to become one of the leading mental disorders by 2030. While audiovisual cues have garnered significant attention in depression recognition research owing to their non-invasive acquisition and rich emotional expressiveness. However, conventional centralized training paradigms raise substantial privacy concerns for individuals with depression and are further hindered by data heterogeneity and label inconsistency across datasets. To overcome these challenges, a hybrid architecture, termed Federated Domain Adversarial with Attention Mechanism (FedDAAM), for privacy preserving multimodal depression assessment, is proposed. FedDAAM introduces a mechanism by differentiating discriminative features into depression-public and depression-private features. Specifically, to extract visual depression-private features from the AVEC2013 and AVEC2014 datasets, a local attention-aware (LAA) architecture is developed. For the depression-public features, action units (AUs), landmarks, head poses, and eye gazes features are adopted. In addition, to consider the transferability and performance of individual client, a dynamic parameter aggregation mechanism, termed FedDyA, is proposed. Extensive validations are performed on the AVEC2013, AVEC2014 and AVEC2017 databases, resulting in root mean square error (RMSE) and mean absolute error (MAE) of 8.61/6.78, 8.59/6.77, and 4.71/3.68, respectively. More importantly, to the best of our knowledge, this is the first study to borrow federated learning (FL) for multimodal depression assessment. The proposed framework offers a novel solution for privacy-aware, distributed clinical diagnosis of depression. Code will be available at: https://github.com/helang818/FedDAAM/.
Weizhao Yang, Junnan Zhao, Dongmei Jiang
IEEE Trans. Circuits Syst. Video Technol.5
2026 Rethinking the Knowledge Gap Between Cloud and Device Models for Effective Co-Adaptation
abstract
By collaboratively updating the cloud (large-scale) and device (small-scale) models, co-adaptation aims to enhance the generalization performance of device models in response to the distribution shifts in the incoming data. Existing methods often rely on low-entropy samples that are selected by thedevice modelfor co-adaptation, which ignores the differences between the predictions of the cloud and device models that are caused by the knowledge gap. As a result, some of the selected samples are redundant and contribute limited value to cloud model updating and knowledge distillation. To this end, we propose a test-time co-adaptation method by Rethinking the Knowledge Gap (RKG) between cloud and device models, which effectively updates the models by informative sample selection and targeted knowledge distillation for image-based classification tasks. Specifically, we design a sample selection module that integrates semantic prediction entropy with object structure cues to identify valuable samples, which effectively alleviates the redundancy problem. Based on these selected samples, we further construct a reweighting module that measures the prediction consistency between the two models and assigns greater emphasis to samples with larger prediction discrepancies, i.e., larger knowledge gaps, to improve knowledge distillation. Furthermore, by jointly leveraging these two modules, RKG enables efficient and effective co-adaptation, thereby achieving robust model generalization to continuously changing data in classification scenarios. Extensive experiments demonstrate that RKG outperforms state-of-the-art methods while requiring fewer uploaded samples.
Yingjian Li 0001, Yushi Zeng, Dongmei Jiang, Yaowei Wang 0001, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2026 DreamAssemble: Complex Multi-Object Text-to-3D Generation via Multi-Density Neural Fields
Bin Huang 0016, Jinbao Wang 0001, Dongmei Jiang, Hongjuan Pei, Qiulu Li, Jian Xue 0002, Ke Lu 0002
IEEE Trans. Image Process.3
2025 Unsupervised Degradation Representation Aware Transform for Real-World Blind Image Super-Resolution
abstract
Blind image super-resolution (blind SR) aims to restore a high-resolution (HR) image from a low-resolution (LR) image with unknown degradation. Many existing methods explicitly estimate degradation information from various LR images. However, in most cases, image degradations are independent of image content. Their estimations may be influenced by the image content resulting in inaccuracy. Unlike existing works, we design a dual-encoder for degradation representation (DEDR) to preclude the influence of image content from LR images. This benefits in extracting the intrinsic degradation representation more accurately. To the best of our knowledge, this paper is the first work that estimates degradation representation through filtering out image content. Based on the degradation representation extracted by DEDR, we present a novel framework, named degradation representation aware transform network (DRAT) for blind SR. We propose global degradation aware (GDA) blocks to propagate degradation information across spatial and channel dimensions, in which a degradation representation transform module (DRT) is introduced to render features degradation-aware, thereby enhancing the restoration of LR images. Extensive experiments are conducted on three benchmark datasets (including Gaussian 8, DIV2KRK, and real-world datasets) under large scaling factors with complex degradations. The experimental results demonstrate that DRAT surpasses state-of-the-art supervised kernel estimation and unsupervised degradation representation methods.
Hongying Liu 0001, Chaowei Fang, Fanhua Shang, Yuanyuan Liu 0001, Dongmei Jiang
AAAI7
2025 Transferable Adversarial Face Attack with Text Controlled Attribute
abstract
Traditional adversarial attacks typically produce adversarial examples under norm-constrained conditions, whereas unrestricted adversarial examples are free-form with semantically meaningful perturbations. Current unrestricted adversarial impersonation attacks exhibit limited control over adversarial face attributes and often suffer from low transferability. In this paper, we propose a novel Text Controlled Attribute Attack (TCA2) to generate photorealistic adversarial impersonation faces guided by natural language. Specifically, the category-level personal softmax vector is employed to precisely guide the impersonation attacks. Additionally, we propose both data and model augmentation strategies to achieve transferable attacks on unknown target models. Finally, a generative model, i.e, Style-GAN, is utilized to synthesize impersonated faces with desired attributes. Extensive experiments on two high-resolution face recognition datasets validate that our TCA2 method can generate natural text-guided adversarial impersonation faces with high transferability. We also evaluate our method on real-world face recognition systems, i.e, Face++ and Aliyun, further demonstrating the practical potential of our approach.
Wenyun Li 0001, Zheng Zhang 0006, Xiangyuan Lan, Dongmei Jiang
AAAI4
2025 Learning Hierarchical Continuous Dynamics for Facial Action Unit Intensity Estimation
abstract
Dynamic facial action recognition is key to understanding human emotions and behaviors, yet estimating facial action units (AUs) intensities in videos is difficult due to subtle muscle motions and complex spatial-temporal dependencies. Existing methods often use fixed or coarse graphs, limiting the ability to capture intricate AU relations and long-range dynamics. This paper presents a hierarchical framework CDAU to effectively capture Continuous Dynamics for AU intensity estimation task. Our approach dynamically constructs multiscale graphs for fine-grained spatiotemporal AU interactions and adaptively fusing information across levels. A bidirectional state-space module further captures long-range temporal dependencies. Extensive experiments on FEAFA and DISFA show that CD-AU outperforms existing methods in both ICC and MAE metrics, validating its generalization capability and stability across subjects and expressions.
Ke Lu 0002, Yan Li 0121, Menghao Hu, Guohong Hu, Dongmei Jiang, Jian Xue 0002
BIBM6
2025 Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues
abstract
Understanding human behavior and environmental information in egocentric videos is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has shown promising results. However, existing visual-to-visual and visual-to-textual Ego-Exo video alignment methods struggle with the issue that some activities may have non-visual overlap. To address this, we propose using sound as a bridge, as audio is often consistent across Ego-Exo videos. However, direct audio-to-audio alignment lacks context. Thus, we introduce two context-aware sound modules: one aligns audio with vision via a visual-audio cross-attention module, and another aligns text with sound closed caption generated by LLM. Experimental results on two Ego-Exo video association benchmarks show that each of the proposed modules enhances the state-of-the-art methods. Moreover, the proposed sound-aware egocentric or exocentric representation boosts the performance of downstream tasks, such as action recognition of exocentric videos and scene recognition of egocentric videos. The code and models can be accessed at https://github.com/shhuangcoder/SoundBridge.
Sihong Huang, Jiaxin Wu 0001, Xiaoyong Wei, Yi Cai 0001, Dongmei Jiang, Yaowei Wang 0001
CVPR5
2025 AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
abstract
Cross-Modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-Based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-sequence or large-scale data. Although recent Mamba-Based approaches achieve linear complexity, their sequential scanning mechanism poses fundamental challenges in comprehensively modeling cross-modal relationships. To address this limitation, we propose Align-Mamba, an efficient and effective method for multimodal fusion. Specifically, grounded in Optimal Transport, we introduce a local cross-modal alignment module that explicitly learns token-level correspondences between different modalities. Moreover, we propose a global cross-modal alignment loss based on Maximum Mean Discrepancy to implicitly enforce the consistency between different modal distributions. Finally, the unimodal representations after local and global alignment are passed to the Mamba backbone for further cross-modal interaction and multimodal fusion. Extensive experiments on complete and incomplete multimodal fusion tasks demonstrate the effectiveness and efficiency of the proposed method. For instance, on the CMU-MOSI dataset, AlignMamba improves classification accuracy by 0.9%, reduces GPU memory usage by 20.3%, and decreases inference time by 83.3%.
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
CVPR6
2025 Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy
abstract
Building an agent that can mimic human behavior patterns to accomplish various open-world tasks is a longterm goal. To enable agents to effectively learn behavioral patterns across diverse tasks, a key challenge lies in modeling the intricate relationships among observations, actions, and language. To this end, we propose Optimus-2, a novel Minecraft agent that incorporates a Multimodal Large Language Model (MLLM) for highlevel planning, alongside a Goal-Observation-Action Conditioned Policy (GOAP) for low-level control. GOAP contains (1) an Action-guided Behavior Encoder that models causal relationships between observations and actions at each timestep, then dynamically interacts with the historical observation-action sequence, consolidating it into fixedlength behavior tokens, and (2) an MLLM that aligns behavior tokens with open-ended language instructions to predict actions auto-regressively. Moreover, we introduce a high-quality Minecraft Goal-Observation-Action (MGOA) dataset, which contains 25,000 videos across 8 atomic tasks, providing about 30M goal-observation-action pairs. The automated construction method, along with the MGOA dataset, can contribute to the community’s efforts to train Minecraft agents. Extensive experimental results demonstrate that Optimus-2 exhibits superior performance across atomic tasks, long-horizon tasks, and open-ended instruction tasks in Minecraft. Please see the project page at https://cybertronagent.github.io/Optimus-2.github.io/.
Zaijing Li, Yuquan Xie, Rui Shao 0001, Gongwei Chen, Dongmei Jiang, Liqiang Nie
CVPR5
2025 Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval
abstract
Asymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for pre-determined platforms, using compatible learning. Nevertheless, these methods suffer from limited flexibility for multi-platform deployment. For example, when introducing a new platform into the retrieval systems, developers have to train an additional model at an appropriate capacity that is compatible with existing models via backward-compatible learning. In this paper, we propose a Prunable Network with self-compatibility, which allows developers to generate compatible subnetworks at any desired capacity through post-training pruning. Thus it allows the creation of a sparse subnetwork matching the resources of the new platform without additional training. Specifically, we optimize both the architecture and weight of subnetworks at different capacities within a dense network in compatible learning. We also design a conflict-aware gradient integration scheme to handle the gradient conflicts between the dense network and subnetworks during compatible learning. Extensive experiments on diverse benchmarks and visual backbones demonstrate the effectiveness of our method. The code will be made publicly available.
Yushuai Sun, Zikun Zhou, Dongmei Jiang, Yaowei Wang 0001, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei
CVPR3
2025 Comprehensive Perturbation Consistency for Semi-Supervised Change Detection in Remote Sensing Images
abstract
Currently, many change detection (CD) methods rely on supervised learning, which necessitates extensive manually annotated data, resulting in significant labor and time requirements. Recently, semi-supervised (SS) approaches have emerged in the CD community, which exploit large amounts of unlabeled data by utilizing consistency regularization. However, these methods do not consider the broader perturbation consistency to confer better generalization of the model. In this paper, we propose a novel SS CD framework with a comprehensive perturbation consistency called CPC, which extends perturbation consistency to the entire learning period. Specifically, our CPC combines the input, feature, and network perturbations for comprehensive perturb space. And, we design two distinct structures in practice, decoupled CPC and coupled CPC. Furthermore, we propose a change-aware input perturbation that introduces expensive annotation information to further expand the input perturb space. Extensive experiments conducted on the WHUCD, and GZ-CD datasets demonstrate that the proposal performs favorably against the state-of-the-art methods.
Zan Mao, Ze Luo, Yingjuan Tang, Dongmei Jiang
ICASSP5
2025 EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
abstract
Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanced cross-modal alignment between visual and textural latents, negatively impacting performance on multi-modal tasks. In this work, we propose Empowering Multi-modal Mamba with Structural and Hierarchical Alignment (EMMA), which enables the MLLM to extract fine-grained visual information. Specifically, we propose a pixel-wise alignment module to autoregressively optimize the learning and processing of spatial image-level features along with textual tokens, enabling structural alignment at the image level. In addition, to prevent the degradation of visual information during the cross-model alignment process, we propose a multi-scale feature fusion (MFF) module to combine multi-scale visual features from intermediate layers, enabling hierarchical alignment at the feature level. Extensive experiments are conducted across a variety of multi-modal benchmarks. Our model shows lower latency than other Mamba-based MLLMs and is nearly four times faster than transformer-based MLLMs of similar scale during inference. Due to better cross-modal alignment, our model exhibits lower degrees of hallucination and enhanced sensitivity to visual details, which manifests in superior performance across diverse multi-modal benchmarks. Code provided at https://github.com/xingyifei2016/EMMA.
Yifei Xing 0001, Xiangyuan Lan, Ruiping Wang 0001, Dongmei Jiang, Yaowei Wang 0001
ICLR4
2025 CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models
abstract
Virtual try-on methods based on diffusion models achieve realistic effects but often require additional encoding modules, a large number of training parameters, and complex preprocessing, which increases the burden on training and inference. In this work, we re-evaluate the necessity of additional modules and analyze how to improve training efficiency and reduce redundant steps in the inference process. Based on these insights, we propose CatVTON, a simple and efficient virtual try-on diffusion model that transfers in-shop or worn garments of arbitrary categories to target individuals by concatenating them along spatial dimensions as inputs of the diffusion model. The efficiency of CatVTON is reflected in three aspects: (1) Lightweight network. CatVTON consists only of a VAE and a simplified denoising UNet, removing redundant image and text encoders as well as cross-attentions, and includes just 899.06M parameters. (2) Parameter-efficient training. Through experimental analysis, we identify self-attention modules as crucial for adapting pre-trained diffusion models to the virtual try-on task, enabling high-quality results with only 49.57M training parameters. (3) Simplified inference. CatVTON eliminates unnecessary preprocessing, such as pose estimation, human parsing, and captioning, requiring only a person image and garment reference to guide the virtual try-on process, reducing over 49% memory usage compared to other diffusion-based methods. Extensive experiments demonstrate that CatVTON achieves superior qualitative and quantitative results compared to baseline methods and demonstrates strong generalization performance in in-the-wild scenarios, despite being trained solely on public datasets with 73K samples.
Zheng Chong, Xujie Zhang, Dongmei Jiang, Xiaodan Liang
ICLR8
2025 PolaFormer: Polarity-aware Linear Attention for Vision Transformers
abstract
Linear attention has emerged as a promising alternative to softmax-based attention, leveraging kernelized feature maps to reduce complexity from quadratic to linear in sequence length. However, the non-negative constraint on feature maps and the relaxed exponential function used in approximation lead to significant information loss compared to the original query-key dot products, resulting in less discriminative attention maps with higher entropy. To address the missing interactions driven by negative values in query-key pairs, we propose a polarity-aware linear attention mechanism that explicitly models both same-signed and opposite-signed query-key interactions, ensuring comprehensive coverage of relational information. Furthermore, to restore the spiky properties of attention maps, we provide a theoretical analysis proving the existence of a class of element-wise functions (with positive first and second derivatives) that can reduce entropy in the attention distribution. For simplicity, and recognizing the distinct contributions of each dimension, we employ a learnable power function for rescaling, allowing strong and weak attention signals to be effectively separated. Extensive experiments demonstrate that the proposed PolaFormer improves performance on various vision tasks, enhancing both expressiveness and efficiency by up to 4.6%.
Weikang Meng, Yadan Luo, Xin Li 0003, Dongmei Jiang, Zheng Zhang 0006
ICLR4
2025 DTAD: A Distribution-Transformed Supervised Anomaly Detection Method
abstract
Most anomaly detection (AD) methods adopt an unsupervised approach, relying exclusively on normal samples during training, which limits the model’s discriminative ability. In real world scenarios, only small amounts of anomaly data are typically available, which can still provide valuable insights for model learning. However, supervised anomaly detection methods may be impacted by the scarcity of anomaly samples, leading to significant overlap in the feature distributions of normal and anomaly samples, thereby degrading overall performance. To address this, we propose a distribution-transformed supervised anomaly detection method (DTAD). This method employs a two-stage distribution transformation to progressively reduce the overlap between normal and anomaly distributions, thereby enhancing the model’s discriminative performance. In the first stage, residual calculation is used to initially separate normal and abnormal distributions, with an attention network highlighting critical feature contributions. In the second stage, a contrastive loss function (semi-push-pull loss) is employed to expand the decision boundary between normal and abnormal samples, further improving distribution separation. On the benchmark datasets MVTecAD, AITEX, ELPV, and BrainMRI, our method outperforms recent state-of-the-art approaches, demonstrating its effectiveness.
Lingxing Chen, Yang Gu 0001, Jianqi Chen, Yingting Zhu, Yehong Zhuo, Dongmei Jiang, Yiqiang Chen 0001
ICME7
2025 Open-Det: An Efficient Learning Framework for Open-Ended Detection
abstract
Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training, suffer from slow convergence, and exhibit limited performance. To address these issues, we present a novel and efficient Open-Det framework, consisting of four collaborative parts. Specifically, Open-Det accelerates model training in both the bounding box and object name generation process by reconstructing the Object Detector and the Object Name Generator. To bridge the semantic gap between Vision and Language modalities, we propose a Vision-Language Aligner with V-to-L and L-to-V alignment mechanisms, incorporating with the Prompts Distiller to transfer knowledge from the VLM into VL-prompts, enabling accurate object name generation for the LLM. In addition, we design a Masked Alignment Loss to eliminate contradictory supervision and introduce a Joint Loss to enhance classification, resulting in more efficient training. Compared to GenerateU, Open-Det, using only 1.5% of the training data (0.077M vs. 5.077M), 20.8% of the training epochs (31 vs. 149), and fewer GPU resources (4 V100 vs. 16 A100), achieves even higher performance (+1.0% in APr). The source codes are available at: https://github.com/Med-Process/Open-Det.
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang
ICML6
2025 DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
abstract
Popular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g., content query and positional query) are still underexplored. These queries are generally predefined with a fixed number (fixed-query), which limits their flexibility. We find that the learning of these fixed-query is impaired by Recurrent Opposing in Teractions (ROT) between two attention operations: Self-Attention (query-to-query) and Cross-Attention (query-to-encoder), thereby degrading decoder efficiency. Furthermore, "query ambiguity" arises when shared-weight decoder layers are processed with both one-to-one and one-to-many label assignments during training, violating DETR's one-to-one matching principle. To address these challenges, we propose DS-Det, a more efficient detector capable of detecting a flexible number of objects in images. Specifically, we reformulate and introduce a new unified Single-Query paradigm for decoder modeling, transforming the fixed-query into flexible. Furthermore, we propose a simplified decoder framework through attention disentangled learning: locating boxes with Cross-Attention (one-to-many process), deduplicating predictions with Self-Attention (one-to-one process), addressing ''query ambiguity'' and ''ROT'' issues directly, and enhancing decoder efficiency. We further introduce a unified PoCoo loss that leverages box size priors to prioritize query learning on hard samples such as small objects. Extensive experiments across five different backbone models on COCO2017 and WiderPerson datasets demonstrate the general effectiveness and superiority of DS-Det. The source codes are available at https://github.com/Med-Process/DS-Det/.
Guiping Cao, Xiangyuan Lan, Wenjian Huang 0001, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
ACM Multimedia5
2025 GraphATC: advancing multilevel and multi-label anatomical therapeutic chemical classification via atom-level graph learning
abstract
The accurate categorization of compounds within the anatomical therapeutic chemical (ATC) system is fundamental for drug development and fundamental research. Although this area has garnered significant research focus for over a decade, the majority of prior studies have concentrated solely on the Level 1 labels defined by the World Health Organization (WHO), neglecting the labels of the remaining four levels. This narrow focus fails to address the true nature of the task as a multilevel, multi-label classification challenge. Moreover, existing benchmarks like Chen-2012 and ATC-SMILES have become outdated, lacking the incorporation of new drugs or updated properties of existing ones that have emerged in recent years and have been integrated into the WHO ATC system. To tackle these shortcomings, we present a comprehensive approach in this paper. Firstly, we systematically cleanse and enhance the drug dataset, expanding it to encompass all five levels through a rigorous cross-resource validation process involving KEGG, PubChem, ChEMBL, ChemSpider, and ChemicalBook. This effort culminates in the creation of a novel benchmark termed ATC-GRAPH. Secondly, we extend the classification task to encompass Level 2 and introduce graph-based learning techniques to provide more accurate representations of drug molecular structures. This approach not only facilitates the modeling of Polymers, Macromolecules, and Multi-Component drugs more precisely but also enhances the overall fidelity of the classification process. The efficacy of our proposed framework is validated through extensive experiments, establishing a new state-of-the-art methodology. To facilitate the replication of this study, we have made the benchmark dataset, source code, and web server openly accessible.
Wengyu Zhang, Qi Tian 0001, Wenqi Fan, Dongmei Jiang, Yaowei Wang 0001, Qing Li 0001, Xiaoyong Wei
Briefings Bioinform.5
2025 DDC: Dynamic distribution calibration for few-shot learning under multi-scale representation
Lingxing Chen, Yang Gu 0001, Dongmei Jiang, Yiqiang Chen 0001
Knowl. Based Syst.5
2025 AVES: An Audio-Visual Emotion Stream Dataset for Temporal Emotion Detection
abstract
Human emotions vary over time, which can be vividly described as a stream of emotions. Observing the emotion stream in daily life provides valuable insights into an individual's mental state. However, existing research in emotion understanding has mainly focused on classification tasks, assigning an emotion category to a well-trimmed segment or each frame within a continuous signal. In contrast, the task of temporal emotion detection, which involveslocatingthe boundaries of emotion segments andrecognizingtheir categories in untrimmed signals, has not been fully explored. To advance research in this area, this paper introduces an in-the-wild Audio-Visual Emotion Stream (AVES) dataset, which is reliably annotated with the time boundaries and emotion category for each emotion segment in the videos. Thus, AVES can serve as a solid benchmark for temporal emotion detection tasks. Moreover, considering the flexible boundaries and varying durations of emotion segments, we propose a Boundary Combination Network (BoCoNet) for temporal emotion detection, which leverages short-term temporal context information to first predict the boundaries of emotion segments and then locate the entire emotion segments. Extensive experiments conducted on various representative unimodal and multimodal representations demonstrate that BoCoNet achieves state-of-the-art results. The AVES dataset will be released to the research community. We expect that this paper can advance the research on emotion stream and temporal emotion detection.
Yan Li 0121, Ke Lu 0002, Dongmei Jiang, Ramesh Jain 0001
IEEE Trans. Affect. Comput.4
2025 Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
abstract
Small Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promising performance, their potential for SOD remains largely unexplored. In typical DETR-like frameworks, the CNN backbone network, specialized in aggregating local information, struggles to capture the necessary contextual information for SOD. The multiple attention layers in the Transformer Encoder face difficulties in effectively attending to small objects and can also lead to blurring of features. Furthermore, the model's lower class prediction score of small objects compared to large objects further increases the difficulty of SOD. To address these challenges, we introduce a novel approach calledCross-DINO. This approach incorporates the deep MLP network to aggregate initial feature representations with both short and long range information for SOD. Then, a new Cross Coding Twice Module (CCTM) is applied to integrate these initial representations to the Transformer Encoder feature, enhancing the details of small objects. Additionally, we introduce a new kind of soft label named Category-Size (CS), integrating the Category and Size of objects. By treating CS as new ground truth, we propose a new loss function called Boost Loss to improve the class prediction score of the model. Extensive experimental results on COCO, WiderPerson, VisDrone, AI-TOD, and SODA-D datasets demonstrate that Cross-DINO efficiently improves the performance of DETR-like models on SOD. Specifically, our model achieves36.4%AP$_{S}$on COCO for SOD with only 45M parameters, outperforming the DINO by+4.4%AP$_{S}$(36.4% vs. 32.0%) with fewer parameters and FLOPs, under 12 epochs training setting.
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
IEEE Trans. Multim.5
2025 ExpLLM: Towards Chain of Thought for Facial Expression Recognition
abstract
Facial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is essential for accurately recognizing them. Current approaches, such as those based on facial action units (AUs), typically provide AU names and intensities but lack insight into the interactions and relationships between AUs and the overall expression. In this paper, we propose a novel method called ExpLLM, which leverages large language models to generate an accurate chain of thought (CoT) for facial expression recognition. Specifically, we have designed the CoT mechanism from three key perspectives: key observations, overall emotional interpretation, and conclusion. The key observations describe the AU's name, intensity, and associated emotions. The overall emotional interpretation provides an analysis based on multiple AUs and their interactions, identifying the dominant emotions and their relationships. Finally, the conclusion presents the final expression label derived from the preceding analysis. Furthermore, we also introduce the Exp-CoT Engine, designed to construct this expression CoT and generate instruction-description data for training our ExpLLM. Extensive experiments on the RAF-DB and AffectNet datasets demonstrate that ExpLLM outperforms current state-of-the-art FER methods. ExpLLM also surpasses the latest GPT-4o in expression CoT generation, particularly in recognizing micro-expressions where GPT-4o frequently fails.
Xing Lan, Jian Xue 0002, Ji Qi 0003, Dongmei Jiang, Ke Lu 0002, Tat-Seng Chua
IEEE Trans. Multim.4
2024 Deep Homography Estimation for Visual Place Recognition
abstract
Visual place recognition (VPR) is a fundamental task for many applications such as robot localization and augmented reality. Recently, the hierarchical VPR methods have received considerable attention due to the trade-off between accuracy and efficiency. They usually first use global features to retrieve the candidate images, then verify the spatial consistency of matched local features for re-ranking. However, the latter typically relies on the RANSAC algorithm for fitting homography, which is time-consuming and non-differentiable. This makes existing methods compromise to train the network only in global feature extraction. Here, we propose a transformer-based deep homography estimation (DHE) network that takes the dense feature map extracted by a backbone network as input and fits homography for fast and learnable geometric verification. Moreover, we design a re-projection error of inliers loss to train the DHE network without additional homography labels, which can also be jointly trained with the backbone network to help it extract the features that are more suitable for local matching. Extensive experiments on benchmark datasets show that our method can outperform several state-of-the-art methods. And it is more than one order of magnitude faster than the mainstream hierarchical VPR methods using RANSAC. The code is released at https://github.com/Lu-Feng/DHE-VPR.
Shuting Dong, Bingxi Liu 0001, Xiangyuan Lan, Dongmei Jiang, Chun Yuan 0003
AAAI6
2024 CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place Recognition
abstract
Over the past decade, most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and neglect the cross-image variations (e.g. viewpoint and illumination), which limits their robustness in challenging scenes. In this paper, we propose a robust global representation method with cross-image correlation awareness for VPR, named CricaVPR. Our method uses the attention mechanism to correlate multiple images within a batch. These images can be taken in the same place with different conditions or viewpoints, or even captured from different places. Therefore, our method can utilize the cross-image variations as a cue to guide the representation learning, which ensures more robust features are produced. To further facilitate the robustness, we propose a multi-scale convolution-enhanced adaptation method to adapt pre-trained visual foundation models to the VPR task, which introduces the multi-scale local information to further enhance the cross-image correlation-aware representation. Experimental results show that our method out-performs state-of-the-art methods by a large margin with significantly less training time. The code is released at https://github.com/Lu-Feng/CricaVPR.
Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Chun Yuan 0003
CVPR4
2024 Facial Action Unit Detection with the Semantic Prompt
abstract
Facial action unit (AU) detection is an essential technique for fine-grained facial expression analysis. To improve the detection performance, the associations among different action units within the detection network should be exploited. In light of this, we propose to exploit the semantic corrections between AUs and improve the detection accuracy via a novel AU prompt framework. Specifically, we incorporate a pre-trained text encoder to extract the textual embeddings for AU descriptions. Then, we treat these embeddings as semantic prompts and feed them into a vision-language cross-attention module to capture the relations among AUs. The cross-attention module will adaptively aggregate the spatial features of a face image encoder, and finally generate discriminative features for each AU. Extensive experiments on BP4D, DISFA, and GFT datasets demonstrate that the proposed framework outperforms state-of-the-art methods in both within-dataset and cross-dataset settings.
Chenyue Liang, Jiabei Zeng, Dongmei Jiang, Shiguang Shan
ICME4
2024 MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
IJCAI5
2024 Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
abstract
Building a general-purpose agent is a long-standing vision in the field of artificial intelligence. Existing agents have made remarkable progress in many domains, yet they still struggle to complete long-horizon tasks in an open world. We attribute this to the lack of necessary world knowledge and multimodal experience that can guide agents through a variety of long-horizon tasks. In this paper, we propose a Hybrid Multimodal Memory module to address the above challenges. It 1) transforms knowledge into Hierarchical Directed Knowledge Graph that allows agents to explicitly represent and learn world knowledge, and 2) summarises historical information into Abstracted Multimodal Experience Pool that provide agents with rich references for in-context learning. On top of the Hybrid Multimodal Memory module, a multimodal agent, Optimus-1, is constructed with dedicated Knowledge-guided Planner and Experience-Driven Reflector, contributing to a better planning and reflection in the face of long-horizon tasks in Minecraft. Extensive experimental results show that Optimus-1 significantly outperforms all existing agents on challenging long-horizon task benchmarks, and exhibits near human-level performance on many tasks. In addition, we introduce various Multimodal Large Language Models (MLLMs) as the backbone of Optimus-1. Experimental results show that Optimus-1 exhibits strong generalization with the help of the Hybrid Multimodal Memory module, outperforming the GPT-4V baseline on many tasks.
Zaijing Li, Yuquan Xie, Rui Shao 0001, Gongwei Chen, Dongmei Jiang, Liqiang Nie
NeurIPS5
2024 A single frame and multi-frame joint network for 360-degree panorama video super-resolution
Hongying Liu 0001, Wanhao Ma, Zhubo Ruan, Chaowei Fang, Fanhua Shang, Yuanyuan Liu 0001, Chaoli Wang 0001, Dongmei Jiang
Eng. Appl. Artif. Intell.9
2024 Facial Action Unit Representation Based on Self-Supervised Learning With Ensembled Priori Constraints
abstract
Facial action units (AUs) focus on a comprehensive set of atomic facial muscle movements for human expression understanding. Based on supervised learning, discriminative AU representation can be achieved from local patches where the AUs are located. Unfortunately, accurate AU localization and characterization are challenged by the tremendous manual annotations, which limits the performance of AU recognition in realistic scenarios. In this study, we propose an end-to-end self-supervised AU representation learning model (SsupAU) to learn AU representations from unlabeled facial videos. Specifically, the input face is decomposed into six components using auto-encoders: five photo-geometric meaningful components, together with 2D flow field AUs. By constructing the canonical neutral face, posed neutral face, and posed expressional face gradually, these components can be disentangled without supervision, therefore the AU representations can be learned. To construct the canonical neutral face without manually labeled ground truth of emotion state or AU intensity, two priori knowledge based assumptions are proposed: 1) identity consistency, which explores the identical albedos and depths of different frames in a face video, and helps to learn the camera color mode as an extra cue for canonical neutral face recovery. 2) average face, which enables the model to discover a 'neutral facial expression' of the canonical neutral face and decouple the AUs in representation learning. To the best of our knowledge, this is the first attempt to design self-supervised AU representation learning method based on the definition of AUs. Substantial experiments on benchmark datasets have demonstrated the superior performance of the proposed work in comparison to other state-of-the-art approaches, as well as an outstanding capability of decomposing input face into meaningful factors for its reconstruction. The code is made available at https://github.com/Sunner4nwpu/SsupAU.
Peng Zhang 0005, Chujia Guo, Ke Lu 0002, Dongmei Jiang
IEEE Trans. Image Process.5
2023 Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based Network
abstract
Electroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDANet) to classify the correct auditory stimulus evoking the EEG signal. It adopts the Attention-based Correlation Module (ACM) to discover the connection between auditory speech and EEG from global aspect, and the Shallow-Deep Similarity Classification Module (SDSCM) to decide the classification result via the embeddings learned from the shallow and deep layers. Moreover, various training strategies and data augmentation are used to boost the model robustness. Experiments are conducted on the dataset provided by Auditory EEG challenge (ICASSP Signal Processing Grand Challenge 2023). Results show that the proposed model has a significant gain over the baseline on the match-mismatch track.
Fan Cui, Liyong Guo, Jiyao Liu, Ercheng Pei, Dongmei Jiang
ICASSP7
2023 Strip-MLP: Efficient Token Interaction for Vision MLP
abstract
Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the model’s expressive ability, especially in deep layers where the feature are down-sampled to a small spatial size. To address this issue, we present a novel method called Strip-MLP to enrich the token interaction power in three ways. Firstly, we introduce a new MLP paradigm called Strip MLP layer that allows the token to interact with other tokens in a cross-strip manner, enabling the tokens in a row (or column) to contribute to the information aggregations in adjacent but different strips of rows (or columns). Secondly, a Cascade Group Strip Mixing Module (CGSMM) is proposed to overcome the performance degradation caused by small spatial feature size. The module allows tokens to interact more effectively in the manners of within-patch and cross-patch, which is independent to the feature spatial size. Finally, based on the Strip MLP layer, we propose a novel Local Strip Mixing Module (LSMM) to boost the token interaction power in the local region. Extensive experiments demonstrate that Strip-MLP significantly improves the performance of MLP-based models on small datasets and obtains comparable or even better results on ImageNet. In particular, Strip-MLP models achieve higher average Top-1 accuracy than existing MLP-based models by +2.44% on Caltech-101 and +2.16% on CIFAR-100. The source codes will be available at https://github.com/Med-Process/Strip_MLP.
Guiping Cao, Shengda Luo, Wenjian Huang 0001, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Jianguo Zhang 0001
ICCV5
2023 Semi-Supervised Multimodal Emotion Recognition with Class-Balanced Pseudo-labeling
abstract
This paper presents our solution for the Semi-Supervised Multimodal Emotion Recognition Challenge (MER2023-SEMI), addressing the issue of limited annotated data in emotion recognition. Recently, the self-training-based Semi-Supervised Learning~(SSL) method has demonstrated its effectiveness in various tasks, including emotion recognition. However, previous studies focused on reducing the confirmation bias of data without adequately considering the issue of data imbalance, which is of great importance in emotion recognition. Additionally, previous methods have primarily focused on unimodal tasks and have not considered the inherent multimodal information in emotion recognition tasks. We propose a simple yet effective semi-supervised multimodal emotion recognition method to address the above issues. We assume that the pseudo-labeled samples with consistent results across unimodal and multimodal classifiers have a more negligible confirmation bias. Based on this assumption, we suggest using a class-balanced strategy to select top-k high-confidence pseudo-labeled samples from each class. The proposed method is validated to be effective on the MER2023-SEMI Grand Challenge, with the weighted F1 score reaching 88.53% on the test set.
Chujia Guo, Yan Li 0121, Peng Zhang 0005, Dongmei Jiang
ACM Multimedia5
2023 Towards Adaptable Graph Representation Learning: An Adaptive Multi-Graph Contrastive Transformer
abstract
Significant progress has been made in graph representation learning in recent years. However, most of these methods model spatial relationships via predefined graphs or decouple spatial-temporal representations, which limits the generalization and effectiveness of the model. To address these issues, we introduce an adaptive multi-graph contrastive transformer (AMGCT) for general spatial-temporal graph representation learning. Specifically, we first propose adaptive multi-graph contrastive learning (AMGCL). Without any expert knowledge, AMGCL can gradually generate adaptive spatial graphs with different topologies to learn spatial representations from different views. Cross-graph contrastive learning further explores potential correlations between different views, making each view's features more discriminative. In addition, to avoid insufficient interaction caused by decoupling spatial-temporal information in existing methods, we design a coupled graph transformer (CGT) to consider spatial relationships at each stage of temporal modeling, explore complementary information between spatial and temporal domains, and obtain more compact spatial-temporal representations. Experimental results on two different spatial-temporal graph datasets and tasks demonstrate that the proposed method achieves excellent performance.
Yan Li 0121, Liang Zhang 0042, Xiangyuan Lan, Dongmei Jiang
ACM Multimedia4
2023 Benign Shortcut for Debiasing: Fair Visual Recognition via Intervention with Shortcut Features
abstract
Machine learning models often learn to make predictions that rely on sensitive social attributes like gender and race, which poses significant fairness risks, especially in societal applications, such as hiring, banking, and criminal justice. Existing work tackles this issue by minimizing the employed information about social attributes in models for debiasing. However, the high correlation between target task and these social attributes makes learning on the target task incompatible with debiasing. Given that model bias arises due to the learning of bias features (i.e., gender) that help target task optimization, we explore the following research question: Can we leverage shortcut features to replace the role of bias feature in target task optimization for debiasing? To this end, we propose Shortcut Debiasing, to first transfer the target task's learning of bias attributes from bias features to shortcut features, and then employ causal intervention to eliminate shortcut features during inference. The key idea of Shortcut Debiasing is to design controllable shortcut features to on one hand replace bias features in contributing to the target task during the training stage, and on the other hand be easily removed by intervention during the inference stage. This guarantees the learning of the target task does not hinder the elimination of bias features. We apply Shortcut Debiasing to several benchmark datasets, and achieve significant improvements over the state-of-the-art debiasing methods in both accuracy and fairness.
Yi Zhang 0101, Jitao Sang 0001, Junyang Wang 0001, Dongmei Jiang, Yaowei Wang 0001
ACM Multimedia4
2023 Efficient spatiotemporal context modeling for action recognition
Congqi Cao, Yue Lu 0008, Yifan Zhang 0001, Dongmei Jiang, Yanning Zhang 0001
Neurocomputing4
2023 HiT-MST: Dynamic facial expression recognition with hierarchical transformers and multi-scale spatiotemporal aggregation
Xiaohan Xia, Dongmei Jiang
Inf. Sci.2
2023 Cross-view adaptive graph attention network for dynamic facial expression recognition
Min Xia 0002, Dongmei Jiang
Multim. Syst.3
2023 Region Attentive Action Unit Intensity Estimation With Uncertainty Weighted Multi-Task Learning
abstract
Facial action units (AUs) refer to a comprehensive set of atomic facial muscle movements. Recent works have focused on exploring complementary information by learning the relationships among AUs. Most existing approaches process AU co-occurrence and enhance AU recognition by learning the dependencies among AUs from labels, however, the complementary information among features of different AUs are ignored. Moreover, ground truth annotations suffer from a large intra-class variance and their associated intensity levels may vary depending on the annotators’ experience. In this paper, we propose the Region Attentive AU intensity estimation method with Uncertainty Weighted Multi-task Learning (RA-UWML). A RoI-Net is first used to extract features from the pre-defined facial patches where the AUs locate. Then, we use the co-occurrence of AUs using both within patch and between patches representation learning. Within a given patch, we propose sharing representation learning in a multi-task manner. To achieve complementarity and avoid redundancy between different image patches, we propose to use a multi-head self-attention mechanism to adaptively and attentively encode each patch specific representation. Moreover, the AU intensity is represented as a Gaussian distribution, instead of a single value, where the mean value indicates the most likely AU intensity and the variance indicates the uncertainty of the estimated AU intensity. The estimated variances are leveraged to automatically weight the loss of each AU in the multitask learning model. In extensive experiments on the Disfa, Fera2015 and Feafa benchmarks, it is shown that the proposed AU intensity estimation model achieves better results compared to the state-of-the-art models.
Dongmei Jiang, Xiaoyong Wei, Ke Lu 0002, Hichem Sahli
IEEE Trans. Affect. Comput.2
2023 A Bayesian Filtering Framework for Continuous Affect Recognition From Facial Images
abstract
Continuous affective state estimation from facial information is a task which requires the prediction of time series of emotional state outputs from a facial image sequence. Modeling the spatial-temporal evolution of facial information plays an important role in affective state estimation. One of the most widely used methods is Recurrent Neural Networks (RNN). RNNs provide an attractive framework for propagating information over a sequence using a continuous-valued hidden layer representation. In this work, we propose to instead learn rich affective state dynamics. We model human affect as a dynamical system and define the affective state in terms of valence, arousal and their higher-order derivatives. We then pose the affective state estimation problem as a jointly trained state estimator for high-dimensional input images, combining an RNN and a Bayesian Filter, i.e. Kalman filters (KF) and Extended Kalman filters (EKF), so that all weights in the resulting network can be trained using backpropagation. We use a recently proposed general framework for designing and learning discriminative state estimators framed as computational graphs. Such approach can handle high dimensional observations and efficiently optimize, in an end-to-end fashion, the state estimator. In addition, to deal with the asynchrony between emotion labels and input images, caused by the inherent reaction lag of the annotators, we introduce a convolutional layer that aligns features with emotion labels. Experimental results, on the RECOLA and SEMAINE datasets for continuous emotion prediction, illustrate the potential of the proposed framework compared to recent state-of-the-art models.
Ercheng Pei, Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli
IEEE Trans. Multim.4
2022 Uncertainty-Aware Semi-Supervised Learning of 3D Face Rigging from Single Image
abstract
We present a method to rig 3D faces via Action Units (AUs), viewpoint and light direction, from single input image. Existing 3D methods for face synthesis and animation rely heavily on 3D morphable model (3DMM), which was built on 3D data and cannot provide intuitive expression parameters, while AU-driven 2D methods cannot handle head pose and lighting effect. We bridge the gap by integrating a recent 3D reconstruction method with 2D AU-driven method in a semi-supervised fashion. Built upon the auto-encoding 3D face reconstruction model that decouples depth, albedo, viewpoint and light without any supervision, we further decouple expression from identity for depth and albedo with a novel conditional feature translation module and pretrained critics for AU intensity estimation and image classification. Novel objective functions are designed using unlabeled in-the-wild images and in-door images with AU labels. We also leverage uncertainty losses to model the probably changing AU region of images as input noise for synthesis, and model the noisy AU intensity labels for intensity estimation of the AU critic. Experiments with face editing and animation on four datasets show that, compared with six state-of-the-art methods, our proposed method is superior and effective on expression consistency, identity similarity and pose similarity.
Hichem Sahli, Ke Lu 0002, Dongmei Jiang
ACM Multimedia5
2022 A multi-scale multi-attention network for dynamic facial expression recognition
Xiaohan Xia, Le Yang 0009, Xiaoyong Wei, Hichem Sahli, Dongmei Jiang
Multim. Syst.5
2022 Leveraging the Deep Learning Paradigm for Continuous Affect Estimation from Facial Expressions
abstract
Continuous affect estimation from facial expressions has attracted increased attention in the affective computing research community. This paper presents a principled framework for estimating continuous affect from video sequences. Based on recent developments, we address the problem of continuous affect estimation by leveraging the Bayesian filtering paradigm, i.e., considering affect as a latent dynamical system corresponding to a general feeling of pleasure with a degree of arousal, and recursively estimating its state using a sequence of visual observations. To this end, we advance the state-of-the-art as follows: (i) Canonical face representation (CFR): a novel algorithm for two-dimensional face frontalization, (ii) Convex unsupervised representation learning (CURL): a novel frequency-domain convex optimization algorithm for unsupervised training of deep convolutional neural networks (CNN)s, and (iii) Deep extended Kalman filtering (DEKF): an extended Kalman filtering-based algorithm for affect estimation from a sequence of CNN observations. The performance of the resulting CFR-CURL-DEKF algorithmic framework is empirically evaluated on publicly available benchmark datasets for facial expression recognition (CK+) and continuous affect estimation (AVEC 2012 and 2014).
Meshia Cédric Oveneke, Ercheng Pei, Abel Díaz Berenguer, Dongmei Jiang, Hichem Sahli
IEEE Trans. Affect. Comput.5
2021 Action Unit Driven Facial Expression Synthesis from a Single Image with Patch Attentive GAN
abstract
Abstract Recent advances in generative adversarial networks (GANs) have shown tremendous success for facial expression generation tasks. However, generating vivid and expressive facial expressions at Action Units (AUs) level is still challenging, due to the fact that automatic facial expression analysis for AU intensity itself is an unsolved difficult task. In this paper, we propose a novel synthesis‐by‐analysis approach by leveraging the power of GAN framework and state‐of‐the‐art AU detection model to achieve better results for AU‐driven facial expression generation. Specifically, we design a novel discriminator architecture by modifying the patch‐attentive AU detection network for AU intensity estimation and combine it with a global image encoder for adversarial learning to force the generator to produce more expressive and realistic facial images. We also introduce a balanced sampling approach to alleviate the imbalanced learning problem for AU synthesis. Extensive experimental results on DISFA and DISFA+ show that our approach outperforms the state‐of‐the‐art in terms of photo‐realism and expressiveness of the facial expression quantitatively and qualitatively.
Le Yang 0009, Ercheng Pei, Meshia Cédric Oveneke, Mitchel Alioscha-Pérez, Dongmei Jiang, Hichem Sahli
Comput. Graph. Forum7
2021 Integrating Deep and Shallow Models for Multi-Modal Depression Analysis - Hybrid Architectures
abstract
At present, although great progress has been made in automatic depression assessment, most of the recent works only concern the audio and video paralinguistic information, rather than the linguistic information from the spoken content. In this work, we argue that beside developing good audio and video features, to build reliable depression detection systems, text-based content features are also of importance to analyse depression-related textual indicators. Furthermore, to improve the performance of automatic depression assessment systems, powerful models, capable of modelling the characteristics of depression embedded in the audio, visual and text descriptors, are also required. This paper proposes new text and video features and hybridizes deep and shallow models for depression estimation and classification from audio, video and text descriptors. The proposed hybrid framework consists of three main parts: 1) A Deep Convolutional Neural Network (DCNN) and Deep Neural Network (DNN) based audio-visual multi-modal depression recognition model for estimating the Patient Health Questionnaire depression scale (PHQ-8); 2) A Paragraph Vector (PV) and Support Vector Machine (SVM) based model for inferring the physical and mental conditions of the individual from the transcripts of the interview; 3) A Random Forest (RF) model for depression classification from the estimated PHQ-8 score and the inferred conditions of the individual. In the PV-SVM model, PV embedding is used to obtain fixed-length feature vectors from transcripts of the answers to the questions associated with psychoanalytic aspects of depression, which are subsequently fed into the SVM classifiers for detecting the presence/absence of the considered psychoanalytic symptoms. To our best knowledge, this approach is the first attempt to apply PV for depression analysis. Besides, we propose a new visual descriptor - Histogram of Displacement Range (HDR) to characterize the displacement and velocity of the facial landmarks in the video segment. Experiments have been carried out on the Audio Visual Emotion Challenge (AVEC2016) depression dataset, they demonstrate that: 1) The proposed hybrid framework effectively improves the accuracies of both depression estimation and depression classification, with an average F1 measure up to 0.746, which is higher than the best result (0.724) of the depression sub-challenge of AVEC2016. 2) HDR obtains better depression recognition performance than Bag-of-Words (BoW) and Motion History Histogram (MHH) features.
Le Yang 0009, Dongmei Jiang, Hichem Sahli
IEEE Trans. Affect. Comput.2
2021 Transformer Encoder With Multi-Modal Multi-Head Attention for Continuous Affect Recognition
abstract
Continuous affect recognition is becoming an increasingly attractive research topic in affective computing. Previous works mainly focused on modelling the temporal dependency within a sensor modality, or adopting early or late fusion for multi-modal affective state recognition. However, early fusion suffers from the curse of dimensionality, and late fusion ignores the complementarity and redundancy between multiple modal streams. In this paper, we first introduce the transformer-encoder with a self-attention mechanism and propose a Convolutional Neural Network-Transformer Encoder (CNN-TE) framework to model the temporal dependency for single modal affect recognition. Further, to effectively consider the complementarity and redundancy between multiple streams we propose a Transformer Encoder with Multi-modal Multi-head Attention (TEMMA) for multi-modal affect recognition. TEMMA allows to progressively and simultaneously refine the inter-modality interactions and intra-modality temporal dependency. The learned multi-modal representations are fed to an Inference Sub-network with fully connected layers to estimate the affective state. The proposed framework is trained in a nutshell and demonstrates its effectiveness on the AVEC2016 and AVEC2019 datasets. Compared to state-of-the-art models, our approach obtains remarkable improvements on both arousal and valence in terms of concordance correlation coefficient (CCC) reaching 0.583 for arousal and 0.564 for valence on the AVEC2019 test set.
Dongmei Jiang, Hichem Sahli
IEEE Trans. Multim.2
2021 Monocular 3D Facial Expression Features for Continuous Affect Recognition
abstract
Automated facial expression analysis from image sequences for continuous emotion recognition is a very challenging task due to the loss of the three-dimensional information during the image formation process. State-of-the-art relied on estimating dynamic textures features and convolutional neural network features to derive spatio-temporal features. Despite their great success, such features are insensitive to micro facial muscle deformations and are affected by identity, face pose, illumination variation, and self-occlusion. In this work, we argue that retrieving, from image sequences, 3D facial spatio-temporal information, which describes the natural facial muscle deformation, provides a semantical and efficient way of representation and is useful for emotion recognition. In this paper, we propose a framework for extracting three-dimensional facial spatio-temporal features from monocular image sequences using an extended 3D Morphable Model (3DMM) which disentangles the identity factor from the facial expressions of a specific person. An LSTM model is used to evaluate the effectiveness of the proposed spatio-temporal features on video-based facial expression recognition task and continuous affect recognition task. Experimental results, on the AFEW6.0 datasets for facial expression recognition, and the RECOLA and SEMAINE datasets for continuous emotion prediction, illustrate the potential of the proposed 3D spatio-temporal features for facial expressions analysis and continuous affect recognition, as well as their efficiency compared to recent state-of-the-art features.
Ercheng Pei, Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli
IEEE Trans. Multim.4
2020 Emotion recognition from spatiotemporal EEG representations with hybrid convolutional recurrent neural networks via wearable multi-channel headset
Jingxia Chen, Dongmei Jiang, Yanning Zhang 0001
Comput. Commun.2
2020 An efficient model-level fusion approach for continuous affect recognition from audiovisual signals
Ercheng Pei, Dongmei Jiang, Hichem Sahli
Neurocomputing2
2019 FACS3D-Net: 3D Convolution based Spatiotemporal Representation for Action Unit Detection
abstract
Most approaches to automatic facial action unit (AU) detection consider only spatial information and ignore AU dynamics. For humans, dynamics improves AU perception. Is same true for algorithms? To make use of AU dynamics, recent work in automated AU detection has proposed a sequential spatiotemporal approach: Model spatial information using a 2D CNN and then model temporal information using LSTM (Long-Short-Term Memory). Inspired by the experience of human FACS coders, we hypothesized that combining spatial and temporal information simultaneously would yield more powerful AU detection. To achieve this, we propose FACS3D-Net that simultaneously integrates 3D and 2D CNN. Evaluation was on the Expanded BP4D+ database of 200 participants. FACS3D-Net outperformed both 2D CNN and 2D CNN-LSTM approaches. Visualizations of learnt representations suggest that FACS3D-Net is consistent with the spatiotemporal dynamics attended to by human FACS coders. To the best of our knowledge, this is the first work to apply 3D CNN to the problem of AU detection.
Le Yang 0009, Itir Önal, Jeffrey F. Cohn, Zakia Hammal, Dongmei Jiang, Hichem Sahli
ACII5
2019 Continuous affect recognition with weakly supervised learning
Ercheng Pei, Dongmei Jiang, Mitchel Alioscha-Pérez, Hichem Sahli
Multim. Tools Appl.2
2019 A video prediction approach for animating single face image
Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli
Multim. Tools Appl.3
2019 Automatic Depression Analysis Using Dynamic Facial Appearance Descriptor and Dirichlet Process Fisher Encoding
abstract
Depression causes mood disorders with noticeable problems in day-to-day activities. Current methods of assessing depression depend almost entirely on clinical interviews or questionnaires. They lack systematic and efficient ways of incorporating behavioral observations that are strong indicators of a psychological disorder. To help clinicians effectively and efficiently diagnose depression severity, automated systems, using objective and quantifiable data for depression assessment, are being developed. This paper presents a framework toward estimating a clinical depression-specific score, namely the Beck Depression Inventory-II (BDI-II) score, based on the analysis of facial expressions features. To extract facial dynamic features, we propose a novel dynamic feature descriptor denoted as median robust local binary patterns from three orthogonal planes (MRLBP-TOP), which can capture both the microstructure and macrostructure of facial appearance and dynamics. To aggregate the MRLBP-TOP over an image sequence, we propose a variant to the Fisher vector (FV) encoding scheme, denoted as the Dirichlet process FV (DPFV). DPFV adopts Dirichlet process Gaussian mixture models (DPGMM) to automatically learn the number of GMM mixtures and model parameters. Experimental results on the AVEC2013 and AVEC2014 depression databases have demonstrated the effectiveness of the proposed method.
Dongmei Jiang, Hichem Sahli
IEEE Trans. Multim.2
2018 An Improved Camouflage Target Detection Using Hyperspectral Image Based on Block-Diagonal and Low-Rank Representation
Fei Li 0011, Xiuwei Zhang 0001, Lei Zhang 0054, Yanning Zhang 0001, Dongmei Jiang, Genping Zhao
PRCV (4)5
2018 Hierarchical sparse coding framework for speech emotion recognition
Diana Torres, Meshia Cédric Oveneke, Fengna Wang, Dongmei Jiang, Werner Verhelst, Hichem Sahli
Speech Commun.4
2018 Leveraging the Bayesian Filtering Paradigm for Vision-Based Facial Affective State Estimation
abstract
Estimating a person's affective state from facial information is an essential capability for social interaction. Automatizing such a capability has therefore increasingly driven multidisciplinary research for the past decades. At the heart of this issue are very challenging signal processing and artificial intelligence problems driven by the inherent complexity of human affect. We therefore propose a principled framework for designing automated systems capable of continuously estimating the human affective state from an incoming stream of images. First, we model human affect as a dynamical system and define the affective state in terms of valence, arousal and their higher-order derivatives. We then pose the affective state estimation problem as a Bayesian filtering problem and provide a solution based on Kalman filtering (KF) for probabilistic reasoning overtime, combined with multiple instance sparse Gaussian processes (MI-SGP) for inferring affect-related measurements from image sequences. We quantitatively and qualitatively evaluate our proposed framework on the AVEC 2012 and AVEC 2014 benchmark datasets and obtain state-of-the-art results using the baseline features as input to our MI-SGP-KF model. We therefore believe that leveraging the Bayesian filtering paradigm can pave the way for further enhancing the design of automated systems for affective state estimation.
Meshia Cédric Oveneke, Isabel Gonzalez, Valentin Enescu, Dongmei Jiang, Hichem Sahli
IEEE Trans. Affect. Comput.4
2018 Exploiting Structured Sparsity for Hyperspectral Anomaly Detection
abstract
Sparse representation-based background modeling facilitates much recent progress in hyperspectral anomaly detection (AD). The sparse representation of background often exhibits underlying structure, which is crucial to distinguish between background and anomaly. However, how to exploit such underlying structure is still challenging. To address this problem, we present a novel hyperspectral AD method, which can exploit the structured sparsity in modeling the background more accurately. With the plausible background area detected by a local RX detector, a robust background spectrum dictionary is learned in a principal component analysis way. A reweighted Laplace prior-based structured sparse representation model is then employed to reconstruct the spectrum of each pixel. With considering the structured sparsity in representation, the background pixels can be reconstructed more accurately than the anomaly ones, which thus can be detected based on the reconstruction error. To further improve the detection performance, an intracluster reconstruction model is developed to exploit the spatial similarity among the background pixels in the same cluster. The anomaly pixels can then be detected based on the cost of intracluster reconstruction error. By linearly combining these two detection results, improvement is obviously achieved on detection accuracy. Experimental results on both simulated and real-world data sets demonstrate that the proposed method outperforms several state-of-the-art hyperspectral AD methods.
Fei Li 0011, Xiuwei Zhang 0001, Lei Zhang 0054, Dongmei Jiang, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2017 DCNN and DNN based multi-modal depression recognition
abstract
In this paper, we propose an audio visual multimodal depression recognition framework composed of deep convolutional neural network (DCNN) and deep neural network (DNN) models. For each modality, corresponding feature descriptors are input into a DCNN to learn high-level global features with compact dynamic information, which are then fed into a DNN to predict the PHQ-8 score. For multi-modal depression recognition, the predicted PHQ-8 scores from each modality are integrated in a DNN for the final prediction. In addition, we propose the Histogram of Displacement Range as a novel global visual descriptor to quantify the range and speed of the facial landmarks' displacements. Experiments have been carried out on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset for the Depression Sub-challenge of the Audio-Visual Emotion Challenge (AVEC 2016), results show that the proposed multi-modal depression recognition framework obtains very promising results on both the development set and test set, which outperforms the state-of-the-art results.
Le Yang 0009, Dongmei Jiang, Wenjing Han, Hichem Sahli
ACII2
2016 Hyperspectral anomaly detection using background learning and structured sparse representation
abstract
A novel background dictionary learning and structured sparse representation based anomaly detection method is proposed for hyperspectral imagery. First, a robust PCA spectrum dictionary is learned from the plausible background area detected by the local RX detector. With the learned dictionary, the reweighted Laplace prior based structured sparse representation model is then employed to reconstruct the spectrum of each pixel in the image. Due to considering the structured sparsity in representation, the background spectra can be reconstructed more accurately than anomaly ones. Thus, reconstruction error is utilized to separate the anomaly pixels and background ones. Experimental results on both simulated and real-world datasets demonstrate that the proposed method outperforms several state-of-the-art hyperspectral anomaly detection methods.
Fei Li 0011, Yanning Zhang 0001, Lei Zhang 0054, Xiuwei Zhang 0001, Dongmei Jiang
IGARSS5
2016 A multiCell visual tracking algorithm using multi-task particle swarm optimization for low-contrast image sequences
Yayun Ren, Benlian Xu, Peiyi Zhu, Mingli Lu, Dongmei Jiang
Appl. Intell.5
2015 Framework for combination aware AU intensity recognition
abstract
We present a framework for combination aware AU intensity recognition. It includes a feature extraction approach that can handle small head movements which does not require face alignment. A three layered structure is used for the AU classification. The first layer is dedicated to independent AU recognition, and the second layer incorporates AU combination knowledge. At a third layer, AU dynamics are handled based on variable duration semi-Markov model. The first two layers are modeled using extreme learning machines (ELMs). ELMs have equal performance to support vector machines but are computationally more efficient, and can handle multi-class classification directly. Moreover, they include feature selection via manifold regularization. We show that the proposed layered classification scheme can improve results by considering AU combinations as well as intensity recognition.
Isabel Gonzalez, Werner Verhelst, Meshia Cédric Oveneke, Hichem Sahli, Dongmei Jiang
ACII5
2015 Multimodal depression recognition with dynamic visual and audio cues
abstract
In this paper, we present our system design for audio visual multi-modal depression recognition. To improve the estimation accuracy of the Beck Depression Inventory (BDI) score, besides the Low Level Descriptors (LLD) features and the Local Gabor Binary Pattern-Three Orthogonal Planes (LGBP-TOP) features provided by the 2014 Audio/Visual Emotion Challenge and Workshop (AVEC2014), we extract extra features to capture key behavioural changes associated with depression. From audio we extract the speaking rate, and from video, the head pose features, the Space-Temporal Interesting Point (STIP) features, and local kinematic features via the Divergence-Curl-Shear descriptors. These features describe body movements, and spatio-temporal changes within the image sequence. We also consider global dynamic features, obtained using motion history histogram (MHH), bag of words (BOW) features and vector of local aggregated descriptors (VLAD). To capture the complementary information within the used features, we evaluate two fusion systems - the feature fusion scheme, and the model fusion scheme via local linear regression (LLR). Experiments are carried out on the training set and development set of the Depression Recognition Sub-Challenge (DSC) of AVEC2014, we obtain root mean square error (RMSE) of 7.6697, and mean absolute error (MAE) of 6.1683 on the development set, which are better or comparable with the state of the art results of the AVEC2014 challenge.
Dongmei Jiang, Hichem Sahli
ACII2
2015 Monocular 3D facial information retrieval for automated facial expression analysis
abstract
Understanding social signals is a very important aspect of human communication and interaction and has therefore attracted increased attention from various research areas. Among the different types of social signals, particular attention has been paid to facial expression of emotions and its automated analysis from image sequences. Automated facial expression analysis is a very challenging task due to the complex three-dimensional deformation and motion of the face associated to the facial expressions and the loss of 3D information during the image formation process. As a consequence, retrieving 3D spatio-temporal facial information from image sequences is essential for automated facial expression analysis. In this paper, we propose a framework for retrieving three-dimensional facial structure, motion and spatio-temporal features from monocular image sequences. First, we estimate monocular 3D scene flow by retrieving the facial structure using shape-from-shading (SFS) and combine it with 2D optical flow. Secondly, based on the retrieved structure and motion of the face, we extract spatio-temporal features for automated facial expression analysis. Experimental results illustrate the potential of the proposed 3D facial information retrieval framework for facial expression analysis, i.e. facial expression recognition and facial action-unit recognition on a benchmark dataset. This paves the way for future research on monocular 3D facial expression analysis.
Meshia Cédric Oveneke, Isabel Gonzalez, Dongmei Jiang, Hichem Sahli
ACII4
2015 Multimodal dimensional affect recognition using deep bidirectional long short-term memory recurrent neural networks
abstract
In this paper we propose the deep bidirectional long short-term memory recurrent neural network (DBLSTM-RNN) based single modal and multi-modal affect recognition frameworks. In the single modal framework DBLSTM with moving average (MA), audio or visual features are input into the DBLSTM-RNN model, whose output estimations of a dimension are smoothed by the moving average filter. After the smoothed estimations are expanded to the frame rate of the ground truth labels, another MA is adopted for smoothing the final results. In the multi-modal framework DBLSTM-DBLSTM-MA, the initial estimations from the audio and visual modalities via the first layer of DBLSTM-RNNs are input into a second layer of DBLSTM-RNN, whose outputs are smoothed by MA. The smoothed estimations are then expanded to the frame rate of the ground truth labels and smoothed again by another MA. Affect recognition experiments are carried out on the training set and development set of the AVEC2014 database, results show that the proposed DBLSTM-MA framework outperforms linear regression, support vector regression (SVR), and BLSTM for single modal dimension estimation. For audio visual multi-modal affect recognition, DBLSTM-DBLSTM-MA obtains better or comparable performance than the state of the art results in the competition of AVEC2014, with the average correlation coefficient (COR) reaches 0.599 on the Freeform database, 0.630 on the Northwind database, and 0.615 on the Freeform-Northwind database.
Ercheng Pei, Le Yang 0009, Dongmei Jiang, Hichem Sahli
ACII3
2015 3D emotional facial animation synthesis with factored conditional Restricted Boltzmann Machines
abstract
This paper presents a 3D emotional facial animation synthesis approach based on the Factored Conditional Restricted Boltzmann Machines (FCRBM). Facial Action Parameters (FAPs) extracted from 2D face image sequences, are adopted to train the FCRBM model parameters. Based on the trained model, given an emotion label sequence and several initial frames of FAPs, the corresponding FAP sequence is generated via the Gibbs sampling, and then used to construct the MPEG-4 compliant 3D facial animation. Emotion recognition and subjective evaluation on the synthesized animations show that the proposed method can obtain natural facial animations representing well the dynamic process of emotions. Besides, facial animation with smooth emotion transitions can be obtained by blending the emotion labels.
Dongmei Jiang, Hichem Sahli
ACII2
2015 Relevance units machine based dimensional and continuous speech emotion prediction
Fengna Wang, Hichem Sahli, Junbin Gao, Dongmei Jiang, Werner Verhelst
Multim. Tools Appl.4
2014 Speech-driven head motion synthesis using neural networks
abstract
This paper presents a neural network approach for speech-driven head motion synthesis, which can automatically predict a speaker’s head movement from his/her speech. Specifically, we realize speech-to-head-motion mapping by learning a multi-layer perceptron from audio-visual broadcast news data. First, we show that a generatively pre-trained neural network significantly outperforms a randomly initialized network and the hidden Markov model (HMM) approach. Second, we demonstrate that the feature combination of log Mel-scale filter-bank (FBank), energy and fundamental frequency (F0) performs best in head motion prediction. Third, we discover that using long context acoustic information can further improve the performance. Finally, extra unlabeled training data used in the pre-training stage can achieve more performance gain. The proposed speech-driven head motion synthesis approach increases the CCA from 0.299 (the HMM approach) to 0.565 and it can be effectively used in expressive talking avatar animation. Index Terms: head motion synthesis, neural network, deep neural network, talking avatar
Chuang Ding, Pengcheng Zhu 0004, Lei Xie 0001, Dongmei Jiang, Zhong-Hua Fu
INTERSPEECH4
2014 Speech driven photo realistic facial animation based on an articulatory DBN model and AAM features
Dongmei Jiang, Hichem Sahli, Yanning Zhang 0001
Multim. Tools Appl.1
2013 Hybrid Deep Neural Network-Hidden Markov Model (DNN-HMM) Based Speech Emotion Recognition
abstract
Deep Neural Network Hidden Markov Models, or DNN-HMMs, are recently very promising acoustic models achieving good speech recognition results over Gaussian mixture model based HMMs (GMM-HMMs). In this paper, for emotion recognition from speech, we investigate DNN-HMMs with restricted Boltzmann Machine (RBM) based unsupervised pre-training, and DNN-HMMs with discriminative pre-training. Emotion recognition experiments are carried out on these two models on the eNTERFACE'05 database and Berlin database, respectively, and results are compared with those from the GMM-HMMs, the shallow-NN-HMMs with two layers, as well as the Multi-layer Perceptrons HMMs (MLP-HMMs). Experimental results show that when the numbers of the hidden layers as well hidden units are properly set, the DNN could extend the labeling ability of GMM-HMM. Among all the models, the DNN-HMMs with discriminative pre-training obtain the best results. For example, for the eNTERFACE'05 database, the recognition accuracy improves 12.22% from the DNN-HMMs with unsupervised pre-training, 11.67% from the GMM-HMMs, 10.56% from the MLP-HMMs, and even 17.22% from the shallow-NN-HMMs, respectively.
Dongmei Jiang, Yanning Zhang 0001, Fengna Wang, Isabel Gonzalez, Valentin Enescu, Hichem Sahli
ACII3
2013 Multiuser two-way relay processing and power control methods for cognitive radio networks
abstract
ABSTRACT We consider a cognitive radio system where a secondary network shares the spectrum band with a primary network. Aiming at improving the frequency efficiency of the secondary network, we set a multiantenna relay station in the secondary network to perform two‐way relaying. Three linear processing schemes at the relay station based on zero forcing, zero forcing‐maximum ratio transmission, and minimum mean square error criteria are derived to guarantee the quality of service of primary users and to suppress the intrapair and interpair interference among secondary users (SUs). In addition, the transmit power of SUs is optimized to maximize the sum rate of SUs and to limit the interference brought to PUs. Numerical results show that the proposed multiuser two‐way relay processing schemes and the optimal power control policies can efficiently limit the interference caused by the secondary network to primary users, and the sum rate of SUs can also be greatly improved. Copyright © 2011 John Wiley & Sons, Ltd.
Dongmei Jiang, Haixia Zhang 0001, Dongfeng Yuan
Wirel. Commun. Mob. Comput.1
2012 Power-efficient resource allocation with QoS guarantees for TDMA fading channels
abstract
ABSTRACT This paper proposes two power‐efficient resource allocation policies with statistical delay Quality of Service (QoS) guarantees for uplink time‐division multiple access (TDMA) communication links. Specifically, the first policy aims at maximizing the system throughput while fulfilling the delay QoS and average power constraints, and the second policy is devised as an effort to minimize the total average power subject to individual delay QoS constraints. Convex optimization problems associated with the resource allocation policies are formulated based on a cross‐layer framework, where the queue at the data link layer is served by the resource allocation policy. By employing the Lagrangian duality theory and the dual decomposition theory, two subgradient iteration algorithms are developed to obtain the globally optimal solutions. The aforementioned resource allocation policies have been shown to be deterministic functions of delay QoS requirements and channel fading states. Moreover, numerical results are provided to demonstrate the performance of the proposed resource allocation policies. Copyright © 2010 John Wiley & Sons, Ltd.
Yanbo Ma, Haixia Zhang 0001, Dongfeng Yuan, Dongmei Jiang
Wirel. Commun. Mob. Comput.4
2011 Kalman Filter-Based Facial Emotional Expression Recognition
Isabel Gonzalez, Valentin Enescu, Hichem Sahli, Dongmei Jiang
ACII (1)5
2011 Audio Visual Emotion Recognition Based on Triple-Stream Dynamic Bayesian Network Models
Dongmei Jiang, Yulu Cui, Isabel Gonzalez, Hichem Sahli
ACII (1)1
2010 Realistic mouth animation based on an articulatory DBN model with constrained asynchrony
abstract
In this paper, we propose an approach to convert acoustic speech to video realistic mouth animation based on an articulatory dynamic Bayesian network model with constrained asynchrony (AF_AVDBN). Conditional probability distributions are defined to control the asynchronies between the articulators such as lips, tongue and glottis/velum. An EM-based conversion algorithm is also presented to learn the optimal visual features given an auditory input and the trained AF_AVDBN parameters. In the training of the AF_AVDBN models, downsampled YUV spatial frequency features of the interpolated mouth image sequences are extracted as visual features. For reproducing the mouth animation sequence, from the learned visual features, a spatial upsampling and a temporal downsampling are applied. Both qualitative and quantitative results show that the proposed method is capable of producing more natural and realistic mouth animations, and the accuracy is further improved compared to the state of the art multi-stream Hidden Markov Model (MSHMM) and articulatory DBN model without asynchrony constraint (AF_DBN).
Dongmei Jiang, Ilse Ravyse, Peizhen Liu, Hichem Sahli, Werner Verhelst
ICASSP1
2009 Audio-Visual Emotion Recognition Based on a DBN Model with Constrained Asynchrony
abstract
This paper presents an audio visual multi-stream DBN model (Asy_DBN) for emotion recognition with constraint asynchrony, in which audio state and visual state transit individually in their corresponding stream but the transition is constrained by the allowed maximum audio visual asynchrony. Emotion recognition experiments of Asy_DBN with different asynchrony constraints are carried out on an audio visual speech database of four emotions, and compared with the single stream HMM, state synchronous HMM (Syn_HMM) and state synchronous DBN model, as well the state asynchronous DBN model without asynchrony constraint. Results show that by setting the appropriate maximum asynchrony constraint between audio and visual streams, the proposed audio visual asynchronous DBN model gets the highest emotion recognition performance, with an improvement of 15% over Syn_HMM.
Danqi Chen 0001, Dongmei Jiang, Ilse Ravyse, Hichem Sahli
ICIG2
2009 A Visual Silence Detector Constraining Speech Source Separation
abstract
We propose an audiovisual source separation algorithm for speech signals. In our proposed algorithm we first extract the time segments with low activity of the mouth region from synchronous video recordings. An automatically selected optimal classifier is used to detect silent intervals in these instants of low visual mouth activity. Then, the source separation problem is formulated and solved for the entire signal duration. Our approach was tested on two challenging speech corpora with two speakers and two microphones, namely in the first corpus separate source signals were mixed in a simulated room, and the second corpus contains recorded conversations. The results are promising on both corpora: with the visual silence detector the performance of the source separation algorithm, measured by the signal to noise inference ratio increases.
Isabel Gonzalez, Ilse Ravyse, Henk Brouckxon, Werner Verhelst, Dongmei Jiang, Hichem Sahli
ICIG5
2009 Video Realistic Mouth Animation Based on an Audio Visual DBN Model with Articulatory Features and Constrained Asynchrony
abstract
This paper presents a mouth animation construction method based on the DBN models with articulatory features (AF_AVDBN), in which the articulatory features of lips, tongue, glottis/velum can be asynchronous within a maximum asynchrony constraint to describe the speech production process more reasonably. Given an audio input and the trained AF_AVDBN models, the optimal visual feature learning algorithm is deduced based on the Maximum Likelihood Estimation criterion. The learned visual features are then used to construct the mouth images for the input speech. Objective and subjective evaluations on the mouth animations of 110 speech sentences show that the learned visual features from the AF_AVDBN models track the real visual features very closely, and the constructed mouth images from the AF_AVDBN models are very much like the real ones.
Dongmei Jiang, Peizhen Liu, Ilse Ravyse, Hichem Sahli, Werner Verhelst
ICIG1
2009 Manifold Analysis for Subject Independent Dynamic Emotion Recognition in Video Sequences
abstract
This paper proposes subject independent manifold features for dynamic emotion recognition. Facial action features, based on FACs, are firstly embedded into a low-dimensional manifold space using the ISOMAP algorithm, then the manifold features from different subjects are aligned into a global coordinate space by the supervised ISOMAP algorithm for recognition. To validate and evaluate the proposed manifold representation for emotion recognition, experiments with GMMs are presented. Given a new expression sequence, and tracked facial features, we are able to pin-point the actual occurrence of specific expressions, while characterizing its intensity by considering different expression temporal transition characteristics. Finally, experimental results show that our approach is able to separate different expressions successfully.
Dongmei Jiang, Fengna Wang, Ilse Ravyse, Hichem Sahli
ICIG2
2008 Accurate visual speech synthesis based on diviseme unit selection and concatenation
abstract
This paper presents a novel speech driven accurate realistic visual speech synthesis approach. Firstly, an audio visual instance database is built for different viseme context combinations, i.e. diviseme units, using 100 audio visual speech sentences of a female speaker. Then a diviseme instance selection algorithm is introduced to choose the optimal diviseme instances for the viseme contexts in the input speech, considering both the concatenation smoothness of the image sequences, and matching of the mouth movements to the acoustic pronunciation process, as well the intensity of the input speech. Finally mouth image sequences of corresponding viseme segments in the selected diviseme instances are time warped and blended to construct the mouth images of the final animation. Visual speech synthesis experiments and subjective evaluation results show that mouth animations can be obtained which are not only realistic with clear and smooth mouth images, but also in good accordance with the acoustic pronunciation and intensity of the input speech.
Dongmei Jiang, Ilse Ravyse, Hichem Sahli, Yanning Zhang 0001
MMSP1
2007 Multi-stream Asynchrony Modeling for Audio-Visual Speech Recognition
abstract
In this paper, two multi-stream asynchrony Dynamic Bayesian Network models (MS-ADBN model and MM-ADBN model) are proposed for audio-visual speech recognition (AVSR). The proposed models, with different topology structures, loose the asynchrony of audio and visual streams to word level. For MS-ADBN model, both in audio stream and in visual stream, each word is composed of its corresponding phones, and each phone is associated with observation vector. MM- ADBN model is an augmentation of MS-ADBN model, a level of hidden nodes--state level, is added between the phone level and the observation node level, to describe the dynamic process of phones. Essentially, MS-ADBN model is a word model, while MM-ADBN model is a phone model. Speech recognition experiments are done on a digit audio-visual (A-V) database, as well as on a continuous A-V database. The results demonstrate that the asynchrony description between audio and visual stream is important for AVSR system, and MM-ADBN model has the best performance for the task of continuous A-V speech recognition.
Guoyun Lv, Dongmei Jiang, Rongchun Zhao, Yunshu Hou
ISM2
2006 Personalization of internet telephony services for presence with SIP and extended CPL
Dongmei Jiang, Ramiro Liscano, Luigi Logrippo
Comput. Commun.1