EDBT 2026 Demo / reviewers in the wild / expert
Zixing Zhang 0001
dblp:02/10648-1
· DBLP profile ↗
104ranked-venue papers
23as first author
53since 2021 · last 2026
0000-0001-8487-0561ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 14 first-author · 31 since 2021Artificial intelligence and machine learning · 39 · 11 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 7Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech CodecsabstractNeural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strategically decouples the quantization of Semantic and Acoustic details. The semantic anchoring is achieved via a lightweight projector that aligns acoustic features with a frozen, large-scale mHuBERT codebook, injecting linguistic priors while guaranteeing full codebook utilization. Sequentially, for acoustic details, a residual activation module with SimVQ enables a single-layer quantizer (acoustic path) to faithfully recover fine-grained information. At just 1.5 kbps, SACodec establishes a new state of the art by excelling in both fidelity and semantics: subjective listening tests confirm that its reconstruction quality is perceptually highly comparable to ground-truth audio, while its tokens demonstrate substantially improved semantic richness in downstream tasks. This work suggests that assigning specialized semantic quantizers to distinct information streams offers an effective path to reconcile the long-standing trade-off between fidelity, semantics, and modeling simplicity in low-bitrate speech tokenization. Zhongren Dong, Jing Han 0010, Xiaojun Mo, Yimin Cao, Zixing Zhang 0001 |
AAAI | 7 |
| 2026 | PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation NetworkabstractExisting multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive modeling tailored to specific task requirements. To address these limitations, we propose a Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network (PLUM-Net). It first employs a multilevel semantic alignment module to synchronize global and local semantics across audio, visual and textual streams. On this aligned foundation, a prototype-based single-modal label generation module derives modality-specific hard and soft-labels that subtly steer the network toward a cleaner split between shared and private cues. Guided by these labels, the task-conditioned feature bifurcator module channels information through the most beneficial common or private pathway for the given task, after which a private refinement module polishes and fuses each modality’s idiosyncratic signals. Extensive experiments show that PLUM-Net delivers strong performance on datasets such as CMU-MOSI, CMU-MOSEI and UR-FUNNY, achieving an ACC-2 of 90.3% on CMU-MOSI, representing a 2%–4% improvement over previous SOTA models. Huan Zhao 0003, Xupeng Zha, Guanghui Ye, Zixing Zhang 0001 |
AAAI | 8 |
| 2026 | DyGIN: Modality-Unified Dynamic Graph Inception Network for Context-Aware Emotion RecognitionabstractGraph representation learning has attracted considerable attention for its ability to generate node representations by aggregating information from neighboring nodes. However, existing graph-based methods for sequential data merely focus on a static graph constructed on an entire unimodal sequence, largely ignoring their dynamic evolution. To overcome this limitation, we introduce a modality-unifiedDynamicGraphInceptionNetwork (DyGIN) that models dynamic evolutionary graphs for different modalities. DyGIN constructs dynamic graphs from temporal subsequences using a sliding window, and incorporates a Temporal Graph Evolution Gated Recurrent Unit (TGE-GRU) to update graph weights at each segment. Additionally, we propose a context-aware similarity matrix, which updates node representations based on the degree of neighboring nodes, replacing the traditional averaging method. Our model is optimized with a combination of graph structure loss, classification loss, and a learnable pooling function. We validate DyGIN on emotion recognition tasks across three public datasets—RML, IEMOCAP, and DEAP—where it outperforms state-of-the-art models by 1.6% on RML, 1.81% and 2.33% on IEMOCAP, and 2.78% and 6.81% on DEAP, based on accuracy and weighted F1 score. Our code is available at https://github.com/G22-web/DyGIN. Yingxue Gao, Huan Zhao 0003, Zixing Zhang 0001 |
IEEE Internet Things J. | 3 |
| 2026 | TTT-Dance: A Diffusion Framework Leveraging Test-Time Training and LayerScale for Adaptive Music-Driven Dance GenerationabstractMusic-driven dance generation requires precise motion modeling tightly aligned with musical rhythms. However, existing methods, while producing perceptually convincing motions, oftenfail to generalize to out-of-distribution music and unseen dance styles. We propose TTT-Dance, a diffusion-based framework that integratesTest-Time Training (TTT)into the denoising network, enabling rapid, sample-specific adaptation during inference toenhance robustness and generalization. Additionally, we incorporateLayerScalemodules into the decoder to stabilize the optimization of long temporal sequences and cross-modal interactions, thereby accelerating convergence and improving motion coherence. Extensive experiments on public benchmarks show that TTT-Dance consistently surpasses stateof- the-art baselines across multiple metrics, including physical plausibility, rhythm alignment, and motion diversity, demonstrating its effectiveness in music-dance generation. Zixing Zhang 0001, Hubo Liao, Jing Han 0010 |
IEEE Signal Process. Lett. | 2 |
| 2025 | ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech SynthesisabstractProsody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks when synthesizing long sentences with complex structures but also produce unnatural intonation. We propose ProsodyFM, a prosody-aware text-to-speech synthesis (TTS) model with a flow-matching (FM) backbone that aims to enhance the phrasing and intonation aspects of prosody. ProsodyFM introduces two key components: a Phrase Break Encoder to capture initial phrase break locations, followed by a Duration Predictor for the flexible adjustment of break durations; and a Terminal Intonation Encoder which learns a bank of intonation shape tokens combined with a novel Pitch Processor for more robust modeling of human-perceived intonation change. ProsodyFM is trained with no explicit prosodic labels and yet can uncover a broad spectrum of break durations and intonation patterns. Experimental results demonstrate that ProsodyFM can effectively improve the phrasing and intonation aspects of prosody, thereby enhancing the overall intelligibility compared to four state-of-the-art (SOTA) models. Out-of-distribution experiments show that this prosody improvement can further bring ProsodyFM superior generalizability for unseen complex sentences and speakers. Our case study intuitively illustrates the powerful and fine-grained controllability of ProsodyFM over phrasing and intonation. Xiangheng He, Zixing Zhang 0001, Björn W. Schuller |
AAAI | 3 |
| 2025 | Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift ModelingabstractConversational Emotion Recognition (CER) has recently been explored through conversational context modeling to learn the emotion distribution, i.e., the likelihood over emotion categories associated with each utterance. While these methods have shown promising results in emotion classification, they often focus on the interactions between utterances (utterance-view) and overlook shifts in the speaker's emotions (emotion-view). This emphasis on homogeneous view modeling limits their overall effectiveness. To address this limitation, we propose DVL-CER, a novel Dual-View Learning approach for CER. DVL-CER integrates both the utterance-view and emotion-view using two projection heads, enabling cross-view projection of emotion distributions. Our approach offers several key advantages: (1) We introduce an emotion-view that captures shifts in a speaker's emotions from initial to subsequent states within a conversation. This view enriches the conversation modeling and supports seamless integration with various CER baseline models. (2) Our dual-view projection learning strategy flexibly balances consistency and independence between the two heterogeneous views, promoting view-specific adaptation learning and incorporating the emotion verification capability within CER. We validate DVL-CER through extensive experiments on two widely-used datasets, IEMOCAP and EmoryNLP. The results demonstrate that DVL-CER achieves state-of-the-art performance, delivering robust and high-quality emotion distributions compared with existing CER methods and other dual-view learning strategies. Xupeng Zha, Huan Zhao 0003, Guanghui Ye, Zixing Zhang 0001 |
AAAI | 4 |
| 2025 | XDGesture: An xLSTM-based Diffusion Model for Co-speech Gesture GenerationabstractIn multimodal human-computer interaction, generating co-speech gestures is crucial for enhancing interaction naturalness and user experience. However, achieving synchronized and natural gesture sequences remains a significant challenge due to the complexity of modeling temporal dependencies across different modalities. Existing methods often rely on simple concatenation techniques, which are limited in effectively handling multimodal information. To address this issue, we propose XDGesture, a diffusion-based framework that integrates a Cross-Modal Fusion module and xLSTM. The Cross-Modal Fusion module efficiently merges information from different modalities, providing the model with rich contextual conditions. Meanwhile, xLSTM, with its enhanced memory structure and exponential gating mechanism, processes the fused multimodal data, capturing long-range dependencies between speech and gestures. This enables the generation of high-quality gesture sequences that are naturally synchronized with speech. Experimental results demonstrate that XDGesture remarkably outperforms existing baselines on multiple datasets, particularly in terms of gesture quality, naturalness, and synchronization with speech. Zixing Zhang 0001, Huan Zhao 0003, Björn W. Schuller |
ICASSP | 1 |
| 2025 | Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph PropagationabstractMultimodal Emotion Recognition in Conversations (ERC) plays a crucial role in understanding human language and behavior in real-world scenarios. However, existing research tends to simply concatenate multimodal representations, failing to capture the complex relationships between modalities. Recent advances have shown that Graph Neural Networks (GNNs) are effective in capturing complex data relationships, offering a promising solution for multimodal ERC. Despite this, current GNN-based methods still face challenges, including weak interactions between modalities, neglecting the information entropy of utterances, and erasure of high-frequency signals that capture key variations and discrepancies between closely related nodes. To address these limitations, we propose a GNNs-based multi-frequency propagation method enhanced by contextual filtering for multimodal ERC. Our approach introduces a context filtering module that combines a similarity matrix and an information entropy matrix, enabling GNNs to effectively capture the inherent relationships among utterances and provide sufficient multimodal and contextual modeling. Additionally, our method explores multivariate relationships by recognizing the varying importance of emotional discrepancies and commonalities through multi-frequency signals. Experimental results on two benchmark datasets, IEMOCAP and MELD, demonstrate that our method outperforms the latest (non-)graph-based works. Our method is available at https://github.com/G22-web/ConFilMER. Huan Zhao 0003, Yingxue Gao, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001 |
ICASSP | 6 |
| 2025 | Parameter-Efficient Federal-Tuning Enhances Privacy Preserving for Speech Emotion RecognitionabstractThe Pre-trained Speech Models (PSMs) generate universal speech representations using self-supervised or weakly-supervised learning from large-scale datasets. It achieves promising performance when fine-tuned for specific tasks such as Speech Emotion Recognition (SER). However, fine-tuning on various datasets requires storing the entire model’s weight parameters, complicating real-world deployment. Additionally, centralized fine-tuning relies on user data, posing significant privacy risks. To address these challenges, we propose employing Federated Learning (FL) for fine-tuning PSMs with a Parameter-Efficient Fine-Tuning (PEFT) method. By embedding trainable layers in the feed-forward layers of the pre-trained model, we keep the backbone model frozen and only update the trainable layer parameters during federated training, significantly reducing parameter transmission. Specifically, we evaluated the performance of downstream model fine-tuning, adapter tuning, embedding prompt tuning, and LoRA within a federated fine-tuning framework for PSMs, demonstrating the framework’s feasibility and effectiveness. Furthermore, attribute inference attack tests showed that gender inference results on three datasets were at chance levels. Haijiao Chen, Huan Zhao 0003, Yingxue Gao, Zixing Zhang 0001 |
ICASSP | 5 |
| 2025 | MHSDB: A Comprehensive Benchmark for Multimodal Humor and Sarcasm Detection Leveraging Foundation ModelsabstractUnderstanding multimodal humor and sarcasm detection remains a key challenge in artificial intelligence. Despite recent advances, inconsistencies in feature extraction, evaluation methods, and experimental setups have hindered fair comparisons across different approaches. To address this issue, we propose the Multimodal Humor and Sarcasm Detection Benchmark (MHSDB), the first unified evaluation platform specifically designed for these tasks. MHSDB combines four datasets in English and Hindi and standardizes feature extraction and evaluation processes to facilitate consistent comparisons. We systematically evaluate mainstream foundation models across audio, video, and text modalities. Unimodal representations are assessed using self-attention mechanisms, while multimodal representations are evaluated through mainstream fusion strategies, including utterance-level and sequence-level approaches. Our experimental results reveal that multimodal approaches outperform unimodal ones in capturing complex contexts and multi-layered semantics. Additionally, specific fusion strategies excel at integrating cross-modal information, achieving state-of-the-art performance, and paving the way for future research on optimizing feature representation and multimodal fusion. Zhongren Dong, Donghao Wang, Ciqiang Chen, Dong-Yan Huang, Zixing Zhang 0001 |
ICASSP | 5 |
| 2025 | Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-LabelingabstractThe lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Fréchet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines. Yuanchao Li, Zixing Zhang 0001, Jing Han 0010, Peter Bell 0001, Catherine Lai |
ICASSP | 2 |
| 2025 | DSSM: Dual State Space Model For Human Motions GenerationabstractText-driven human motion generation has attracted considerable critical attention in recent years. The task requires generating movements that are diverse, natural, and comfortable in accordance with the text description. However, while generating the human motion, there is a significant gap in the amount of feature information contained in word-based text description modality and joint-based human motion modality. This results in an extreme imbalance of feature information in the latent space across the modalities, which seriously affects the effect of feature fusion. To alleviate the imbalance between modalities, we propose the Dual State Space Model (DSSM), which reconstruction the fused feature from coarse-to-fine. The DSSM contains two unit structures: the Masked State Space Model (MSSM) and the Hierarchical State Space Block (HSSB). At the same time, in order to make better use of timing information and reduce the computational complexity of the model, the DSSM is also the first method to introduce the state space model (SSM) into the text-driven motion sequence generation. We evaluated the DSSM on the HumanML3D and KITML benchmark datasets, and the experimental results show that our approach achieves state-of-the-art performance. The source code is available on GitHub at https://github.com/yimingliu123/DSSM. Huan Zhao 0003, Yaqian Liu, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001 |
ICASSP | 7 |
| 2025 | GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion RecognitionabstractIn recent years, multimodal emotion recognition (MER) has gained significant attention due to its potential to integrate information from diverse signals. However, existing methods often struggle to effectively capture complex interactions and contextual information both inter- and intra-modalities, and even to extract the salient representations from pre-trained models. To address these issues, we propose a novel model, gated Mixture of Multimodal Experts (MoME) and Mixtral of Experts (MixMoE) models, namely GateM2Former. The gate mechanism is used to select the most relevant representations from pre-trained models. The MoME and MixMoE expert modules respectively learn the individual characteristics of each modality and the intrinsic alignment and potential interactions between modalities. Besides, we design a hierarchical merge structure to better suit the long sequence scenario (i. e., speech in our case). To verify the effectiveness of the introduced model, we conducted extensive experiments on the IEMOCAP and MELD datasets. The results show that GateM2Former, with a universal multimodal structure, is able to achieve the best results on IEMOCAP and MELD compared with other latest approaches. Zhongren Dong, Runming Wang, Xinzhou Xu, Zixing Zhang 0001 |
ICASSP | 5 |
| 2025 | SSE: A Speaking Style Extractor Based on Fine-Grained Contrastive Learning between Speech and Descriptive TextabstractEffective extraction of paralinguistic features from speech, such as emotion, accent, and age, remains a challenging task in speech processing. Traditional methods typically address each type of paralinguistic information with separate classification or regression tasks—e. g., emotion recognition, accent detection, and age estimation. These approaches can result in fragmented and incomplete descriptions of speech style, failing to capture the full range of paralinguistic attributes in an integrated way. To address these limitations, we propose the Speech Style Extractor (SSE), a novel approach that aims to provide a more comprehensive extraction of speech style features. SSE leverages an optimized fine-grained contrastive learning scheme, enhancing the extraction of diverse paralinguistic features. The optimization approach outperforms the baseline on 11 out of 13 datasets, with an average improvement of 2.1% in terms of the R@1 evaluation metric. We also propose a complete data generation framework using large language models to create 199k text samples across 9 paralinguistic categories, such as emotion, deception, stuttering, and accent, to fill the gap in speaking style descriptions in public datasets. This extensive dataset not only facilitates our research but also serves as a valuable resource for the speech processing community. Zixing Zhang 0001, Yimeng Wu, Zhongren Dong, Wulong Xiang, Shengfan Shen, Björn W. Schuller |
ICASSP | 1 |
| 2025 | Rethinking Removal Attack and Fingerprinting Defense for Model Intellectual Property Protection: A Frequency PerspectiveabstractTraining deep neural networks is resource-intensive, making it crucial to protect their intellectual property from infringement. However, current model ownership resolution (MOR) methods predominantly address general removal attacks that involve weight modifications, with limited research considering alternative attack perspectives. In this work, we propose a frequency-based model ownership removal attack, grounded in a key observation: modifying a model's high-frequency coefficients does not significantly impact its performance but does alter its weights and decision boundary. This change invalidates the existing MOR methods. We further propose a frequency-based fingerprinting technique as a defense mechanism. By extracting frequency-domain characteristics instead of decision boundary or model weights, our fingerprinting defense effectively against the proposed frequency-based removal attack and demonstrates robustness against existing general removal attacks. The experimental results show that the frequency-based removal attack can easily defeat state-of-the-art white-box watermarking and fingerprinting schemes while preserving model performance, and the proposed defense method is also effective. Our code is released at: https://github.com/huangtingqiao/RRA-IJCAI25. Cheng Zhang 0035, Yang Xu 0013, Tingqiao Huang, Zixing Zhang 0001 |
IJCAI | 4 |
| 2025 | Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment AnalysisabstractUnderstanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) Challenge invited participants to tackle two demanding subtasks: (1) extracting a comprehensive sentiment sextuple-including holder, target, aspect, opinion, sentiment, and rationale-from multi-speaker dialogues, and (2) detecting sentiment flipping, which detects dynamic sentiment shifts and their underlying triggers. For Subtask-I, in the present paper, we designed a structured prompting pipeline that guided large language models (LLMs) to sequentially extract sentiment components with refined contextual understanding. For Subtask-II, we further leveraged the complementary strengths of three LLMs through ensembling to robustly identify sentiment transitions and their triggers. Our system achieved a 47.38% average score on Subtask-I and a 74.12% exact match F1 on Subtask-II, showing the effectiveness of step-wise refinement and ensemble strategies in rich, multimodal sentiment analysis tasks. Shihao Gao, Zixing Zhang 0001, Jing Han 0010 |
ACM Multimedia | 3 |
| 2025 | PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentabstractSpeech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high lip-synchronization accuracy, existing methods largely overlook the intricate nuances of individual speaking styles, which limits personalization and realism. In this work, we present a novel framework for personalized 3D talking head animation, namely ''PTalker''. This framework preserves speaking style through style disentanglement from audio and facial motion sequences and enhances lip-synchronization accuracy through a three-level alignment mechanism between audio and mesh modalities. Specifically, to effectively disentangle style and content, we design disentanglement constraints that encode driven audio and motion sequences into distinct style and content spaces to enhance speaking style representation. To improve lip-synchronization accuracy, we adopt a modality alignment mechanism incorporating three aspects: spatial alignment using Graph Attention Networks to capture vertex connectivity in the 3D mesh structure, temporal alignment using cross-attention to capture and synchronize temporal dependencies, and feature alignment by top-k bidirectional contrastive losses and KL divergence constraints to ensure consistency between speech and mesh modalities. Extensive qualitative and quantitative experiments on public datasets demonstrate that PTalker effectively generates realistic, stylized 3D talking heads that accurately match identity-specific speaking styles, outperforming state-of-the-art methods. The source code and supplementary videos are available at: PTalker. Yang Xu 0013, Huan Zhao 0003, Hao Zhang 0139, Zixing Zhang 0001 |
ACM Multimedia | 5 |
| 2025 | Tyee: A Unified, Modular, and Fully-Integrated Configurable Toolkit for Intelligent Physiological Health Care
Lingyu Shu, Zixing Zhang 0001, Jing Han 0010 |
ACM Multimedia | 3 |
| 2025 | AudioFab: Building A General and Intelligent Audio Factory through Tool LearningabstractCurrently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient framework to unlock their full potential. Existing audio agent frameworks often suffer from complex environment configurations and inefficient tool collaboration. To address these limitations, we introduce AudioFab, an open-source agent framework aimed at establishing an open and intelligent audio-processing ecosystem. Compared to existing solutions, AudioFab's modular design resolves dependency conflicts, simplifying tool integration and extension. It also optimizes tool learning through intelligent selection and few-shot learning, improving efficiency and accuracy in complex audio tasks. Furthermore, AudioFab provides a user-friendly natural language interface tailored for non-expert users. As a foundational framework, AudioFab's core contribution lies in offering a stable and extensible platform for future research and development in audio and multimodal AI. The code is available at https://github.com/SmileHnu/AudioFab. Jing Han 0010, Qianshuai Xue, Huan Zhao 0003, Zixing Zhang 0001 |
ACM Multimedia | 6 |
| 2025 | Federal parameter-efficient fine-tuning for speech emotion recognition
Haijiao Chen, Huan Zhao 0003, Zixing Zhang 0001, Keqin Li 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Reparameterization of Lightweight Transformer for On-Device Speech Emotion RecognitionabstractWith the increasing implementation of machine learning models on the edge or Internet of Things (IoT) devices, deploying advanced models on resource-constrained IoT devices remains challenging. Transformer models, a currently dominant neural architecture, have achieved great success in broad domains but their complexity hinders its deployment on IoT devices with limited computation capability and storage size. Although many model compression approaches have been explored, they often suffer from notorious performance degradation. To address this issue, we introduce a new method, namely Transformer reparameterization, to boost the performance of lightweight Transformer models. It consists of two processes: 1) the high-rank factorization (HRF) process in the training stage and 2) the de-HRF (deHRF) process in the inference stage. In the former process, we insert an additional linear layer before the feed-forward network (FFN) of the lightweight Transformer. It is supposed that the inserted HRF layers can enhance the model learning capability. In the later process, the auxiliary HRF layer will be merged together with the following FFN layer into one linear layer and thus recover the original structure of the lightweight model. To examine the effectiveness of the proposed method, we evaluate it on three widely used Transformer variants, i.e., ConvTransformer, Conformer, and SpeechFormer networks, in the application of speech emotion recognition on the IEMOCAP, M3ED, and DAIC-WOZ datasets. Experimental results show that our proposed method consistently improves the performance of lightweight Transformers, even making them comparable to large models. The proposed reparameterization approach enables advanced Transformer models to be deployed on resource-constrained IoT devices. Zixing Zhang 0001, Zhongren Dong, Jing Han 0010 |
IEEE Internet Things J. | 1 |
| 2025 | UniDE: A multi-level and low-resource framework for automatic dialogue evaluation via LLM-based data augmentation and multitask learning
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Zhihua Jiang |
Inf. Process. Manag. | 3 |
| 2025 | STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion RecognitionabstractSpeech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models. Yi Chang 0004, Zhao Ren, Zixing Zhang 0001, Xin Jing 0001, Kun Qian 0003, Xi Shao, Bin Hu 0001, Tanja Schultz, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics Over Acoustic Foundation ModelsabstractComputational paralinguistics (ComParal) aims to develop algorithms and models to automatically detect, analyze, and interpret non-verbal information from speech communication, e. g., emotion, health state, age, and gender. Despite its rapid progress, it heavily depends on sophisticatedly designed models given specific paralinguistic tasks. Thus, the heterogeneity and diversity of ComParal models largely prevent the realistic implementation of ComParal models. Recently, with the advent of acoustic foundation models because of self-supervised learning, developing more generic models that can efficiently perceive a plethora of paralinguistic information has become an active topic in speech processing. However, it lacks a unified evaluation framework for a fair and consistent performance comparison. To bridge this gap, we conduct a large-scale benchmark, namelyParaLBench, which concentrates on standardizing the evaluation process of diverse paralinguistic tasks, including critical aspects of affective computing such as emotion recognition and emotion dimensions prediction, over different acoustic foundation models. This benchmark contains ten datasets with thirteen distinct paralinguistic tasks, covering short-, medium- and long-term characteristics. Each task is carried out on 14 acoustic foundation models under a unified evaluation framework, which allows for an unbiased methodological comparison and offers a grounded reference for the ComParal community. Based on the insights gained from ParaLBench, we also point out potential research directions, i. e., the cross-corpus generalizability, to propel ComParal research in the future. The code associated with this study will be available to foster the transparency and replicability of this work for succeeding researchers. Zixing Zhang 0001, Zhongren Dong, Kanglin Wang, Yimeng Wu, Runming Wang, Dong-Yan Huang |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Robust Hashing With Bilinear Drift for Image-Text RetrievalabstractSupervised hashing models for image-text retrieval are fundamental and versatile in social media analysis and cross-lingual web search. Among them, supervised bilinear drift hashing is one of the most popular approaches. However, it still faces several challenges. For instance, how to leverage the power of bilinear drift hashing to distinguish similar and dissimilar data samples effectively; how to strengthen the semantic relationship between similar data and supervision. To solve these problems, we propose Robust Hashing with Bilinear Drift (RHBD) to improve the accuracy and robustness of the supervised model. The key idea of this work is to generate effective hash codes between image-text feature representations by combining robust data distributions and multiple supervision information. The benefits of bilinear drift with robust hashing, which enhance the discrimination of hash binary, are manifested mainly in two ways: (1) RHBD employs a semantic autoencoder with a linear drift to get a discriminative common feature representation between image and text modalities; (2) RHBD explores iteration quantization with a linear drift to well generate similarity-preserving hash codes. Moreover, we introduce multiple supervision learning to promote the consistency between data information and supervision knowledge for semantic complementarity. Results on three public datasets show that RHBD is effective in image-text retrieval, consistently outperforming other state-of-the-art models with comparable training efficiency to competitive baselines. Huan Zhao 0003, Song Wang 0016, Zixing Zhang 0001, Keqin Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Modality Imbalance? Dynamic Multi-Modal Knowledge Distillation in Automatic Alzheimer's Disease RecognitionabstractAlzheimer's disease (AD), as the most prevalent form of dementia, necessitates early identification and treatment for the critical enhancement of patients' quality of life. Recent studies strive to explore advanced machine learning approaches with multiple information cues, such as speech and text, to automatically and precisely detect this disease from conversations. However, these multi-modality-based approaches often suffer from a modality-imbalance challenge that leads to performance degradation. That is, the multi-modal model performs worse than the best mono-modal model, although the former contains more information. To address this issue, we propose a Dynamic Multi-Modal Knowledge Distillation (DMMKD) approach, which dynamically identify the dominant modality and the weak modality, and opt to conduct an inter(cross)-modal or intra-modal knowledge distillation. The core idea is to balance the individual learning speed in the multi-modal learning process by boosting the weak modality with the dominant modality. To evaluate the effectiveness of the introduced DMMKD algorithm, we conducted extensive experiments on two publicly available and widely used AD datasets, i. e., ADReSSo and ADReSS-M. Compared to the multi-modal approaches without dealing with the modality imbalance issue, the introduced DMMKD indicates substantial performance improvements by 15.4% and 10.9% in terms of relative accuracy on the ADReSSo and ADReSS-M datasets, respectively. Moreover, when compared to the state-of-the-art models for automatic AD detection, the DMMKD achieves the best performance of 91.5% and 87.0% accuracies on the two datasets, respectively. Zhongren Dong, Xinzhou Xu, Zixing Zhang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | BAHBench: A Unified Benchmark for Evaluating Bio-Acoustic Health With Acoustic Foundation ModelsabstractAcoustic foundation models, through self-supervised learning on large amounts of unlabeled speech data, can acquire rich acoustic representations. In recent years, these models have demonstrated substantial potential in audio-based health-related tasks, remarkably enhancing the efficiency and quality of healthcare services and contributing to the advancement of smart healthcare. However, there is currently a lack of systematic research and exploration on the performance of acoustic foundation models in health-related tasks. Furthermore, inconsistencies in evaluation methods and experimental setups hinder fair comparisons between different methods, severely impeding progress in this field. To address these challenges, we establish a unified Benchmark for evaluating Bio-Acoustic health via acoustic foundation models, namely BAHBench. BAHBench encompasses 6 distinct health-related tasks and evaluates 12 acoustic foundation models within a unified evaluation framework and parameter settings, enabling fair comparisons across different models. Our objective is to explore the effectiveness of current acoustic foundation models in health-related tasks. Thus, we discuss the impact of model size and data diversity on performance, and investigate feature selection and efficient fine-tuning strategy. Experimental results show that different health-related tasks benefit from features from different layers of the foundation model, while LoRA fine-tuning further enhances the model's performance on downstream tasks. Our goal is to provide clear and comprehensive guidance for future researchers. The code related to this study will be available to the research community to promote transparency and reproducibility. Zhongren Dong, Runming Wang, Zixing Zhang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty EstimationabstractHeart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learning algorithms and the uncertainty of predictions. This paper introduces a heart murmur detection method based on a parallel-attentive model, which consists of two branches: One is based on a self-attention module and the other one is based on a convolutional network. Unlike traditional approaches, this structure is better equipped to handle long-term dependencies in sequential data, and thus effectively captures the local and global features of heart murmurs. Additionally, we acknowledge the significance of understanding the uncertainty of model predictions in the medical field for clinical decision-making. Therefore, we have incorporated an effective uncertainty estimation method based on Monte Carlo Dropout into our model. Furthermore, we have employed temperature scaling to calibrate the predictions of our probabilistic model, enhancing its reliability. In experiments conducted on the CirCor Digiscope dataset for heart murmur detection, our proposed method achieves a weighted accuracy of 79.8 % and an F1 of 65.1 %, representing state-of-the-art results. Zixing Zhang 0001, Jing Han 0010, Björn W. Schuller |
ICASSP | 1 |
| 2024 | HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous SpeechabstractAutomatically detecting Alzheimer’s Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational complexity associated with self-attention and the length of audio poses a challenge when deploying such models on edge devices. In this context, we construct a novel framework, namely Hierarchical Attention-Free Transformer (HAFFormer), to better deal with long speech for AD detection. Specifically, we employ an attention-free module of Multi-Scale Depthwise Convolution to replace the self-attention and thus avoid the expensive computation, and a GELU-based Gated Linear Unit to replace the feedforward layer, aiming to automatically filter out the redundant information. Moreover, we design a hierarchical structure to force it to learn a variety of information grains, from the frame level to the dialogue level. By conducting extensive experiments on the ADReSS-M dataset, the introduced HAFFormer can achieve competitive results (82.6% accuracy) with other recent work, but with significant computational complexity and model size reduction compared to the standard Transformer. This shows the efficiency of HAFFormer in dealing with long audio for AD detection. Zhongren Dong, Zixing Zhang 0001, Jing Han 0010, Jianjun Ou, Björn W. Schuller |
ICASSP | 2 |
| 2024 | Adaptive Speech Emotion Representation Learning Based On Dynamic GraphabstractGraph representation learning has become a hot research topic due to its powerful nonlinear fitting capability in extracting representative node embeddings. However, for sequential data such as speech signals, most traditional methods merely focus on the static graph created within a sequence, and largely overlook the intrinsic evolving patterns of these data. This may reduce the efficiency of graph representation learning for sequential data. For this reason, we propose an adaptive graph representation learning method based on dynamically evolved graphs, which are consecutively constructed on a series of subsequences segmented by a sliding window. In doing this, it is better to capture local and global context information within a long sequence. Moreover, we introduce a weighted approach to update the node representation rather than the conventional average one, where the weights are calculated by a novel matrix computation based on the degree of neighboring nodes. Finally, we construct a learnable graph convolutional layer that combines the graph structure loss and classification loss to optimize the graph structure. To verify the effectiveness of the proposed method, we conducted experiments for speech emotion recognition on the IEMOCAP and RAVDESS datasets. Experimental results show that the proposed method outperforms the latest (non-)graph-based models. Yingxue Gao, Huan Zhao 0003, Zixing Zhang 0001 |
ICASSP | 3 |
| 2024 | Customising General Large Language Models for Specialised Emotion Recognition TasksabstractThe advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning, and others. Leveraging the capability of LLMs inevitably becomes an essential solution for emotion recognition. To this end, we further comprehensively investigate how LLMs perform in linguistic emotion recognition if we concentrate on this specific task. Specifically, we exemplify a publicly available and widely used LLM – Chat General Language Model, and customise it for our target by using two different model adaptation techniques, i.e., deep prompt tuning and low-rank adaptation. The experimental results obtained on six widely used datasets present that the adapted LLM can easily outperform other state-of-the-art but specialised deep models. This indicates the strong transferability and feasibility of LLMs in the field of emotion recognition. Liyizhe Peng, Zixing Zhang 0001, Jing Han 0010, Huan Zhao 0003, Björn W. Schuller |
ICASSP | 2 |
| 2024 | Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion RecognitionabstractConversational Emotion Recognition (CER) aims to predict the emotion expressed by an utterance (referred to as an "event") during a conversation. Existing graph-based methods mainly focus on event interactions to comprehend the conversational context, while overlooking the direct influence of the speaker’s emotional state on the events. In addition, real-time modeling of the conversation is crucial for real-world applications but is rarely considered. Toward this end, we propose a novel graph-based approach, namely Event-State Interactions infused Heterogeneous Graph Neural Network (ESIHGNN), which incorporates the speaker’s emotional state and constructs a heterogeneous event-state interaction graph to model the conversation. Specifically, a heterogeneous directed acyclic graph neural network is employed to dynamically update and enhance the representations of events and emotional states at each turn, thereby improving conversational coherence and consistency. Furthermore, to further improve the performance of CER, we enrich the graph’s edges with external knowledge. Experimental results on four publicly available CER datasets show the superiority of our approach and the effectiveness of the introduced heterogeneous event-state interaction graph. Xupeng Zha, Huan Zhao 0003, Zixing Zhang 0001 |
ICASSP | 3 |
| 2024 | LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement FeedbackabstractGuanghui Ye, Huan Zhao, Zixing Zhang, Xupeng Zha, Zhihua Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Xupeng Zha, Zhihua Jiang |
NAACL-HLT | 3 |
| 2024 | Individual mapping and asymmetric dual supervision for discrete cross-modal hashing
Song Wang 0016, Huan Zhao 0003, Zixing Zhang 0001, Keqin Li 0001 |
Expert Syst. Appl. | 3 |
| 2024 | An Angle-Oriented Approach to Transferring Speech to Gesture for Highly Anthropomorphized Embodied Conversational AgentsabstractRealistic co-speech gestures are important to anthropomorphize ECAs, as nonverbal behavior improves expressiveness of their speech greatly. However, the existing approaches to generating co-speech gestures with sufficient details (including fingers, etc.) in 3D scenarios are indeed rare. Additionally, they hardly address the problem of abnormal gestures, temporal–spatial coherence and diversity of gesture sequences comprehensively. To handle abnormal gesture issues, we put forward an angle conversion method to remove body part length from the original in-the-wild video dataset via transferring coordinates of human upper body key points into relative deflection angles and pitch angles. We also propose a neural network called HARP with encoder–decoder architecture to transfer MFCC featured speech audio into aforementioned angles on the basis of CNN and LSTM. The angles then can be rendered as corresponding co-speech gestures. Compared with the other latest approaches, the co-speech gestures generated by HARP are proved to be almost as good as the real person, i.e., they have strong temporal–spatial coherence, diversity, persuasiveness and credibility. Our approach puts finer control on co-speech gestures than most of the existing works by handling all key points of the human upper body. It is more feasible for industrial application, since HARP can be adaptive to any human upper body model. All related code and evidence videos of HARP can be accessed at https://github.com/drrobincroft/HARP . Zheng Qin 0001, Zixing Zhang 0001, Jixin Zhang |
Int. J. Comput. Intell. Appl. | 3 |
| 2024 | Discriminative Feature Learning-Based Federated Lightweight Distillation Against Multiple AttacksabstractThanks to the advantages of cloud and edge computing, federated learning (FL)–based speech emotion recognition (SER) tasks can be well-scaled to cloud-edge-terminal ecosystems. It aims to characterize emotions while protecting data privacy. However, catastrophic forgetting caused by data heterogeneity, potential system attacks, and possible privacy leakage and communication overhead from parameter sharing have constrained its breakthrough. Some schemes that attempt to tackle the FL bottleneck do not consider these issues comprehensively. We propose a federated distillation-based multiple defense approach (FedMud), which simultaneously considers how to balance system performance, privacy security, and communication overhead. First, It employs a server-side lightweight generator to learn global view knowledge and guides client-side updates through distillation, further mitigating catastrophic forgetting and improving system performance. In addition, we design a multi-path integrated defense paradigm to counter potential system attacks, with a data perturbation technique based on gradient modification, a dynamically weighted selection method, and a privacy-enhanced strategy by capturing discriminative features. Moreover, to minimize parameter leakage, the parameter-decoupled hierarchical sharing mechanism is utilized, which also significantly reduces the communication overhead. The experimental results show that our approach is effective, with gender predictions down to chance levels while maintaining SER performance enhancements. Haijiao Chen, Huan Zhao 0003, Zixing Zhang 0001, Keqin Li 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Gradient-Level Differential Privacy Against Attribute Inference Attack for Speech Emotion RecognitionabstractThe Federated Learning (FL) paradigm for distributed privacy preservation is valued for its ability to collaboratively train Speech Emotion Recognition (SER) models while keeping data localized. However, recent studies reveal privacy leakage in the model sharing process. Existing differential privacy schemes face increasing inference attack risks as clients expose more model updates. To address these challenges, we propose aGradient-levelHierarchicalDifferentialPrivacy (GHDP) strategy to mitigate attribute inference attacks. GHDP employs normalization to distinguish gradient importance, clipping significant gradients and filtering out sensitive information that may lead to privacy leaks. Additionally, increased random perturbations are applied to early model layers during backpropagation, achieving hierarchical differential privacy through layered noise addition. This theoretically grounded approach offers enhanced protection for critical information. Our experiments show that GHDP maintains stable SER performance while providing robust privacy protection, unaffected by the number of model updates. Haijiao Chen, Huan Zhao 0003, Zixing Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Refashioning Emotion Recognition Modeling: The Advent of Generalized Large ModelsabstractAfter its inception, emotion recognition or affective computing has increasingly become an active research topic due to its broad applications. The corresponding computational models have gradually migrated from statistically shallow models to neural-network-based deep models, which can significantly boost the performance of emotion recognition and consistently achieve the best results on different benchmarks, and thus has been considered the first option for emotion recognition. However, the debut of large language models (LLMs), such as ChatGPT and GPT4, has remarkably astonished the world due to their emerged capabilities of zero/few-shot learning, in-context learning (ICL), chain-of-thought, and others that are never shown in previous deep models. In the present article, we comprehensively investigate how the LLMs perform in emotion recognition in terms of diverse aspects, including ICL, few-shot prompting, accuracy, generalization, and explanation. Moreover, we offer some insights and pose other potential challenges, hoping to ignite broader discussions about enhancing emotion recognition in the new era of advanced and more generalized models. Zixing Zhang 0001, Liyizhe Peng, Jing Han 0010, Huan Zhao 0003, Björn W. Schuller |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed PrototypesabstractZero-shot Speech Emotion Recognition (SER) enables machines to perceive unseen-emotional speech without knowing any samples from these emotional states, which is helpful in audio-based autonomous affective computing. However, existing works on zero-shot SER directly employ original prototypes and only consider inter-domain knowledge transfer through learning unseen-emotional classifiers. In this regard, we propose a zero-shot SER approach using generative learning with reconstructed prototypes in this paper. Within the proposed approach, we first reconstruct prototypes using the alignment from paralinguistic features to semantic prototypes. Then, generative learning is performed to build the connection from the reconstructed prototypes to the features. Afterwards, zero-shot experiments on emotional-speech data demonstrate that the proposed approach achieves better performance compared with the state-of-the-art approaches. Xinzhou Xu, Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 3 |
| 2023 | Privacy-Enhanced Federated Learning Against Attribute Inference Attack for Speech Emotion RecognitionabstractFederal learning-based (FL) Speech Emotion Recognition (SER) framework aims to protect data privacy when characterizing emotions. However, previous studies have shown that the framework is vulnerable, because curious servers can indirectly infer user private information. To address this challenge, we propose a novel privacy- enhanced SER approach against attribute inference attack. It helps filter sensitive information and attends to highlight emotion features before uploading the shared model updates under the FL. Firstly, a bi-directional recurrent neural network captures the latent representations in sequences to discard partial redundant features. Then, a feature attention mechanism is applied to focus on the salient regions in the latent representations, further hiding emotion-irrelevant attributes. The experimental results show that the introduced model is effective. The attack capability of a gender prediction model is reduced to a chance level while retaining SER performance. Huan Zhao 0003, Haijiao Chen, Yufeng Xiao, Zixing Zhang 0001 |
ICASSP | 4 |
| 2023 | Speaker-aware Cross-modal Fusion Architecture for Conversational Emotion Recognition
Huan Zhao 0003, Zixing Zhang 0001 |
INTERSPEECH | 3 |
| 2023 | Frequency Domain Feature Learning with Wavelet Transform for Image Translation
Huan Zhao 0003, Yujiang Wang 0005, Song Wang 0016, Lixuan Li, Xupeng Zha, Zixing Zhang 0001 |
PRICAI (3) | 7 |
| 2022 | Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional NetworkabstractAutomated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classification tasks, which tend to be smaller and imbalanced. In this paper, we propose to explore the effectiveness of a multi-branch Temporal Convolutional Network (TCN) architecture integrated with Squeeze-and-Excitation Network (SEnet), a system denoted herein as MBTCNSE, for respiratory sound classification. To the best of the authors’ knowledge, this is the first time that such a hybrid architecture has been employed for respiratory sounds classification. Experiments based on the ICBHI challenge respiratory sound dataset demonstrate the effectiveness of our method. Ziping Zhao 0001, Mingyue Niu, Haishuai Wang, Zixing Zhang 0001, Ya Li 0001 |
ICASSP | 6 |
| 2022 | Deliberation Selector for Knowledge-Grounded Conversation Generation
Huan Zhao 0003, Song Wang 0016, Zixing Zhang 0001, Xupeng Zha |
PRICAI (3) | 5 |
| 2022 | Rethinking Auditory Affective Descriptors Through Zero-Shot Emotion Recognition in SpeechabstractZero-shot speech emotion recognition (SER) endows machines with the ability of sensing unseen-emotional states in speech, compared with conventional SER endeavors on supervised cases. On addressing the zero-shot SER task, auditory affective descriptors (AADs) are typically employed to transfer affective knowledge from seen- to unseen-emotional states. However, it remains unknown which types of AADs can well describe emotional states in speech during the transfer. In this regard, we define and research on three types of AADs, namely, per-emotion semantic-embedding, per-emotion manually annotated, and per-sample manually annotated AADs, through zero-shot emotion recognition in speech. This leads to a systematic design including prototype- and annotation-based zero-shot SER modules, relying on the input from per-emotion and per-sample AADs, respectively. We then perform extensive experimental comparisons between human and machines’ AADs on the French emotional speech corpus CINEMO for positive-negative (PN) and within-negative (WN) tasks. The experimental results indicate that semantic-embedding prototypes from pretrained models can outperform manually annotated emotional dimensions in zero-shot SER. The results further demonstrate that it is possible for machines to understand and describe affective information in speech better than human beings, with the help of sufficient pretrained models. Xinzhou Xu, Zixing Zhang 0001, Xijian Fan, Li Zhao 0003, Laurence Devillers, Björn W. Schuller |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding PrototypesabstractSpeech Emotion Recognition (SER) makes it possible for machines to perceive affective information. Our previous research differed from conventional SER endeavours in that it focused on recognising unseen emotions in speech autonomously through machine learning. Such a step would enable the automatic leaning of unknown emerging emotional states. This type of learning framework, however, still relied on manual annotations to obtain multiple samples of each emotion. In order to reduce this additional workload, herein, we propose a zero-shot SER framework employing a per-emotion semantic-embedding paradigm to describe emotions in zero-shot SER, instead of using the sample-wise descriptors. Aiming to optimise the relationship between emotions, prototypes, and speech samples, this framework includes two types of learning strategies: Sample-wise learning and emotion-wise learning. These strategies apply a novel learning process to speech samples and emotions, respectively, via specifically designed semantic-embedding prototypes. We verify the utility of these approaches by performing an extensive experimental evaluation on two corpora on three aspects, namely the influence of different types of learning strategies, emotional-pair comparison, and the selections of semantic-embedding prototypes and paralinguistic features. The experimental results indicate that it is applicable to use semantic-embedding prototypes for zero-shot emotion recognition in speech, despite the influence of choosing optimal strategies and prototypes. Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
IEEE Trans. Multim. | 4 |
| 2021 | Learning audio sequence representations for acoustic event classification
Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Björn W. Schuller |
Expert Syst. Appl. | 1 |
| 2021 | Internet of emotional people: Towards continual affective computing cross cultures via audiovisual signals
Jing Han 0010, Zixing Zhang 0001, Maja Pantic, Björn W. Schuller |
Future Gener. Comput. Syst. | 2 |
| 2021 | Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19abstractComputer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose. Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE Internet Things J. | 13 |
| 2021 | Combining a parallel 2D CNN with a self-attention Dilated Residual Network for CTC-based discrete speech emotion recognition
Ziping Zhao 0001, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
Neural Networks | 3 |
| 2021 | EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion EmbeddingsabstractDespite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous presence of all involved multiple modalities. Motivated by this, we propose a novel crossmodal emotion embedding framework called EmoBed, which aims to leverage the knowledge from other auxiliary modalities to improve the performance of an emotion recognition system at hand. The framework generally includes two main learning components, i.e., joint multimodal training and crossmodal training. Both of them tend to explore the underlying semantic emotion information but with a shared recognition network or with a shared emotion embedding space, respectively. In doing this, the enhanced system trained with this approach can efficiently make use of the complementary information from other modalities. Nevertheless, the presence of these auxiliary modalities is not demanded during inference. To empirically investigate the effectiveness and robustness of the proposed framework, we perform extensive experiments on the two benchmark databases RECOLA and OMG-Emotion for the tasks of dimensional emotion regression and categorical emotion classification, respectively. The obtained results show that the proposed framework significantly outperforms related baselines in monomodal inference, and are also competitive or superior to the recently reported systems, which emphasises the importance of the proposed crossmodal learning for emotion recognition. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Can Machine Learning Assist Locating the Excitation of Snore Sound? A ReviewabstractIn the past three decades, snoring (affecting more than 30 % adults of the UK population) has been increasingly studied in the transdisciplinary research community involving medicine and engineering. Early work demonstrated that, the snore sound can carry important information about the status of the upper airway, which facilitates the development of non-invasive acoustic based approaches for diagnosing and screening of obstructive sleep apnoea and other sleep disorders. Nonetheless, there are more demands from clinical practice on finding methods to localise the snore sound's excitation rather than only detecting sleep disorders. In order to further the relevant studies and attract more attention, we provide a comprehensive review on the state-of-the-art techniques from machine learning to automatically classify snore sounds. First, we introduce the background and definition of the problem. Second, we illustrate the current work in detail and explain potential applications. Finally, we discuss the limitations and challenges in the snore sound classification task. Overall, our review provides a comprehensive guidance for researchers to contribute to this area. Kun Qian 0003, Christoph Janott, Maximilian Schmitt, Zixing Zhang 0001, Clemens Heiser, Werner Hemmert, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Self-attention transfer networks for speech emotion recognitionabstractA crucial element of human–machine interaction, the automatic detection of emotional states from human speech has long been regarded as a challenging task for machine learning models. One vital challenge in speech emotion recognition (SER) is how to learn robust and discriminative representations from speech. Meanwhile, although machine learning methods have been widely applied in SER research, the inadequate amount of available annotated data has become a bottleneck that impedes the extended application of techniques (e.g., deep neural networks). To address this issue, we present a deep learning method that combines knowledge transfer and self-attention for SER tasks. Here, we apply the log-Mel spectrogram with deltas and delta-deltas as input. Moreover, given that emotions are time-dependent, we apply Temporal Convolutional Neural Networks (TCNs) to model the variations in emotions. We further introduce an attention transfer mechanism, which is based on a self-attention algorithm in order to learn long-term dependencies. The Self-Attention Transfer Network (SATN) in our proposed approach, takes advantage of attention autoencoders to learn attention from a source task, and then from speech recognition, followed by transferring this knowledge into SER. Evaluation built on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) demonstrates the effectiveness of the novel model. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Shihuang Sun, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition ModelsabstractThe development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this study, we aim to train deep models to defending against non-targeted white-box adversarial attacks. Adversarial data is first generated from the real data using the fast gradient sign method. Then in the research field of speech emotion recognition, adversarial-based training is employed as a method for protecting against adversarial attack. We then train deep convolutional models with both real and adversarial data, and compare the performances of two adversarial training procedures - namely, vanilla adversarial training, and similarity-based adversarial training. In our experiments, through the use of adversarial data augmentation, both of the considered adversarial training procedures can improve the performance when validated on the real data. Additionally, the similarity-based adversarial training learns a more robust model when working with adversarial data. Finally, the considered VGG-16 model performs the best across all models, for both real and generated data. Zhao Ren, Alice Baird, Jing Han 0010, Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 4 |
| 2020 | Hierarchical Attention Transfer Networks for Depression Assessment from SpeechabstractA growing area of mental health research is the search for speech-based objective markers for conditions such as depression. However, when combined with machine learning, this search can be challenging due to a limited amount of annotated training data. In this paper, we propose a novel crosstask approach which transfers attention mechanisms from speech recognition to aid depression severity measurement. This transfer is applied in a two-level hierarchical network which mirrors the natural hierarchical structure of speech. Experiments based on the Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) dataset, as used in the 2017 Audio/Visual Emotion Challenge, demonstrate the effectiveness of our Hierarchical Attention Transfer Network. On the development set, the proposed approach achieves a root mean square error (RMSE) of 3.85, and a mean absolute error (MAE) of 2.99, on a Patient Health Questionnaire (PHQ)-8 scale [0], [24], while on the test set, it achieves an RMSE of 5.66 and an MAE of 4.28. To the best of our knowledge, these scores represent the best-known speech-only results to date on this corpus. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Björn W. Schuller |
ICASSP | 3 |
| 2020 | An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and AnxietyabstractThe COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease. Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
INTERSPEECH | 12 |
| 2020 | Exploiting time-frequency patterns with LSTM-RNNs for low-bitrate audio restoration
Björn W. Schuller, Florian Eyben, Dagmar Schuller, Zixing Zhang 0001, Holly Francois, Eunmi Oh |
Neural Comput. Appl. | 5 |
| 2020 | Snore-GANs: Improving Automatic Snore Sound Classification With Synthesized DataabstractOne of the frontier issues that severely hamper the development of automatic snore sound classification (ASSC) associates to the lack of sufficient supervised training data. To cope with this problem, we propose a novel data augmentation approach based on semi-supervised conditional generative adversarial networks (scGANs), which aims to automatically learn a mapping strategy from a random noise space to original data distribution. The proposed approach has the capability of well synthesizing "realistic" high-dimensional data, while requiring no additional annotation process. To handle the mode collapse problem of GANs, we further introduce an ensemble strategy to enhance the diversity of the generated data. The systematic experiments conducted on a widely used Munich-Passau snore sound corpus demonstrate that the scGANs-based systems can remarkably outperform other classic data augmentation systems, and are also competitive to other recently reported systems for ASSC. Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Christoph Janott, Yanan Guo 0001, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | Audiovisual Analysis for Recognising Frustration during Game-Play: Introducing the Multimodal Game Frustration DatabaseabstractAutomatic recognition of frustration, by analysing facial and vocal expressions, can help user experience designers to identify interaction obstacles. To encourage the development of automated systems such as these, we present a novel audiovisual database: the Multimodal Game Frustration Database (MGFD), consisting of ca. 5 hours of audiovisual data, collected from 67 Chinese students speaking in English. For data collection, we developed ‘Crazy Trophy’, a Wizard-of-Oz voice activated web-game designed with a variety of usability problems and aimed to induce increasing amounts of frustration. We also present a baseline for binary multimodal frustration classification (frustration vs no-frustration). For this, we compare the performance of a conventional method, Support Vector Machine classifier, and a state-of-the-art method utilising Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN), extracting both audio (Mel-frequency Cepstral Coefficients) and video (facial action units) features. Using LSTM-RNN and a feature-based multi-model fusion strategy, the best result acheived for the baseline was 60.3 % UAR. To enable further research in this area, the game (‘Crazy Trophy’), the database (MGFD), and the partitioning considered in the presented baseline, are made accessible to the research community. Meishu Song, Zijiang Yang 0007, Alice Baird, Emilia Parada-Cabaleiro, Zixing Zhang 0001, Ziping Zhao 0001, Björn W. Schuller |
ACII | 5 |
| 2019 | Attention-augmented End-to-end Multi-task Learning for Emotion Prediction from SpeechabstractDespite the increasing research interest in end-to-end learning systems for speech emotion recognition, conventional systems either suffer from the overfitting due in part to the limited training data, or do not explicitly consider the different contributions of automatically learnt representations for a specific task. In this contribution, we propose a novel end-to-end framework which is enhanced by learning other auxiliary tasks and an attention mechanism. That is, we jointly train an end-to-end network with several different but related emotion prediction tasks, i. e., arousal, valence, and dominance predictions, to extract more robust representations shared among various tasks than traditional systems with the hope that it is able to relieve the overfitting problem. Meanwhile, an attention layer is implemented on top of the layers for each task, with the aim to capture the contribution distribution of different segment parts for each individual task. To evaluate the effectiveness of the proposed system, we conducted a set of experiments on the widely used database IEMOCAP. The empirical results show that the proposed systems significantly outperform corresponding baseline systems. Zixing Zhang 0001, Bingwen Wu, Björn W. Schuller |
ICASSP | 1 |
| 2019 | Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono ModalityabstractDespite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the training procedure for either speech or facial emotion recognition. Specifically, the model consists of one modality-specific network per individual modality and one shared network to map both audio and visual cues into final predictions. In the training process, we additionally take the loss from one auxiliary modality into account besides the main modality. To evaluate the effectiveness of the implicit fusion model, we conduct extensive experiments for mono-modal emotion classification and regression, and find that the implicit fusion models outperform the standard mono-modal training process. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Björn W. Schuller |
ICASSP | 2 |
| 2019 | Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion RecognitionabstractDespite the great advances, most of the recently developed automatic speech recognition systems focus on working in a server-client manner, and thus often require a high computational cost, such as the storage size and memory accesses. This, however, does not satisfy the increasing demand for a succinct model that can run smoothly in embedded devices like smartphones. To this end, in this paper we propose a neural network compression method, in the way of quantizing the weights of the neural networks from the original full-precised values into binary values that then can be stored and processed with only one bit per value. In doing this, the traditional neural network-based large-size speech emotion recognition models can be greatly compressed into smaller ones, which demand lower computational cost. To evaluate the feasibility of the proposed approach, we take a state-of-the-art speech emotion recognition model, i. e., convolutional recurrent neural networks, as an example, and conduct experiments on two widely used emotional databases. We find that the proposed binary neural networks are able to yield a remarkable model compression rate but at limited expense of model performance. Huan Zhao 0003, Yufeng Xiao, Jing Han 0010, Zixing Zhang 0001 |
ICASSP | 4 |
| 2019 | VCMNet: Weakly Supervised Learning for Automatic Infant Vocalisation Maturity AnalysisabstractUsing neural networks to classify infant vocalisations into important subclasses (such as crying versus speech) is an emergent task in speech technology. One of the biggest roadblocks standing in the way of progress lies in the datasets: The performance of a learning model is affected by the labelling quality and size of the dataset used, and infant vocalisation datasets with good quality labels tend to be small. In this paper, we assess the performance of three models for infant VoCalisation Maturity (VCM) trained with a large dataset annotated automatically using a purpose-built classifier and a small dataset annotated by highly trained human coders. The two datasets are used in three different training strategies, whose performance is compared against a baseline model. The first training strategy investigates adversarial training, while the second exploits multi-task learning as the neural network trains on both datasets simultaneously. In the final strategy, we integrate adversarial training and multi-task learning. All of the training strategies outperform the baseline, with the adversarial training strategy yielding the best results on the development set. Najla Al Futaisi, Zixing Zhang 0001, Alejandrina Cristià, Anne S. Warlaumont, Björn W. Schuller |
ICMI | 2 |
| 2019 | Autonomous Emotion Learning in Speech: A View of Zero-Shot Speech Emotion RecognitionabstractConventionally, speech emotion recognition is achieved using passive learning approaches.Differing from such approaches, we herein propose and develop a dynamic method of autonomous emotion learning based on zero-shot learning.The proposed methodology employs emotional dimensions as the attributes in the zero-shot learning paradigm, resulting in two phases of learning, namely attribute learning and label learning.Attribute learning connects the paralinguistic features and attributes utilising speech with known emotional labels, while label learning aims at defining unseen emotions through the attributes.The experimental results achieved on the CINEMO corpus indicate that zero-shot learning is a useful technique for autonomous speech-based emotion learning, achieving accuracies considerably better than chance level and an attribute-based gold-standard setup.Furthermore, different emotion recognition tasks, emotional attributes, and employed approaches strongly influence system performance. Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
INTERSPEECH | 4 |
| 2019 | Attention-Enhanced Connectionist Temporal Classification for Discrete Speech Emotion RecognitionabstractDiscrete speech emotion recognition (SER), the assignment of a single emotion label to an entire speech utterance, is typically performed as a sequence-to-label task.This approach, however, is limited, in that it can result in models that do not capture temporal changes in the speech signal, including those indicative of a particular emotion.One potential solution to overcome this limitation is to model SER as a sequence-to-sequence task instead.In this regard, we have developed an attention-based bidirectional long short-term memory (BLSTM) neural network in combination with a connectionist temporal classification (CTC) objective function (Attention-BLSTM-CTC) for SER.We also assessed the benefits of incorporating two contemporary attention mechanisms, namely component attention and quantum attention, into the CTC framework.To the best of the authors' knowledge, this is the first time that such a hybrid architecture has been employed for SER.We demonstrated the effectiveness of our approach on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and FAU-Aibo Emotion corpora.The experimental results demonstrate that our proposed model outperforms current state-of-the-art approaches. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Björn W. Schuller |
INTERSPEECH | 3 |
| 2019 | Dynamic Difficulty Awareness Training for Continuous Emotion PredictionabstractTime-continuous emotion prediction has become an increasingly compelling task in machine learning. Considerable efforts have been made to advance the performance of these systems. Nonetheless, the main focus has been the development of more sophisticated models and the incorporation of different expressive modalities (e.g., speech, face, and physiology). In this paper, motivated by the benefit of difficulty awareness in a human learning procedure, we propose a novel machine learning framework, namely, dynamic difficulty awareness training (DDAT), which sheds fresh light on the research - directly exploiting the difficulties in learning to boost the machine learning process. The DDAT framework consists of two stages: information retrieval and information exploitation. In the first stage, we make use of the reconstruction error of input features or the annotation uncertainty to estimate the difficulty of learning specific information. The obtained difficulty level is then used in tandem with original features to update the model input in a second learning stage with the expectation that the model can learn to focus on high difficulty regions of the learning process. We perform extensive experiments on a benchmark database REmote COLlaborative and affective to evaluate the effectiveness of the proposed framework. The experimental results show that our approach outperforms related baselines as well as other well-established time-continuous emotion prediction systems, which suggests that dynamically integrating the difficulty information for neural networks can help enhance the learning process. Zixing Zhang 0001, Jing Han 0010, Eduardo Coutinho, Björn W. Schuller |
IEEE Trans. Multim. | 1 |
| 2018 | Towards Conditional Adversarial Training for Predicting Emotions from SpeechabstractMotivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consists of two networks, trained in an adversarial manner: The first network tries to predict emotion from acoustic features, while the second network aims at distinguishing between the predictions provided by the first network and the emotion labels from the database using the acoustic features as conditional information. We evaluate the performance of the proposed conditional adversarial training framework on the widely used emotion database RECOLA. Experimental results show that the proposed training strategy outperforms the conventional training method, and is comparable with, or even superior to other recently reported approaches, including deep and end-to-end learning. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Fabien Ringeval, Björn W. Schuller |
ICASSP | 2 |
| 2018 | Exploring A New Method for Food Likability Rating Based on DT-CWT TheoryabstractIn this paper, we mainly investigate subjects' food likability based on audio-related features as a contribution to EAT ? the ICMI 2018 Eating Analysis and Tracking challenge. Specifically, we conduct 4-level Double Tree Complex Wavelet Transform decomposition of an audio signal, and obtain five sub-audio signals with frequencies ranging from low to high. For each sub-audio signal, not only 'traditional' functional-based features but also deep learning-based features via pretrained CNNs based on SliCQ-nonstationary Gabor transform and a cochleagram map, are calculated. Besides, the original audio signals based Bag-of-Audio-Words features extracted by the openXBOW toolkit are used to enhance the model as well. Finally, the early fusion of all these three kinds of features can lead to promising results, yielding the highest UAR of 79.2 % by means of a leave-one-speaker-out cross-validation, which holds a 12.7 % absolute gain compared with the baseline of 66.5 % UAR. Yanan Guo 0001, Jing Han 0010, Zixing Zhang 0001, Björn W. Schuller, Yide Ma |
ICMI | 3 |
| 2018 | Evolving Learning for Analysing Mood-Related Infant VocalisationabstractInfant vocalisation analysis plays an important role in the study of the development of pre-speech capability of infants, while machine-based approaches nowadays emerge with an aim to advance such an analysis.However, conventional machine learning techniques require heavy feature-engineering and refined architecture designing.In this paper, we present an evolving learning framework to automate the design of neural network structures for infant vocalisation analysis.In contrast to manually searching by trial and error, we aim to automate the search process in a given space with less interference.This framework consists of a controller and its child networks, where the child networks are built according to the controller's estimation.When applying the framework to the Interspeech 2018 Computational Paralinguistics (ComParE) Crying Subchallenge, we discover several deep recurrent neural network structures, which are able to deliver competitive results to the best ComParE baseline method. Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Björn W. Schuller |
INTERSPEECH | 1 |
| 2018 | Bags in Bag: Generating Context-Aware Bags for Tracking Emotions from SpeechabstractInternational audience Jing Han 0010, Zixing Zhang 0001, Maximilian Schmitt, Zhao Ren, Fabien Ringeval, Björn W. Schuller |
INTERSPEECH | 2 |
| 2018 | Automated Classification of Children's Linguistic versus Non-Linguistic VocalisationsabstractA key outstanding task for speech technology involves dealing with non-standard speakers, notably young children.Distinguishing children's linguistic from non-linguistic vocalisations is crucial for a number of applied and fundamental research goals, and yet there are few systems available for such a classification.This paper investigates two large-scale framelevel acoustic feature sets (eGeMAPS and ComParE16) followed by a dynamic model (GRU-RNN), and two kinds of derived static feature sets on the segment level (functional-based and Bag of Audio Words) combined with a static model (SVM), and automatically learnt representations directly from original raw voice signals by using an end-to-end system.These are applied to a large database of children's vocalisations (total N = 6,298) drawn from daylong recordings gathered in Namibia, Bolivia, and Vanuatu.Among these systems, the one implemented with GRU-RNN using ComParE16 features empirically performs best.We further identify promising paths of further research, including the application of a finer-grained classification of children's vocalisations onto these data, and the exploration of other feature systems. Zixing Zhang 0001, Alejandrina Cristià, Anne S. Warlaumont, Björn W. Schuller |
INTERSPEECH | 1 |
| 2018 | Exploring Spatio-Temporal Representations by Integrating Attention-based Bidirectional-LSTM-RNNs and FCNs for Speech Emotion RecognitionabstractAutomatic emotion recognition from speech, which is an important and challenging task in the field of affective computing, heavily relies on the effectiveness of the speech features for classification. Previous approaches to emotion recognition have mostly focused on the extraction of carefully hand-crafted features. How to model spatio-temporal dynamics for speech emotion recognition effectively is still under active investigation. In this paper, we propose a method to tackle the problem of emotional relevant feature extraction from speech by leveraging Attention-based Bidirectional Long Short-Term Memory Recurrent Neural Networks with fully convolutional networks in order to automatically learn the best spatio-temporal representations of speech signals. The learned high-level features are then fed into a deep neural network (DNN) to predict the final emotion. The experimental results on the Chinese Natural Audio-Visual Emotion Database (CHEAVD) and the Interactive Emotional Dyadic Motion Capture (IEMOCAP) corpora show that our method provides more accurate predictions compared with other existing emotion recognition algorithms. Ziping Zhao 0001, Yu Zheng 0013, Zixing Zhang 0001, Haishuai Wang, Yiqin Zhao |
INTERSPEECH | 3 |
| 2018 | Semisupervised Autoencoders for Speech Emotion RecognitionabstractDespite the widespread use of supervised learning methods for speech emotion recognition, they are severely restricted due to the lack of sufficient amount of labelled speech data for the training. Considering the wide availability of unlabelled speech data, therefore, this paper proposes semisupervised autoencoders to improve speech emotion recognition. The aim is to reap the benefit from the combination of the labelled data and unlabelled data. The proposed model extends a popular unsupervised autoencoder by carefully adjoining a supervised learning objective. We extensively evaluate the proposed model on the INTERSPEECH 2009 Emotion Challenge database and other four public databases in different scenarios. Experimental results demonstrate that the proposed model achieves state-of-the-art performance with a very small number of labelled data on the challenge task and other tasks, and significantly outperforms other alternative methods. Xinzhou Xu, Zixing Zhang 0001, Sascha Frühholz, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Deep Learning for Environmentally Robust Speech Recognition: An Overview of Recent DevelopmentsabstractEliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition but still remains an important challenge. Data-driven supervised approaches, especially the ones based on deep neural networks, have recently emerged as potential alternatives to traditional unsupervised approaches and with sufficient training, can alleviate the shortcomings of the unsupervised methods in various real-life acoustic environments. In this light, we review recently developed, representative deep learning approaches for tackling non-stationary additive and convolutional degradation of speech with the aim of providing guidelines for those involved in the development of environmentally robust speech recognition systems. We separately discuss single- and multi-channel techniques developed for the front-end and back-end of speech recognition systems, as well as joint front-end and back-end training frameworks. In the meanwhile, we discuss the pros and cons of these approaches and provide their experimental results on benchmark databases. We expect that this overview can facilitate the development of the robustness of speech recognition systems in acoustic noisy environments. Zixing Zhang 0001, Jürgen T. Geiger, Jouni Pohjalainen, Amr El-Desoky Mousa, Wenyu Jin 0002, Björn W. Schuller |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2017 | Reconstruction-error-based learning for continuous emotion recognition in speechabstractTo advance the performance of continuous emotion recognition from speech, we introduce a reconstruction-error-based (RE-based) learning framework with memory-enhanced Recurrent Neural Networks (RNN). In the framework, two successive RNN models are adopted, where the first model is used as an autoencoder for reconstructing the original features, and the second is employed to perform emotion prediction. The RE of the original features is used as a complementary descriptor, which is merged with the original features and fed to the second model. The assumption of this framework is that the system has the ability to learn its `drawback' which is expressed by the RE. Experimental results on the RECOLA database show that the proposed framework significantly outperforms the baseline systems without any RE information in terms of Concordance Correlation Coefficient (.729 vs .710 for arousal, .360 vs .237 for valence), and also significantly overcomes other state-of-the-art methods. Jing Han 0010, Zixing Zhang 0001, Fabien Ringeval, Björn W. Schuller |
ICASSP | 2 |
| 2017 | Prediction-based learning for continuous emotion recognition in speechabstractIn this paper, a prediction-based learning framework is proposed for a continuous prediction task of emotion recognition from speech, which is one of the key components of affective computing in multimedia. The main goal of this framework is to utmost exploit the individual advantages of different regression models cooperatively. To this end, we take two widely used regression models for example, i. e., support vector regression and bidirectional long short-term memory recurrent neural network. We concatenate the two models in a tandem structure by different ways, forming a united cascaded framework. The outputs predicted by the former model are combined together with the original features as the input of the following model for final predictions. The experimental results on a time- and value-continuous spontaneous emotion database (RECOLA) show that, the prediction-based learning framework significantly outperforms the individual models for both arousal and valence dimensions, and provides significantly better results in comparison to other state-of-the-art methodologies on this corpus. Jing Han 0010, Zixing Zhang 0001, Fabien Ringeval, Björn W. Schuller |
ICASSP | 2 |
| 2017 | Towards intoxicated speech recognitionabstractIn a real-life scenario, the acoustic characteristics of speech often suffer from the variations induced by diverse environmental noises and different speakers. To overcome the speaker-related speech variation problem for Automatic Speech Recognition (ASR), many speaker adaptation techniques have been proposed and studied. Almost all of these studies, however, only considered the speakers' long-term traits, such as age, gender, and dialect. Speakers' short-term states, for example, affect and intoxication, are largely ignored. In this study, we address one particular speaker state, alcohol intoxication, which has rarely been studied in the context of ASR. To do this, empirical experiments are performed on a publicly available database used for the INTERSPEECH 2011 Speaker State Challenge, Intoxication Sub-Challenge. The experimental results show that the intoxicated state of the speaker indeed degrades the performance of ASR systems by a large margin for all of the three considered speech styles (spontaneous speech, tongue twisters, command & control). In addition, this paper further shows that multi-condition training can notably improve the acoustic model. Zixing Zhang 0001, Felix Weninger, Martin Wöllmer, Jing Han 0010, Björn W. Schuller |
IJCNN | 1 |
| 2017 | Towards Intelligent Crowdsourcing for Audio Data Annotation: Integrating Active Learning in the Real WorldabstractIn this contribution, we combine the advantages of traditional crowdsourcing with contemporary machine learning algorithms with the aim of ultimately obtaining reliable training data for audio processing in a faster, cheaper and therefore more efficient manner than has been previously possible.We propose a novel crowdsourcing approach, which brings a simulated active learning annotation scenario into a real world environment creating an intelligent and gamified crowdsourcing platform for manual audio annotation.Our platform combines two active learning query strategies with an internally calculated trustability score to efficiently reduce manual labelling efforts.This reduction is achieved in a twofold manner: first our system automatically decides if an instance requires annotation; second, it dynamically decides, depending on the quality of previously gathered annotations, on exactly how many annotations are needed to reliably label an instance.Results presented indicate that our approach drastically reduces the annotation load and is considerably more efficient than conventional methods. Simone Hantke, Zixing Zhang 0001, Björn W. Schuller |
INTERSPEECH | 2 |
| 2017 | From Hard to Soft: Towards more Human-like Emotion Recognition by Modelling the Perception UncertaintyabstractOver the last decade, automatic emotion recognition has become well established. The gold standard target is thereby usually calculated based on multiple annotations from different raters. All related efforts assume that the emotional state of a human subject can be identified by a 'hard' category or a unique value. This assumption tries to ease the human observer's subjectivity when observing patterns such as the emotional state of others. However, as the number of annotators cannot be infinite, uncertainty remains in the emotion target even if calculated from several, yet few human annotators. The common procedure to use this same emotion target in the learning process thus inevitably introduces noise in terms of an uncertain learning target. In this light, we propose a 'soft' prediction framework to provide a more human-like and comprehensive prediction of emotion. In our novel framework, we provide an additional target to indicate the uncertainty of human perception based on the inter-rater disagreement level, in contrast to the traditional framework which is merely producing one single prediction (category or value). To exploit the dependency between the emotional state and the newly introduced perception uncertainty, we implement a multi-task learning strategy. To evaluate the feasibility and effectiveness of the proposed soft prediction framework, we perform extensive experiments on a time- and value-continuous spontaneous audiovisual emotion database including late fusion results. We show that the soft prediction framework with multi-task learning of the emotional state and its perception uncertainty significantly outperforms the individual tasks in both the arousal and valence dimensions. Jing Han 0010, Zixing Zhang 0001, Maximilian Schmitt, Maja Pantic, Björn W. Schuller |
ACM Multimedia | 2 |
| 2017 | Strength modelling for real-worldautomatic continuous affect recognition from audiovisual signals
Jing Han 0010, Zixing Zhang 0001, Nicholas Cummins, Fabien Ringeval, Björn W. Schuller |
Image Vis. Comput. | 2 |
| 2017 | Universum Autoencoder-Based Domain Adaptation for Speech Emotion RecognitionabstractOne of the serious obstacles to the applications of speech emotion recognition systems in real-life settings is the lack of generalization of the emotion classifiers. Many recognition systems often present a dramatic drop in performance when tested on speech data obtained from different speakers, acoustic environments, linguistic content, and domain conditions. In this letter, we propose a novel unsupervised domain adaptation model, called Universum autoencoders, to improve the performance of the systems evaluated in mismatched training and test conditions. To address the mismatch, our proposed model not only learns discriminative information from labeled data, but also learns to incorporate the prior knowledge from unlabeled data into the learning. Experimental results on the labeled Geneva Whispered Emotion Corpus database plus other three unlabeled databases demonstrate the effectiveness of the proposed method when compared to other domain adaptation methods. Xinzhou Xu, Zixing Zhang 0001, Sascha Frühholz, Björn W. Schuller |
IEEE Signal Process. Lett. | 3 |
| 2017 | A Two-Dimensional Framework of Multiple Kernel Subspace Learning for Recognizing Emotion in SpeechabstractAs a highly active topic in computational paralinguistics, speech emotion recognition (SER) aims to explore ideal representations for emotional factors in speech. In order to improve the performance of SER, multiple kernel learning (MKL) dimensionality reduction has been utilized to obtain effective information for recognizing emotions. However, the solution of MKL usually provides only one nonnegative mapping direction for multiple kernels; this may lead to loss of valuable information. To address this issue, we propose a two-dimensional framework for multiple kernel subspace learning. This framework provides more linear combinations on the basis of MKL without nonnegative constraints, which preserves more information in the learning procedures. It also leverages both of MKL and two-dimensional subspace learning, combining them into a unified structure. To apply the framework to SER, we also propose an algorithm, namely generalised multiple kernel discriminant analysis (GMKDA), by employing discriminant embedding graphs in this framework. GMKDA takes advantage of the additional mapping directions for multiple kernels in the proposed framework. In order to evaluate the performance of the proposed algorithm a wide range of experiments is carried out on several key emotional corpora. These experimental results demonstrate that the proposed methods can achieve better performance compared with some conventional and subspace learning methods in dealing with SER. Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Wavelet features for classification of vote snore soundsabstractLocation and form of the upper airway obstruction is essential for a targeted therapy of obstructive sleep apnea (OSA). Utilizing snore sounds (SnS) to reveal the pathological characters of OSA patients has been the subject of scientific research for several decades. Fewer studies exist on the evaluation of SnS to identify the corresponding obstruction types in the upper airway. In this study, we propose a novel feature set based on wavelet transform with a support vector machine classifier to discriminate VOTE (velum, oropharyngeal lateral walls, tongue base and epiglottis) snore sounds labelled during drug-induced sleep endoscopy (DISE). Based on snore sound data collected from 24 snoring subjects, processed by a subject-independent 2-fold cross validation experiment, we can show that our wavelet features outperform the frequently-used acoustic features (formants, MFCC, power ratio, crest factor, fundamental frequency) at an WAR (weighted average recall) of 78.2 % and an UAR (unweighted average recall) of 71.2%, with an enhancement ranging from 5.1 % to 24.4% and 12.2% to 46.4% in WAR and UAR, respectively. Kun Qian 0003, Christoph Janott, Zixing Zhang 0001, Clemens Heiser, Björn W. Schuller |
ICASSP | 3 |
| 2016 | Enhanced semi-supervised learning for multimodal emotion recognitionabstractSemi-Supervised Learning (SSL) techniques have found many applications where labeled data is scarce and/or expensive to obtain. However, SSL suffers from various inherent limitations that limit its performance in practical applications. A central problem is that the low performance that a classifier can deliver on challenging recognition tasks reduces the trustability of the automatically labeled data. Another related issue is the noise accumulation problem - instances that are misclassified by the system are still used to train it in future iterations. In this paper, we propose to address both issues in the context of emotion recognition. Initially, we exploit the complementarity between audio-visual features to improve the performance of the classifier during the supervised phase. Then, we iteratively re-evaluate the automatically labeled instances to correct possibly mislabeled data and this enhances the overall confidence of the system's predictions. Experimental results performed on the RECOLA database demonstrate that our methodology delivers a strong performance in the classification of high/low emotional arousal (UAR = 76.5%), and significantly outperforms traditional SSL methods by at least 5.0% (absolute gain). Zixing Zhang 0001, Fabien Ringeval, Eduardo Coutinho, Erik Marchi, Björn W. Schuller |
ICASSP | 1 |
| 2016 | Multiscale kernel locally penalised discriminant analysis exemplified by emotion recognition in speechabstractWe propose a novel method to learn multiscale kernels with locally penalised discriminant analysis, namely Multiscale-Kernel Locally Penalised Discriminant Analysis (MS-KLPDA). As an exemplary use-case, we apply it to recognise emotions in speech. Specifically, we employ the term of locally penalised discriminant analysis by controlling the weights of marginal sample pairs, while the method learns kernels with multiple scales. Evaluated in a series of experiments on emotional speech corpora, our proposed MS-KLPDA is able to outperform the previous research of Multiscale-Kernel Fisher Discriminant Analysis and some conventional methods in solving speech emotion recognition. Xinzhou Xu, Maryna Gavryukova, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller |
ICMI | 4 |
| 2016 | Facing Realism in Spontaneous Emotion Recognition from Speech: Feature Enhancement by Autoencoder with LSTM Neural NetworksabstractInternational audience Zixing Zhang 0001, Fabien Ringeval, Jing Han 0010, Erik Marchi, Björn W. Schuller |
INTERSPEECH | 1 |
| 2016 | Spectral and Cepstral Audio Noise Reduction Techniques in Speech Emotion RecognitionabstractSignal noise reduction can improve the performance of machine learning systems dealing with time signals such as audio. Real-life applicability of these recognition technologies requires the system to uphold its performance level in variable, challenging conditions such as noisy environments. In this contribution, we investigate audio signal denoising methods in cepstral and log-spectral domains and compare them with common implementations of standard techniques. The different approaches are first compared generally using averaged acoustic distance metrics. They are then applied to automatic recognition of spontaneous and natural emotions under simulated smartphone-recorded noisy conditions. Emotion recognition is implemented as support vector regression for continuous-valued prediction of arousal and valence on a realistic multimodal database. In the experiments, the proposed methods are found to generally outperform standard noise reduction algorithms. Jouni Pohjalainen, Fabien Ringeval, Zixing Zhang 0001, Björn W. Schuller |
ACM Multimedia | 3 |
| 2015 | On rater reliability and agreement based dynamic active learningabstractIn this paper, we propose two novel Dynamic Active Learning (DAL) methods with the aim of ultimately reducing the costly human labelling work for subjective tasks such as speech emotion recognition. Compared to conventional Active Learning (AL) algorithms, the proposed DAL approaches employ a highly efficient adaptive query strategy that minimises the number of annotations through three advancements. First, we shift from the standard majority voting procedure, in which unlabelled instances are annotated by a fixed number of raters, to an agreement-based annotation technique that dynamically determines how many human annotators are required to label a selected instance. Second, we introduce the concept of the order-based DAL algorithm by considering rater reliability and inter-rater agreement. Third, a highly dynamic development trend is successfully implemented by upgrading the agreement levels depending on the prediction uncertainty. In extensive experiments on standardised test-beds, we show that the new dynamic methods significantly improve the efficiency of the existing AL algorithms by reducing human labelling effort up to 85.41%, while achieving the same classification accuracy. Thus, the enhanced DAL derivations opens up high-potential research directions for the utmost exploitation of unlabelled data. Yue Zhang 0014, Eduardo Coutinho, Björn W. Schuller, Zixing Zhang 0001, Michael G. Adam |
ACII | 4 |
| 2015 | Dynamic Active Learning Based on Agreement and Applied to Emotion Recognition in Spoken InteractionsabstractIn this contribution, we propose a novel method for Active Learning (AL) - Dynamic Active Learning (DAL) - which targets the reduction of the costly human labelling work necessary for modelling subjective tasks such as emotion recognition in spoken interactions. The method implements an adaptive query strategy that minimises the amount of human labelling work by deciding for each instance whether it should automatically be labelled by machine or manually by human, as well as how many human annotators are required. Extensive experiments on standardised test-beds show that DAL significantly improves the efficiency of conventional AL. In particular, DAL achieves the same classification accuracy obtained with AL with up to 79.17% less human annotation effort. Yue Zhang 0014, Eduardo Coutinho, Zixing Zhang 0001, Caijiao Quan, Björn W. Schuller |
ICMI | 3 |
| 2015 | Cooperative Learning and its Application to Emotion Recognition from SpeechabstractIn this paper, we propose a novel method for highly efficient exploitation of unlabeled data-Cooperative Learning. Our approach consists of combining Active Learning and Semi-Supervised Learning techniques, with the aim of reducing the costly effects of human annotation. The core underlying idea of Cooperative Learning is to share the labeling work between human and machine efficiently in such a way that instances predicted with insufficient confidence value are subject to human labeling, and those with high confidence values are machine labeled. We conducted various test runs on two emotion recognition tasks with a variable number of initial supervised training instances and two different feature sets. The results show that Cooperative Learning consistently outperforms individual Active and Semi-Supervised Learning techniques in all test cases. In particular, we show that our method based on the combination of Active Learning and Co-Training leads to the same performance of a model trained on the whole training set, but using 75% fewer labeled instances. Therefore, our method efficiently and robustly reduces the need for human annotations. Zixing Zhang 0001, Eduardo Coutinho, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Introducing shared-hidden-layer autoencoders for transfer learning and their application in acoustic emotion recognitionabstractThis study addresses a situation in practice where training and test samples come from different corpora - here in acoustic emotion recognition. In this situation, a model is trained on one database while tested on another disjoint one. The typical inherent mismatch between the corpora and by that between test and training set usually leads to significant performance degradation. To cope with this problem when no training data from the target domain exists, we propose a `shared-hidden-layer autoencoder' (SHLA) approach for learning common feature representations shared across the training and test set in order to reduce the discrepancy in them. To exemplify effectiveness of our approach, we select the Interspeech Emotion Challenge's FAU Aibo Emotion Corpus as test database and two other publicly available databases as training set for extensive evaluation. The experimental results show that our SHLA method significantly improves over the baseline performance and outperforms today's state-of-the-art domain adaptation methods. Zixing Zhang 0001, Yang Liu 0004, Björn W. Schuller |
ICASSP | 3 |
| 2014 | Linked Source and Target Domain Subspace Feature Transfer Learning - Exemplified by Speech Emotion RecognitionabstractThe typical inherent mismatch between the test and training corpora and by that between 'target' and 'source' sets usually leads to significant performance downgrades. To cope with this, this study presents a feature transfer learning method using Denoising Auto encoders (DAEs) to build high order subspaces of the source and target corpora, where features in the source domain are transferred to the target domain by an additional neural network. To exemplify effectiveness of our approach, we select the INTERSPEECH Emotion Challenge's FAU Aibo Emotion Corpus as target corpus and further two publicly available databases as source corpora for extensive and reproducible evaluation. The experimental results show that our method significantly improves over the baseline performance and outperforms today's state-of-the-art domain adaptation methods. Zixing Zhang 0001, Björn W. Schuller |
ICPR | 2 |
| 2014 | Robust speech recognition using long short-term memory recurrent neural networks for hybrid acoustic modellingabstractOne method to achieve robust speech recognition in adverse conditions including noise and reverberation is to employ acoustic modelling techniques involving neural networks. Using long short-term memory (LSTM) recurrent neural networks proved to be efficient for this task in a setup for phoneme prediction in a multi-stream GMM-HMM framework. These networks exploit a self-learnt amount of temporal context, which makes them especially suited for a noisy speech recognition task. One shortcoming of this approach is the necessity of a GMM acoustic model in the multi-stream framework. Furthermore, potential modelling power of the network is lost when predicting phonemes, compared to the classical hybrid setup where the network predicts HMM states. In this work, we propose to use LSTM networks in a hybrid HMM setup, in order to overcome these drawbacks. Experiments are performed using the medium-vocabulary recognition track of the 2nd CHiME challenge, containing speech utterances in a reverberant and noisy environment. A comparison of different network topologies for phoneme or state prediction used either in the hybrid or double-stream setup shows that state prediction networks perform better than networks predicting phonemes, leading to stateof-the-art results for this database. Index Terms: acoustic modelling, robust speech recognition, neural networks, long short-term memory Jürgen T. Geiger, Zixing Zhang 0001, Felix Weninger, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 2 |
| 2014 | Autoencoder-based Unsupervised Domain Adaptation for Speech Emotion RecognitionabstractWith the availability of speech data obtained from different devices and varied acquisition conditions, we are often faced with scenarios, where the intrinsic discrepancy between the training and the test data has an adverse impact on affective speech analysis. To address this issue, this letter introduces an Adaptive Denoising Autoencoder based on an unsupervised domain adaptation method, where prior knowledge learned from a target set is used to regularize the training on a source set. Our goal is to achieve a matched feature space representation for the target and source sets while ensuring target domain knowledge transfer. The method has been successfully evaluated on the 2009 INTERSPEECH Emotion Challenge's FAU Aibo Emotion Corpus as target corpus and two other publicly available speech emotion corpora as sources. The experimental results show that our method significantly improves over the baseline performance and outperforms related feature domain adaptation methods. Zixing Zhang 0001, Florian Eyben, Björn W. Schuller |
IEEE Signal Process. Lett. | 2 |
| 2014 | Distributing Recognition in Computational ParalinguisticsabstractIn this paper, we propose and evaluate a distributed system for multiple Computational Paralinguistics tasks in a client-server architecture. The client side deals with feature extraction, compression, and bit-stream formatting, while the server side performs the reverse process, plus model training, and classification. The proposed architecture favors large-scale data collection and continuous model updating, personal information protection, and transmission bandwidth optimization. In order to preliminarily investigate the feasibility and reliability of the proposed system, we focus on the trade-off between transmission bandwidth and recognition accuracy. We conduct large-scale evaluations of some key functions, namely, feature compression/decompression, model training and classification, on five common paralinguistic tasks related to emotion, intoxication, pathology, age and gender. We show that, for most tasks, with compression ratios up to 40 (bandwidth savings up to 97.5 percent), the recognition accuracies are very close to the baselines. Our results encourage future exploitation of the system proposed in this paper, and demonstrate that we are not far from the creation of robust distributed multi-task paralinguistic recognition systems which can be applied to a myriad of everyday life scenarios. Zixing Zhang 0001, Eduardo Coutinho, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 1 |
| 2013 | Sparse Autoencoder-Based Feature Transfer Learning for Speech Emotion RecognitionabstractIn speech emotion recognition, training and test data used for system development usually tend to fit each other perfectly, but further 'similar' data may be available. Transfer learning helps to exploit such similar data for training despite the inherent dissimilarities in order to boost a recogniser's performance. In this context, this paper presents a sparse auto encoder method for feature transfer learning for speech emotion recognition. In our proposed method, a common emotion-specific mapping rule is learnt from a small set of labelled data in a target domain. Then, newly reconstructed data are obtained by applying this rule on the emotion-specific data in a different domain. The experimental results evaluated on six standard databases show that our approach significantly improves the performance relative to learning each source domain independently. Zixing Zhang 0001, Erik Marchi, Björn W. Schuller |
ACII | 2 |
| 2013 | Feature enhancement by bidirectional LSTM networks for conversational speech recognition in highly non-stationary noiseabstractThe recognition of spontaneous speech in highly variable noise is known to be a challenge, especially at low signal-to-noise ratios (SNR). In this paper, we investigate the effect of applying bidirectional Long Short-Term Memory (BLSTM) recurrent neural networks for speech feature enhancement in noisy conditions. BLSTM networks tend to prevail over conventional neural network architectures, whenever the recognition or regression task relies on an intelligent exploitation of temporal context information. We show that BLSTM networks are well-suited for mapping from noisy to clean speech features and that the obtained recognition performance gain is partly complementary to improvements via additional techniques such as speech enhancement by non-negative matrix factorization and probabilistic feature generation by Bottleneck-BLSTM networks. Compared to simple multi-condition training or feature enhancement via standard recurrent neural networks, our BLSTM-based feature enhancement approach leads to remarkable gains in word accuracy in a highly challenging task of recognizing spontaneous speech at SNR levels between -6 and 9 dB. Martin Wöllmer, Zixing Zhang 0001, Felix Weninger, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 2 |
| 2013 | Co-training succeeds in Computational ParalinguisticsabstractData sparsity is one of the major bottlenecks in the field of Computational Paralinguistics. Partially supervised learning approaches can help leverage this problem without the need of cost-intensive human labelling efforts. We thus investigate the feasibility of cotraining for exemplary paralinguistic speech analysis tasks spanning along the time-continuum: from short-term-related emotion to mid-term-related sleepiness and finally to long-term trait of gender. By dividing the acoustic feature space with two views as independent and sufficient as possible, the semi-supervised learning approach of co-training selects instances with high confidence scores in each view, and agglomerates them along with their predictions into initial training sets per iteration. Our experimental results on official Interspeech Computational Paralinguistics Challenge tasks effectively demonstrate co-training's superiority over the baseline formed by single-view self-training, especially for the short- and medium-term tasks emotion and sleepiness recognition. Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 1 |
| 2013 | Active learning by label uncertainty for acoustic emotion recognitionabstractSpeech data is in principle available in large amounts for the training of acoustic emotion recognisers.However, emotional labelling is usually not given and the distribution is heavily unbalanced, as most data is 'rather neutral' than truly 'emotional'.In the 'hay stack' of speech data, Active Learning automatically identifies the 'needles', i.e., the more informative instances to reduce human labelling effort when building a classifier, e.g., for acoustic emotion recognition.The critical issue thus is the determination and quantification of informativeness.To this end, we suggest to exploit the reliability of the usual ambiguity of emotional labels, i.e., we propose a novel approach based on label uncertainty.By building a certainty model and predicting the candidate instances, informativeness is thus based on labeller agreement.In addition, we consider class sparseness.The results of extensive test runs under well standardised conditions show the method's great potential in reducing labelling costs while boosting performance. Zixing Zhang 0001, Erik Marchi, Björn W. Schuller |
INTERSPEECH | 1 |
| 2012 | Automatic recognition of emotion evoked by general sound eventsabstractWithout a doubt there is emotion in sound. So far, however, research efforts have focused on emotion in speech and music despite many applications in emotion-sensitive sound retrieval. This paper is an attempt at automatic emotion recognition of general sounds. We selected sound clips from different areas of the daily human environment and model them using the increasingly popular dimensional approach in the emotional arousal and valence space. To establish a reliable ground truth, we compare mean and median of four annotators with their evaluator weighted estimator. We discuss human labelers' consistency, feature relevance, and automatic regression. Results reach correlation coefficients of .61 (arousal) and .49 (valence). Björn W. Schuller, Simone Hantke, Felix Weninger, Wenjing Han, Zixing Zhang 0001, Shri Narayanan |
ICASSP | 5 |
| 2012 | Semi-supervised learning helps in sound event classificationabstractWe investigate the suitability of semi-supervised learning in sound event classification on a large database of 17 k sound clips. Seven categories are chosen based on the findsounds.com schema: animals, people, nature, vehicles, noisemakers, office, and musical instruments. Our results show that adding unlabelled sound event data to the training set based on sufficient classifier confidence level after its automatic labelling level can significantly enhance classification performance. Furthermore, combined with optimal re-sampling of originally labelled instances and iteratively learning in semi-supervised manner, the expected gain can reach approximately half the one achieved by using the originally manually labelled data. Overall, maximum performance of 71.7% can be reported for the automatic classification of sound in a large-scale archive. Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 1 |
| 2012 | Active Learning by Sparse Instance Tracking and Classifier Confidence in Acoustic Emotion RecognitionabstractData scarcity is an ever crucial problem in the field of acoustic emotion recognition.How to get the most informative data from a huge amount of data by least human work and at the same time to obtain the highest performance is quite important.In this paper, we propose and investigate two active learning strategies in acoustic emotion recognition: Based on sparse instances or based on classifier confidence scores.The first strategy focuses on the problem of unbalanced binary or multiple classes.The latter strategy pays more attention on clearing up the boundary confusion between different classes.Our experimental results show that by using active learning aiming at sparse instances or based on classifier confidence, the amount of transcribed data needed is significantly reduced and the unweigted accuracy boosts greatly as well. Zixing Zhang 0001, Björn W. Schuller |
INTERSPEECH | 1 |
| 2011 | Unsupervised learning in cross-corpus acoustic emotion recognitionabstractOne of the ever-present bottlenecks in Automatic Emotion Recognition is data sparseness. We therefore investigate the suitability of unsupervised learning in cross-corpus acoustic emotion recognition through a large-scale study with six commonly used databases, including acted and natural emotion speech, and covering a variety of application scenarios and acoustic conditions. We show that adding unlabeled emotional speech to agglomerated multi-corpus training sets can enhance recognition performance even in a challenging cross-corpus setting; furthermore, we show that the expected gain by adding unlabeled data on average is approximately half the one achieved by additional manually labeled data in leave-one-corpus-out validation. Zixing Zhang 0001, Felix Weninger, Martin Wöllmer, Björn W. Schuller |
ASRU | 1 |
| 2011 | Using Multiple Databases for Training in Emotion Recognition: To Unite or to Vote?abstractWe present an extensive study on the performance of data agglomeration and decision-level fusion for robust cross-corpus emotion recognition.We compare joint training with multiple databases and late fusion of classifiers trained on single databases, employing six frequently used corpora of natural or elicited emotion, namely ABC, AVIC, DES, eNTERFACE, SAL, VAM, and three classifiers i. e. SVM, Random Forests, Naïve Bayes to best cover for singular effects.On average over classifier and database, data agglomeration and majority voting deliver relative improvements of unweighted accuracy by 9.0 % and 4.8 %, respectively, over single-database cross-corpus classification of arousal, while majority voting performs best for valence recognition. Björn W. Schuller, Zixing Zhang 0001, Felix Weninger, Gerhard Rigoll |
INTERSPEECH | 2 |