EDBT 2026 Demo / reviewers in the wild / expert
Jing Han 0010
dblp:90/2562-10
· DBLP profile ↗
47ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0001-5776-6849ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech CodecsabstractNeural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strategically decouples the quantization of Semantic and Acoustic details. The semantic anchoring is achieved via a lightweight projector that aligns acoustic features with a frozen, large-scale mHuBERT codebook, injecting linguistic priors while guaranteeing full codebook utilization. Sequentially, for acoustic details, a residual activation module with SimVQ enables a single-layer quantizer (acoustic path) to faithfully recover fine-grained information. At just 1.5 kbps, SACodec establishes a new state of the art by excelling in both fidelity and semantics: subjective listening tests confirm that its reconstruction quality is perceptually highly comparable to ground-truth audio, while its tokens demonstrate substantially improved semantic richness in downstream tasks. This work suggests that assigning specialized semantic quantizers to distinct information streams offers an effective path to reconcile the long-standing trade-off between fidelity, semantics, and modeling simplicity in low-bitrate speech tokenization. Zhongren Dong, Jing Han 0010, Xiaojun Mo, Yimin Cao, Zixing Zhang 0001 |
AAAI | 3 |
| 2026 | TTT-Dance: A Diffusion Framework Leveraging Test-Time Training and LayerScale for Adaptive Music-Driven Dance GenerationabstractMusic-driven dance generation requires precise motion modeling tightly aligned with musical rhythms. However, existing methods, while producing perceptually convincing motions, oftenfail to generalize to out-of-distribution music and unseen dance styles. We propose TTT-Dance, a diffusion-based framework that integratesTest-Time Training (TTT)into the denoising network, enabling rapid, sample-specific adaptation during inference toenhance robustness and generalization. Additionally, we incorporateLayerScalemodules into the decoder to stabilize the optimization of long temporal sequences and cross-modal interactions, thereby accelerating convergence and improving motion coherence. Extensive experiments on public benchmarks show that TTT-Dance consistently surpasses stateof- the-art baselines across multiple metrics, including physical plausibility, rhythm alignment, and motion diversity, demonstrating its effectiveness in music-dance generation. Zixing Zhang 0001, Hubo Liao, Jing Han 0010 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Heart Sounds for High Blood Pressure PredictionabstractHypertension, a major risk factor for cardiovascular diseases, often goes undetected due to its asymptomatic nature. This study explores a novel approach to detecting elevated blood pressure using heart sounds, aiming to provide a non-invasive, potentially continuous monitoring solution. We evaluated our approach on a new dataset of 260 participants, employing patient-independent cross-validation to ensure generalisability. Our methodology utilises a convolutional neural network-based Hidden Semi-Markov Model for heart sound segmentation, followed by extraction of hand-crafted amplitude, duration, and frequency features. A random forest model was implemented for the binary classification of hypertension, achieving a promising 70% accuracy with 72% sensitivity. We conducted comprehensive analyses, including auscultation location and feature importance evaluation, and the investigation of the heart rate – blood pressure relationship. Our findings demonstrate the feasibility of this approach, providing a robust foundation for further research and development in this domain. Erika Bondareva, Jing Han 0010, Katarzyna Szczurek, Dawid Szczepanek, Tomasz Jadczyk, Cecilia Mascolo |
ICASSP | 2 |
| 2025 | Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-LabelingabstractThe lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Fréchet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines. Yuanchao Li, Zixing Zhang 0001, Jing Han 0010, Peter Bell 0001, Catherine Lai |
ICASSP | 3 |
| 2025 | Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment AnalysisabstractUnderstanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) Challenge invited participants to tackle two demanding subtasks: (1) extracting a comprehensive sentiment sextuple-including holder, target, aspect, opinion, sentiment, and rationale-from multi-speaker dialogues, and (2) detecting sentiment flipping, which detects dynamic sentiment shifts and their underlying triggers. For Subtask-I, in the present paper, we designed a structured prompting pipeline that guided large language models (LLMs) to sequentially extract sentiment components with refined contextual understanding. For Subtask-II, we further leveraged the complementary strengths of three LLMs through ensembling to robustly identify sentiment transitions and their triggers. Our system achieved a 47.38% average score on Subtask-I and a 74.12% exact match F1 on Subtask-II, showing the effectiveness of step-wise refinement and ensemble strategies in rich, multimodal sentiment analysis tasks. Shihao Gao, Zixing Zhang 0001, Jing Han 0010 |
ACM Multimedia | 6 |
| 2025 | Tyee: A Unified, Modular, and Fully-Integrated Configurable Toolkit for Intelligent Physiological Health Care
Lingyu Shu, Zixing Zhang 0001, Jing Han 0010 |
ACM Multimedia | 4 |
| 2025 | AudioFab: Building A General and Intelligent Audio Factory through Tool LearningabstractCurrently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient framework to unlock their full potential. Existing audio agent frameworks often suffer from complex environment configurations and inefficient tool collaboration. To address these limitations, we introduce AudioFab, an open-source agent framework aimed at establishing an open and intelligent audio-processing ecosystem. Compared to existing solutions, AudioFab's modular design resolves dependency conflicts, simplifying tool integration and extension. It also optimizes tool learning through intelligent selection and few-shot learning, improving efficiency and accuracy in complex audio tasks. Furthermore, AudioFab provides a user-friendly natural language interface tailored for non-expert users. As a foundational framework, AudioFab's core contribution lies in offering a stable and extensible platform for future research and development in audio and multimodal AI. The code is available at https://github.com/SmileHnu/AudioFab. Jing Han 0010, Qianshuai Xue, Huan Zhao 0003, Zixing Zhang 0001 |
ACM Multimedia | 2 |
| 2025 | Reparameterization of Lightweight Transformer for On-Device Speech Emotion RecognitionabstractWith the increasing implementation of machine learning models on the edge or Internet of Things (IoT) devices, deploying advanced models on resource-constrained IoT devices remains challenging. Transformer models, a currently dominant neural architecture, have achieved great success in broad domains but their complexity hinders its deployment on IoT devices with limited computation capability and storage size. Although many model compression approaches have been explored, they often suffer from notorious performance degradation. To address this issue, we introduce a new method, namely Transformer reparameterization, to boost the performance of lightweight Transformer models. It consists of two processes: 1) the high-rank factorization (HRF) process in the training stage and 2) the de-HRF (deHRF) process in the inference stage. In the former process, we insert an additional linear layer before the feed-forward network (FFN) of the lightweight Transformer. It is supposed that the inserted HRF layers can enhance the model learning capability. In the later process, the auxiliary HRF layer will be merged together with the following FFN layer into one linear layer and thus recover the original structure of the lightweight model. To examine the effectiveness of the proposed method, we evaluate it on three widely used Transformer variants, i.e., ConvTransformer, Conformer, and SpeechFormer networks, in the application of speech emotion recognition on the IEMOCAP, M3ED, and DAIC-WOZ datasets. Experimental results show that our proposed method consistently improves the performance of lightweight Transformers, even making them comparable to large models. The proposed reparameterization approach enables advanced Transformer models to be deployed on resource-constrained IoT devices. Zixing Zhang 0001, Zhongren Dong, Jing Han 0010 |
IEEE Internet Things J. | 4 |
| 2024 | Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty EstimationabstractHeart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learning algorithms and the uncertainty of predictions. This paper introduces a heart murmur detection method based on a parallel-attentive model, which consists of two branches: One is based on a self-attention module and the other one is based on a convolutional network. Unlike traditional approaches, this structure is better equipped to handle long-term dependencies in sequential data, and thus effectively captures the local and global features of heart murmurs. Additionally, we acknowledge the significance of understanding the uncertainty of model predictions in the medical field for clinical decision-making. Therefore, we have incorporated an effective uncertainty estimation method based on Monte Carlo Dropout into our model. Furthermore, we have employed temperature scaling to calibrate the predictions of our probabilistic model, enhancing its reliability. In experiments conducted on the CirCor Digiscope dataset for heart murmur detection, our proposed method achieves a weighted accuracy of 79.8 % and an F1 of 65.1 %, representing state-of-the-art results. Zixing Zhang 0001, Jing Han 0010, Björn W. Schuller |
ICASSP | 3 |
| 2024 | HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous SpeechabstractAutomatically detecting Alzheimer’s Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational complexity associated with self-attention and the length of audio poses a challenge when deploying such models on edge devices. In this context, we construct a novel framework, namely Hierarchical Attention-Free Transformer (HAFFormer), to better deal with long speech for AD detection. Specifically, we employ an attention-free module of Multi-Scale Depthwise Convolution to replace the self-attention and thus avoid the expensive computation, and a GELU-based Gated Linear Unit to replace the feedforward layer, aiming to automatically filter out the redundant information. Moreover, we design a hierarchical structure to force it to learn a variety of information grains, from the frame level to the dialogue level. By conducting extensive experiments on the ADReSS-M dataset, the introduced HAFFormer can achieve competitive results (82.6% accuracy) with other recent work, but with significant computational complexity and model size reduction compared to the standard Transformer. This shows the efficiency of HAFFormer in dealing with long audio for AD detection. Zhongren Dong, Zixing Zhang 0001, Jing Han 0010, Jianjun Ou, Björn W. Schuller |
ICASSP | 4 |
| 2024 | Customising General Large Language Models for Specialised Emotion Recognition TasksabstractThe advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning, and others. Leveraging the capability of LLMs inevitably becomes an essential solution for emotion recognition. To this end, we further comprehensively investigate how LLMs perform in linguistic emotion recognition if we concentrate on this specific task. Specifically, we exemplify a publicly available and widely used LLM – Chat General Language Model, and customise it for our target by using two different model adaptation techniques, i.e., deep prompt tuning and low-rank adaptation. The experimental results obtained on six widely used datasets present that the adapted LLM can easily outperform other state-of-the-art but specialised deep models. This indicates the strong transferability and feasibility of LLMs in the field of emotion recognition. Liyizhe Peng, Zixing Zhang 0001, Jing Han 0010, Huan Zhao 0003, Björn W. Schuller |
ICASSP | 4 |
| 2024 | Towards Open Respiratory Acoustic Foundation Models: Pretraining and BenchmarkingabstractRespiratory audio, such as coughing and breathing sounds, has predictive power for a wide range of healthcare applications, yet is currently under-explored. The main problem for those applications arises from the difficulty in collecting large labeled task-specific data for model development. Generalizable respiratory acoustic foundation models pretrained with unlabeled data would offer appealing advantages and possibly unlock this impasse. However, given the safety-critical nature of healthcare applications, it is pivotal to also ensure openness and replicability for any proposed foundation model solution. To this end, we introduce OPERA, an OPEn Respiratory Acoustic foundation model pretraining and benchmarking system, as the first approach answering this need. We curate large-scale respiratory audio datasets ($\sim$136K samples, over 400 hours), pretrain three pioneering foundation models, and build a benchmark consisting of 19 downstream respiratory health tasks for evaluation. Our pretrained models demonstrate superior performance (against existing acoustic models pretrained with general audio on 16 out of 19 tasks) and generalizability (to unseen datasets and new respiratory audio modalities). This highlights the great promise of respiratory acoustic foundation models and encourages more studies using OPERA as an open resource to accelerate research on respiratory audio for health. The system is accessible from https://github.com/evelyn0414/OPERA. Yuwei Zhang 0001, Tong Xia, Jing Han 0010, Yu Wu 0021, Georgios Rizos, Yang Liu 0101, Mohammed Mosuily, Jagmohan Chauhan, Cecilia Mascolo |
NeurIPS | 3 |
| 2024 | Refashioning Emotion Recognition Modeling: The Advent of Generalized Large ModelsabstractAfter its inception, emotion recognition or affective computing has increasingly become an active research topic due to its broad applications. The corresponding computational models have gradually migrated from statistically shallow models to neural-network-based deep models, which can significantly boost the performance of emotion recognition and consistently achieve the best results on different benchmarks, and thus has been considered the first option for emotion recognition. However, the debut of large language models (LLMs), such as ChatGPT and GPT4, has remarkably astonished the world due to their emerged capabilities of zero/few-shot learning, in-context learning (ICL), chain-of-thought, and others that are never shown in previous deep models. In the present article, we comprehensively investigate how the LLMs perform in emotion recognition in terms of diverse aspects, including ICL, few-shot prompting, accuracy, generalization, and explanation. Moreover, we offer some insights and pose other potential challenges, hoping to ignite broader discussions about enhancing emotion recognition in the new era of advanced and more generalized models. Zixing Zhang 0001, Liyizhe Peng, Jing Han 0010, Huan Zhao 0003, Björn W. Schuller |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Uncertainty-Aware Health Diagnostics via Class-Balanced Evidential Deep LearningabstractUncertainty quantification is critical for ensuring the safety of deep learning-enabled health diagnostics, as it helps the model account for unknown factors and reduces the risk of misdiagnosis. However, existing uncertainty quantification studies often overlook the significant issue of class imbalance, which is common in medical data. In this paper, we propose a class-balanced evidential deep learning framework to achieve fair and reliable uncertainty estimates for health diagnostic models. This framework advances the state-of-the-art uncertainty quantification method of evidential deep learning with two novel mechanisms to address the challenges posed by class imbalance. Specifically, we introduce a pooling loss that enables the model to learn less biased evidence among classes and a learnable prior to regularize the posterior distribution that accounts for the quality of uncertainty estimates. Extensive experiments using benchmark data with varying degrees of imbalance and various naturally imbalanced health data demonstrate the effectiveness and superiority of our method. Our work pushes the envelope of uncertainty quantification from theoretical studies to realistic healthcare application scenarios. By enhancing uncertainty estimation for class-imbalanced data, we contribute to the development of more reliable and practical deep learning-enabled health diagnostic systems. Tong Xia, Ting Dang, Jing Han 0010, Lorena Qendro, Cecilia Mascolo |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Cross-Device Federated Learning for Mobile Health Diagnostics: A First Study on COVID-19 DetectionabstractFederated learning (FL) aided health diagnostic models can incorporate data from a large number of personal edge devices (e.g., mobile phones) while keeping the data local to the originating devices, largely ensuring privacy. However, such a cross-device FL approach for health diagnostics still imposes many challenges due to both local data imbalance (as extreme as local data consists of a single disease class) and global data imbalance (the disease prevalence is generally low in a population). Since the federated server has no access to data distribution information, it is not trivial to solve the imbalance issue towards an unbiased model. In this paper, we propose FedLoss, a novel cross-device FL framework for health diagnostics. Here the federated server averages the models trained on edge devices according to the predictive loss on the local data, rather than using only the number of samples as weights. As the predictive loss better quantifies the data distribution at a device, FedLoss alleviates the impact of data imbalance. Through a real-world dataset on respiratory sound and symptom-based COVID-19 detection task, we validate the superiority of FedLoss. It achieves competitive COVID-19 detection performance compared to a centralised model with an AUC-ROC of 79%. It also outperforms the state-of-the-art FL baselines in sensitivity and convergence speed. Our work not only demonstrates the promise of federated COVID-19 detection but also paves the way to a plethora of mobile health model development in a privacy-preserving fashion. Tong Xia, Jing Han 0010, Abhirup Ghosh, Cecilia Mascolo |
ICASSP | 2 |
| 2023 | Conditional Neural ODE Processes for Individual Disease Progression Forecasting: A Case Study on COVID-19abstractTime series forecasting, as one of the fundamental machine learning areas, has attracted tremendous attentions over recent years. The solutions have evolved from statistical machine learning (ML) methods to deep learning techniques. One emerging sub-field of time series forecasting is individual disease progression forecasting, e.g., predicting individuals' disease development over a few days (e.g., deteriorating trends, recovery speed) based on few past observations. Despite the promises in the existing ML techniques, a variety of unique challenges emerge for disease progression forecasting, such as irregularly-sampled time series, data sparsity, and individual heterogeneity in disease progression. To tackle these challenges, we propose novel Conditional Neural Ordinary Differential Equations Processes (CNDPs), and validate it in a COVID-19 disease progression forecasting task using audio data. CNDPs allow for irregularly-sampled time series modelling, enable accurate forecasting with sparse past observations, and achieve individual-level progression forecasting. CNDPs show strong performance with an Unweighted Average Recall (UAR) of 78.1%, outperforming a variety of commonly used Recurrent Neural Networks based models. With the proposed label-enhancing mechanism (i.e., including the initial health status as input) and the customised individual-level loss, CNDPs further boost the performance reaching a UAR of 93.6%. Additional analysis also reveals the model's capability in tracking individual-specific recovery trend, implying the potential usage of the model for remote disease progression monitoring. In general, CNDPs pave new pathways for time series forecasting, and provide considerable advantages for disease progression monitoring. Ting Dang, Jing Han 0010, Tong Xia, Erika Bondareva, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Dimitris Spathis, Pietro Cicuta, Cecilia Mascolo |
KDD | 2 |
| 2022 | Fitbeat: COVID-19 estimation based on wristband heart rate using a contrastive convolutional auto-encoder
Shuo Liu 0012, Jing Han 0010, Estela Laporta Puyal, Spyridon Kontaxis, Shaoxiong Sun, Patrick Locatelli, Judith Dineley, Florian B. Pokorny, Gloria Dalla Costa, Letizia Leocani, Ana Isabel Guerrero, Carlos Nos, Ana Zabalza, Per Soelberg Sørensen, Mathias Buron, Melinda Magyari, Yatharth Ranjan, Zulqarnain Rashid, Pauline Conde, Callum L. Stewart, Amos Folarin, Richard J. B. Dobson, Raquel Bailón, Srinivasan Vairavan, Nicholas Cummins, Vaibhav A. Narayan, Matthew Hotopf, Giancarlo Comi, Björn W. Schuller |
Pattern Recognit. | 2 |
| 2021 | Exploring Automatic COVID-19 Diagnosis via Voice and Symptoms from Crowdsourced DataabstractThe development of fast and accurate screening tools, which could facilitate testing and prevent more costly clinical tests, is key to the current pandemic of COVID-19. In this context, some initial work shows promise in detecting diagnostic signals of COVID-19 from audio sounds. In this paper, we propose a voice-based framework to automatically detect individuals who have tested positive for COVID-19. We evaluate the performance of the proposed framework on a subset of data crowdsourced from our app, containing 828 samples from 343 participants. By combining voice signals and reported symptoms, an AUC of 0.79 has been attained, with a sensitivity of 0.68 and a specificity of 0.82. We hope that this study opens the door to rapid, low-cost, and convenient pre-screening tools to automatically detect the disease. Jing Han 0010, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Cecilia Mascolo |
ICASSP | 1 |
| 2021 | Towards Sonification in Multimodal and User-friendlyExplainable Artificial IntelligenceabstractWe are largely used to hearing explanations. For example, if someone thinks you are sad today, they might reply to your “why?” with “because you were so Hmmmmm-mmm-mmm”. Today’s Artificial Intelligence (AI), however, is – if at all – largely providing explanations of decisions in a visual or textual manner. While such approaches are good for communication via visual media such as in research papers or screens of intelligent devices, they may not always be the best way to explain; especially when the end user is not an expert. In particular, when the AI’s task is about Audio Intelligence, visual explanations appear less intuitive than audible, sonified ones. Sonification has also great potential for explainable AI (XAI) in systems that deal with non-audio data – for example, because it does not require visual contact or active attention of a user. Hence, sonified explanations of AI decisions face a challenging, yet highly promising and pioneering task. That involves incorporating innovative XAI algorithms to allow pointing back at the learning data responsible for decisions made by an AI, and to include decomposition of the data to identify salient aspects. It further aims to identify the components of the preprocessing, feature representation, and learnt attention patterns that are responsible for the decisions. Finally, it targets decision-making at the model-level, to provide a holistic explanation of the chain of processing in typical pattern recognition problems from end-to-end. Sonified AI explanations will need to unite methods for sonification of the identified aspects that benefit decisions, decomposition and recomposition of audio to sonify which parts in the audio were responsible for the decision, and rendering attention patterns and salient feature representations audible. Benchmarking sonified XAI is challenging, as it will require a comparison against a backdrop of existing, state-of-the-art visual and textual alternatives, as well as synergistic complementation of all modalities in user evaluations. Sonified AI explanations will need to target different user groups to allow personalisation of the sonification experience for different user needs, to lead to a major breakthrough in comprehensibility of AI via hearing how decisions are made, hence supporting tomorrow’s humane AI’s trustability. Here, we introduce and motivate the general idea, and provide accompanying considerations including milestones of realisation of sonifed XAI and foreseeable risks. Björn W. Schuller, Tuomas Virtanen, Maria Riveiro 0001, Georgios Rizos, Jing Han 0010, Annamaria Mesaros, Konstantinos Drossos |
ICMI | 5 |
| 2021 | The INTERSPEECH 2021 Computational Paralinguistics Challenge: COVID-19 Cough, COVID-19 Speech, Escalation & PrimatesabstractThe INTERSPEECH 2021 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the COVID-19 Cough and COVID-19 Speech Sub-Challenges, a binary classification on COVID-19 infection has to be made based on coughing sounds and speech; in the Escalation SubChallenge, a three-way assessment of the level of escalation in a dialogue is featured; and in the Primates Sub-Challenge, four species vs background need to be classified. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AuDeep toolkit, and deep feature extraction from pre-trained CNNs using the Deep Spectrum toolkit; in addition, we add deep end-to-end sequential modelling, and partially linguistic analysis. Björn W. Schuller, Anton Batliner, Christian Bergler, Cecilia Mascolo, Jing Han 0010, Iulia Lefter, Heysem Kaya, Shahin Amiriparian, Alice Baird, Lukas Stappen, Sandra Ottl, Maurice Gerczuk, Panagiotis Tzirakis, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Léon J. M. Rothkrantz, Joeri A. Zwerts, Jelle Treep, Casper S. Kaandorp |
Interspeech | 5 |
| 2021 | Uncertainty-Aware COVID-19 Detection from Imbalanced Sound DataabstractRecently, sound-based COVID-19 detection studies have shown great promise to achieve scalable and prompt digital prescreening.However, there are still two unsolved issues hindering the practice.First, collected datasets for model training are often imbalanced, with a considerably smaller proportion of users tested positive, making it harder to learn representative and robust features.Second, deep learning models are generally overconfident in their predictions.Clinically, false predictions aggravate healthcare costs.Estimation of the uncertainty of screening would aid this.To handle these issues, we propose an ensemble framework where multiple deep learning models for sound-based COVID-19 detection are developed from different but balanced subsets from original data.As such, data are utilized more effectively compared to traditional up-sampling and down-sampling approaches: an AUC of 0.74 with a sensitivity of 0.68 and a specificity of 0.69 is achieved.Simultaneously, we estimate uncertainty from the disagreement across multiple models.It is shown that false predictions often yield higher uncertainty, enabling us to suggest the users with certainty higher than a threshold to repeat the audio test on their phones or to take clinical tests if digital diagnosis still fails.This study paves the way for a more robust sound-based COVID-19 automated screening system. Tong Xia, Jing Han 0010, Lorena Qendro, Ting Dang, Cecilia Mascolo |
Interspeech | 2 |
| 2021 | Learning audio sequence representations for acoustic event classification
Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Björn W. Schuller |
Expert Syst. Appl. | 3 |
| 2021 | Internet of emotional people: Towards continual affective computing cross cultures via audiovisual signals
Jing Han 0010, Zixing Zhang 0001, Maja Pantic, Björn W. Schuller |
Future Gener. Comput. Syst. | 1 |
| 2021 | Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19abstractComputer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose. Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE Internet Things J. | 5 |
| 2021 | SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the WildabstractNatural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our life. Accurately annotated real-world data are the crux in devising such systems. However, existing databases usually consider controlled settings, low demographic variability, and a single task. In this paper, we introduce the SEWA database of more than 2,000 minutes of audio-visual data of 398 people coming from six cultures, 50 percent female, and uniformly spanning the age range of 18 to 65 years old. Subjects were recorded in two different contexts: while watching adverts and while discussing adverts in a video chat. The database includes rich annotations of the recordings in terms of facial landmarks, facial action units (FAU), various vocalisations, mirroring, and continuously valued valence, arousal, liking, agreement, and prototypic examples of (dis)liking. This database aims to be an extremely valuable resource for researchers in affective computing and automatic human sensing and is expected to push forward the research in human behaviour analysis, including cultural studies. Along with the database, we provide extensive baseline experiments for automatic FAU detection and automatic valence, arousal, and (dis)liking intensity estimation. Jean Kossaifi, Robert Walecki, Yannis Panagakis, Jie Shen 0008, Maximilian Schmitt, Fabien Ringeval, Jing Han 0010, Vedhas Pandit, Antoine Toisoul, Björn W. Schuller, Kam Star, Elnar Hajiyev, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion EmbeddingsabstractDespite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous presence of all involved multiple modalities. Motivated by this, we propose a novel crossmodal emotion embedding framework called EmoBed, which aims to leverage the knowledge from other auxiliary modalities to improve the performance of an emotion recognition system at hand. The framework generally includes two main learning components, i.e., joint multimodal training and crossmodal training. Both of them tend to explore the underlying semantic emotion information but with a shared recognition network or with a shared emotion embedding space, respectively. In doing this, the enhanced system trained with this approach can efficiently make use of the complementary information from other modalities. Nevertheless, the presence of these auxiliary modalities is not demanded during inference. To empirically investigate the effectiveness and robustness of the proposed framework, we perform extensive experiments on the two benchmark databases RECOLA and OMG-Emotion for the tasks of dimensional emotion regression and categorical emotion classification, respectively. The obtained results show that the proposed framework significantly outperforms related baselines in monomodal inference, and are also competitive or superior to the recently reported systems, which emphasises the importance of the proposed crossmodal learning for emotion recognition. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | CAA-Net: Conditional Atrous CNNs With Attention for Explainable Device-Robust Acoustic Scene ClassificationabstractAcoustic Scene Classification (ASC) aims to classify the environment in which the audio signals are recorded. Recently, Convolutional Neural Networks (CNNs) have been successfully applied to ASC. However, the data distributions of the audio signals recorded with multiple devices are different. There has been little research on the training of robust neural networks on acoustic scene datasets recorded with multiple devices, and on explaining the operation of the internal layers of the neural networks. In this article, we focus on training and explaining device-robust CNNs on multi-device acoustic scene data. We propose conditional atrous CNNs with attention for multi-device ASC. Our proposed system contains an ASC branch and a device classification branch, both modelled by CNNs. We visualise and analyse the intermediate layers of the atrous CNNs. A time-frequency attention mechanism is employed to analyse the contribution of each time-frequency bin of the feature maps in the CNNs. On the Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 ASC dataset, recorded with three devices, our proposed model performs significantly better than CNNs trained on single-device data. Zhao Ren, Qiuqiang Kong, Jing Han 0010, Mark D. Plumbley, Björn W. Schuller |
IEEE Trans. Multim. | 3 |
| 2020 | Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition ModelsabstractThe development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this study, we aim to train deep models to defending against non-targeted white-box adversarial attacks. Adversarial data is first generated from the real data using the fast gradient sign method. Then in the research field of speech emotion recognition, adversarial-based training is employed as a method for protecting against adversarial attack. We then train deep convolutional models with both real and adversarial data, and compare the performances of two adversarial training procedures - namely, vanilla adversarial training, and similarity-based adversarial training. In our experiments, through the use of adversarial data augmentation, both of the considered adversarial training procedures can improve the performance when validated on the real data. Additionally, the similarity-based adversarial training learns a more robust model when working with adversarial data. Finally, the considered VGG-16 model performs the best across all models, for both real and generated data. Zhao Ren, Alice Baird, Jing Han 0010, Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 3 |
| 2020 | An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and AnxietyabstractThe COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease. Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
INTERSPEECH | 1 |
| 2020 | Enhancing Transferability of Black-Box Adversarial Attacks via Lifelong Learning for Speech Emotion Recognition ModelsabstractWell-designed adversarial examples can easily fool deep speech emotion recognition models into misclassifications.The transferability of adversarial attacks is a crucial evaluation indicator when generating adversarial examples to fool a new target model or multiple models.Herein, we propose a method to improve the transferability of black-box adversarial attacks using lifelong learning.First, black-box adversarial examples are generated by an atrous Convolutional Neural Network (CNN) model.This initial model is trained to attack a CNN target model.Then, we adapt the trained atrous CNN attacker to a new CNN target model using lifelong learning.We use this paradigm, as it enables multi-task sequential learning, which saves more memory space than conventional multi-task learning.We verify this property on an emotional speech database, by demonstrating that the updated atrous CNN model can attack all target models which have been learnt, and can better attack a new target model than an attack model trained on one target model only. Zhao Ren, Jing Han 0010, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 2 |
| 2020 | Exploring Automatic Diagnosis of COVID-19 from Crowdsourced Respiratory Sound DataabstractAudio signals generated by the human body (e.g., sighs, breathing, heart, digestion, vibration sounds) have routinely been used by clinicians as indicators to diagnose disease or assess disease progression. Until recently, such signals were usually collected through manual auscultation at scheduled visits. Research has now started to use digital technology to gather bodily sounds (e.g., from digital stethoscopes) for cardiovascular or respiratory examination, which could then be used for automatic analysis. Some initial work shows promise in detecting diagnostic signals of COVID-19 from voice and coughs. In this paper we describe our data analysis over a large-scale crowdsourced dataset of respiratory sounds collected to aid diagnosis of COVID-19. We use coughs and breathing to understand how discernible COVID-19 sounds are from those in asthma or healthy controls. Our results show that even a simple binary machine learning classifier is able to classify correctly healthy and COVID-19 sounds. We also show how we distinguish a user who tested positive for COVID-19 and has a cough from a healthy user with a cough, and users who tested positive for COVID-19 and have a cough from users with asthma and a cough. Our models achieve an AUC of above 80% across all tasks. These results are preliminary and only scratch the surface of the potential of this type of data and audio-based machine learning. This work opens the door to further investigation of how automatically analysed respiratory patterns could be used as pre-screening signals to aid COVID-19 diagnosis. Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Jing Han 0010, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Cecilia Mascolo |
KDD | 4 |
| 2020 | Snore-GANs: Improving Automatic Snore Sound Classification With Synthesized DataabstractOne of the frontier issues that severely hamper the development of automatic snore sound classification (ASSC) associates to the lack of sufficient supervised training data. To cope with this problem, we propose a novel data augmentation approach based on semi-supervised conditional generative adversarial networks (scGANs), which aims to automatically learn a mapping strategy from a random noise space to original data distribution. The proposed approach has the capability of well synthesizing "realistic" high-dimensional data, while requiring no additional annotation process. To handle the mode collapse problem of GANs, we further introduce an ensemble strategy to enhance the diversity of the generated data. The systematic experiments conducted on a widely used Munich-Passau snore sound corpus demonstrate that the scGANs-based systems can remarkably outperform other classic data augmentation systems, and are also competitive to other recently reported systems for ASSC. Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Christoph Janott, Yanan Guo 0001, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono ModalityabstractDespite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the training procedure for either speech or facial emotion recognition. Specifically, the model consists of one modality-specific network per individual modality and one shared network to map both audio and visual cues into final predictions. In the training process, we additionally take the loss from one auxiliary modality into account besides the main modality. To evaluate the effectiveness of the implicit fusion model, we conduct extensive experiments for mono-modal emotion classification and regression, and find that the implicit fusion models outperform the standard mono-modal training process. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Björn W. Schuller |
ICASSP | 1 |
| 2019 | Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic ScenesabstractThe goal of Acoustic Scene Classification (ASC) is to recognise the environment in which an audio waveform has been recorded. Recently, deep neural networks have been applied to ASC and have achieved state-of-the-art performance. However, few works have investigated how to visualise and understand what a neural network has learnt from acoustic scenes. Previous work applied local pooling after each convolutional layer, therefore reduced the size of the feature maps. In this paper, we suggest that local pooling is not necessary, but the size of the receptive field is important. We apply atrous Convolutional Neural Networks (CNNs) with global attention pooling as the classification model. The internal feature maps of the attention model can be visualised and explained. On the Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 dataset, our proposed method achieves an accuracy of 72.7 %, significantly outperforming the CNNs without dilation at 60.4 %. Furthermore, our results demonstrate that the learnt feature maps contain rich information on acoustic scenes in the time-frequency domain. Zhao Ren, Qiuqiang Kong, Jing Han 0010, Mark D. Plumbley, Björn W. Schuller |
ICASSP | 3 |
| 2019 | Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion RecognitionabstractDespite the great advances, most of the recently developed automatic speech recognition systems focus on working in a server-client manner, and thus often require a high computational cost, such as the storage size and memory accesses. This, however, does not satisfy the increasing demand for a succinct model that can run smoothly in embedded devices like smartphones. To this end, in this paper we propose a neural network compression method, in the way of quantizing the weights of the neural networks from the original full-precised values into binary values that then can be stored and processed with only one bit per value. In doing this, the traditional neural network-based large-size speech emotion recognition models can be greatly compressed into smaller ones, which demand lower computational cost. To evaluate the feasibility of the proposed approach, we take a state-of-the-art speech emotion recognition model, i. e., convolutional recurrent neural networks, as an example, and conduct experiments on two widely used emotional databases. We find that the proposed binary neural networks are able to yield a remarkable model compression rate but at limited expense of model performance. Huan Zhao 0003, Yufeng Xiao, Jing Han 0010, Zixing Zhang 0001 |
ICASSP | 3 |
| 2019 | Dynamic Difficulty Awareness Training for Continuous Emotion PredictionabstractTime-continuous emotion prediction has become an increasingly compelling task in machine learning. Considerable efforts have been made to advance the performance of these systems. Nonetheless, the main focus has been the development of more sophisticated models and the incorporation of different expressive modalities (e.g., speech, face, and physiology). In this paper, motivated by the benefit of difficulty awareness in a human learning procedure, we propose a novel machine learning framework, namely, dynamic difficulty awareness training (DDAT), which sheds fresh light on the research - directly exploiting the difficulties in learning to boost the machine learning process. The DDAT framework consists of two stages: information retrieval and information exploitation. In the first stage, we make use of the reconstruction error of input features or the annotation uncertainty to estimate the difficulty of learning specific information. The obtained difficulty level is then used in tandem with original features to update the model input in a second learning stage with the expectation that the model can learn to focus on high difficulty regions of the learning process. We perform extensive experiments on a benchmark database REmote COLlaborative and affective to evaluate the effectiveness of the proposed framework. The experimental results show that our approach outperforms related baselines as well as other well-established time-continuous emotion prediction systems, which suggests that dynamically integrating the difficulty information for neural networks can help enhance the learning process. Zixing Zhang 0001, Jing Han 0010, Eduardo Coutinho, Björn W. Schuller |
IEEE Trans. Multim. | 2 |
| 2018 | Towards Conditional Adversarial Training for Predicting Emotions from SpeechabstractMotivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consists of two networks, trained in an adversarial manner: The first network tries to predict emotion from acoustic features, while the second network aims at distinguishing between the predictions provided by the first network and the emotion labels from the database using the acoustic features as conditional information. We evaluate the performance of the proposed conditional adversarial training framework on the widely used emotion database RECOLA. Experimental results show that the proposed training strategy outperforms the conventional training method, and is comparable with, or even superior to other recently reported approaches, including deep and end-to-end learning. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Fabien Ringeval, Björn W. Schuller |
ICASSP | 1 |
| 2018 | Exploring A New Method for Food Likability Rating Based on DT-CWT TheoryabstractIn this paper, we mainly investigate subjects' food likability based on audio-related features as a contribution to EAT ? the ICMI 2018 Eating Analysis and Tracking challenge. Specifically, we conduct 4-level Double Tree Complex Wavelet Transform decomposition of an audio signal, and obtain five sub-audio signals with frequencies ranging from low to high. For each sub-audio signal, not only 'traditional' functional-based features but also deep learning-based features via pretrained CNNs based on SliCQ-nonstationary Gabor transform and a cochleagram map, are calculated. Besides, the original audio signals based Bag-of-Audio-Words features extracted by the openXBOW toolkit are used to enhance the model as well. Finally, the early fusion of all these three kinds of features can lead to promising results, yielding the highest UAR of 79.2 % by means of a leave-one-speaker-out cross-validation, which holds a 12.7 % absolute gain compared with the baseline of 66.5 % UAR. Yanan Guo 0001, Jing Han 0010, Zixing Zhang 0001, Björn W. Schuller, Yide Ma |
ICMI | 2 |
| 2018 | Evolving Learning for Analysing Mood-Related Infant VocalisationabstractInfant vocalisation analysis plays an important role in the study of the development of pre-speech capability of infants, while machine-based approaches nowadays emerge with an aim to advance such an analysis.However, conventional machine learning techniques require heavy feature-engineering and refined architecture designing.In this paper, we present an evolving learning framework to automate the design of neural network structures for infant vocalisation analysis.In contrast to manually searching by trial and error, we aim to automate the search process in a given space with less interference.This framework consists of a controller and its child networks, where the child networks are built according to the controller's estimation.When applying the framework to the Interspeech 2018 Computational Paralinguistics (ComParE) Crying Subchallenge, we discover several deep recurrent neural network structures, which are able to deliver competitive results to the best ComParE baseline method. Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Björn W. Schuller |
INTERSPEECH | 2 |
| 2018 | Bags in Bag: Generating Context-Aware Bags for Tracking Emotions from SpeechabstractInternational audience Jing Han 0010, Zixing Zhang 0001, Maximilian Schmitt, Zhao Ren, Fabien Ringeval, Björn W. Schuller |
INTERSPEECH | 1 |
| 2017 | Reconstruction-error-based learning for continuous emotion recognition in speechabstractTo advance the performance of continuous emotion recognition from speech, we introduce a reconstruction-error-based (RE-based) learning framework with memory-enhanced Recurrent Neural Networks (RNN). In the framework, two successive RNN models are adopted, where the first model is used as an autoencoder for reconstructing the original features, and the second is employed to perform emotion prediction. The RE of the original features is used as a complementary descriptor, which is merged with the original features and fed to the second model. The assumption of this framework is that the system has the ability to learn its `drawback' which is expressed by the RE. Experimental results on the RECOLA database show that the proposed framework significantly outperforms the baseline systems without any RE information in terms of Concordance Correlation Coefficient (.729 vs .710 for arousal, .360 vs .237 for valence), and also significantly overcomes other state-of-the-art methods. Jing Han 0010, Zixing Zhang 0001, Fabien Ringeval, Björn W. Schuller |
ICASSP | 1 |
| 2017 | Prediction-based learning for continuous emotion recognition in speechabstractIn this paper, a prediction-based learning framework is proposed for a continuous prediction task of emotion recognition from speech, which is one of the key components of affective computing in multimedia. The main goal of this framework is to utmost exploit the individual advantages of different regression models cooperatively. To this end, we take two widely used regression models for example, i. e., support vector regression and bidirectional long short-term memory recurrent neural network. We concatenate the two models in a tandem structure by different ways, forming a united cascaded framework. The outputs predicted by the former model are combined together with the original features as the input of the following model for final predictions. The experimental results on a time- and value-continuous spontaneous emotion database (RECOLA) show that, the prediction-based learning framework significantly outperforms the individual models for both arousal and valence dimensions, and provides significantly better results in comparison to other state-of-the-art methodologies on this corpus. Jing Han 0010, Zixing Zhang 0001, Fabien Ringeval, Björn W. Schuller |
ICASSP | 1 |
| 2017 | Towards intoxicated speech recognitionabstractIn a real-life scenario, the acoustic characteristics of speech often suffer from the variations induced by diverse environmental noises and different speakers. To overcome the speaker-related speech variation problem for Automatic Speech Recognition (ASR), many speaker adaptation techniques have been proposed and studied. Almost all of these studies, however, only considered the speakers' long-term traits, such as age, gender, and dialect. Speakers' short-term states, for example, affect and intoxication, are largely ignored. In this study, we address one particular speaker state, alcohol intoxication, which has rarely been studied in the context of ASR. To do this, empirical experiments are performed on a publicly available database used for the INTERSPEECH 2011 Speaker State Challenge, Intoxication Sub-Challenge. The experimental results show that the intoxicated state of the speaker indeed degrades the performance of ASR systems by a large margin for all of the three considered speech styles (spontaneous speech, tongue twisters, command & control). In addition, this paper further shows that multi-condition training can notably improve the acoustic model. Zixing Zhang 0001, Felix Weninger, Martin Wöllmer, Jing Han 0010, Björn W. Schuller |
IJCNN | 4 |
| 2017 | From Hard to Soft: Towards more Human-like Emotion Recognition by Modelling the Perception UncertaintyabstractOver the last decade, automatic emotion recognition has become well established. The gold standard target is thereby usually calculated based on multiple annotations from different raters. All related efforts assume that the emotional state of a human subject can be identified by a 'hard' category or a unique value. This assumption tries to ease the human observer's subjectivity when observing patterns such as the emotional state of others. However, as the number of annotators cannot be infinite, uncertainty remains in the emotion target even if calculated from several, yet few human annotators. The common procedure to use this same emotion target in the learning process thus inevitably introduces noise in terms of an uncertain learning target. In this light, we propose a 'soft' prediction framework to provide a more human-like and comprehensive prediction of emotion. In our novel framework, we provide an additional target to indicate the uncertainty of human perception based on the inter-rater disagreement level, in contrast to the traditional framework which is merely producing one single prediction (category or value). To exploit the dependency between the emotional state and the newly introduced perception uncertainty, we implement a multi-task learning strategy. To evaluate the feasibility and effectiveness of the proposed soft prediction framework, we perform extensive experiments on a time- and value-continuous spontaneous audiovisual emotion database including late fusion results. We show that the soft prediction framework with multi-task learning of the emotional state and its perception uncertainty significantly outperforms the individual tasks in both the arousal and valence dimensions. Jing Han 0010, Zixing Zhang 0001, Maximilian Schmitt, Maja Pantic, Björn W. Schuller |
ACM Multimedia | 1 |
| 2017 | Strength modelling for real-worldautomatic continuous affect recognition from audiovisual signals
Jing Han 0010, Zixing Zhang 0001, Nicholas Cummins, Fabien Ringeval, Björn W. Schuller |
Image Vis. Comput. | 1 |
| 2016 | Cross lingual speech emotion recognition using canonical correlation analysis on principal component subspaceabstractThis paper proposes an analytical approach based on Kernel Canonical Correlation Analysis (KCCA) for domain adaptation. To generate paired instances for KCCA, we mapped source and target data onto both source and target principal components. We performed pair-wise domain adaptation between four emotional speech corpora with different languages (English, German, Italian, and Polish) to validate the approach. We compared our approach with the Shared-Hidden-Layer Auto-Encoder (SHLA) and kernel based principal components. On average, the proposed approach yields higher classification performance. Hesam Sagha, Maryna Gavryukova, Jing Han 0010, Björn W. Schuller |
ICASSP | 4 |
| 2016 | Facing Realism in Spontaneous Emotion Recognition from Speech: Feature Enhancement by Autoencoder with LSTM Neural NetworksabstractInternational audience Zixing Zhang 0001, Fabien Ringeval, Jing Han 0010, Erik Marchi, Björn W. Schuller |
INTERSPEECH | 3 |