EDBT 2026 Demo / reviewers in the wild / expert
Ziping Zhao 0001
dblp:13/3015-1
· DBLP profile ↗
33ranked-venue papers
15as first author
17since 2021 · last 2025
0000-0002-8719-6389ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 11 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SeeNet: A Soft Emotion Expert and Data Augmentation Method to Enhance Speech Emotion RecognitionabstractSpeech emotion recognition (SER) systems are designed to enable machines to recognize emotional states in human speech during human-computer interactions, enhancing the interactive experience. While considerable progress has been achieved in this field recently, SER systems still encounter challenges related to performance and robustness, primarily stemming from the limited labeled data. To this end, we propose a novel multitask learning framework to learn a distinctive and robust emotional representation by our “Soft Emotion Expert Network (SeeNet)”. SeeNet consists of three components: a pretrained model, an auxiliary task soft emotion expert (SEE) module and an energy-based mixup (EBM) data augmentation module. The pretrained model and EBM module are employed to mitigate the challenges arising from limited labeled data, thereby enhancing the model performance and bolstering robustness. The SEE module as an auxiliary task is designed to assist the main task of SER by enhancing the distinction between samples exhibiting high similarity across categories. This aims to further improve the performance and robustness of the system. Comprehensive experiments on three different settings and multiple datasets are conducted to evaluate the performance and robustness of our proposed method. The experimental results demonstrate that SeeNet surpasses the state-of-the-art (SOTA) methods in both performance and robustness. Yingming Gao, Yuhua Wen, Ziping Zhao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | A Knowledge Distillation-Based Approach to Speech Emotion RecognitionabstractDue to rapid advancements in deep learning, Transformer-based architectures have proven effective in speech emotion recognition (SER), largely due to their ability to model long-term dependencies more effectively than recurrent networks. The current Transformer architecture is not well-suited for SER because its large parameter number demands significant computational resources, making it less feasible in environments with limited resources. Furthermore, its application to SER is limited because human emotions, which are expressed in long segments of continuous speech, are inherently complex and ambiguous. Therefore, designing specialized Transformer models tailored for SER is essential. To address these challenges, we propose a novel knowledge distillation framework that combines meta-knowledge and curriculum-based distillation. Specifically, we fine-tune the teacher model to optimize it for the SER task. For the student model, we embed individual sequence time points into variable tokens, which are used to aggregate the global speech representation. Additionally, we combine supervised contrastive and cross-entropy loss to increase the inter-class distance between learnable features. Finally, we optimize the student model using both meta-knowledge and the curriculum-based distillation framework. Experimental results on two benchmark datasets, IEMOCAP and MELD, demonstrate that our method performs competitively with state-of-the-art approaches in SER. Ziping Zhao 0001, Haishuai Wang, Danushka Bandara, Jianhua Tao 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | Dense Coordinate Channel Attention Network for Depression Level Estimation from Speech
Ziping Zhao 0001, Shizhao Liu, Mingyue Niu, Haishuai Wang, Björn W. Schuller |
ICPR (13) | 1 |
| 2024 | MFDR: Multiple-stage Fusion and Dynamically Refined Network for Multimodal Emotion Recognition
Ziping Zhao 0001, Haishuai Wang, Björn W. Schuller |
INTERSPEECH | 1 |
| 2024 | MTDAN: A Lightweight Multi-Scale Temporal Difference Attention Networks for Automated Video Depression DetectionabstractDeep learning based video depression analysis has been recently an interesting and challenging topic. Most of existing works focus on learning single-scale facial dynamics of participants for depression detection. Besides, they usually adopt expensive deep learning models with high computational complexity, resulting in difficulty in real-time clinical applications. To address these two issues, this work proposes a lightweight Multi-scale Temporal Difference Attention Networks (MTDAN) integrating the temporal difference and attention mechanism to model both short-term and long-term temporal facial behaviors for automated video depression detection. Initially, two simple yet effective sub-branches, i.e., a Short-term Temporal Difference Attention Network (ST-TDAN), and a Long-term Temporal Difference Attention Network (LT-TDAN), are designed to perform individually short-term and long-term depressive behavior modeling. Then, a simple Interactive Multi-head Attention Fusion (IMHAF) strategy is employed for integrating short-term and long-term spatiotemporal features, followed by a linear fully-collected layer for depression score prediction. Experiments on two public AVEC2013 and AVEC2014 datasets show that our proposed method not only achieves highly competitive performance to state-of-the-art methods, but also has much smaller computational complexity than them on video depression detection tasks. Shiqing Zhang, Xingnan Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Mingyue Niu, Ziping Zhao 0001, Jun Yu 0002, Qi Tian 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Attention-Based Temporal Graph Representation Learning for EEG-Based Emotion RecognitionabstractDue to the objectivity of emotional expression in the central nervous system, EEG-based emotion recognition can effectively reflect humans' internal emotional states. In recent years, convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have made significant strides in extracting local features and temporal dependencies from EEG signals. However, CNNs ignore spatial distribution information from EEG electrodes; moreover, RNNs may encounter issues such as exploding/vanishing gradients and high time consumption. To address these limitations, we propose an attention-based temporal graph representation network (ATGRNet) for EEG-based emotion recognition. Firstly, a hierarchical attention mechanism is introduced to integrate feature representations from both frequency bands and channels ordered by priority in EEG signals. Second, a graph convolutional neural network with top-k operation is utilized to capture internal relationships between EEG electrodes under different emotion patterns. Next, a residual-based graph readout mechanism is applied to accumulate the EEG feature node-level representations into graph-level representations. Finally, the obtained graph-level representations are fed into a temporal convolutional network (TCN) to extract the temporal dependencies between EEG frames. We evaluated our proposed ATGRNet on the SEED, DEAP and FACED datasets. The experimental findings show that the proposed ATGRNet surpasses the state-of-the-art graph-based mehtods for EEG-based emotion recognition. Feng Wang 0075, Ziping Zhao 0001, Haishuai Wang, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Hierarchical Network with Decoupled Knowledge Distillation for Speech Emotion RecognitionabstractThe goal of Speech Emotion Recognition (SER) is to enable computers to recognize the emotion category of a given utterance in the same way that humans do. The accuracy of SER is strongly dependent on the validity of the utterance-level representation obtained by the model. Nevertheless, the "dark knowledge" carried by non-target classes is always ignored by previous studies. In this paper, we propose a hierarchical network, called DKDFMH, which employs decoupled knowledge distillation in a deep convolutional neural network with a fused multi-head attention mechanism. Our approach applies logit distillation to obtain higher-level semantic features from different scales of attention sets and delve into the knowledge carried by non-target classes, thus guiding the model to focus more on the differences between sentiment features. To validate the effectiveness of our model, we conducted experiments on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. We achieved competitive performance, with 79.1 % weighted accuracy (WA) and 77.1 % unweighted accuracy (UA). To the best of our knowledge, this is the first time since 2015 that logit distillation has been returned to state-of-the-art status. Ziping Zhao 0001, Haishuai Wang, Björn W. Schuller |
ICASSP | 1 |
| 2023 | SWRR: Feature Map Classifier Based on Sliding Window Attention and High-Response Feature Reuse for Multimodal Emotion Recognition
Ziping Zhao 0001, Haishuai Wang, Björn W. Schuller |
INTERSPEECH | 1 |
| 2023 | Dual Attention and Element Recalibration Networks for Automatic Depression Level PredictionabstractPhysiological studies have identified that facial dynamics can be considered as biomarkers to analyze depression severity. This paper accordingly develops a Dual Attention and Element Recalibration (DAER) network to extract facial changes to predict the depression level. In this model, we propose two blocks: a Dual Attention (DA) block and Element Recalibration (ER) block. The DA block uses the self-attention to investigate the dynamic changes in the representation sequence of a facial video segment. It further examines the influence of feature components of the representation sequence on depression level prediction through bilinear-attention. Moreover, to improve the representation ability of network, the ER block is used to obtain the global information to recalibrate each element of the tensor. Adopting this approach, for the depression level prediction task, we first divide the long-term video into fixed-length segments and use the trained ResNet50 to encode each frame to generate the representation sequences of video segments. Second, the representation sequences are input into DAER network to obtain the depression level scores. Finally, the average of these scores yields the prediction result corresponding to the long-term video. Experiments on publicly available AVEC 2013 and AVEC 2014 depression databases illustrate the effectiveness of our method. Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Automatic Depression Level Assessment from Speech By Long-Term Global Information EmbeddingabstractDepression is a serious mood disorder which brings negative effects on people's social activities. Therefore, growing attention has been paid to automatic depression assessment, especially from speech. However, most of the previous work uses hand-crafted features or deep neural network-based feature extractors to obtain deep features and then feed them into a classifier or a regression, which ignores the temporal relation of these features. To address this issue, this paper proposes a global information embedding (GIE) to make use of the long-term global information of depression and re-weight the LSTM output sequence. The short-term features are then pooled into long-term features by LASSO optimization to further improve the accuracy of depression recognition. Experiments on AVEC 2013 and AVEC 2014 verified the proposed method, and the RMSEs are 9.63 and 9.40, respectively. Ya Li 0001, Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001 |
ICASSP | 3 |
| 2022 | Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional NetworkabstractAutomated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classification tasks, which tend to be smaller and imbalanced. In this paper, we propose to explore the effectiveness of a multi-branch Temporal Convolutional Network (TCN) architecture integrated with Squeeze-and-Excitation Network (SEnet), a system denoted herein as MBTCNSE, for respiratory sound classification. To the best of the authors’ knowledge, this is the first time that such a hybrid architecture has been employed for respiratory sounds classification. Experiments based on the ICBHI challenge respiratory sound dataset demonstrate the effectiveness of our method. Ziping Zhao 0001, Mingyue Niu, Haishuai Wang, Zixing Zhang 0001, Ya Li 0001 |
ICASSP | 1 |
| 2022 | Adversarial Training for Predicting the Trend of the COVID-19 PandemicabstractIt is significant to accurately predict the epidemic trend of COVID-19 due to its detrimental impact on the global health and economy. Although machine learning based approaches have been applied to predict epidemic trend, standard models have shown low accuracy for long-term prediction due to a high level of uncertainty and lack of essential training data. This paper proposes an improved machine learning framework employing Generative Adversarial Network (GAN) and Long Short-Term Memory (LSTM) for adversarial training to forecast the potential threat of COVID-19 in countries where COVID-19 is rapidly spreading. It also investigates the most updated COVID-19 epidemiological data before October 18, 2020 and model the epidemic trend as time series that can be fed into the proposed model for data augmentation and trend prediction of the epidemic. The proposed model is trained to predict daily numbers of cumulative confirmed cases of COVID-19 in Italy, USA, China, Germany, UK, and across the world. Paper further analyzes and suggests which populations are at risk of contracting COVID-19. Haishuai Wang, Ziping Zhao 0001, Zhenyi Jia, Zhenyan Ji, Jun Wu 0007 |
J. Database Manag. | 3 |
| 2022 | Selective Element and Two Orders Vectorization Networks for Automatic Depression Severity Diagnosis via Facial ChangesabstractPhysiological studies have shown that healthy and depressed individuals present different facial changes. Thus, many researchers have attempted to use Convolutional Neural Networks (CNNs) to extract high-level facial dynamic representations for predicting depression severity. However, the max-pooling (or average-pooling) layers in the CNN lead to the loss of subtle depression cues. Without pooling layers, the CNN cannot extract multi-scale information and has difficulties for tensor vectorization. To this end, we propose a Selective Element and Two Orders Vectorization (SE-TOV) network. For the SE-TOV network, an SE block is constructed to adaptively select the effective elements from the tensors obtained by receptive fields of different sizes. Moreover, we propose a TOV block for vectorizing a high-dimensional tensor. On the one hand, TOV block inputs a tensor into the Global Average Pooling layer to obtain the first-order vectorization result. On the other hand, it takes principal components of the correlation matrix of channels in a tensor as the second-order vectorization result. Experimental results on AVEC 2013 (RMSE$=7.42$, MAE$=6.09$) and AVEC 2014 (RMSE$=7.39$, MAE$=5.87$) depression databases illustrate the superiority of our approach over previous works. Mingyue Niu, Ziping Zhao 0001, Jianhua Tao 0001, Ya Li 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Hierarchical Attention-Based Temporal Convolutional Networks for Eeg-Based Emotion RecognitionabstractEEG-based emotion recognition is an effective way to infer the inner emotional state of human beings. Recently, deep learning methods, particularly long short-term memory recurrent neural networks (LSTM-RNNs), have made encouraging progress for in the field of emotion recognition. However, the LSTM-RNNs are time-consuming and have difficulty avoiding the problem of exploding/vanishing gradients when during training. In addition, EEG-based emotion recognition often suffers due to the existence of silent and emotional irrelevant frames from intra-channel. Not all channels carry the same emotional discriminative information. In order to tackle these problems, a hierarchical attention-based temporal convolutional networks (HATCN) for efficient EEG-based emotion recognition is proposed. Firstly, a spectrogram representation is generated from raw EEG signals in each channel to capture their time and frequency information. Secondly, temporal convolutional networks (TCNs) are utilised to automatically learn more robust/intrinsic long-term dynamic characters in emotion response. Next, a hierarchical attention mechanism is investigated that aggregates the emotional information at both the frame and channel level. The experimental results on the DEAP dataset show that our method achieves an average recognition accuracy of 0.716 and an F1-score of 0.642 over four emotional dimensions and outperforms other state-of-the-art methods in a user-independent scenario. Ziping Zhao 0001, Nicholas Cummins, Björn W. Schuller |
ICASSP | 3 |
| 2021 | Combining a parallel 2D CNN with a self-attention Dilated Residual Network for CTC-based discrete speech emotion recognition
Ziping Zhao 0001, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
Neural Networks | 1 |
| 2021 | Frustration recognition from speech during game interaction using wide residual networksabstractAlthough frustration is a common emotional reaction while playing games, an excessive level of frustration can negatively impact a user's experience, discouraging them from further game interactions. The automatic detection of frustration can enable the development of adaptive systems that can adapt a game to a user's specific needs through real-time difficulty adjustment, thereby optimizing the player's experience and guaranteeing game success. To this end, we present a speech-based approach for the automatic detection of frustration during game interactions, a specific task that remains underexplored in research. The experiments were performed on the Multimodal Game Frustration Database (MGFD), an audiovisual dataset—collected within the Wizard-of-Oz framework—that is specially tailored to investigate verbal and facial expressions of frustration during game interactions. We explored the performance of a variety of acoustic feature sets, including Mel-Spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and the low-dimensional knowledge-based acoustic feature set eGeMAPS. Because of the continual improvements in speech recognition tasks achieved by the use of convolutional neural networks (CNNs), unlike the MGFD baseline, which is based on the Long Short-Term Memory (LSTM) architecture and Support Vector Machine (SVM) classifier—in the present work, we consider typical CNNs, including ResNet, VGG, and AlexNet. Furthermore, given the unresolved debate on the suitability of shallow and deep networks, we also examine the performance of two of the latest deep CNNs: WideResNet and EfficientNet. Our best result, achieved with WideResNet and Mel-Spectrogram features, increases the system performance from 58.8% unweighted average recall (UAR) to 93.1% UAR for speech-based automatic frustration recognition. Meishu Song, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Zijiang Yang 0007, Shuo Liu 0012, Zhao Ren, Ziping Zhao 0001, Björn W. Schuller |
Virtual Real. Intell. Hardw. | 7 |
| 2021 | Self-attention transfer networks for speech emotion recognitionabstractA crucial element of human–machine interaction, the automatic detection of emotional states from human speech has long been regarded as a challenging task for machine learning models. One vital challenge in speech emotion recognition (SER) is how to learn robust and discriminative representations from speech. Meanwhile, although machine learning methods have been widely applied in SER research, the inadequate amount of available annotated data has become a bottleneck that impedes the extended application of techniques (e.g., deep neural networks). To address this issue, we present a deep learning method that combines knowledge transfer and self-attention for SER tasks. Here, we apply the log-Mel spectrogram with deltas and delta-deltas as input. Moreover, given that emotions are time-dependent, we apply Temporal Convolutional Neural Networks (TCNs) to model the variations in emotions. We further introduce an attention transfer mechanism, which is based on a self-attention algorithm in order to learn long-term dependencies. The Self-Attention Transfer Network (SATN) in our proposed approach, takes advantage of attention autoencoders to learn attention from a source task, and then from speech recognition, followed by transferring this knowledge into SER. Evaluation built on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) demonstrates the effectiveness of the novel model. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Shihuang Sun, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
Virtual Real. Intell. Hardw. | 1 |
| 2020 | Hierarchical Attention Transfer Networks for Depression Assessment from SpeechabstractA growing area of mental health research is the search for speech-based objective markers for conditions such as depression. However, when combined with machine learning, this search can be challenging due to a limited amount of annotated training data. In this paper, we propose a novel crosstask approach which transfers attention mechanisms from speech recognition to aid depression severity measurement. This transfer is applied in a two-level hierarchical network which mirrors the natural hierarchical structure of speech. Experiments based on the Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) dataset, as used in the 2017 Audio/Visual Emotion Challenge, demonstrate the effectiveness of our Hierarchical Attention Transfer Network. On the development set, the proposed approach achieves a root mean square error (RMSE) of 3.85, and a mean absolute error (MAE) of 2.99, on a Patient Health Questionnaire (PHQ)-8 scale [0], [24], while on the test set, it achieves an RMSE of 5.66 and an MAE of 4.28. To the best of our knowledge, these scores represent the best-known speech-only results to date on this corpus. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Björn W. Schuller |
ICASSP | 1 |
| 2020 | Learning Higher Representations from Bioacoustics: A Sequence-to-Sequence Deep Learning Approach for Bird Sound Classification
Kun Qian 0003, Ziping Zhao 0001 |
ICONIP (5) | 3 |
| 2020 | Exploring Spatial-Temporal Representations for fNIRS-based Intimacy Detection via an Attention-enhanced Cascade Convolutional Recurrent Neural NetworkabstractThe detection of intimacy plays a crucial role in the improvement of intimate relationship, which contributes to promote the family and social harmony. Previous studies have shown that different degrees of intimacy have significant differences in brain imaging. Recently, work has emerged to recognise intimacy automatically by using machine learning techniques. Moreover, considering the temporal dynamic characteristics of intimacy relationship on neural mechanism, how to model spatiotemporal dynamics for intimacy prediction effectively is still a challenge. In this paper, we propose a novel method to explore deep spatial-temporal representations for intimacy prediction by anAttention-enhancedCascadeConvolutionalRecurrentNeuralNetwork(ACCRNN). Given the advantages of time-frequency resolution in complex neuronal activities analysis, this paper utilizesfunctionalnear-infraredspectroscopy(fNIRS) to analyse and infer intimate relationship. We collected fNIRS-based dataset for the analysis of intimate relationship. Forty-two-channel fNIRS signals are recorded from the 44 subjects' prefrontal cortex when they watched a total of 18 photos of lovers, friends and strangers for 30 seconds per photo. The experimental results show that our proposed method outperforms the others in terms of accuracy with the precision of 96.5%. To the best of our knowledge, this is the first time that such a hybrid deep architecture has been employed for fNIRS-based intimacy prediction. Ziping Zhao 0001, Li Gu, Björn W. Schuller |
ICPR | 3 |
| 2020 | Hierarchical Component-attention Based Speaker Turn Embedding for Emotion RecognitionabstractTraditional discrete-time Speech Emotion Recognition (SER) modelling techniques typically assume that an entire speaker chunk or turn is indicative of its corresponding label. An alternative approach is to assume emotional saliency varies over the course of a speaker turn and use modelling techniques capable of identifying and utilising the most emotionally salient segments, such as those with higher emotional intensity. This strategy has the potential to improve the accuracy of SER systems. Towards this goal, we developed a novel hierarchical recurrent neural network model that produces turn level embeddings for SER. Specifically, we apply two levels of attention to learn to identify salient emotional words in a turn as well as the more informative frames within these words. In a set of experiments on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, we demonstrate that component-attention is more effective within our hierarchical framework than both standard soft-attention and conventional local-attention. Our best network, a hierarchical component-attention network with an attention scope of seven, achieved an Unweighted Average Recall (UAR) of 65.0 % and a Weighted Average Recall (WAR) of 66.1 %, outperforming other baseline attention approaches on the IEMOCAP database. Shuo Liu 0012, Jinlong Jiao, Ziping Zhao 0001, Judith Dineley, Nicholas Cummins, Björn W. Schuller |
IJCNN | 3 |
| 2020 | Hybrid Network Feature Extraction for Depression Assessment from SpeechabstractA fast-growing area of mental health research is the search for speech-based objective markers for conditions such as depression.One vital challenge in the development of speech-based depression severity assessment systems is the extraction of depression-relevant features from speech signals.In order to deliver more comprehensive feature representation, we herein explore the benefits of a hybrid network that encodes depressionrelated characteristics in speech for the task of depression severity assessment.The proposed network leverages self-attention networks (SAN) trained on low-level acoustic features and deep convolutional neural networks (DCNN) trained on 3D Log-Mel spectrograms.The feature representations learnt in the SAN and DCNN are concatenated and average pooling is exploited to aggregate complementary segment-level features.Finally, support vector regression is applied to predict a speaker's Beck Depression Inventory-II score.Experiments based on a subset of the Audio-Visual Depressive Language Corpus, as used in the 2013 and 2014 Audio/Visual Emotion Challenges, demonstrate the effectiveness of our proposed hybrid approach. Ziping Zhao 0001, Nicholas Cummins, Bin Liu 0041, Haishuai Wang, Jianhua Tao 0001, Björn W. Schuller |
INTERSPEECH | 1 |
| 2020 | Exploring temporal representations by leveraging attention-based bidirectional LSTM-RNNs for multi-modal emotion recognition
Zhongtian Bao, Linhao Li, Ziping Zhao 0001 |
Inf. Process. Manag. | 4 |
| 2019 | Audiovisual Analysis for Recognising Frustration during Game-Play: Introducing the Multimodal Game Frustration DatabaseabstractAutomatic recognition of frustration, by analysing facial and vocal expressions, can help user experience designers to identify interaction obstacles. To encourage the development of automated systems such as these, we present a novel audiovisual database: the Multimodal Game Frustration Database (MGFD), consisting of ca. 5 hours of audiovisual data, collected from 67 Chinese students speaking in English. For data collection, we developed ‘Crazy Trophy’, a Wizard-of-Oz voice activated web-game designed with a variety of usability problems and aimed to induce increasing amounts of frustration. We also present a baseline for binary multimodal frustration classification (frustration vs no-frustration). For this, we compare the performance of a conventional method, Support Vector Machine classifier, and a state-of-the-art method utilising Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN), extracting both audio (Mel-frequency Cepstral Coefficients) and video (facial action units) features. Using LSTM-RNN and a feature-based multi-model fusion strategy, the best result acheived for the baseline was 60.3 % UAR. To enable further research in this area, the game (‘Crazy Trophy’), the database (MGFD), and the partitioning considered in the presented baseline, are made accessible to the research community. Meishu Song, Zijiang Yang 0007, Alice Baird, Emilia Parada-Cabaleiro, Zixing Zhang 0001, Ziping Zhao 0001, Björn W. Schuller |
ACII | 6 |
| 2019 | Analysing and Inferring of Intimacy Based on fNIRS Signals and Peripheral Physiological SignalsabstractIntimacy refers to a relatively long-lasting affinity relationship between individuals, which involves complex neuronal activities and physiological changes in the body. Recent advancements in the field of neuroimaging have demonstrated that functional near-infrared spectroscopy (fNIRS) has excellent potential for intimate relationship analysis. Signals such as fNIRS and physiological signals are increasingly utilised in this regard due to their consistency and complementarity. In this paper, first, we apply fNIRS and physiological database collected from 26 subjects when viewing lover, friend and stranger pictures to analyse and infer the intimacy. Then, the time domain information from both the fNIRS and physiological signals are utilised to exploit the representation of intimacy by General Linear Model (GLM) and Complex Brain Network Analysis (CBNA) methods. Based on these two methods, the intimacy can be analysed with different brain activation patterns. Finally, different machine learning techniques are utilised to predict the intimate relationship. The results demonstrate that multi-modal features are more efficient for intimacy research. Moreover, the average classification accuracy of ensemble learning is 98.72% whereas for KNN it is 91.03%. Ziping Zhao 0001, Li Gu, Nicholas Cummins, Björn W. Schuller |
IJCNN | 3 |
| 2019 | Speech Augmentation via Speaker-Specific Noise in Unseen EnvironmentabstractSpeech augmentation is a common and effective strategy to avoid overfitting and improve on the robustness of an emotion recognition model.In this paper, we investigate for the first time the intrinsic attributes in a speech signal using the multi-resolution analysis theory and the Hilbert-Huang Spectrum, with the goal of developing a robust speech augmentation approach from raw speech data.Specifically, speech decomposition in a double tree complex wavelet transform domain is realized, to obtain sub-speech signals; then, the Hilbert Spectrum using Hilbert-Huang Transform is calculated for each sub-band to capture the noise content in unseen environments with the voice restriction to 100-4000 Hz; finally, the speechspecific noise that varies with the speaker individual, scenarios, environment, and voice recording equipment, can be reconstructed from the top two high-frequency sub-bands to enhance the raw signal.Our proposed speech augmentation is demonstrated using five robust machine learning architectures based on the RAVDESS database, achieving up to 9.3 % higher accuracy compared to the performance on raw data for an emotion recognition task. Yanan Guo 0001, Ziping Zhao 0001, Yide Ma, Björn W. Schuller |
INTERSPEECH | 2 |
| 2019 | A Hierarchical Attention Network-Based Approach for Depression Detection from Transcribed Clinical InterviewsabstractThe high prevalence of depression in society has given rise to a need for new digital tools that can aid its early detection. Among other effects, depression impacts the use of language. Seeking to exploit this, this work focuses on the detection of depressed and non-depressed individuals through the analysis of linguistic information extracted from transcripts of clinical interviews with a virtual agent. Specifically, we investigated the advantages of employing hierarchical attention-based networks for this task. Using Global Vectors (GloVe) pretrained word embedding models to extract low-level representations of the words, we compared hierarchical local-global attention networks and hierarchical contextual attention networks. We performed our experiments on the Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WoZ) dataset, which contains audio, visual, and linguistic information acquired from participants during a clinical session. Our results using the DAIC-WoZ test set indicate that hierarchical contextual attention networks are the most suitable configuration to detect depression from transcripts. The configuration achieves an Unweighted Average Recall (UAR) of .66 using the test set, surpassing our baseline, a Recurrent Neural Network that does not use attention. Adria Mallol-Ragolta, Ziping Zhao 0001, Lukas Stappen, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 2 |
| 2019 | Attention-Enhanced Connectionist Temporal Classification for Discrete Speech Emotion RecognitionabstractDiscrete speech emotion recognition (SER), the assignment of a single emotion label to an entire speech utterance, is typically performed as a sequence-to-label task.This approach, however, is limited, in that it can result in models that do not capture temporal changes in the speech signal, including those indicative of a particular emotion.One potential solution to overcome this limitation is to model SER as a sequence-to-sequence task instead.In this regard, we have developed an attention-based bidirectional long short-term memory (BLSTM) neural network in combination with a connectionist temporal classification (CTC) objective function (Attention-BLSTM-CTC) for SER.We also assessed the benefits of incorporating two contemporary attention mechanisms, namely component attention and quantum attention, into the CTC framework.To the best of the authors' knowledge, this is the first time that such a hybrid architecture has been employed for SER.We demonstrated the effectiveness of our approach on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and FAU-Aibo Emotion corpora.The experimental results demonstrate that our proposed model outperforms current state-of-the-art approaches. Ziping Zhao 0001, Zhongtian Bao, Zixing Zhang 0001, Nicholas Cummins, Haishuai Wang, Björn W. Schuller |
INTERSPEECH | 1 |
| 2019 | Design and Implementation of Credit Evaluation System for Healthy Aged ServiceabstractAs the problem of global aging intensifies, the quality of credit for healthy aged service has increasingly received more attention. The traditional evaluation of credit quality of healthy aged service still uses the manual entry way to evaluate the service quality, which cannot well deal with huge data in the evaluation for healthy aged service quality. Therefore, in this paper, a credit quality evaluation system is designed and implemented to predict and evaluate user's credit ratings by a standard operation flow for healthy aged service. In this system, we establish a predictive model to automatically evaluate healthy aged service credit quality by machine learning technique. The credit data is quantified and processed for extracting credit ratings-related features. Machine learning technique is utilized to explore the latent relationship between the credit data and its rating. The credit evaluation model for healthy aged service can be built by supervised learning method for predicting user's credit ratings. The design and implementation of this system provides a reasonable solution for auto-evaluation of credit quality of healthy aged service, which can save more manpower and improve service quality in healthy aged service. Yiqin Zhao, Ziping Zhao 0001 |
SERVICES | 5 |
| 2018 | Exploring Spatio-Temporal Representations by Integrating Attention-based Bidirectional-LSTM-RNNs and FCNs for Speech Emotion RecognitionabstractAutomatic emotion recognition from speech, which is an important and challenging task in the field of affective computing, heavily relies on the effectiveness of the speech features for classification. Previous approaches to emotion recognition have mostly focused on the extraction of carefully hand-crafted features. How to model spatio-temporal dynamics for speech emotion recognition effectively is still under active investigation. In this paper, we propose a method to tackle the problem of emotional relevant feature extraction from speech by leveraging Attention-based Bidirectional Long Short-Term Memory Recurrent Neural Networks with fully convolutional networks in order to automatically learn the best spatio-temporal representations of speech signals. The learned high-level features are then fed into a deep neural network (DNN) to predict the final emotion. The experimental results on the Chinese Natural Audio-Visual Emotion Database (CHEAVD) and the Interactive Emotional Dyadic Motion Capture (IEMOCAP) corpora show that our method provides more accurate predictions compared with other existing emotion recognition algorithms. Ziping Zhao 0001, Yu Zheng 0013, Zixing Zhang 0001, Haishuai Wang, Yiqin Zhao |
INTERSPEECH | 1 |
| 2015 | Active learning for the prediction of prosodic phrase boundaries in Chinese speech synthesis systems using conditional random fieldsabstractProsodic structure contributes to speech production and comprehension. One of the crucial problems in achieving natural-sounding synthesized speech is the prediction of appropriate phrase boundaries. Unfortunately, obtaining human annotations of prosodic phrases to train a supervised system can be laborious and costly. Active learning has been proven effective in reducing labeling efforts for supervised learning. This study explores active learning techniques with the objective to reduce the amount of human-annotated data needed to attain a given level of performance. It presents an approach based on active learning to predict the Chinese prosodic phrase boundaries in unrestricted Chinese text. Experiments show that for most of the cases considered, the active selection strategies for labeling the prosodic phrase boundaries are as good as or exceed the performance of random data selection. Ziping Zhao 0001, Xirong Ma |
SNPD | 1 |
| 2013 | Active Learning for Speech Emotion Recognition Using Conditional Random FieldsabstractWith the increasing demand for spoken language interfaces in human-computer interactions, automatic recognition of emotional states from human speeches has become of increasing importance. Unfortunately, obtaining human annotations of emotion corpus to train a supervised system can become a laborious and costly effort. To address this, we explore active learning techniques with the objective of reducing the amount of human-annotated data needed to attain a given level of performance. In this paper we proposed an approach for speech emotion recognition based on Active Conditional Random Fields. Experiments show that for most of the cases considered, active selection strategies when recognizing speech emotion are as good as or exceed the performance of random data selection. Ziping Zhao 0001, Xirong Ma |
SNPD | 1 |
| 2011 | Semi Supervised Learning for Prediction of Prosodic Phrase Boundaries in Chinese TTS Using Conditional Random Fields
Ziping Zhao 0001, Xirong Ma, Weidong Pei |
ISNN (2) | 1 |