Shuo Liu 0012

dblp:07/6773-12 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
14since 2021 · last 2024
0000-0001-8133-8588ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Automatic Bird Sound Source Separation Based on Passive Acoustic Devices in Wild Environment
abstract
The Internet of Things (IoT)-based passive acoustic monitoring (PAM) has shown great potential in large-scale remote bird monitoring. However, field recordings often contain overlapping signals, making precise bird information extraction challenging. To solve this challenge, first, the inter-channel spatial feature is chosen as complementary information to the spectral feature to obtain additional spatial correlations between the sources. Then, an end-to-end model named BACPPNet is built based on Deeplabv3plus and enhanced with the polarized self-attention mechanism to estimate the spectral amplitude mask (SMM) for separating bird vocalizations. Finally, the separated bird vocalizations are recovered from SMMs and the spectrogram of mixed audio using the inverse short Fourier transform (ISTFT). We evaluate our proposed method utilizing the generated mixed dataset. Experiments have shown that our method can separate bird vocalizations from mixed audio with RMSE, SDR, SIR, SAR, and STOI values of 2.82, 10.00dB, 29.90 dB, 11.08 dB, and 0.66, respectively, which are better than existing methods. Furthermore, the average classification accuracy of the separated bird vocalizations drops the least. This indicates that our method outperforms other compared separation methods in bird sound separation and preserves the fidelity of the separated sound sources, which might help us better understand wild bird sound recordings.
Jiangjian Xie, Yuwei Shi, Dongming Ni, Manuel Milling, Shuo Liu 0012, Junguo Zhang, Kun Qian 0003, Björn W. Schuller
IEEE Internet Things J.5
2024 Audio Enhancement for Computer Audition - An Iterative Training Paradigm Using Sample Importance
Manuel Milling, Shuo Liu 0012, Andreas Triantafyllopoulos, Ilhan Aslan, Björn W. Schuller
J. Comput. Sci. Technol.2
2023 COVID-19 Detection from Speech in Noisy Conditions
abstract
We explore the integration of audio enhancement into a speech-based COVID-19 detection system in an attempt to make speech captured in noisy environments from everyday life useful for the detection of the virus. For this purpose, two multi-task learning approaches are exploited to jointly optimise a front-end speech enhancement model and a subsequent COVID-19 detection model. In comparison to several baseline methods, such as noisy data augmentation, cold cascade of speech enhancement, and COVID-19 models, our proposed solutions are able to recover a substantial percentage of the performance reduction caused by real-world noises. Our best-performing model, which is trained using the synthetic data of the DiCOVA speech corpus and AudioSet environmental backgrounds, can achieve an average AUC of 76.87 % on the test data covering a wide range of noise intensities, which is over 10 % better than a COVID-19 model trained with clean audio.
Shuo Liu 0012, Adria Mallol-Ragolta, Björn W. Schuller
ICASSP1
2023 The MASCFLICHT Corpus: Face Mask Type and Coverage Area Recognition from Speech
Adria Mallol-Ragolta, Nils Urbach, Shuo Liu 0012, Anton Batliner, Björn W. Schuller
INTERSPEECH3
2023 An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
abstract
Speech is the fundamental mode of human communication, and its synthesis has long been a core priority in human–computer interaction research. In recent years, machines have managed to master the art of generating speech that is understandable by humans. However, the linguistic content of an utterance encompasses only a part of its meaning. Affect, or expressivity, has the capacity to turn speech into a medium capable of conveying intimate thoughts, feelings, and emotions—aspects that are essential for engaging and naturalistic interpersonal communication. While the goal of imparting expressivity to synthesized utterances has so far remained elusive, following recent advances in text-to-speech synthesis, a paradigm shift is well under way in the fields of affective speech synthesis and conversion as well. Deep learning, as the technology that underlies most of the recent advances in artificial intelligence, is spearheading these efforts. In this overview, we outline ongoing trends and summarize state-of-the-art approaches in an attempt to provide a broad overview of this exciting field.
Andreas Triantafyllopoulos, Björn W. Schuller, Gökçe Iymen, Tevfik Metin Sezgin, Xiangheng He, Zijiang Yang 0007, Panagiotis Tzirakis, Shuo Liu 0012, Silvan Mertes, Elisabeth André, Ruibo Fu, Jianhua Tao 0001
Proc. IEEE8
2022 Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough Sounds
abstract
Covid-19 has caused a huge health crisis worldwide in the past two years. Although an early detection of the virus through nucleic acid screening can considerably reduce its spread, the efficiency of this diagnostic process is limited by its complexity and costs. Hence, an effective and inexpensive way to early detect Covid-19 is still needed. Considering that the cough of an infected person contains a large amount of information, we propose an algorithm for the automatic recognition of Covid-19 from cough signals. Our approach generates static log-Mel spectrograms with deltas and delta-deltas from the cough signal and subsequently extracts feature maps through a Convolutional Neural Network (CNN). Following the advances on transformers in the realm of deep learning, our proposed architecture exploits a novel adaptive position embedding structure which can learn the position information of the features from the CNN output. This make the transformer structure rapidly lock the attention feature location by overlaying with the CNN output, which yields better classification. The efficiency of the proposed architecture is shown by the improvement, w. r. t. the baseline, of our experimental results on the INTERPSEECH 2021 Computational Paralinguistics Challenge CCS (Coughing Sub Challenge) database, which reached 72.6 % UAR (Unweighted Average Recall).
Tianhao Yan, Shuo Liu 0012, Emilia Parada-Cabaleiro, Zhao Ren, Björn W. Schuller
ICASSP3
2022 Fitbeat: COVID-19 estimation based on wristband heart rate using a contrastive convolutional auto-encoder
Shuo Liu 0012, Jing Han 0010, Estela Laporta Puyal, Spyridon Kontaxis, Shaoxiong Sun, Patrick Locatelli, Judith Dineley, Florian B. Pokorny, Gloria Dalla Costa, Letizia Leocani, Ana Isabel Guerrero, Carlos Nos, Ana Zabalza, Per Soelberg Sørensen, Mathias Buron, Melinda Magyari, Yatharth Ranjan, Zulqarnain Rashid, Pauline Conde, Callum L. Stewart, Amos Folarin, Richard J. B. Dobson, Raquel Bailón, Srinivasan Vairavan, Nicholas Cummins, Vaibhav A. Narayan, Matthew Hotopf, Giancarlo Comi, Björn W. Schuller
Pattern Recognit.1
2022 Capturing Time Dynamics From Speech Using Neural Networks for Surgical Mask Detection
abstract
The importance of detecting whether a person wears a face mask while speaking has tremendously increased since the outbreak of SARS-CoV-2 (COVID-19), as wearing a mask can help to reduce the spread of the virus and mitigate the public health crisis. Besides affecting human speech characteristics related to frequency, face masks cause temporal interferences in speech, altering the pace, rhythm, and pronunciation speed. In this regard, this paper presents two effective neural network models to detect surgical masks from audio. The proposed architectures are both based on Convolutional Neural Networks (CNNs), chosen as an optimal approach for the spatial processing of the audio signals. One architecture applies a Long Short-Term Memory (LSTM) network to model the time-dependencies. Through an additional attention mechanism, the LSTM-based architecture enables the extraction of more salient temporal information. The other architecture (named ConvTx) retrieves the relative position of a sequence through the positional encoder of a transformer module. In order to assess to which extent both architectures can complement each other when modelling temporal dynamics, we also explore the combination of LSTM and Transformers in three hybrid models. Finally, we also investigate whether data augmentation techniques, such as, using transitions between audio frames and considering gender-dependent frameworks might impact the performance of the proposed architectures. Our experimental results show that one of the hybrid models achieves the best performance, surpassing existing state-of-the-art results for the task at hand.
Shuo Liu 0012, Adria Mallol-Ragolta, Tianhao Yan, Kun Qian 0003, Emilia Parada-Cabaleiro, Bin Hu 0001, Björn W. Schuller
IEEE J. Biomed. Health Informatics1
2021 A Novel Attention-Based Gated Recurrent Unit and its Efficacy in Speech Emotion Recognition
abstract
Notwithstanding the significant advancements in the field of deep learning, the basic long short-term memory (LSTM) or Gated Recurrent Unit (GRU) units have largely remained unchanged and unexplored. There are several possibilities in advancing the state-of-art by rightly adapting and enhancing the various elements of these units. Activation functions are one such key element. In this work, we explore using diverse activation functions within GRU and bi-directional GRU (BiGRU) cells in the context of speech emotion recognition (SER). We also propose a novel Attention ReLU GRU (AR-GRU) that employs attention-based Rectified Linear Unit (AReLU) activation within GRU and BiGRU cells. We demonstrate the effectiveness of AR-GRU on one exemplary application using the recently proposed network for SER namely Interaction-Aware Attention Network (IAAN). Our proposed method utilising AR-GRU within this network yields significant performance gain and achieves an unweighted accuracy of 68.3% (2% over the baseline) and weighted accuracy of 66.9 % (2.2 % absolute over the baseline) in four class emotion recognition on the IEMOCAP database.
Srividya Tirunellai Rajamani, Kumar T. Rajamani, Adria Mallol-Ragolta, Shuo Liu 0012, Björn W. Schuller
ICASSP4
2021 Deep speaker conditioning for speech emotion recognition
abstract
In this work, we explore the use of speaker conditioning sub-networks for speaker adaptation in a deep neural network (DNN) based speech emotion recognition (SER) system. We use a ResNet architecture trained on log spectrogram features, and augment this architecture with an auxiliary network providing speaker embeddings, which conditions multiple layers of the primary classification network on a single neutral speech sample of the target speaker. The whole system is trained end-to-end using a standard cross-entropy loss for utterance-level SER. Relative to the same architecture without the auxiliary embedding sub-network, we are able to improve by 8.3% on IEMOCAP, and by 5.0% and 30.9% on the 2-class and 5-class SER tasks on FAU-AIBO, respectively.
Andreas Triantafyllopoulos, Shuo Liu 0012, Björn W. Schuller
ICME2
2021 Coughing-Based Recognition of Covid-19 with Spatial Attentive ConvLSTM Recurrent Neural Networks
abstract
The rapid emergence of COVID-19 has become a major public health threat around the world.Although early detection is crucial to reduce its spread, the existing diagnostic methods are still insufficient in bringing the pandemic under control.Thus, more sophisticated systems, able to easily identify the infection from a larger variety of symptoms, such as cough, are urgently needed.Deep learning models can indeed convey numerous signal features relevant to fight against the disease; yet, the performance of state-of-the-art approaches is still severely restricted by the feature information loss typically due to the high number of layers.To mitigate this phenomenon, identifying the most relevant feature areas by drawing into attention mechanisms becomes essential.In this paper, we introduce Spatial Attentive ConvLSTM-RNN (SACRNN), a novel algorithm that is using Convolutional Long-Short Term Memory Recurrent Neural Networks with embedded attention that has the ability to identify the most valuable features.The promising results achieved by the fusion between the proposed model and a conventional Attentive Convolutional Recurrent Neural Network, on the automatic recognition of COVID-19 coughing (73.2 % of Unweighted Average Recall) show the great potential of the presented approach in developing efficient solutions to defeat the pandemic.
Tianhao Yan, Emilia Parada-Cabaleiro, Shuo Liu 0012, Meishu Song, Björn W. Schuller
Interspeech4
2021 Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19
abstract
Computer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose.
Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE Internet Things J.12
2021 N-HANS: A neural network-based toolkit for in-the-wild audio enhancement
abstract
Abstract The unprecedented growth of noise pollution over the last decades has raised an always increasing need for developing efficient audio enhancement technologies. Yet, the variety of difficulties related to processing audio sources in-the-wild, such as handling unseen noises or suppressing specific interferences, makes audio enhancement a still open challenge. In this regard, we present (the Neuro-Holistic Audio-eNhancement System), a Python toolkit for in-the-wild audio enhancement that includes functionalities for audio denoising, source separation, and —for the first time in such a toolkit—selective noise suppression. The architecture is specially developed to automatically adapt to different environmental backgrounds and speakers. This is achieved by the use of two identical neural networks comprised of stacks of residual blocks, each conditioned on additional speech- and noise-based recordings through auxiliary sub-networks. Along to a Python API, a command line interface is provided to researchers and developers, both of them carefully documented. Experimental results indicate that achieves great performance w. r. t. existing methods, preserving also the audio quality at a high level; thus, ensuring a reliable usage in real-life application, e. g., for in-the-wild speech processing, which encourages the development of speech-based intelligent technology.
Shuo Liu 0012, Gil Keren, Emilia Parada-Cabaleiro, Björn W. Schuller
Multim. Tools Appl.1
2021 Frustration recognition from speech during game interaction using wide residual networks
abstract
Although frustration is a common emotional reaction while playing games, an excessive level of frustration can negatively impact a user's experience, discouraging them from further game interactions. The automatic detection of frustration can enable the development of adaptive systems that can adapt a game to a user's specific needs through real-time difficulty adjustment, thereby optimizing the player's experience and guaranteeing game success. To this end, we present a speech-based approach for the automatic detection of frustration during game interactions, a specific task that remains underexplored in research. The experiments were performed on the Multimodal Game Frustration Database (MGFD), an audiovisual dataset—collected within the Wizard-of-Oz framework—that is specially tailored to investigate verbal and facial expressions of frustration during game interactions. We explored the performance of a variety of acoustic feature sets, including Mel-Spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and the low-dimensional knowledge-based acoustic feature set eGeMAPS. Because of the continual improvements in speech recognition tasks achieved by the use of convolutional neural networks (CNNs), unlike the MGFD baseline, which is based on the Long Short-Term Memory (LSTM) architecture and Support Vector Machine (SVM) classifier—in the present work, we consider typical CNNs, including ResNet, VGG, and AlexNet. Furthermore, given the unresolved debate on the suitability of shallow and deep networks, we also examine the performance of two of the latest deep CNNs: WideResNet and EfficientNet. Our best result, achieved with WideResNet and Mel-Spectrogram features, increases the system performance from 58.8% unweighted average recall (UAR) to 93.1% UAR for speech-based automatic frustration recognition.
Meishu Song, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Zijiang Yang 0007, Shuo Liu 0012, Zhao Ren, Ziping Zhao 0001, Björn W. Schuller
Virtual Real. Intell. Hardw.5
2020 A Curriculum Learning Approach for Pain Intensity Recognition from Facial Expressions
abstract
The high prevalence of chronic pain in society raises the need to develop new digital tools that can automatically and objectively assess pain intensity in individuals. These tools can contribute to an optimisation of clinical resources, as they offer cost-effective solutions for early detection, continuous monitoring, and treatment personalisation by utilising Artificial Intelligence techniques. In this work, we present our contribution to the Pain Intensity Estimation from Facial Expressions task of the EMOPAIN 2020 Challenge. Specifically, we compare the performance of Recurrent Neural Networks trained with standard or Curriculum Learning (CL) approaches to predict the pain intensity level of individuals reported in an 11-point scale from facial expressions. The results obtained using the test partition support the use of CL-based approaches in the automatic prediction of pain from facial features. The best model trained using a CL approach achieved a Concordance Correlation Coefficient (CCC) of 0.196 in the test partition, while the model trained using a standard approach, without CL, achieved a CCC of 0.174. In terms of CCC, these results respectively represent an improvement of 0.136 and 0.114 on the best results of the baseline system reported by the Challenge organisers using the test partition.
Adria Mallol-Ragolta, Shuo Liu 0012, Nicholas Cummins, Björn W. Schuller
FG2
2020 Hierarchical Component-attention Based Speaker Turn Embedding for Emotion Recognition
abstract
Traditional discrete-time Speech Emotion Recognition (SER) modelling techniques typically assume that an entire speaker chunk or turn is indicative of its corresponding label. An alternative approach is to assume emotional saliency varies over the course of a speaker turn and use modelling techniques capable of identifying and utilising the most emotionally salient segments, such as those with higher emotional intensity. This strategy has the potential to improve the accuracy of SER systems. Towards this goal, we developed a novel hierarchical recurrent neural network model that produces turn level embeddings for SER. Specifically, we apply two levels of attention to learn to identify salient emotional words in a turn as well as the more informative frames within these words. In a set of experiments on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, we demonstrate that component-attention is more effective within our hierarchical framework than both standard soft-attention and conventional local-attention. Our best network, a hierarchical component-attention network with an attention scope of seven, achieved an Unweighted Average Recall (UAR) of 65.0 % and a Weighted Average Recall (WAR) of 66.1 %, outperforming other baseline attention approaches on the IEMOCAP database.
Shuo Liu 0012, Jinlong Jiao, Ziping Zhao 0001, Judith Dineley, Nicholas Cummins, Björn W. Schuller
IJCNN1
2020 An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety
abstract
The COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease.
Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller
INTERSPEECH6
2020 Towards Speech Robustness for Acoustic Scene Classification
abstract
This work discusses the impact of human voice on acoustic scene classification (ASC) systems.Typically, such systems are trained and evaluated on data sets lacking human speech.We show experimentally that the addition of speech can be detrimental to system performance.Furthermore, we propose two alternative solutions to mitigate that effect in the context of deep neural networks (DNNs).We first utilise data augmentation to make the algorithm robust against the presence of human speech in the data.We also introduce a voice-suppression algorithm that removes human speech from audio recordings, and test the DNN classifier on those denoised samples.Experimental results show that both approaches reduce the negative effects of human voice in ASC systems.Compared to using data augmentation, applying voice suppression achieved better classification accuracy and managed to perform more stably for different speech intensity.
Shuo Liu 0012, Andreas Triantafyllopoulos, Zhao Ren, Björn W. Schuller
INTERSPEECH1
2020 Adventitious Respiratory Classification Using Attentive Residual Neural Networks
abstract
Every year, respiratory diseases affect millions of people worldwide, becoming one of the main causes of death in nowadays society.Currently, the COVID-19-known as a novel respiratory illness-has triggered a global health crisis, which has been identified as the greatest challenge of our time since the Second World War.COVID-19 and many other respiratory diseases present often common symptoms, which impairs their early diagnosis; thus, restricting their prevention and treatment.In this regard, in order to encourage a faster and more accurate detection of these kinds of diseases, the automatic identification of respiratory illness through the application of machine learning methods is a very promising area of research aimed to support clinicians.With this in mind, we apply attention-based Convolutional Neural Networks for the recognition of adventitious respiratory cycles on the International Conference on Biomedical Health Informatics 2017 challenge database.Experimental results indicate that the architecture of residual networks with attention mechanism achieves a significant improvement w. r. t. the baseline models.
Zijiang Yang 0007, Shuo Liu 0012, Meishu Song, Emilia Parada-Cabaleiro, Björn W. Schuller
INTERSPEECH2