Meishu Song

dblp:255/0014 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-0023-415XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 MADUV: The 1st INTERSPEECH Mice Autism Detection via Ultrasound Vocalization Challenge
Zijiang Yang 0007, Meishu Song, Xin Jing 0001, Kun Qian 0003, Bin Hu 0001, Kota Tamada, Toru Takumi, Björn W. Schuller, Yoshiharu Yamamoto
INTERSPEECH2
2023 Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis
abstract
Translating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the Japanese Daily Speech Dataset (JDSD), a large in-the-wild daily speech emotion dataset consisting of 20,827 speech samples from 342 speakers and 54 hours of total duration. The data is annotated on the Depression and Anxiety Mood Scale (DAMS) – 9 self-reported emotions to evaluate mood state including "vigorous", "gloomy", "concerned", "happy", "unpleasant", "anxious", "cheerful", "depressed", and "worried". Our dataset possesses emotional states, activity, and time diversity, making it useful for training models to track daily emotional states for healthcare purposes. We partition our corpus and provide a multi-task benchmark across nine emotions, demonstrating that mental health states can be predicted reliably from self-reports with a Concordance Correlation Coefficient value of .547 on average. We hope that JDSD will become a valuable resource to further the development of daily emotional healthcare tracking.
Meishu Song, Andreas Triantafyllopoulos, Zijiang Yang 0007, Hiroki Takeuchi, Toru Nakamura, Akifumi Kishi, Tetsuro Ishizawa, Kazuhiro Yoshiuchi, Xin Jing 0001, Vincent Karas, Zhonghao Zhao, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
ICASSP1
2022 An Overview & Analysis of Sequence-to-Sequence Emotional Voice Conversion
abstract
Emotional voice conversion (EVC) focuses on converting a speech utterance from a source to a target emotion; it can thus be a key enabling technology for human-computer interaction applications and beyond. However, EVC remains an unsolved research problem with several challenges. In particular, as speech rate and rhythm are two key factors of emotional conversion, models have to generate output sequences of differing length. Sequence-to-sequence modelling is recently emerging as a competitive paradigm for models that can overcome those challenges. In an attempt to stimulate further research in this promising new direction, recent sequence-to-sequence EVC papers were systematically investigated and reviewed from six perspectives: their motivation, training strategies, model architectures, datasets, model inputs, and evaluation methods. This information is organised to provide the research community with an easily digestible overview of the current state-of-the-art. Finally, we discuss existing challenges of sequence-to-sequence EVC.
Zijiang Yang 0007, Xin Jing 0001, Andreas Triantafyllopoulos, Meishu Song, Ilhan Aslan, Björn W. Schuller
INTERSPEECH4
2021 Coughing-Based Recognition of Covid-19 with Spatial Attentive ConvLSTM Recurrent Neural Networks
abstract
The rapid emergence of COVID-19 has become a major public health threat around the world.Although early detection is crucial to reduce its spread, the existing diagnostic methods are still insufficient in bringing the pandemic under control.Thus, more sophisticated systems, able to easily identify the infection from a larger variety of symptoms, such as cough, are urgently needed.Deep learning models can indeed convey numerous signal features relevant to fight against the disease; yet, the performance of state-of-the-art approaches is still severely restricted by the feature information loss typically due to the high number of layers.To mitigate this phenomenon, identifying the most relevant feature areas by drawing into attention mechanisms becomes essential.In this paper, we introduce Spatial Attentive ConvLSTM-RNN (SACRNN), a novel algorithm that is using Convolutional Long-Short Term Memory Recurrent Neural Networks with embedded attention that has the ability to identify the most valuable features.The promising results achieved by the fusion between the proposed model and a conventional Attentive Convolutional Recurrent Neural Network, on the automatic recognition of COVID-19 coughing (73.2 % of Unweighted Average Recall) show the great potential of the presented approach in developing efficient solutions to defeat the pandemic.
Tianhao Yan, Emilia Parada-Cabaleiro, Shuo Liu 0012, Meishu Song, Björn W. Schuller
Interspeech5
2021 Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19
abstract
Computer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose.
Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE Internet Things J.9
2021 Frustration recognition from speech during game interaction using wide residual networks
abstract
Although frustration is a common emotional reaction while playing games, an excessive level of frustration can negatively impact a user's experience, discouraging them from further game interactions. The automatic detection of frustration can enable the development of adaptive systems that can adapt a game to a user's specific needs through real-time difficulty adjustment, thereby optimizing the player's experience and guaranteeing game success. To this end, we present a speech-based approach for the automatic detection of frustration during game interactions, a specific task that remains underexplored in research. The experiments were performed on the Multimodal Game Frustration Database (MGFD), an audiovisual dataset—collected within the Wizard-of-Oz framework—that is specially tailored to investigate verbal and facial expressions of frustration during game interactions. We explored the performance of a variety of acoustic feature sets, including Mel-Spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and the low-dimensional knowledge-based acoustic feature set eGeMAPS. Because of the continual improvements in speech recognition tasks achieved by the use of convolutional neural networks (CNNs), unlike the MGFD baseline, which is based on the Long Short-Term Memory (LSTM) architecture and Support Vector Machine (SVM) classifier—in the present work, we consider typical CNNs, including ResNet, VGG, and AlexNet. Furthermore, given the unresolved debate on the suitability of shallow and deep networks, we also examine the performance of two of the latest deep CNNs: WideResNet and EfficientNet. Our best result, achieved with WideResNet and Mel-Spectrogram features, increases the system performance from 58.8% unweighted average recall (UAR) to 93.1% UAR for speech-based automatic frustration recognition.
Meishu Song, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Zijiang Yang 0007, Shuo Liu 0012, Zhao Ren, Ziping Zhao 0001, Björn W. Schuller
Virtual Real. Intell. Hardw.1
2020 An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety
abstract
The COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease.
Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller
INTERSPEECH3
2020 Adventitious Respiratory Classification Using Attentive Residual Neural Networks
abstract
Every year, respiratory diseases affect millions of people worldwide, becoming one of the main causes of death in nowadays society.Currently, the COVID-19-known as a novel respiratory illness-has triggered a global health crisis, which has been identified as the greatest challenge of our time since the Second World War.COVID-19 and many other respiratory diseases present often common symptoms, which impairs their early diagnosis; thus, restricting their prevention and treatment.In this regard, in order to encourage a faster and more accurate detection of these kinds of diseases, the automatic identification of respiratory illness through the application of machine learning methods is a very promising area of research aimed to support clinicians.With this in mind, we apply attention-based Convolutional Neural Networks for the recognition of adventitious respiratory cycles on the International Conference on Biomedical Health Informatics 2017 challenge database.Experimental results indicate that the architecture of residual networks with attention mechanism achieves a significant improvement w. r. t. the baseline models.
Zijiang Yang 0007, Shuo Liu 0012, Meishu Song, Emilia Parada-Cabaleiro, Björn W. Schuller
INTERSPEECH3
2019 Audiovisual Analysis for Recognising Frustration during Game-Play: Introducing the Multimodal Game Frustration Database
abstract
Automatic recognition of frustration, by analysing facial and vocal expressions, can help user experience designers to identify interaction obstacles. To encourage the development of automated systems such as these, we present a novel audiovisual database: the Multimodal Game Frustration Database (MGFD), consisting of ca. 5 hours of audiovisual data, collected from 67 Chinese students speaking in English. For data collection, we developed ‘Crazy Trophy’, a Wizard-of-Oz voice activated web-game designed with a variety of usability problems and aimed to induce increasing amounts of frustration. We also present a baseline for binary multimodal frustration classification (frustration vs no-frustration). For this, we compare the performance of a conventional method, Support Vector Machine classifier, and a state-of-the-art method utilising Long Short-Term Memory Recurrent Neural Networks (LSTM-RNN), extracting both audio (Mel-frequency Cepstral Coefficients) and video (facial action units) features. Using LSTM-RNN and a feature-based multi-model fusion strategy, the best result acheived for the baseline was 60.3 % UAR. To enable further research in this area, the game (‘Crazy Trophy’), the database (MGFD), and the partitioning considered in the presented baseline, are made accessible to the research community.
Meishu Song, Zijiang Yang 0007, Alice Baird, Emilia Parada-Cabaleiro, Zixing Zhang 0001, Ziping Zhao 0001, Björn W. Schuller
ACII1