Yoshiharu Yamamoto

dblp:14/7919 · DBLP profile ↗
← Back
31ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0002-1132-0355ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Computer networks · 4 · 4 since 2021
YearPublicationVenuePosition
2026 BBANet: Bilateral biological auditory-inspired neural network for heart sound classification
Yang Tan 0003, Hanhan Wu, Kun Qian 0003, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
Eng. Appl. Artif. Intell.7
2026 Can Information Representations Inspired by the Human Auditory Perception Benefit Computer Audition-Based Disease Detection? An Interpretable Comparative Study
abstract
Computer audition-based methods have attracted a great deal of attention in the field of disease detection due to their significant advantages, e.g., non-invasive and convenient operation. Among them, the introduction of information representations inspired by human auditory perception, e.g., Mel-frequency transformation, gives it great potential to approach and even exceed the limits of the human auditory system. However, according to previous research, it remains challenging to fairly assess whether information representations inspired by human auditory perception have a significant positive effect on disease detection. Moreover, performance differences among various information representations and their underlying causes are yet to be thoroughly investigated and analyzed. To this end, we propose an interpretable comparative study on information representations inspired by human auditory perception for disease detection. First, the detection accuracy of different information representations are investigated on two sound datasets (a psychological and a physiological disease) based on the classical model and the proposed Temporal-Spatial Multi-Scale Perception Network. Then, the noise robustness of these information representations are compared by introducing Gaussian noise with varying signal-to-noise ratios (SNRs). Finally, by combining the human auditory perception mechanism and explainable AI techniques, we analyze the reasons for performance differences among various information representations from qualitative and quantitative perspectives. Experimental results demonstrate that information representations inspired by human auditory perception can improve the performance of disease detection with statistical significance. Furthermore, Gammatone Frequency Cepstral Coefficients (GFCCs) outperform other information representations by achieving the highest accuracy, particularly under noisy conditions. The interpretable results further reveal the underlying reasons for GFCC's superior performance, highlighting its ability to capture critical auditory features robustly across varying noise levels.These findings emphasize the potential of auditory perception-inspired representations in advancing computer audition-based disease detection systems and provide a solid foundation for future research in this domain.
Yang Tan 0003, Rui Wang 0198, Kun Qian 0003, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE J. Biomed. Health Informatics7
2025 Exploring the Power of Empirical Mode Decomposition for Sensing the Sound of Silence: A Pilot Study on Mice Autism Detection via Ultrasonic Vocalisation
Chenhao Wu 0004, Xiangjun Cai, Tianrui Jia, Yilu Deng, Kun Qian 0003, Björn W. Schuller, Yoshiharu Yamamoto, Jiang Liu 0005
INTERSPEECH8
2025 MADUV: The 1st INTERSPEECH Mice Autism Detection via Ultrasound Vocalization Challenge
Zijiang Yang 0007, Meishu Song, Xin Jing 0001, Kun Qian 0003, Bin Hu 0001, Kota Tamada, Toru Takumi, Björn W. Schuller, Yoshiharu Yamamoto
INTERSPEECH10
2025 Toward Practical Colorectal Cancer Diagnosis: A Bowel-Sound-Based System With Portable Sensor and On-Board Lightweight AI Model
abstract
Colorectal Cancer (CRC) is one of the leading causes of cancer-related deaths worldwide, and early screening plays a crucial role in improving patient outcomes. In this study, we present a novel AI-assisted CRC diagnostic system using Bowel Sound (BS) signals. We first develop two portable BS acquisition devices with distinct form factors for high-fidelity signal capture in both clinical and home-care scenarios. A total of 221 recordings were collected under expert-guided protocol, with 144 CRC recordings and 59 Non-CRC healthy controls using the developed device. To enable low-resource deployment, we design a lightweight deep learning model optimized for real-time, on-board inference. The model incorporates multiple training strategies, including transfer learning on a large-scale public BS dataset, self-supervised temporal feature learning, and a hybrid semi-and weakly-supervised approach that leverages both unlabeled and real-noise data. Furthermore, a Sound Event Detection (SED) attention mechanism and iterative consistency learning are introduced to enhance the model’s sensitivity to BS activity. The proposed model comprises only 264.7 K parameters and 253.2 M Floating-Point Operations (FLOPs), requiring 1.57 MB of RAM and 1.03 MB of FLASH when deployed on microcontroller. It performs inference in approximately 3.4 s with low power consumption, making it well-suited for low-resource environments. Despite its compact design, the model achieves 93.06% classification accuracy, 96.46% sensitivity, and 86.99% specificity for binary-classes in CRC diagnosis. These results demonstrate the system’s potential for accessible and cost-effective CRC screening in community, home, and rural healthcare scenarios.
Fuze Tian, Yang Tan 0003, Enze Li, Jiedong Ma, Jingyu Liu 0002, Kun Qian 0003, Jing Li 0046, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE Internet Things J.11
2025 Enhancing Emotion Regulation in Mental Disorder Treatment: An AIGC-Based Closed-Loop Music Intervention System
abstract
Mental disorders have increased rapidly and have emerged as a serious social health issue in the recent decade. Undoubtedly, the timely treatment of mental disorders is crucial. Emotion regulation has been proven to be an effective method for treating mental disorders. Music therapy as one of the methods that can achieve emotional regulation has gained increasing attention in the field of mental disorder treatment. However, traditional music therapy methods still face some unresolved issues, such as the lack of real-time capability and the inability to form closed-loop systems. With the advancement of artificial intelligence (AI), especially AI-generated content (AIGC), AI-based music therapy holds promise in addressing these issues. In this paper, an AIGC-based closed-loop music intervention system demonstration is proposed to regulate emotions for mental disorder treatment. This system demonstration consists of an emotion recognition model and a music generation model. The emotion recognition model can assess mental states, while the music generation model generates the corresponding emotional music for regulation. The system continuously performs recognition and regulation, thus forming a closed-loop process. In the experiment, we first conduct experiments on both the emotion recognition model and the music generation model to validate the accuracy of the recognition model and the music quality generated by the music generation models. In conclusion, we conducted comprehensive tests on the entire system to verify its feasibility and effectiveness.
Cuiping Zhu, Ruobing Li, Kun Qian 0003, Fuze Tian, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
IEEE Trans. Affect. Comput.9
2025 FedKDC: Consensus-Driven Knowledge Distillation for Personalized Federated Learning in EEG-Based Emotion Recognition
abstract
Federated learning (FL) has gained prominence in electroencephalogram (EEG)-based emotion recognition because of its ability to enable secure collaborative training without centralized data. However, traditional FL faces challenges due to model and data heterogeneity in smart healthcare settings. For example, medical institutions have varying computational resources, which creates a need for personalized local models. Moreover, EEG data from medical institutions typically face data heterogeneity issues stemming from limitations in participant availability, ethical constraints, and cultural differences among subjects, which can slow model convergence and degrade model performance. To address these challenges, we propose FedKDC, a novel FL framework that incorporates clustered knowledge distillation (CKD). This method introduces a consensus-based distributed learning mechanism to facilitate the clustering process. It then enhances the convergence speed through intraclass distillation and reduces the negative impact of heterogeneity through interclass distillation. Additionally, we introduce a DriftGuard mechanism to mitigate client drift, along with an entropy reducer to decrease the entropy of aggregated knowledge. The framework is validated on the SEED, SEED-IV, SEED-FRA, and SEED-GER datasets, demonstrating its effectiveness in scenarios where both the data and the models are heterogeneous. Experimental results show that FedKDC outperforms other FL frameworks in emotion recognition, achieving a maximum average accuracy of 85.2%, and in convergence efficiency, with faster and more stable convergence.
Xihang Qiu, Wanyong Qiu, Ye Zhang 0017, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
IEEE J. Biomed. Health Informatics8
2025 An On-Board Executable Multi-Feature Transfer-Enhanced Fusion Model for Three-Lead EEG Sensor-Assisted Depression Diagnosis
abstract
The development of affective computing and medical electronic technologies has led to the emergence of Artificial Intelligence (AI)-based methods for the early detection of depression. However, previous studies have often overlooked the necessity for the AI-assisted diagnosis system to be wearable and accessible in practical scenarios for depression recognition. In this work, we present an on-board executable multi-feature transfer-enhanced fusion model for our custom-designed wearable three-lead Electroencephalogram (EEG) sensor, based on EEG data collected from 73 depressed patients and 108 healthy controls. Experimental results show that the proposed model exhibits low-computational complexity (65.0 K parameters), promising Floating-Point Operations (FLOPs) performance (25.6 M), real-time processing (1.5 s/execution), and low power consumption (320.8 mW). Furthermore, it requires only 202.0 KB of Random Access Memory (RAM) and 279.6 KB of Read-Only Memory (ROM) when deployed on the EEG sensor. Despite its low computational and spatial complexity, the model achieves a notable classification accuracy of 95.2%, specificity of 94.0%, and sensitivity of 96.9% under independent test conditions. These results underscore the potential of deploying the model on the wearable three-lead EEG sensor for assisting in the diagnosis of depression.
Fuze Tian, Yang Tan 0003, Lixian Zhu, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
IEEE J. Biomed. Health Informatics9
2025 An AI-Assisted All-in-One Integrated Coronary Artery Disease Diagnosis System Using a Portable Heart Sound Sensor With an On-Board Executable Lightweight Model
abstract
Heart sounds play a crucial role in assessing Coronary Artery Disease (CAD). The advancement of Artificial Intelligence (AI) technologies has given rise to Computer Audition (CA)-based methods for CAD detection. However, previous research has focused primarily on analyzing and modeling heart sound data, overlooking practical application scenarios. In this work, we design a pervasive heart sound collection device used for high-quality heart sound data acquisition. Moreover, we introduce an on-board executable lightweight network tailored for the designed portable device, referred to as TYKDModel. Further, heart sound data from 41 CAD patients and 22 non-CAD healthy controls are collected using the developed device. Experimental results show that the TYKDModel exhibits low-computational complexity, with 52.16 K parameters and 5.03 M Floating-Point Operations (FLOPs). When deployed on the board, it requires only 1.10 MB of Random Access Memory (RAM) and 236.27 KB of Read-Only Memory (ROM), and takes around 1.72 seconds to perform a classification. Despite the low computational and spatial complexity, the TYKDModel achieves a notable classification accuracy of 85.2%, specificity of 88.6%, and sensitivity of 82.8% on the board. These results indicate the promising potential of AI-assisted all-in-one integrated system for the diagnosis of heart sound-assisted CAD.
Fuze Tian, Yang Tan 0003, Jingyu Liu 0002, Kun Qian 0003, Yalei Han, Gong Su, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
IEEE Trans. Mob. Comput.12
2024 Clearer Lub-Dub: A Novel Approach in Heart Sound Denoising Based on Transfer Learning
abstract
Cardiovascular diseases (CVDs) constitute the primary cause of human mortality globally in recent decades. To effectively detect CVDs, heart auscultation plays an important role in early diagnosis. With the development of artificial intelligence (AI), many studies have designed varying AI-assisted diagnosis systems helping people discriminate abnormal heart sounds. Yet, a robust system usually requires a noise-less input signal, which is critical as heart sounds are often affected by some unavoidable noise. Therefore, many heart sound classification models use filters or other methods to obtain the clean signals. However, these classic techniques are not adaptable enough to distinguish the meaningful murmurs and real noises. Thus, we propose a novel approach to transfer an audio source separation model to denoise the heart sound. In this paper, we test different denoisers on synthesis heart sound with additive white Gaussian noises. Our method performs well on the noise reduction metrics. Meanwhile, we evaluate the classification performance of each denoiser with some classifiers on the PhysioNet dataset. Experimental results demonstrate that our method can outperform other denoising techniques by achieving the highest unweighted average recall (UAR) at 95.7% with the smallest standard deviation. The results confirm that our method is robust and adaptable in improving audio's denoising.
Jiang Liu 0005, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
HealthCom9
2024 Inspiration of Prototype Knowledge: Introducing a Meta-Learning Approach to Heart Sound Classification
abstract
Cardiovascular diseases (CVDs) stand as the primary reason of fatalities globally, especially in low- and middle-income countries. In recent years, with the leverage of computer audition technologies, the diagnosis of CVDs through heart sounds become a popular topic. Current models and techniques are trained, validated, and tested on the same dataset, which need to be retrained when encountering new data. To make the best use of sparse data, we propose a Prototypical Network framework with heuristic weight for heart sound recognition. After extracting two different features (Mel Spectrogram and Mel Frequency Cepstral Coefficients) and encoding the features, we calculate the distance between two categories (normal and abnormal), then, a heuristic weight is assigned to the distance that makes the blurred boundaries more distinct. By considering the subject independence, the Unweighted Average Recall (UAR) on the PhysioNet/CinC Challenge 2016 is 68.2 % and 67.7 % on two features, respectively. The capability of our model to work on different datasets is proved by a UAR of 66.4 %, which exceeds the baseline UAR of 58.6 % under a single model.
Qingrong Jackie Wu, Mengkai Sun, Boyang Meng, Kun Qian 0003, Bin Hu 0001, Toru Nakamura, Taishin Nomura, Björn W. Schuller, Yoshiharu Yamamoto
HealthCom12
2024 Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation
abstract
Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation. In this work, we propose a novel hybrid network that involves three main branches: the Multi-Layer Perception (MLP) branch, the Convolutional Neural Network (CNN) branch, and a Fusion branch. The MLP and CNN branches aim to learn global and local features, respectively. To fuse these, the fusion branch introduces a novel hierarchical fusion that performs multi-layered fusions that generate high-level representations to enhance segmentation. Our evaluation with two datasets shows strong performance of the proposed method compared to state-of-the-art baselines.
Chengyu Yuan, Hao Xiong 0001, Guoqing Shangguan, Hualei Shen, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto, Shlomo Berkovsky
ICASSP11
2024 E-ODN: An Emotion Open Deep Network for Generalised and Adaptive Speech Emotion Recognition
Liuxian Ma, Ruobing Li, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
INTERSPEECH8
2024 Study Selectively: An Adaptive Knowledge Distillation based on a Voting Network for Heart Sound Classification
abstract
Phonocardiogram classification methods using deep neural networks have been widely applied to the early detection of cardiovascular diseases recently.Despite their excellent recognition rate, the sizeable computational complexity limits their further development.Nowadays, knowledge distillation (KD) is an established paradigm for model compression.While current research on multi-teacher KD has shown potential to impart more comprehensive knowledge to the student than single-teacher KD, this approach is not suitable for all scenarios.This paper proposes a novel KD strategy to realise an adaptive multi-teacher instruction mechanism.We design a teacher selection strategy called voting network to tell the contribution of different teachers on each distillation points, so that the student can choose the useful information and renounce the redundant one.An evaluation demonstrates that our method reaches excellent accuracy (92.8 %) while maintaining a low computational complexity (0.7 M).
Xihang Qiu, Lixian Zhu, Zikai Song, Kun Qian 0003, Ye Zhang 0017, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
INTERSPEECH9
2024 Fed-MStacking: Heterogeneous Federated Learning With Stacking Misaligned Labels for Abnormal Heart Sound Detection
abstract
Ubiquitous sensing has been widely applied in smart healthcare, providing an opportunity for intelligent heart sound auscultation. However, smart devices contain sensitive information, raising user privacy concerns. To this end, federated learning (FL) has been adopted as an effective solution, enabling decentralised learning without data sharing, thus preserving data privacy in the Internet of Health Things (IoHT). Nevertheless, traditional FL requires the same architectural models to be trained across local clients and global servers, leading to a lack of model heterogeneity and client personalisation. For medical institutions with private data clients, this study proposes Fed-MStacking, a heterogeneous FL framework that incorporates a stacking ensemble learning strategy to support clients in building their own models. The secondary objective of this study is to address scenarios involving local clients with data characterised by inconsistent labelling. Specifically, the local client contains only one case type, and the data cannot be shared within or outside the institution. To train a global multi-class classifier, we aggregate missing class information from all clients at each institution and build meta-data, which then participates in FL training via a meta-learner. We apply the proposed framework to a multi-institutional heart sound database. The experiments utilise random forests (RFs), feedforward neural networks (FNNs), and convolutional neural networks (CNNs) as base classifiers. The results show that the heterogeneous stacking of local models performs better compared to homogeneous stacking.
Wanyong Qiu, Yuying Li 0006, Yi Chang 0004, Kun Qian 0003, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE J. Biomed. Health Informatics7
2024 Automated Cough Sound Analysis for Detecting Childhood Pneumonia
abstract
Pneumonia is one of the leading causes of death in children. Prompt diagnosis and treatment can help prevent these deaths, particularly in resource poor regions where deaths due to pneumonia are highest. Clinical symptom-based screening of childhood pneumonia yields excessive false positives, highlighting the necessity for additional rapid diagnostic tests. Cough is a prevalent symptom of acute respiratory illnesses and the sound of a cough can indicate the underlying pathological changes resulting from respiratory infections. In this study, we propose a fully automated approach to evaluate cough sounds to distinguish pneumonia from other acute respiratory diseases in children. The proposed method involves cough sound denoising, cough sound segmentation, and cough sound classification. The denoising algorithm utilizes multi-conditional spectral mapping with a multilayer perceptron network while the segmentation algorithm detects cough sounds directly from the denoised audio waveform. From the segmented cough signal, we extract various handcrafted features and feature embeddings from a pretrained deep learning network. A multilayer perceptron is trained on the combined feature set for detecting pneumonia. The method we propose is evaluated using a dataset comprising cough sounds from 173 children diagnosed with either pneumonia or other acute respiratory diseases. On average, the denoising algorithm improved the signal-to-noise ratio by 44%. Furthermore, a sensitivity and specificity of 91% and 86%, respectively, is achieved in cough segmentation and 82% and 71%, respectively, in detecting childhood pneumonia using cough sounds alone. This demonstrates its potential as a rapid diagnostic tool, such as using smartphone technology.
Roneel V. Sharan, Kun Qian 0003, Yoshiharu Yamamoto
IEEE J. Biomed. Health Informatics3
2023 AMNet: Introducing an Adaptive Mel-Spectrogram End-to-End Neural Network for Heart Sound Classification
abstract
The cardiovascular diseases (CVDs) cause tremendous deaths yearly. The Mel-spectrogram is widely used as a tool to analyse the heart sound, which facilitate a cheap and efficient diagnosis of CVDs. Nevertheless, the amplitude and frequency responses of the Mel filter banks remain constant, limiting its function to frequency selection. We propose an adaptive Melspectrogram end-to-end neural network (AMNet) for a better characterisation and classification of heart sound in the work. The core of the adaptive Mel-spectrograms (AMel) lies in an adaptive Mel filter banks whose frequency characteristics remain the same as the original Mel-spectrogram (OMel) and amplitude is learnt by the backropagation algorithm. The AMNet learns the raw audio representation directly and outputs the classification results. It reaches 43.5% Unweighted Average Recall (UAR) and surpasses the model with the OMel and the baseline by 6% UAR. It is demonstrated that the AMel characterises the heart sound more effectively.
Yang Tan 0003, Kun Qian 0003, Zhihao Bao, Zheyu Cao, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller
HealthCom7
2023 Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis
abstract
Translating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the Japanese Daily Speech Dataset (JDSD), a large in-the-wild daily speech emotion dataset consisting of 20,827 speech samples from 342 speakers and 54 hours of total duration. The data is annotated on the Depression and Anxiety Mood Scale (DAMS) – 9 self-reported emotions to evaluate mood state including "vigorous", "gloomy", "concerned", "happy", "unpleasant", "anxious", "cheerful", "depressed", and "worried". Our dataset possesses emotional states, activity, and time diversity, making it useful for training models to track daily emotional states for healthcare purposes. We partition our corpus and provide a multi-task benchmark across nine emotions, demonstrating that mental health states can be predicted reliably from self-reports with a Concordance Correlation Coefficient value of .547 on average. We hope that JDSD will become a valuable resource to further the development of daily emotional healthcare tracking.
Meishu Song, Andreas Triantafyllopoulos, Zijiang Yang 0007, Hiroki Takeuchi, Toru Nakamura, Akifumi Kishi, Tetsuro Ishizawa, Kazuhiro Yoshiuchi, Xin Jing 0001, Vincent Karas, Zhonghao Zhao, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
ICASSP15
2023 Federated Intelligent Terminals Facilitate Stuttering Monitoring
abstract
Stuttering is a complicated language disorder. The most common form of stuttering is developmental stuttering, which begins in childhood. Early monitoring and intervention are essential for the treatment of children with stuttering. Automatic speech recognition technology has shown its great potential for non-fluent disorder identification, whereas the previous work has not considered the privacy of users’ data. To this end, we propose federated intelligent terminals for automatic monitoring of stuttering speech in different contexts. Experimental results demonstrate that the proposed federated intelligent terminals model can analyze symptoms of stammering speech by taking personal privacy protection into account. Furthermore, the study has explored that the Shapley value approach in the federated learning setting has comparable performance to data-centralised learning.
Yongzi Yu, Wanyong Qiu, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
ICASSP9
2023 Explainable Stuttering Recognition Using Axial Attention
Kaixiang Yuan, Guangzhe Xuan, Yongzi Yu, Hengrui Zhong, Rui Li 0105, Jian Shen 0004, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto
ICIC (3)12
2022 Psychological Field Versus Physiological Field: From Qualitative Analysis to Quantitative Modeling of the Mental Status
abstract
Welcome to the fifth issue of IEEE Transactions on Computational Social Systems (TCSS) in 2022. After the usual introduction of our 24 regular articles, we would like to discuss the topic of “Psychological Field Versus Physiological Field: From Qualitative Analysis to Quantitative Modelling of the Mental Status.”
Bin Hu 0001, Kun Qian 0003, Qunxi Dong, Yuejia Luo, Yoshiharu Yamamoto, Björn W. Schuller
IEEE Trans. Comput. Soc. Syst.5
2022 Learning Multimodal Representations for Drowsiness Detection
abstract
Drowsiness detection is a crucial step for safe driving. A plethora of efforts has been invested on using pervasive sensor data (e.g., video, physiology) empowered by machine learning to build an automatic drowsiness detection system. Nevertheless, most of the existing methods are based on complicated wearables (e.g., electroencephalogram) or computer vision algorithms (e.g., eye state analysis), which makes the relevant systems hardly applicable in the wild. Furthermore, data based on these methods are insufficient in nature due to limited simulation experiments. In this light, we propose a novel and easily implemented method based on full non-invasive multimodal machine learning analysis for the driver drowsiness detection task. The drowsiness level was estimated by self-reported questionnaire in pre-designed protocols. First, we consider involving environmental data (e.g., temperature, humidity, illuminance, and further more), which can be regarded as complementary information for the human activity data recorded via accelerometers or actigraphs. Second, we demonstrate that the models trained by daily life data can still be efficient to make predictions for the subject performing in a simulator, which may benefit the future data collection methods. Finally, we make a comprehensive study on investigating different machine learning methods including classic ‘shallow’ models and recent deep models. Experimental results show that, our proposed methods can reach 64.6% unweighted average recall for drowsiness detection in a subject-independent scenario.
Kun Qian 0003, Tomoya Koike, Toru Nakamura, Björn W. Schuller, Yoshiharu Yamamoto
IEEE Trans. Intell. Transp. Syst.5
2021 Can Appliances Understand the Behavior of Elderly Via Machine Learning? A Feasibility Study
abstract
Over the last half decade, fast development of the Internet of Things and machine learning (ML) made it feasible to leverage the power of artificial intelligence to facilitate a variety of intelligent systems in smart home. Nevertheless, the studies on designing specific computing technologies for helping elderly to enjoy a comfortable, convenient, and independent daily life are extremely limited. On the one hand, there are increasingly growing demands from the ageing society to implement the cutting edge technology enabling a better life quality for the elderly. On the other hand, there is still a lack on fundamental investigations, applicable infrastructures, and advanced data-driven frameworks. To this end, we propose a novel machine framework for analyzing the daily life behavior of elderly-all in this study are living alone-by the data collected from their home appliances, i.e., television and refrigerator. First, the interevent intervals for the use of the appliances collected in one month from 76 elderly are the raw data to describe the behaviors. Then, three ML paradigms are investigated and compared, which include “classic” ML methods and the state-of-the-art deep learning approaches. Finally, we indicate the current findings and limitations in this feasibility study. Experimental results demonstrate that, our proposed method can reach performance peak at an unweighted average recall of 58.7% (chance level: 50.0%) in a subject-independent test for classifying symptom/nonsymptom days.
Kun Qian 0003, Tomoya Koike, Kazuhiro Yoshiuchi, Björn W. Schuller, Yoshiharu Yamamoto
IEEE Internet Things J.5
2021 Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19
abstract
Computer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose.
Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller
IEEE Internet Things J.14
2021 Can Machine Learning Assist Locating the Excitation of Snore Sound? A Review
abstract
In the past three decades, snoring (affecting more than 30 % adults of the UK population) has been increasingly studied in the transdisciplinary research community involving medicine and engineering. Early work demonstrated that, the snore sound can carry important information about the status of the upper airway, which facilitates the development of non-invasive acoustic based approaches for diagnosing and screening of obstructive sleep apnoea and other sleep disorders. Nonetheless, there are more demands from clinical practice on finding methods to localise the snore sound's excitation rather than only detecting sleep disorders. In order to further the relevant studies and attract more attention, we provide a comprehensive review on the state-of-the-art techniques from machine learning to automatically classify snore sounds. First, we introduce the background and definition of the problem. Second, we illustrate the current work in detail and explain potential applications. Finally, we discuss the limitations and challenges in the snore sound classification task. Overall, our review provides a comprehensive guidance for researchers to contribute to this area.
Kun Qian 0003, Christoph Janott, Maximilian Schmitt, Zixing Zhang 0001, Clemens Heiser, Werner Hemmert, Yoshiharu Yamamoto, Björn W. Schuller
IEEE J. Biomed. Health Informatics7
2020 An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and Anxiety
abstract
The COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease.
Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller
INTERSPEECH13
2020 Learning Higher Representations from Pre-Trained Deep Models with Data Augmentation for the COMPARE 2020 Challenge Mask Task
abstract
Human hand-crafted features are always regarded as expensive, time-consuming, and difficult in almost all of the machinelearning-related tasks.First, those well-designed features extremely rely on human expert domain knowledge, which may restrain the collaboration work across fields.Second, the features extracted in such a brute-force scenario may not be easy to be transferred to another task, which means a series of new features should be designed.To this end, we introduce a method based on a transfer learning strategy combined with data augmentation techniques for the COMPARE 2020 Challenge Mask Sub-Challenge.Unlike the previous studies mainly based on pre-trained models by image data, we use a pre-trained model based on large scale audio data, i. e., AudioSet.In addition, the SpecAugment and mixup methods are used to improve the generalisation of the deep models.Experimental results demonstrate that the best-proposed model can significantly (p < .001,by one-tailed z-test) improve the unweighted average recall (UAR) from 71.8 % (baseline) to 76.2 % on the test set.Finally, the best result, i. e., 77.5 % of the UAR on the test set, is achieved by a late fusion of the two best proposed models and the best single model in the baseline.
Tomoya Koike, Kun Qian 0003, Björn W. Schuller, Yoshiharu Yamamoto
INTERSPEECH4
2020 Machine Listening for Heart Status Monitoring: Introducing and Benchmarking HSS - The Heart Sounds Shenzhen Corpus
abstract
Auscultation of the heart is a widely studied technique, which requires precise hearing from practitioners as a means of distinguishing subtle differences in heart-beat rhythm. This technique is popular due to its non-invasive nature, and can be an early diagnosis aid for a range of cardiac conditions. Machine listening approaches can support this process, monitoring continuously and allowing for a representation of both mild and chronic heart conditions. Despite this potential, relevant databases and benchmark studies are scarce. In this paper, we introduce our publicly accessible database, the Heart Sounds Shenzhen Corpus (HSS), which was first released during the recent INTERSPEECH 2018 ComParE Heart Sound sub-challenge. Additionally, we provide a survey of machine learning work in the area of heart sound recognition, as well as a benchmark for HSS utilising standard acoustic features and machine learning models. At best our support vector machine with Log Mel features achieves 49.7% unweighted average recall on a three category task (normal, mild, moderate/severe).
Fengquan Dong, Kun Qian 0003, Zhao Ren, Alice Baird, Zhenyu Dai, Florian Metze, Yoshiharu Yamamoto, Björn W. Schuller
IEEE J. Biomed. Health Informatics9
2018 Data-driven teaching assessment in inquiry-based learning by topic modeling
Hiroyuki Kuromiya, Ichiro Hidaka, Yoshiharu Yamamoto
ICCE3
2016 Multiscale Analysis of Intensive Longitudinal Biomedical Signals and Its Clinical Applications
abstract
Recent advances in wearable and/or biomedical sensing technologies have made it possible to record very long-term, continuous biomedical signals, referred to as biomedical intensive longitudinal data (ILD). To link ILD to clinical applications, such as personalized healthcare and disease prevention, the development of robust and reliable data analysis techniques is considered important. In this review, we introduce multiscale analysis methods for and the applications to two types of intensive longitudinal biomedical signals, heart rate variability (HRV) and spontaneous physical activity (SPA) time series. It has been shown that these ILD have robust characteristics unique to various multiscale complex systems, and some parameters characterizing the multiscale complexity are in fact altered in pathological states, showing potential usability as a new type of ambient diagnostic and/or prognostic tools. For example, parameters characterizing increased intermittency of HRV are found to be potentially useful in detecting abnormality in the state of the autonomic nervous system, in particular the sympathetic hyperactivity, and intermittency parameters of SPA might also be useful in evaluating symptoms of psychiatric patients with depressive as well as manic episodes, all in the daily settings. Therefore, multiscale analysis might be a useful tool to extract information on clinical events occurring at multiple time scales during daily life and the underlying physiological control mechanisms from biomedical ILD.
Toru Nakamura, Ken Kiyono, Herwig Wendt, Patrice Abry, Yoshiharu Yamamoto
Proc. IEEE5
2015 Covariation of Depressive Mood and Spontaneous Physical Activity in Major Depressive Disorder: Toward Continuous Monitoring of Depressive Mood
abstract
The objective evaluation of depressive mood is considered to be useful for the diagnosis and treatment of depressive disorders. Thus, we investigated psychobehavioral correlates, particularly the statistical associations between momentary depressive mood and behavioral dynamics measured objectively, in patients with major depressive disorder (MDD) and healthy subjects. Patients with MDD ( n = 14) and healthy subjects ( n = 43) wore a watch-type computer device and rated their momentary symptoms using ecological momentary assessment. Spontaneous physical activity in daily life, referred to as locomotor activity, was also continuously measured by an activity monitor built into the device. A multilevel modeling approach was used to model the associations between changes in depressive mood scores and the local statistics of locomotor activity simultaneously measured. We further examined the cross validity of such associations across groups. The statistical model established indicated that worsening of the depressive mood was associated with the increased intermittency of locomotor activity, as characterized by a lower mean and higher skewness. The model was cross validated across groups, suggesting that the same psychobehavioral correlates are shared by both healthy subjects and patients, although the latter had significantly higher mean levels of depressive mood scores. Our findings suggest the presence of robust as well as common associations between momentary depressive mood and behavioral dynamics in healthy individuals and patients with depression, which may lead to the continuous monitoring of the pathogenic processes (from healthy states) and pathological states of MDD.
Jinhyuk Kim, Toru Nakamura, Hiroe Kikuchi, Kazuhiro Yoshiuchi, Tsukasa Sasaki, Yoshiharu Yamamoto
IEEE J. Biomed. Health Informatics6