VLDB 2026 Research / reviewers in the wild / expert
Kun Qian 0003
dblp:77/2062-3
· DBLP profile ↗
68ranked-venue papers
8as first author
57since 2021 · last 2026
0000-0002-1918-6453ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 3 first-author · 26 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 12 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BBANet: Bilateral biological auditory-inspired neural network for heart sound classification
Yang Tan 0003, Hanhan Wu, Kun Qian 0003, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Multi-scale cross-domain and class-wise kernel discriminative alignment for EEG-based emotion recognition
Chengcheng Zheng, Lixian Zhu, Tianqi Fan, Fuze Tian, Dixin Wang, Kun Qian 0003, Bin Hu 0001 |
Pattern Recognit. | 7 |
| 2026 | Bayesian knowledge-guided confidence aware model based on multimodal pre-trained LLMs for depression detection
Haotian Zhai, Tianrui Jia, Kun Qian 0003, Wen Qi 0005, Bin Hu 0001 |
Pattern Recognit. | 6 |
| 2026 | Hybrid Source Selection Fusion Domain-Invariant Attention for Cross-Subject Emotion RecognitionabstractElectroencephalogram (EEG) has been widely used for emotion recognition due to its portability and high temporal resolution. It makes success in subject-dependent scenario but faces significant challenges in cross-subject emotion recognition because of non-stationarity of EEG and individual differences. Most previous studies treat all individuals as a single source domain for transferring emotional knowledge, which may introduce irrelevant information and lead to negative transfer. Besides, there is a potential risk that some important information of common emotional features might be ignored. To deal with the issues, we propose a framework called hybrid source selection fusion domain-invariant attention (HSSFDA) for cross-subject emotion recognition. First, source domains are selected by leveraging local and global similarity for knowledge transfer. Then, a specialized attention mechanism is employed to focus on important emotional information extracted from the domain-invariant features. Finally, domain-invariant and domain-specific features are fused to enhance emotion recognition performance. To evaluate the proposed method, experiments are conducted on several public datasets including SEED, SEED_IV, DREAMER and DEAP. The results demonstrate that HSSFDA achieves accuracies of 85.07 %, 72.11 %, 62.36 %, 77.17 %, 58.51 %, and 63.55 % on SEED, SEED_IV, valence and arousal of DREAMER, and valence and arousal of DEAP datasets, respectively, demonstrating competitive performance compared to popular and state-of-the-art methods. Furthermore, we apply the HSSFDA to a self-recorded dataset collected by self-developed three-channel device and validate its effectiveness in practical applications. In conclusion, HSSFDA is a feasible method for cross-subject emotion recognition and has the potential to broaden the application of EEG in the field of affective computing. Shuaiyi Xu, Wei Zhang 0386, Lixian Zhu, Fuze Tian, Na Chu, Kun Qian 0003, Xiaowei Li 0005, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 8 |
| 2026 | Discriminative Knowledge Fuzzy Transfer Learning Guided by Resting-State EEG for Cross-Subject Emotion RecognitionabstractCross-subject emotion recognition remains a challenge due to inter-subject variability, which limits the generalized ability of models to unseen subjects. Existing studies commonly rely on tasking-state EEG data from the target subject for adaptation, which requires additional emotion-elicitation experiments and limits practical deployment. Motivated by findings that resting-state EEG can reflect individual-specific neural characteristics, this study proposes a Discriminative Knowledge Fuzzy Transfer Learning Guided by Resting-state EEG (DKFTL-R) for cross-subject emotion recognition without requiring tasking-state EEG data from the target subject. First, resting-state EEG is leveraged to characterize subject-specific neural signatures, by which source-domain selection is informed. Second, an emotion-style projection alignment module is introduced, in which discriminative emotional knowledge and domain-specific style are integrated via adaptive weighting so that a more transferable representation is obtained. Finally, a Takagi-Sugeno-Kang fuzzy classifier is employed to perform fuzzy inference on the transferable representation. Experiments are conducted on DEAP and DENS datasets, where accuracies of 58.79%, 55.89%, 62.91%, and 60.42% are achieved, respectively, demonstrating competitive performance compared with popular and recent baseline methods. To evaluate practical applicability and deployability, the proposed method is conducted on a self-constructed emotion EEG dataset (BHE-EMO), and it achieves 67.00% accuracy for two-class classification and 44.92% for three-class classification tasks, further demonstrating its effectiveness and engineering potential in real-world settings. In conclusion, we propose a new perspective on cross-subject emotion recognition by integrating resting-state EEG information with fuzzy modeling. This study also introduces a new calibration paradigm for affective brain-computer interface systems. Na Chu, Lixian Zhu, Chengcheng Zheng, Dixin Wang, Kun Qian 0003, Xiaowei Li 0005, Bin Hu 0001 |
IEEE Trans. Fuzzy Syst. | 6 |
| 2026 | Can Information Representations Inspired by the Human Auditory Perception Benefit Computer Audition-Based Disease Detection? An Interpretable Comparative StudyabstractComputer audition-based methods have attracted a great deal of attention in the field of disease detection due to their significant advantages, e.g., non-invasive and convenient operation. Among them, the introduction of information representations inspired by human auditory perception, e.g., Mel-frequency transformation, gives it great potential to approach and even exceed the limits of the human auditory system. However, according to previous research, it remains challenging to fairly assess whether information representations inspired by human auditory perception have a significant positive effect on disease detection. Moreover, performance differences among various information representations and their underlying causes are yet to be thoroughly investigated and analyzed. To this end, we propose an interpretable comparative study on information representations inspired by human auditory perception for disease detection. First, the detection accuracy of different information representations are investigated on two sound datasets (a psychological and a physiological disease) based on the classical model and the proposed Temporal-Spatial Multi-Scale Perception Network. Then, the noise robustness of these information representations are compared by introducing Gaussian noise with varying signal-to-noise ratios (SNRs). Finally, by combining the human auditory perception mechanism and explainable AI techniques, we analyze the reasons for performance differences among various information representations from qualitative and quantitative perspectives. Experimental results demonstrate that information representations inspired by human auditory perception can improve the performance of disease detection with statistical significance. Furthermore, Gammatone Frequency Cepstral Coefficients (GFCCs) outperform other information representations by achieving the highest accuracy, particularly under noisy conditions. The interpretable results further reveal the underlying reasons for GFCC's superior performance, highlighting its ability to capture critical auditory features robustly across varying noise levels.These findings emphasize the potential of auditory perception-inspired representations in advancing computer audition-based disease detection systems and provide a solid foundation for future research in this domain. Yang Tan 0003, Rui Wang 0198, Kun Qian 0003, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Exploring the Power of Empirical Mode Decomposition for Sensing the Sound of Silence: A Pilot Study on Mice Autism Detection via Ultrasonic Vocalisation
Chenhao Wu 0004, Xiangjun Cai, Tianrui Jia, Yilu Deng, Kun Qian 0003, Björn W. Schuller, Yoshiharu Yamamoto, Jiang Liu 0005 |
INTERSPEECH | 6 |
| 2025 | Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
Yi Chang 0004, Zhao Ren, Zhonghao Zhao, Thanh Tam Nguyen, Kun Qian 0003, Tanja Schultz, Björn W. Schuller |
INTERSPEECH | 5 |
| 2025 | MADUV: The 1st INTERSPEECH Mice Autism Detection via Ultrasound Vocalization Challenge
Zijiang Yang 0007, Meishu Song, Xin Jing 0001, Kun Qian 0003, Bin Hu 0001, Kota Tamada, Toru Takumi, Björn W. Schuller, Yoshiharu Yamamoto |
INTERSPEECH | 5 |
| 2025 | Toward Practical Colorectal Cancer Diagnosis: A Bowel-Sound-Based System With Portable Sensor and On-Board Lightweight AI ModelabstractColorectal Cancer (CRC) is one of the leading causes of cancer-related deaths worldwide, and early screening plays a crucial role in improving patient outcomes. In this study, we present a novel AI-assisted CRC diagnostic system using Bowel Sound (BS) signals. We first develop two portable BS acquisition devices with distinct form factors for high-fidelity signal capture in both clinical and home-care scenarios. A total of 221 recordings were collected under expert-guided protocol, with 144 CRC recordings and 59 Non-CRC healthy controls using the developed device. To enable low-resource deployment, we design a lightweight deep learning model optimized for real-time, on-board inference. The model incorporates multiple training strategies, including transfer learning on a large-scale public BS dataset, self-supervised temporal feature learning, and a hybrid semi-and weakly-supervised approach that leverages both unlabeled and real-noise data. Furthermore, a Sound Event Detection (SED) attention mechanism and iterative consistency learning are introduced to enhance the model’s sensitivity to BS activity. The proposed model comprises only 264.7 K parameters and 253.2 M Floating-Point Operations (FLOPs), requiring 1.57 MB of RAM and 1.03 MB of FLASH when deployed on microcontroller. It performs inference in approximately 3.4 s with low power consumption, making it well-suited for low-resource environments. Despite its compact design, the model achieves 93.06% classification accuracy, 96.46% sensitivity, and 86.99% specificity for binary-classes in CRC diagnosis. These results demonstrate the system’s potential for accessible and cost-effective CRC screening in community, home, and rural healthcare scenarios. Fuze Tian, Yang Tan 0003, Enze Li, Jiedong Ma, Jingyu Liu 0002, Kun Qian 0003, Jing Li 0046, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE Internet Things J. | 8 |
| 2025 | GCD-JFSE: Graph-based class-domain knowledge joint feature selection and ensemble learning for EEG-based emotion recognition
Yutong Han, Weichu Xie, Fuze Tian, Lixian Zhu, Kun Qian 0003, Xiaowei Li 0005, Bin Hu 0001 |
Knowl. Based Syst. | 6 |
| 2025 | STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion RecognitionabstractSpeech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models. Yi Chang 0004, Zhao Ren, Zixing Zhang 0001, Xin Jing 0001, Kun Qian 0003, Xi Shao, Bin Hu 0001, Tanja Schultz, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | IMGWOFS: A Feature Selector With Trade-Off Between Conflict Objectives for EEG-Based Emotion RecognitionabstractFeature selection is a crucial step in EEG emotion recognition. However, it was often used as a single objective problem to either reduce the number of features or maximize classification accuracy, while neglecting their balance. To address the issue, we proposed Improved Multi-objective Grey Wolf Optimization Feature Selection (IMGWOFS). First, we designed a population initialization operator via discriminability and independence of features to accelerate search speed. Second, we employed a two-stage update strategy to improve the global search capabilities of the EEG feature subsets. Finally, we incorporated an adaptive mutation operator to escape the local optima. We conducted experiments on SEED and DEAP datasets, and the accuracy were 86.87$\pm$1.62 % and 60.65$\pm$1.51 % in the beta band using a smaller number of EEG features. In addition, the frontal lobe was related to emotion processing. In conclusion, IMGWOFS is an effective and feasible feature selection method for EEG-based emotion recognition. Chang Yan, Shanshan Qu, Dixin Wang, Na Chu, Fuze Tian, Kun Qian 0003, Xiaowei Li 0005, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 9 |
| 2025 | Enhancing Emotion Regulation in Mental Disorder Treatment: An AIGC-Based Closed-Loop Music Intervention SystemabstractMental disorders have increased rapidly and have emerged as a serious social health issue in the recent decade. Undoubtedly, the timely treatment of mental disorders is crucial. Emotion regulation has been proven to be an effective method for treating mental disorders. Music therapy as one of the methods that can achieve emotional regulation has gained increasing attention in the field of mental disorder treatment. However, traditional music therapy methods still face some unresolved issues, such as the lack of real-time capability and the inability to form closed-loop systems. With the advancement of artificial intelligence (AI), especially AI-generated content (AIGC), AI-based music therapy holds promise in addressing these issues. In this paper, an AIGC-based closed-loop music intervention system demonstration is proposed to regulate emotions for mental disorder treatment. This system demonstration consists of an emotion recognition model and a music generation model. The emotion recognition model can assess mental states, while the music generation model generates the corresponding emotional music for regulation. The system continuously performs recognition and regulation, thus forming a closed-loop process. In the experiment, we first conduct experiments on both the emotion recognition model and the music generation model to validate the accuracy of the recognition model and the music quality generated by the music generation models. In conclusion, we conducted comprehensive tests on the entire system to verify its feasibility and effectiveness. Cuiping Zhu, Ruobing Li, Kun Qian 0003, Fuze Tian, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Semantic Disentangling for Audiovisual Induced EmotionabstractEmotions regulation play an important role in human behavior, but exhibit considerable heterogeneity among individuals, which attenuates the generalization ability of emotion models. In this work, we aim to achieve robust emotion prediction through efficient disentanglement of affective semantic representations. In detail, the data generation mechanism behind observations from different perspectives is causally set, where latent variables that relate to emotion are explicitly separate into three parts: the intrinsic-related part, the extrinsic-related part, and the spurious-related part. Affective semantic features consist of the first two parts, with the understanding that spurious latent variables generate the inherent biases in the data. Furthermore, a variational autoencoder with a reformulated objective function is proposed to learn such disentangled latent variables, and only adopts semantic representations to perform the final classification task, avoiding the interference of spurious variables. In addition, for electroencephalography (EEG) data used in this article, a space-frequency mapping method is introduced to improve information utilization. Comprehensive experiments on popular emotion datasets show that the proposed method can achieve competitive intersubject generalization performance. Our results highlight the potential of efficient latent representation disentanglement in addressing the complexity challenges of emotion recognition. Qunxi Dong, Fuze Tian, Lixian Zhu, Kun Qian 0003, Jingyu Liu 0002 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Creating Healthier Living Environments: The Role of Soundscapes in Promoting Mental Health and Well-Being
Jian Kang 0002, Kun Qian 0003, Björn W. Schuller, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Information Diffusion Prediction With Augmented Diffusion Dependency and Multigranularity Temporal InfluenceabstractInformation diffusion prediction plays a pivotal role in the analysis of information propagation across social networks. Many existing methods rely on learning social homophily solely from users’ social connections as a single diffusion dependency to drive information diffusion. Moreover, these approaches often capture temporal influence from cascades within discrete time intervals, which might be inadequate in describing complex diffusion processes and can limit prediction performance. To overcome these limitations, we propose a novel approach with augmented diffusion dependency and multigranularity temporal influence (ADDMT) for information diffusion prediction. Our method strategically leverages the interactive regularity implicit in historical diffusion cascades. This information is integrated with social homophily through a cross-graph convolution network (GCN) to augment the diffusion dependency among users. Furthermore, we introduce multiple overlapping sliding windows to partition diffusion cascades. Adjacent cascade slices exhibit 50% overlap, enhancing semantic and structural coherence. In addition, we employ the combination of hypergraph convolution networks (HGCNs) and temporal convolution networks (TCNs) to capture multigranularity temporal influence within cascades. This design enables our model to further discern evolutionary trends and ephemeral fluctuations in users’ preferences across time intervals. The experimental results, obtained from comprehensive evaluations on four realistic datasets, demonstrate the superior performance of our proposed model. In particular, our model surpasses previous state-of-the-art diffusion prediction models, as evidenced by improved metrics such as Hits@K and MAP@K. These results underscore the effectiveness and robustness of ADDMT in predicting information diffusion in social networks. Zekun Tao, Kele Xu, Tao Sun 0005, Kun Qian 0003, Yanru Bai, Shanshan Li 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | MF$^{2}$-Net: Exploring a Meta-Fuzzy Multimodal Fusion Network for Depression RecognitionabstractDepression is a prevalent mental illness that significantly impacts the well-being of individuals and the development of society. The current diagnostic methods are largely subjective and time-consuming. Moreover, the existing machine learning-based depression recognition methods struggle to fully exploit the collaborative benefits between modalities, lack interpretability in their fusion processes, and perform inadequately in few-shot depression recognition tasks. To address these challenges, we propose a meta-fuzzy multimodal fusion network (MF$^{2}$-Net) for depression recognition. This innovative approach integrates physiological signals and behavioral data, employs multiple MLPs to learn the fuzzy measures of single base learners and complementary increments, then constructs all fuzzy measures, and finally achieves an interpretable decision-level fusion process through fuzzy integrals. Furthermore, we incorporate model-agnostic meta-learning for conducting few-shot domain-adaptive training, mitigating the issues related to high individual variability levels and the scarcity of multimodal depression data. Our method demonstrates exceptional classification performance in subject-independent experiments implemented on public datasets; offers a reliable solution for objectively, effectively, and conveniently recognizing depression; and has the potential to promote the clinical applications of rapid intelligent depression diagnosis. Jian Shen 0004, Jinwen Wu, Kang Wang 0010, Kechen Hou, Kun Qian 0003, Xiaowei Zhang 0001, Bin Hu 0001 |
IEEE Trans. Fuzzy Syst. | 8 |
| 2025 | A Review of AIoT-Based Human Activity Recognition: From Application to TechniqueabstractThis scoping review paper redefines the Artificial Intelligence-based Internet of Things (AIoT) driven Human Activity Recognition (HAR) field by systematically extrapolating from various application domains to deduce potential techniques and algorithms. We distill a general model with adaptive learning and optimization mechanisms by conducting a detailed analysis of human activity types and utilizing contact or non-contact devices. It presents various system integration mathematical paradigms driven by multimodal data fusion, covering predictions of complex behaviors and redefining valuable methods, devices, and systems for HAR. Additionally, this paper establishes benchmarks for behavior recognition across different application requirements, from simple localized actions to group activities. It summarizes open research directions, including data diversity and volume, computational limitations, interoperability, real-time recognition, data security, and privacy concerns. Finally, we aim to serve as a comprehensive and foundational resource for researchers delving into the complex and burgeoning realm of AIoT-enhanced HAR, providing insights and guidance for future innovations and developments. Wen Qi 0005, Xiangmin Xu 0001, Kun Qian 0003, Björn W. Schuller, Giancarlo Fortino, Andrea Aliverti |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | FedKDC: Consensus-Driven Knowledge Distillation for Personalized Federated Learning in EEG-Based Emotion RecognitionabstractFederated learning (FL) has gained prominence in electroencephalogram (EEG)-based emotion recognition because of its ability to enable secure collaborative training without centralized data. However, traditional FL faces challenges due to model and data heterogeneity in smart healthcare settings. For example, medical institutions have varying computational resources, which creates a need for personalized local models. Moreover, EEG data from medical institutions typically face data heterogeneity issues stemming from limitations in participant availability, ethical constraints, and cultural differences among subjects, which can slow model convergence and degrade model performance. To address these challenges, we propose FedKDC, a novel FL framework that incorporates clustered knowledge distillation (CKD). This method introduces a consensus-based distributed learning mechanism to facilitate the clustering process. It then enhances the convergence speed through intraclass distillation and reduces the negative impact of heterogeneity through interclass distillation. Additionally, we introduce a DriftGuard mechanism to mitigate client drift, along with an entropy reducer to decrease the entropy of aggregated knowledge. The framework is validated on the SEED, SEED-IV, SEED-FRA, and SEED-GER datasets, demonstrating its effectiveness in scenarios where both the data and the models are heterogeneous. Experimental results show that FedKDC outperforms other FL frameworks in emotion recognition, achieving a maximum average accuracy of 85.2%, and in convergence efficiency, with faster and more stable convergence. Xihang Qiu, Wanyong Qiu, Ye Zhang 0017, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | An On-Board Executable Multi-Feature Transfer-Enhanced Fusion Model for Three-Lead EEG Sensor-Assisted Depression DiagnosisabstractThe development of affective computing and medical electronic technologies has led to the emergence of Artificial Intelligence (AI)-based methods for the early detection of depression. However, previous studies have often overlooked the necessity for the AI-assisted diagnosis system to be wearable and accessible in practical scenarios for depression recognition. In this work, we present an on-board executable multi-feature transfer-enhanced fusion model for our custom-designed wearable three-lead Electroencephalogram (EEG) sensor, based on EEG data collected from 73 depressed patients and 108 healthy controls. Experimental results show that the proposed model exhibits low-computational complexity (65.0 K parameters), promising Floating-Point Operations (FLOPs) performance (25.6 M), real-time processing (1.5 s/execution), and low power consumption (320.8 mW). Furthermore, it requires only 202.0 KB of Random Access Memory (RAM) and 279.6 KB of Read-Only Memory (ROM) when deployed on the EEG sensor. Despite its low computational and spatial complexity, the model achieves a notable classification accuracy of 95.2%, specificity of 94.0%, and sensitivity of 96.9% under independent test conditions. These results underscore the potential of deploying the model on the wearable three-lead EEG sensor for assisting in the diagnosis of depression. Fuze Tian, Yang Tan 0003, Lixian Zhu, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Decider: A Dual-System Rule-Controllable Decoding Framework for Language GenerationabstractConstrained decoding approaches aim to control the meaning or style of text generated by a Pre-trained Language Model (PLM) for various task-specific objectives at inference time. However, these methods often guide plausible continuations by greedily and explicitly selecting targets, which, while fulfilling the task requirements, may overlook the natural patterns of human language generation. In this work, we propose a novel decoding framework,Decider, which enables us to program high-level rules on how we might effectively complete tasks to control a PLM. Differing from previous works, our framework transforms the encouragement of concrete target words into the encouragement of all words that satisfy the high-level rules. Specifically,Decideris a dual system in which a PLM is equipped and controlled by a First-Order Logic (FOL) reasoner to express and evaluate the rules, along with a decision function that merges the outputs from both systems to guide the generation. Experiments on CommonGen and PersonaChat demonstrate thatDecidercan effectively follow given rules to guide a PLM in achieving generation tasks in a more human-like manner. Tian Lan 0003, Changlong Yu, Wei Wang 0138, Qunxi Dong, Kun Qian 0003, Piji Li, Wei Bi, Bin Hu 0001 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2025 | An AI-Assisted All-in-One Integrated Coronary Artery Disease Diagnosis System Using a Portable Heart Sound Sensor With an On-Board Executable Lightweight ModelabstractHeart sounds play a crucial role in assessing Coronary Artery Disease (CAD). The advancement of Artificial Intelligence (AI) technologies has given rise to Computer Audition (CA)-based methods for CAD detection. However, previous research has focused primarily on analyzing and modeling heart sound data, overlooking practical application scenarios. In this work, we design a pervasive heart sound collection device used for high-quality heart sound data acquisition. Moreover, we introduce an on-board executable lightweight network tailored for the designed portable device, referred to as TYKDModel. Further, heart sound data from 41 CAD patients and 22 non-CAD healthy controls are collected using the developed device. Experimental results show that the TYKDModel exhibits low-computational complexity, with 52.16 K parameters and 5.03 M Floating-Point Operations (FLOPs). When deployed on the board, it requires only 1.10 MB of Random Access Memory (RAM) and 236.27 KB of Read-Only Memory (ROM), and takes around 1.72 seconds to perform a classification. Despite the low computational and spatial complexity, the TYKDModel achieves a notable classification accuracy of 85.2%, specificity of 88.6%, and sensitivity of 82.8% on the board. These results indicate the promising potential of AI-assisted all-in-one integrated system for the diagnosis of heart sound-assisted CAD. Fuze Tian, Yang Tan 0003, Jingyu Liu 0002, Kun Qian 0003, Yalei Han, Gong Su, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Clearer Lub-Dub: A Novel Approach in Heart Sound Denoising Based on Transfer LearningabstractCardiovascular diseases (CVDs) constitute the primary cause of human mortality globally in recent decades. To effectively detect CVDs, heart auscultation plays an important role in early diagnosis. With the development of artificial intelligence (AI), many studies have designed varying AI-assisted diagnosis systems helping people discriminate abnormal heart sounds. Yet, a robust system usually requires a noise-less input signal, which is critical as heart sounds are often affected by some unavoidable noise. Therefore, many heart sound classification models use filters or other methods to obtain the clean signals. However, these classic techniques are not adaptable enough to distinguish the meaningful murmurs and real noises. Thus, we propose a novel approach to transfer an audio source separation model to denoise the heart sound. In this paper, we test different denoisers on synthesis heart sound with additive white Gaussian noises. Our method performs well on the noise reduction metrics. Meanwhile, we evaluate the classification performance of each denoiser with some classifiers on the PhysioNet dataset. Experimental results demonstrate that our method can outperform other denoising techniques by achieving the highest unweighted average recall (UAR) at 95.7% with the smallest standard deviation. The results confirm that our method is robust and adaptable in improving audio's denoising. Jiang Liu 0005, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
HealthCom | 6 |
| 2024 | Inspiration of Prototype Knowledge: Introducing a Meta-Learning Approach to Heart Sound ClassificationabstractCardiovascular diseases (CVDs) stand as the primary reason of fatalities globally, especially in low- and middle-income countries. In recent years, with the leverage of computer audition technologies, the diagnosis of CVDs through heart sounds become a popular topic. Current models and techniques are trained, validated, and tested on the same dataset, which need to be retrained when encountering new data. To make the best use of sparse data, we propose a Prototypical Network framework with heuristic weight for heart sound recognition. After extracting two different features (Mel Spectrogram and Mel Frequency Cepstral Coefficients) and encoding the features, we calculate the distance between two categories (normal and abnormal), then, a heuristic weight is assigned to the distance that makes the blurred boundaries more distinct. By considering the subject independence, the Unweighted Average Recall (UAR) on the PhysioNet/CinC Challenge 2016 is 68.2 % and 67.7 % on two features, respectively. The capability of our model to work on different datasets is proved by a UAR of 66.4 %, which exceeds the baseline UAR of 58.6 % under a single model. Qingrong Jackie Wu, Mengkai Sun, Boyang Meng, Kun Qian 0003, Bin Hu 0001, Toru Nakamura, Taishin Nomura, Björn W. Schuller, Yoshiharu Yamamoto |
HealthCom | 7 |
| 2024 | Deep Fusion of Shifted MLP and CNN for Medical Image SegmentationabstractMedical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation. In this work, we propose a novel hybrid network that involves three main branches: the Multi-Layer Perception (MLP) branch, the Convolutional Neural Network (CNN) branch, and a Fusion branch. The MLP and CNN branches aim to learn global and local features, respectively. To fuse these, the fusion branch introduces a novel hierarchical fusion that performs multi-layered fusions that generate high-level representations to enhance segmentation. Our evaluation with two datasets shows strong performance of the proposed method compared to state-of-the-art baselines. Chengyu Yuan, Hao Xiong 0001, Guoqing Shangguan, Hualei Shen, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto, Shlomo Berkovsky |
ICASSP | 8 |
| 2024 | E-ODN: An Emotion Open Deep Network for Generalised and Adaptive Speech Emotion Recognition
Liuxian Ma, Ruobing Li, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
INTERSPEECH | 5 |
| 2024 | Study Selectively: An Adaptive Knowledge Distillation based on a Voting Network for Heart Sound ClassificationabstractPhonocardiogram classification methods using deep neural networks have been widely applied to the early detection of cardiovascular diseases recently.Despite their excellent recognition rate, the sizeable computational complexity limits their further development.Nowadays, knowledge distillation (KD) is an established paradigm for model compression.While current research on multi-teacher KD has shown potential to impart more comprehensive knowledge to the student than single-teacher KD, this approach is not suitable for all scenarios.This paper proposes a novel KD strategy to realise an adaptive multi-teacher instruction mechanism.We design a teacher selection strategy called voting network to tell the contribution of different teachers on each distillation points, so that the student can choose the useful information and renounce the redundant one.An evaluation demonstrates that our method reaches excellent accuracy (92.8 %) while maintaining a low computational complexity (0.7 M). Xihang Qiu, Lixian Zhu, Zikai Song, Kun Qian 0003, Ye Zhang 0017, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller |
INTERSPEECH | 6 |
| 2024 | Automatic Bird Sound Source Separation Based on Passive Acoustic Devices in Wild EnvironmentabstractThe Internet of Things (IoT)-based passive acoustic monitoring (PAM) has shown great potential in large-scale remote bird monitoring. However, field recordings often contain overlapping signals, making precise bird information extraction challenging. To solve this challenge, first, the inter-channel spatial feature is chosen as complementary information to the spectral feature to obtain additional spatial correlations between the sources. Then, an end-to-end model named BACPPNet is built based on Deeplabv3plus and enhanced with the polarized self-attention mechanism to estimate the spectral amplitude mask (SMM) for separating bird vocalizations. Finally, the separated bird vocalizations are recovered from SMMs and the spectrogram of mixed audio using the inverse short Fourier transform (ISTFT). We evaluate our proposed method utilizing the generated mixed dataset. Experiments have shown that our method can separate bird vocalizations from mixed audio with RMSE, SDR, SIR, SAR, and STOI values of 2.82, 10.00dB, 29.90 dB, 11.08 dB, and 0.66, respectively, which are better than existing methods. Furthermore, the average classification accuracy of the separated bird vocalizations drops the least. This indicates that our method outperforms other compared separation methods in bird sound separation and preserves the fidelity of the separated sound sources, which might help us better understand wild bird sound recordings. Jiangjian Xie, Yuwei Shi, Dongming Ni, Manuel Milling, Shuo Liu 0012, Junguo Zhang, Kun Qian 0003, Björn W. Schuller |
IEEE Internet Things J. | 7 |
| 2024 | Physiological Electrosignal Asynchronous Acquisition Technology: Insight and PerspectivesabstractWith great pride and enthusiasm, we present the inaugural edition of IEEE Transactions on Computational Social Systems (TCSS) for 2024. Reflecting on the year gone by, 2023 stands as a hallmark of academic excellence and prolific output, wherein our journal has successfully disseminated a substantial volume of scholarly work—301 articles encompassing approximately 3600 pages, distributed across six distinct issues. Bin Hu 0001, Lixian Zhu, Qunxi Dong, Kun Qian 0003, Hanshu Cai, Fuze Tian |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | A Novel Intelligence Evaluation Framework: Exploring the Psychophysiological Patterns of Gifted StudentsabstractIntelligence evaluation is a desirable intelligent application for sensing and interaction in various scenarios, e.g., education, office, and the aviation industry. For example, identifying gifted students, who learn faster and more efficiently than general students due to their neurophysiological advantages, and teaching different students according to their intelligence are urgent requirements in school education. However, current intelligence evaluation mainly relies on intelligence quotient (IQ) tests, which have a problem of decreasing reliability in repeated tests. In addition, no objective assessment criteria are available in the present intelligence evaluation process. Electroencephalogram (EEG) signals, which reflect the neuroelectrical activities of the brain, can be utilized to develop an objective and promising tool for investigating the neurophysiological advantages of gifted groups and augmenting the effects of intelligence evaluation. Consequently, we proposed a novel real-time intelligence evaluation framework based on users’ psychophysiological data. Then, we leveraged the framework to investigate a case study to asses which EEG patterns could be used to effectively characterize gifted students and distinguish them from average students. Experimental results reveal the great differences in the chaos degree of the brain (CDB) between different groups of subjects and the effectiveness of the model in identifying gifted students, thus verifying the practicability and validity of the proposed framework. Jian Shen 0004, Zeguang Zhao, Huajian Liang, Kun Qian 0003, Qunxi Dong |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Fed-MStacking: Heterogeneous Federated Learning With Stacking Misaligned Labels for Abnormal Heart Sound DetectionabstractUbiquitous sensing has been widely applied in smart healthcare, providing an opportunity for intelligent heart sound auscultation. However, smart devices contain sensitive information, raising user privacy concerns. To this end, federated learning (FL) has been adopted as an effective solution, enabling decentralised learning without data sharing, thus preserving data privacy in the Internet of Health Things (IoHT). Nevertheless, traditional FL requires the same architectural models to be trained across local clients and global servers, leading to a lack of model heterogeneity and client personalisation. For medical institutions with private data clients, this study proposes Fed-MStacking, a heterogeneous FL framework that incorporates a stacking ensemble learning strategy to support clients in building their own models. The secondary objective of this study is to address scenarios involving local clients with data characterised by inconsistent labelling. Specifically, the local client contains only one case type, and the data cannot be shared within or outside the institution. To train a global multi-class classifier, we aggregate missing class information from all clients at each institution and build meta-data, which then participates in FL training via a meta-learner. We apply the proposed framework to a multi-institutional heart sound database. The experiments utilise random forests (RFs), feedforward neural networks (FNNs), and convolutional neural networks (CNNs) as base classifiers. The results show that the heterogeneous stacking of local models performs better compared to homogeneous stacking. Wanyong Qiu, Yuying Li 0006, Yi Chang 0004, Kun Qian 0003, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Automated Cough Sound Analysis for Detecting Childhood PneumoniaabstractPneumonia is one of the leading causes of death in children. Prompt diagnosis and treatment can help prevent these deaths, particularly in resource poor regions where deaths due to pneumonia are highest. Clinical symptom-based screening of childhood pneumonia yields excessive false positives, highlighting the necessity for additional rapid diagnostic tests. Cough is a prevalent symptom of acute respiratory illnesses and the sound of a cough can indicate the underlying pathological changes resulting from respiratory infections. In this study, we propose a fully automated approach to evaluate cough sounds to distinguish pneumonia from other acute respiratory diseases in children. The proposed method involves cough sound denoising, cough sound segmentation, and cough sound classification. The denoising algorithm utilizes multi-conditional spectral mapping with a multilayer perceptron network while the segmentation algorithm detects cough sounds directly from the denoised audio waveform. From the segmented cough signal, we extract various handcrafted features and feature embeddings from a pretrained deep learning network. A multilayer perceptron is trained on the combined feature set for detecting pneumonia. The method we propose is evaluated using a dataset comprising cough sounds from 173 children diagnosed with either pneumonia or other acute respiratory diseases. On average, the denoising algorithm improved the signal-to-noise ratio by 44%. Furthermore, a sensitivity and specificity of 91% and 86%, respectively, is achieved in cough segmentation and 82% and 71%, respectively, in detecting childhood pneumonia using cough sounds alone. This demonstrates its potential as a rapid diagnostic tool, such as using smartphone technology. Roneel V. Sharan, Kun Qian 0003, Yoshiharu Yamamoto |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | AMNet: Introducing an Adaptive Mel-Spectrogram End-to-End Neural Network for Heart Sound ClassificationabstractThe cardiovascular diseases (CVDs) cause tremendous deaths yearly. The Mel-spectrogram is widely used as a tool to analyse the heart sound, which facilitate a cheap and efficient diagnosis of CVDs. Nevertheless, the amplitude and frequency responses of the Mel filter banks remain constant, limiting its function to frequency selection. We propose an adaptive Melspectrogram end-to-end neural network (AMNet) for a better characterisation and classification of heart sound in the work. The core of the adaptive Mel-spectrograms (AMel) lies in an adaptive Mel filter banks whose frequency characteristics remain the same as the original Mel-spectrogram (OMel) and amplitude is learnt by the backropagation algorithm. The AMNet learns the raw audio representation directly and outputs the classification results. It reaches 43.5% Unweighted Average Recall (UAR) and surpasses the model with the OMel and the baseline by 6% UAR. It is demonstrated that the AMel characterises the heart sound more effectively. Yang Tan 0003, Kun Qian 0003, Zhihao Bao, Zheyu Cao, Bin Hu 0001, Yoshiharu Yamamoto, Björn W. Schuller |
HealthCom | 3 |
| 2023 | Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured LearningabstractSpeech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep learning has been investigated to improve the performance of SER by training complex models, the memory space and computational capability of edge devices represents a constraint for embedding deep learning models. We propose a neural structured learning (NSL) framework through building synthesized graphs. An SER model is trained on a source dataset and used to build graphs on a target dataset. A relatively lightweight model is then trained with the speech samples and graphs together as the input. Our experiments demonstrate that training a lightweight SER model on the target dataset with speech samples and graphs can not only produce small SER models, but also enhance the model performance compared to models with speech samples only and those using classic transfer learning strategies. Yi Chang 0004, Zhao Ren, Thanh Tam Nguyen, Kun Qian 0003, Björn W. Schuller |
ICASSP | 4 |
| 2023 | Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning AnalysisabstractTranslating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the Japanese Daily Speech Dataset (JDSD), a large in-the-wild daily speech emotion dataset consisting of 20,827 speech samples from 342 speakers and 54 hours of total duration. The data is annotated on the Depression and Anxiety Mood Scale (DAMS) – 9 self-reported emotions to evaluate mood state including "vigorous", "gloomy", "concerned", "happy", "unpleasant", "anxious", "cheerful", "depressed", and "worried". Our dataset possesses emotional states, activity, and time diversity, making it useful for training models to track daily emotional states for healthcare purposes. We partition our corpus and provide a multi-task benchmark across nine emotions, demonstrating that mental health states can be predicted reliably from self-reports with a Concordance Correlation Coefficient value of .547 on average. We hope that JDSD will become a valuable resource to further the development of daily emotional healthcare tracking. Meishu Song, Andreas Triantafyllopoulos, Zijiang Yang 0007, Hiroki Takeuchi, Toru Nakamura, Akifumi Kishi, Tetsuro Ishizawa, Kazuhiro Yoshiuchi, Xin Jing 0001, Vincent Karas, Zhonghao Zhao, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
ICASSP | 12 |
| 2023 | Federated Intelligent Terminals Facilitate Stuttering MonitoringabstractStuttering is a complicated language disorder. The most common form of stuttering is developmental stuttering, which begins in childhood. Early monitoring and intervention are essential for the treatment of children with stuttering. Automatic speech recognition technology has shown its great potential for non-fluent disorder identification, whereas the previous work has not considered the privacy of users’ data. To this end, we propose federated intelligent terminals for automatic monitoring of stuttering speech in different contexts. Experimental results demonstrate that the proposed federated intelligent terminals model can analyze symptoms of stammering speech by taking personal privacy protection into account. Furthermore, the study has explored that the Shapley value approach in the federated learning setting has comparable performance to data-centralised learning. Yongzi Yu, Wanyong Qiu, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
ICASSP | 4 |
| 2023 | Explainable Stuttering Recognition Using Axial Attention
Kaixiang Yuan, Guangzhe Xuan, Yongzi Yu, Hengrui Zhong, Rui Li 0105, Jian Shen 0004, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto |
ICIC (3) | 9 |
| 2023 | Automatic Audio Augmentation for Requests Sub-ChallengeabstractThis paper presents our solution for the Requests Sub-challenge of the ACM Multimedia 2023 Computational Paralinguistics Challenge. Drawing upon the framework of self-supervised learning, we put forth an automated data augmentation technique for audio classification, accompanied by a multi-channel fusion strategy aimed at enhancing overall performance. Specifically, to tackle the issue of imbalanced classes in complaint classification, we propose an audio data augmentation method that generates appropriate augmentation strategies for the challenge dataset. Furthermore, recognizing the distinctive characteristics of the dual-channel HC-C dataset, we individually evaluate the classification performance of the left channel, right channel, channel difference, and channel sum, subsequently selecting the optimal integration approach. Our approach yields a significant improvement in performance when compared to the competitive baselines, particularly in the context of the complaint task. Moreover, our method demonstrates noteworthy cross-task transferability. Yanjie Sun, Kele Xu, Chaorun Liu, Yong Dou, Kun Qian 0003 |
ACM Multimedia | 5 |
| 2023 | Intelligent Music Intervention for Mental Disorders: Insights and PerspectivesabstractWelcome to the first issue of IEEE Transactions on Computational Social Systems (TCSS) of 2023. The past 2022 was again a very productive year, in which we have published 159 articles with about 1850 pages in six issues. We also received much great and exciting news. Kun Qian 0003, Björn W. Schuller, Xiaohong Guan, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Can a Holistic View Facilitate the Development of Intelligent Traditional Chinese Medicine? A SurveyabstractIntelligent traditional Chinese medicine (ITCM) is an emerging interdisciplinary subject. It aims to efficiently and precisely promote the prevention and treatment of diseases and health management in Chinese medicine clinical practice via the combination of traditional Chinese medicine (TCM) fundamentals and artificial intelligence technologies. Presently, it is experiencing dramatic growth in recent years. On the one hand, a holistic view, as a crucial philosophy in the theory of TCM, will be guiding the development of ITCM. On the other hand, a comprehensive discussion of the benefits of such a holistic view of ITCM is lacking. To this end, we conduct this survey by introducing the named holistic view first. Then, adaptive learning and field theory will be presented and discussed with respect to their application in ITCM. Ethical issues of ITCM will then be taken into account by human-centered TCM and potentials based on affective computing of ITCM. In addition, we give our opinions and insights on the challenges and open issues regarding the future of ITCM. We hope that this survey article can be a good guide for experts in the relevant fields. Guihua Tian, Kun Qian 0003, Mengkai Sun, Wanyong Qiu, Xiaoming Xie, Zhonghao Zhao, Liangqing Huang, Siyan Luo, Tianxing Guo, Ran Cai, Björn W. Schuller |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Depression Recognition From EEG Signals Using an Adaptive Channel Fusion Method via Improved Focal LossabstractDepression is a serious and common psychiatric disease characterized by emotional and cognitive dysfunction. In addition, the rates of clinical diagnosis and treatment for depression are low. Therefore, the accurate recognition of depression is important for its effective treatment. Electroencephalogram (EEG) signals, which can objectively reflect the inner states of human brains, are regarded as promising physiological tools that can enable effective and efficient clinical depression diagnosis and recognition. However, one of the challenges regarding EEG-based depression recognition involves sufficiently optimizing the spatial information derived from the multichannel space of EEG signals. Consequently, we propose an adaptive channel fusion method via improved focal loss (FL) functions for depression recognition based on EEG signals to effectively address this challenge. In this method, we propose two improved FL functions that can enhance the separability of hard examples by upweighting their losses as optimization objectives and can optimize the channel weights by a proposed adaptive channel fusion framework. The experimental results obtained on two EEG datasets show that the developed channel fusion method can achieve improved classification performance. The learned channel weights include the individual characteristics of each EEG epoch, which can effectively optimize the spatial information of each EEG epoch via the channel fusion method. In addition, the proposed method performs better than the state-of-the-art channel fusion methods. Jian Shen 0004, Huajian Liang, Zeguang Zhao, Kun Qian 0003, Qunxi Dong, Xiaowei Zhang 0001, Bin Hu 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | An Overview of the FIRST ICASSP Special Session on Computer Audition for HealthcareabstractAudio has been increasingly used as a novel digital phenotype that carries important information of the subject’s health status. We can find tremendous efforts given to this young and promising field, i.e., computer audition for healthcare (CA4H), whereas the application scenarios have not been fully studied as compared to its counterpart in medical areas, computer vision. To this end, the first special session held at ICASSP 2020 was dedicated to the topic. In this overview paper, we at first summarise the invited high-quality contributions from leading scientists from a multi- disciplinary background. Then, we provide a detailed grouping of the contributions to several scenarios such as body sound analysis (e.g., heart sound), human speech analysis (e.g., stress detection), and artificial hearing technologies (e.g., cochlear implants). In addition to the collected works, we will compare them with other recent studies within the topic. Finally, we conclude the limitations and perspectives of the current stage. It is interesting and encouraging to find that the state-of-the-art machine learning and audio signal processing techniques have been successfully applied in the health domain, e.g., to fight with the global challenges of COVID-19 and ageing population. Kun Qian 0003, Tanja Schultz, Björn W. Schuller |
ICASSP | 1 |
| 2022 | A Glance-and-Gaze Network for Respiratory Sound ClassificationabstractA plethora of great successes has been achieved by the existing convolutional neural networks (CNN) for respiratory sound classification. Nevertheless, simultaneously capturing both the local and global features can never be an easy task due to the limitation of a CNN’s structure. In this contribution, we propose a novel glance-and-gaze network to address the aforementioned issue. The glance block aims to learn global information, while the gaze block is responsible for learning local patterns and suppressing the noises that attenuates the final performance. In the proposed method, both the global and local information can be extracted. Moreover, the spectral and temporal representations can be learnt via a feature fusion module. Experimental results on the largest public respiratory sound database demonstrate that the proposed model outperforms the state-of-the-art methods. Shuai Yu 0002, Yiwei Ding, Kun Qian 0003, Bin Hu 0001, Wei Li 0012, Björn W. Schuller |
ICASSP | 3 |
| 2022 | MEDAS: an open-source platform as a service to help break the walls between medicine and informatics
Liang Zhang 0010, Johann Li, Ping Li 0030, Xiaoyuan Lu, Maoguo Gong, Peiyi Shen, Guangming Zhu 0001, Syed Afaq Ali Shah, Mohammed Bennamoun, Kun Qian 0003, Björn W. Schuller |
Neural Comput. Appl. | 10 |
| 2022 | Psychological Field Versus Physiological Field: From Qualitative Analysis to Quantitative Modeling of the Mental StatusabstractWelcome to the fifth issue of IEEE Transactions on Computational Social Systems (TCSS) in 2022. After the usual introduction of our 24 regular articles, we would like to discuss the topic of “Psychological Field Versus Physiological Field: From Qualitative Analysis to Quantitative Modelling of the Mental Status.” Bin Hu 0001, Kun Qian 0003, Qunxi Dong, Yuejia Luo, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | Fundamentals of Computational Psychophysiology: Theory and MethodologyabstractWelcome to the second issue of IEEE Transactions on Computational Social Systems (TCSS) in 2022. In this issue, we are going to present 25 regular articles. After the “scanning the issue,” I would like to share some of my opinions and perspectives on the fundamentals of computational psychophysiology: theory and methodology. Bin Hu 0001, Jian Shen 0004, Lixian Zhu, Qunxi Dong, Hanshu Cai, Kun Qian 0003 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2022 | COVID-19's Impact on Mental Health - The Hour of Computational Aid?abstractWelcome to the fourth issue of IEEE Transactions on Computational Social Systems (TCSS) in 2022. First, we have some exciting news to share. In late June, Clarivate updated the Impact Factor of all journals which are indexed by Web of Science. According to the Journal Citation Reports, the 2021 Journal Impact Factor of IEEE TCSS was 4.727. Many thanks to all for your great effort and support. Björn W. Schuller, Johanna Löchner, Kun Qian 0003, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Digital Mental Health - Breaking a Lance for PreventionabstractWelcome to the last issue of IEEE Transactions On Computational Social Systems (TCSS) of 2022. In this issue, we publish a Special Issue on Advanced Cognitive Computing for Data-Driven Computational Social Systems, which includes 23 articles. Moreover, we also would like to share some of our opinions and perspectives on “Digital Mental Health—Breaking a Lance for Prevention.” Björn W. Schuller, Johanna Löchner, Kun Qian 0003, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Capturing Time Dynamics From Speech Using Neural Networks for Surgical Mask DetectionabstractThe importance of detecting whether a person wears a face mask while speaking has tremendously increased since the outbreak of SARS-CoV-2 (COVID-19), as wearing a mask can help to reduce the spread of the virus and mitigate the public health crisis. Besides affecting human speech characteristics related to frequency, face masks cause temporal interferences in speech, altering the pace, rhythm, and pronunciation speed. In this regard, this paper presents two effective neural network models to detect surgical masks from audio. The proposed architectures are both based on Convolutional Neural Networks (CNNs), chosen as an optimal approach for the spatial processing of the audio signals. One architecture applies a Long Short-Term Memory (LSTM) network to model the time-dependencies. Through an additional attention mechanism, the LSTM-based architecture enables the extraction of more salient temporal information. The other architecture (named ConvTx) retrieves the relative position of a sequence through the positional encoder of a transformer module. In order to assess to which extent both architectures can complement each other when modelling temporal dynamics, we also explore the combination of LSTM and Transformers in three hybrid models. Finally, we also investigate whether data augmentation techniques, such as, using transitions between audio frames and considering gender-dependent frameworks might impact the performance of the proposed architectures. Our experimental results show that one of the hybrid models achieves the best performance, surpassing existing state-of-the-art results for the task at hand. Shuo Liu 0012, Adria Mallol-Ragolta, Tianhao Yan, Kun Qian 0003, Emilia Parada-Cabaleiro, Bin Hu 0001, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Learning Multimodal Representations for Drowsiness DetectionabstractDrowsiness detection is a crucial step for safe driving. A plethora of efforts has been invested on using pervasive sensor data (e.g., video, physiology) empowered by machine learning to build an automatic drowsiness detection system. Nevertheless, most of the existing methods are based on complicated wearables (e.g., electroencephalogram) or computer vision algorithms (e.g., eye state analysis), which makes the relevant systems hardly applicable in the wild. Furthermore, data based on these methods are insufficient in nature due to limited simulation experiments. In this light, we propose a novel and easily implemented method based on full non-invasive multimodal machine learning analysis for the driver drowsiness detection task. The drowsiness level was estimated by self-reported questionnaire in pre-designed protocols. First, we consider involving environmental data (e.g., temperature, humidity, illuminance, and further more), which can be regarded as complementary information for the human activity data recorded via accelerometers or actigraphs. Second, we demonstrate that the models trained by daily life data can still be efficient to make predictions for the subject performing in a simulator, which may benefit the future data collection methods. Finally, we make a comprehensive study on investigating different machine learning methods including classic ‘shallow’ models and recent deep models. Experimental results show that, our proposed methods can reach 64.6% unweighted average recall for drowsiness detection in a subject-independent scenario. Kun Qian 0003, Tomoya Koike, Toru Nakamura, Björn W. Schuller, Yoshiharu Yamamoto |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Multi-label learning with missing and completely unobserved labelsabstractAbstract Multi-label learning deals with data examples which are associated with multiple class labels simultaneously. Despite the success of existing approaches to multi-label learning, there is still a problem neglected by researchers, i.e., not only are some of the values of observed labels missing, but also some of the labels are completely unobserved for the training data. We refer to the problem asmulti-label learning with missing and completely unobserved labels, and argue that it is necessary to discover these completely unobserved labels in order to mine useful knowledge and make a deeper understanding of what is behind the data. In this paper, we propose a new approach named MCUL to solve multi-label learning with Missing and Completely Unobserved Labels. We try to discover the unobserved labels of a multi-label data set with a clustering based regularization term and describe the semantic meanings of them based on the label-specific features learned by MCUL, and overcome the problem of missing labels by exploiting label correlations. The proposed method MCUL can predict both the observed and newly discovered labels simultaneously for unseen data examples. Experimental results validated over ten benchmark datasets demonstrate that the proposed method can outperform other state-of-the-art approaches on observed labels and obtain an acceptable performance on the new discovered labels as well. Jun Huang 0003, Linchuan Xu, Kun Qian 0003, Jing Wang 0023, Kenji Yamanishi |
Data Min. Knowl. Discov. | 3 |
| 2021 | Learning audio sequence representations for acoustic event classification
Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Björn W. Schuller |
Expert Syst. Appl. | 4 |
| 2021 | Can Appliances Understand the Behavior of Elderly Via Machine Learning? A Feasibility StudyabstractOver the last half decade, fast development of the Internet of Things and machine learning (ML) made it feasible to leverage the power of artificial intelligence to facilitate a variety of intelligent systems in smart home. Nevertheless, the studies on designing specific computing technologies for helping elderly to enjoy a comfortable, convenient, and independent daily life are extremely limited. On the one hand, there are increasingly growing demands from the ageing society to implement the cutting edge technology enabling a better life quality for the elderly. On the other hand, there is still a lack on fundamental investigations, applicable infrastructures, and advanced data-driven frameworks. To this end, we propose a novel machine framework for analyzing the daily life behavior of elderly-all in this study are living alone-by the data collected from their home appliances, i.e., television and refrigerator. First, the interevent intervals for the use of the appliances collected in one month from 76 elderly are the raw data to describe the behaviors. Then, three ML paradigms are investigated and compared, which include “classic” ML methods and the state-of-the-art deep learning approaches. Finally, we indicate the current findings and limitations in this feasibility study. Experimental results demonstrate that, our proposed method can reach performance peak at an unweighted average recall of 58.7% (chance level: 50.0%) in a subject-independent test for classifying symptom/nonsymptom days. Kun Qian 0003, Tomoya Koike, Kazuhiro Yoshiuchi, Björn W. Schuller, Yoshiharu Yamamoto |
IEEE Internet Things J. | 1 |
| 2021 | Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19abstractComputer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose. Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE Internet Things J. | 1 |
| 2021 | An Online Robot Collision Detection and Identification Scheme by Supervised Learning and Bayesian Decision TheoryabstractThis article is dedicated to developing an online collision detection and identification (CDI) scheme for human-collaborative robots. The scheme is composed of a signal classifier and an online diagnosor, which monitors the sensory signals of the robot system, detects the occurrence of a physical human–robot interaction, and identifies its type within a short period. In the beginning, we conduct an experiment to construct a data set that contains the segmented physical interaction signals with ground truth. Then, we develop the signal classifier on the data set with the paradigm of supervised learning. To adapt the classifier to the online application with requirements on response time, an auxiliary online diagnosor is designed using the Bayesian decision theory. The diagnosor provides not only a collision identification result but also a confidence index which represents the reliability of the result. Compared to the previous works, the proposed scheme ensures rapid and accurate CDI even in the early stage of a physical interaction. As a result, safety mechanisms can be triggered before further injuries are caused, which is quite valuable and important toward a safe human–robot collaboration. In the end, the proposed scheme is validated on a robot manipulator and applied to a demonstration task with collision reaction strategies. The experimental results reveal that the collisions are detected and classified within 20 ms with an overall accuracy of 99.6%, which confirms the applicability of the scheme to collaborative robots in practice.Note to Practitioners—This article is intended to provide a novel online collision event handling scheme for robots in industrial environments. This scheme is designed to quickly and accurately detect an accidental collision and distinguish it from the intentional human–robot interaction. The method takes the raw signals from external torque sensors and provides a collision diagnosis result with a reliability index. The simple structure makes it easy to be implemented as a regular fault monitoring routine for collaborative robots. Different from the conventional methods, the proposed collision identification scheme in this article especially focuses on overcoming the following two challenges in practice: first, to timely and accurately report a collision within its early stage, and second, to ensure a high identification accuracy in a complicated environment, where ubiquitous disturbance and noise are unneglectable. The experimental validation at the end of this article confirms its promising application value in human–robot collaboration. Zengjie Zhang, Kun Qian 0003, Björn W. Schuller, Dirk Wollherr |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2021 | Can Machine Learning Assist Locating the Excitation of Snore Sound? A ReviewabstractIn the past three decades, snoring (affecting more than 30 % adults of the UK population) has been increasingly studied in the transdisciplinary research community involving medicine and engineering. Early work demonstrated that, the snore sound can carry important information about the status of the upper airway, which facilitates the development of non-invasive acoustic based approaches for diagnosing and screening of obstructive sleep apnoea and other sleep disorders. Nonetheless, there are more demands from clinical practice on finding methods to localise the snore sound's excitation rather than only detecting sleep disorders. In order to further the relevant studies and attract more attention, we provide a comprehensive review on the state-of-the-art techniques from machine learning to automatically classify snore sounds. First, we introduce the background and definition of the problem. Second, we illustrate the current work in detail and explain potential applications. Finally, we discuss the limitations and challenges in the snore sound classification task. Overall, our review provides a comprehensive guidance for researchers to contribute to this area. Kun Qian 0003, Christoph Janott, Maximilian Schmitt, Zixing Zhang 0001, Clemens Heiser, Werner Hemmert, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Learning Higher Representations from Bioacoustics: A Sequence-to-Sequence Deep Learning Approach for Bird Sound Classification
Kun Qian 0003, Ziping Zhao 0001 |
ICONIP (5) | 2 |
| 2020 | An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and AnxietyabstractThe COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease. Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
INTERSPEECH | 2 |
| 2020 | Learning Higher Representations from Pre-Trained Deep Models with Data Augmentation for the COMPARE 2020 Challenge Mask TaskabstractHuman hand-crafted features are always regarded as expensive, time-consuming, and difficult in almost all of the machinelearning-related tasks.First, those well-designed features extremely rely on human expert domain knowledge, which may restrain the collaboration work across fields.Second, the features extracted in such a brute-force scenario may not be easy to be transferred to another task, which means a series of new features should be designed.To this end, we introduce a method based on a transfer learning strategy combined with data augmentation techniques for the COMPARE 2020 Challenge Mask Sub-Challenge.Unlike the previous studies mainly based on pre-trained models by image data, we use a pre-trained model based on large scale audio data, i. e., AudioSet.In addition, the SpecAugment and mixup methods are used to improve the generalisation of the deep models.Experimental results demonstrate that the best-proposed model can significantly (p < .001,by one-tailed z-test) improve the unweighted average recall (UAR) from 71.8 % (baseline) to 76.2 % on the test set.Finally, the best result, i. e., 77.5 % of the UAR on the test set, is achieved by a late fusion of the two best proposed models and the best single model in the baseline. Tomoya Koike, Kun Qian 0003, Björn W. Schuller, Yoshiharu Yamamoto |
INTERSPEECH | 2 |
| 2020 | Machine Listening for Heart Status Monitoring: Introducing and Benchmarking HSS - The Heart Sounds Shenzhen CorpusabstractAuscultation of the heart is a widely studied technique, which requires precise hearing from practitioners as a means of distinguishing subtle differences in heart-beat rhythm. This technique is popular due to its non-invasive nature, and can be an early diagnosis aid for a range of cardiac conditions. Machine listening approaches can support this process, monitoring continuously and allowing for a representation of both mild and chronic heart conditions. Despite this potential, relevant databases and benchmark studies are scarce. In this paper, we introduce our publicly accessible database, the Heart Sounds Shenzhen Corpus (HSS), which was first released during the recent INTERSPEECH 2018 ComParE Heart Sound sub-challenge. Additionally, we provide a survey of machine learning work in the area of heart sound recognition, as well as a benchmark for HSS utilising standard acoustic features and machine learning models. At best our support vector machine with Log Mel features achieves 49.7% unweighted average recall on a three category task (normal, mild, moderate/severe). Fengquan Dong, Kun Qian 0003, Zhao Ren, Alice Baird, Zhenyu Dai, Florian Metze, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Snore-GANs: Improving Automatic Snore Sound Classification With Synthesized DataabstractOne of the frontier issues that severely hamper the development of automatic snore sound classification (ASSC) associates to the lack of sufficient supervised training data. To cope with this problem, we propose a novel data augmentation approach based on semi-supervised conditional generative adversarial networks (scGANs), which aims to automatically learn a mapping strategy from a random noise space to original data distribution. The proposed approach has the capability of well synthesizing "realistic" high-dimensional data, while requiring no additional annotation process. To handle the mode collapse problem of GANs, we further introduce an ensemble strategy to enhance the diversity of the generated data. The systematic experiments conducted on a widely used Munich-Passau snore sound corpus demonstrate that the scGANs-based systems can remarkably outperform other classic data augmentation systems, and are also competitive to other recently reported systems for ASSC. Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Christoph Janott, Yanan Guo 0001, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Evolving Learning for Analysing Mood-Related Infant VocalisationabstractInfant vocalisation analysis plays an important role in the study of the development of pre-speech capability of infants, while machine-based approaches nowadays emerge with an aim to advance such an analysis.However, conventional machine learning techniques require heavy feature-engineering and refined architecture designing.In this paper, we present an evolving learning framework to automate the design of neural network structures for infant vocalisation analysis.In contrast to manually searching by trial and error, we aim to automate the search process in a given space with less interference.This framework consists of a controller and its child networks, where the child networks are built according to the controller's estimation.When applying the framework to the Interspeech 2018 Computational Paralinguistics (ComParE) Crying Subchallenge, we discover several deep recurrent neural network structures, which are able to deliver competitive results to the best ComParE baseline method. Zixing Zhang 0001, Jing Han 0010, Kun Qian 0003, Björn W. Schuller |
INTERSPEECH | 3 |
| 2018 | The INTERSPEECH 2018 Computational Paralinguistics Challenge: Atypical & Self-Assessed Affect, Crying & Heart BeatsabstractThe INTERSPEECH 2018 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Atypical Affect Sub-Challenge, four basic emotions annotated in the speech of handicapped subjects have to be classified; in the Self-Assessed Affect Sub-Challenge, valence scores given by the speakers themselves are used for a three-class classification problem; in the Crying Sub-Challenge, three types of infant vocalisations have to be told apart; and in the Heart Beats Sub-Challenge, three different types of heart beats have to be determined.We describe the Sub-Challenges, their conditions, and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by end-to-end learning, the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit for the first time in the challenge series. Björn W. Schuller, Stefan Steidl, Anton Batliner, Peter B. Marschik, Harald Baumeister, Fengquan Dong, Simone Hantke, Florian B. Pokorny, Eva-Maria Rathner, Katrin D. Bartl-Pokorny, Christa Einspieler, Dajie Zhang, Alice Baird, Shahin Amiriparian, Kun Qian 0003, Zhao Ren, Maximilian Schmitt, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 15 |
| 2017 | The INTERSPEECH 2017 Computational Paralinguistics Challenge: Addressee, Cold & SnoringabstractThe INTERSPEECH 2017 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: In the Addressee sub-challenge, it has to be determined whether speech produced by an adult is directed towards another adult or towards a child; in the Cold sub-challenge, speech under cold has to be told apart from ‘healthy’ speech; and in the Snoring subchallenge, four different types of snoring have to be classified. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, which include data-learnt feature representations by end-to-end learning with convolutional and recurrent neural networks, and bag-of-audiowords for the first time in the challenge series Björn W. Schuller, Stefan Steidl, Anton Batliner, Elika Bergelson, Jarek Krajewski, Christoph Janott, Andrei Amatuni, Marisa Casillas, Amanda Seidl, Melanie Soderstrom, Anne S. Warlaumont, Guillermo Hidalgo, Sebastian Schnieder, Clemens Heiser, Winfried Hohenhorst, Michael Herzog, Maximilian Schmitt, Kun Qian 0003, Yue Zhang 0014, George Trigeorgis, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 18 |
| 2016 | Wavelet features for classification of vote snore soundsabstractLocation and form of the upper airway obstruction is essential for a targeted therapy of obstructive sleep apnea (OSA). Utilizing snore sounds (SnS) to reveal the pathological characters of OSA patients has been the subject of scientific research for several decades. Fewer studies exist on the evaluation of SnS to identify the corresponding obstruction types in the upper airway. In this study, we propose a novel feature set based on wavelet transform with a support vector machine classifier to discriminate VOTE (velum, oropharyngeal lateral walls, tongue base and epiglottis) snore sounds labelled during drug-induced sleep endoscopy (DISE). Based on snore sound data collected from 24 snoring subjects, processed by a subject-independent 2-fold cross validation experiment, we can show that our wavelet features outperform the frequently-used acoustic features (formants, MFCC, power ratio, crest factor, fundamental frequency) at an WAR (weighted average recall) of 78.2 % and an UAR (unweighted average recall) of 71.2%, with an enhancement ranging from 5.1 % to 24.4% and 12.2% to 46.4% in WAR and UAR, respectively. Kun Qian 0003, Christoph Janott, Zixing Zhang 0001, Clemens Heiser, Björn W. Schuller |
ICASSP | 1 |
| 2015 | Automatic detection, segmentation and classification of snore related signals from overnight audio recordingabstractSnore related signals (SRS) have been found to carry important information about the snore source and obstruction site in the upper airway of an Obstructive Sleep Apnea/Hypopnea Syndrome (OSAHS) patient. An overnight audio recording of an individual subject is the preliminary and essential material for further study and diagnosis. Automatic detection, segmentation and classification of SRS from overnight audio recordings are significant in establishing a personal health database and in researching the area on a large scale. In this study, the authors focused on how to implement this intelligent method by combining acoustic signal processing with machine learning techniques. The authors proposed a systematic solution includes SRS events detection, classifier training, automatic segmentation and classification. An overnight audio recording of a severe OSAHS patient is taken as an example to demonstrate the feasibility of their method. Both the experimental data testing and subjective testing of 25 volunteers (17 males and 8 females) demonstrated that their method could be effective in automatic detection, segmentation and classification of the SRS from original audio recordings. Kun Qian 0003, Zhiyong Xu 0002, Huijie Xu, Zhao Zhao 0003 |
IET Signal Process. | 1 |
| 2013 | A Cloud Computing System for Snore Signals Processing
Jian Guo 0004, Kun Qian 0003, Zhaomeng Zhu, Gongxuan Zhang, Huijie Xu |
APPT | 2 |