EDBT 2026 Demo / reviewers in the wild / expert
Zhao Ren
dblp:176/8500
· DBLP profile ↗
44ranked-venue papers
9as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 22 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Federated Graph Recommendation with Semantic Guidance
Thi Minh Chau Nguyen, Hien Trang Nguyen, Van Ho-Long, Zhao Ren |
DAIS | 6 |
| 2026 | Leveraging Semi-Supervised Learning for Multimodal Hate Speech Data Annotation and Detection
Rathi Adarshi Rammohan, Zhao Ren, Dominik Puchala, Aleksandra Swiderska, Dennis Küster, Tanja Schultz |
LREC | 2 |
| 2026 | A review of instruction-guided image editingabstractThe rapid advancement of artificial intelligence (AI) systems such as large language models (LLMs) and multimodal learning frameworks has transformed digital content creation and manipulation. Traditional visual editing tools require significant expertise, limiting accessibility. Recent strides in instruction-guided editing have enabled intuitive interaction with visual content, using natural language as a bridge between user intent and complex editing operations. This survey reviews how AI-implemented instruction-guided image editing models – including approaches rooted in generative adversarial networks and diffusion models – empower users to achieve precise visual modifications without deep technical knowledge. By synthesizing over 100 publications, we examine multimodal integration for fine-grained content control, compare existing literature, and highlight how these AI applications support creative visual storytelling, design workflows, and multimedia production. We also identify key challenges to stimulate further research. Interested readers are encouraged to access our repository at https://github.com/tamlhp/awesome-instructional-editing . Thanh Tam Nguyen, Zhao Ren, Trinh Pham, Phi-Le Nguyen, Nguyen Quoc Viet Hung, Hongzhi Yin |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
Yi Chang 0004, Zhao Ren, Zhonghao Zhao, Thanh Tam Nguyen, Kun Qian 0003, Tanja Schultz, Björn W. Schuller |
INTERSPEECH | 2 |
| 2025 | DiffMV-ETS: Diffusion-based Multi-Voice Electromyography-to-Speech Conversion using Speaker-Independent Speech Training TargetsabstractElectromyography (EMG) signals have been investigated for novel voice prostheses to enable speech communication with silent articulation.In this work, we propose DiffMV-ETS, a multi-voice, diffusion-based EMG-to-speech system that converts EMG signals to speech in selectable voices.We evaluate it for scenarios where no speech of the speaker wearing EMG sensors is used for training.For this purpose, we introduce EMG-VCTK, a dataset containing EMG and audio recordings of sentences from the Voice Conversion Tool Kit corpus.We compare EMG models trained with audio of the same speaker, of auxiliary speakers, and of text-to-speech systems.Experiments indicate that models retain their intelligibility and naturalness when trained with synthetic speech.DiffMV-ETS enhances the speech naturalness and similarity to unseen voices.To the best of our knowledge, this is the first work to train multi-voice EMG-to-speech systems with speaker-independent targets. Kevin Scheck, Tom Dombeck, Zhao Ren, Peter Wu, Michael Wand 0002, Tanja Schultz |
INTERSPEECH | 3 |
| 2025 | Nonparametric Quantile Regression with ReLU-Activated Recurrent Neural NetworksabstractThis paper investigates nonparametric quantile regression using recurrent neural networks (RNNs) and sparse recurrent neural networks (SRNNs) to approximate the conditional quantile function, which is assumed to follow a compositional hierarchical interaction model. We show that RNN- and SRNN-based estimators with rectified linear unit (ReLU) activation and appropriately designed architectures achieve the optimal nonparametric convergence rate, up to a logarithmic factor, under stationary, exponentially $\boldsymbol{\beta}$-mixing processes. To establish this result, we derive sharp approximation error bounds for functions in the hierarchical interaction model using RNNs and SRNNs, exploiting their close connection to sparse feedforward neural networks (SFNNs). Numerical experiments and an empirical study on the Dow Jones Industrial Average (DJIA) further support our theoretical findings. Han Yu 0001, Lyumin Wu, Wenxin Zhou, Zhao Ren |
NeurIPS | 4 |
| 2025 | Privacy-preserving explainable AI: a surveyabstractAbstract As the adoption of explainable AI (XAI) continues to expand, the urgency to address its privacy implications intensifies. Despite a growing corpus of research in AI privacy and explainability, there is little attention on privacy-preserving model explanations. This article presents the first thorough survey about privacy attacks on model explanations and their countermeasures. Our contribution to this field comprises a thorough analysis of research papers with a connected taxonomy that facilitates the categorization of privacy attacks and countermeasures based on the targeted explanations. This work also includes an initial investigation into the causes of privacy leaks. Finally, we discuss unresolved issues and prospective research directions uncovered in our analysis. This survey aims to be a valuable resource for the research community and offers clear insights for those new to this domain. To support ongoing research, we have established an online resource repository, which will be continuously updated with new and relevant findings. Thanh Tam Nguyen, Zhao Ren, Phi-Le Nguyen, Hongzhi Yin, Nguyen Quoc Viet Hung |
Sci. China Inf. Sci. | 3 |
| 2025 | STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion RecognitionabstractSpeech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models. Yi Chang 0004, Zhao Ren, Zixing Zhang 0001, Xin Jing 0001, Kun Qian 0003, Xi Shao, Bin Hu 0001, Tanja Schultz, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | A Survey of Machine UnlearningabstractToday, computer systems hold large amounts of personal data. Yet while such an abundance of data allows breakthroughs in AI, and especially machine learning, its existence can be a threat to user privacy, and it can weaken the bonds of trust between humans and AI. Recent regulations now require that, on request, private information about a user must be removed both from computer systems and from machine learning models—this legislation is more colloquially called “the right to be forgotten.” While removing data from back-end databases should be straightforward, it is not sufficient in the AI context as machine learning models often “remember” the old data. Contemporary adversarial attacks on trained models have proven that we can learn whether an instance or an attribute belonged to the training data. This phenomenon calls for a new paradigm, namely machine unlearning , to make machine learning models forget about particular data. It turns out that recent works on machine unlearning have not been able to completely solve the problem due to the lack of common frameworks and resources. Therefore, this article aspires to present a comprehensive examination of machine unlearning’s concepts, designs, methods, and applications. Specifically, as a category collection of cutting-edge studies, the intention behind this article is to serve as a comprehensive resource for researchers and practitioners seeking an introduction to machine unlearning and its formulations, design criteria, removal requests, algorithms, and applications. In addition, we aim to highlight the key findings, current trends, and new research areas that have not yet featured the use of machine unlearning but could benefit greatly from it. We hope that this survey serves as a valuable resource for machine learning researchers and those seeking to innovate privacy technologies. Our resources are publicly available at https://github.com/tamlhp/awesome-machine-unlearning . Thanh Tam Nguyen, Zhao Ren, Phi-Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, Nguyen Quoc Viet Hung |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
Sheng Li 0010, Yang Cao 0011, Zhao Ren, Tanja Schultz |
MMAsia | 4 |
| 2024 | Portable graph-based rumour detection against multi-modal heterophilyabstractThe propagation of rumours on social media poses an important threat to societies, so that various techniques for graph-based rumour detection have been proposed recently. Existing works, however, are based on homophilic graphs: entities that are connected to each other often have the same label. However, recent studies found that heterophily is more common in real-world social networks, i.e., entities with different labels are also often linked to each other due to ‘innocent’ retweets or camouflage behaviours by malicious users. Especially, the heterophily problem is even more challenging in multi-modal social graphs, in which neighbouring entities might differ in terms of both labels and modalities. To cope with multi-modal homophily in graph-based rumour detection, we propose a Portable Graph Transformer-based Rumour Detection model (PHAROS) with novel multi-modal homophily measures. It integrates label information in the learning process, which enables us to generate discriminative neighbourhoods of entities. Our model can handle multiple modalities (a natural characteristic of social graphs) and is portable to be combined with existing graph-based models. Extensive experiments on real and synthetic data show the superiority, efficiency, robustness, and portability of PHAROS and its heterophily resilience. Thanh Tam Nguyen, Zhao Ren, Jun Jo 0001, Nguyen Quoc Viet Hung, Hongzhi Yin |
Knowl. Based Syst. | 2 |
| 2023 | Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured LearningabstractSpeech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep learning has been investigated to improve the performance of SER by training complex models, the memory space and computational capability of edge devices represents a constraint for embedding deep learning models. We propose a neural structured learning (NSL) framework through building synthesized graphs. An SER model is trained on a source dataset and used to build graphs on a target dataset. A relatively lightweight model is then trained with the speech samples and graphs together as the input. Our experiments demonstrate that training a lightweight SER model on the target dataset with speech samples and graphs can not only produce small SER models, but also enhance the model performance compared to models with speech samples only and those using classic transfer learning strategies. Yi Chang 0004, Zhao Ren, Thanh Tam Nguyen, Kun Qian 0003, Björn W. Schuller |
ICASSP | 2 |
| 2023 | Fast Yet Effective Speech Emotion Recognition with Self-DistillationabstractSpeech emotion recognition (SER) is the task of recognising humans’ emotional states from speech. SER is extremely prevalent in helping dialogue systems to truly understand our emotions and become a trustworthy human conversational partner. Due to the lengthy nature of speech, SER also suffers from the lack of abundant labelled data for powerful models like deep neural networks. Pre-trained complex models on large-scale speech datasets have been successfully applied to SER via transfer learning. However, fine-tuning complex models still requires large memory space and results in low inference efficiency. In this paper, we argue achieving a fast yet effective SER is possible with self-distillation, a method of simultaneously fine-tuning a pretrained model and training shallower versions of itself. The benefits of our self-distillation framework are threefold: (1) the adoption of self-distillation method upon the acoustic modality breaks through the limited ground-truth of speech data, and outperforms the existing models’ performance on an SER dataset; (2) executing powerful models at different depths can achieve adaptive accuracy-efficiency trade-offs on resource-limited edge devices; (3) a new fine-tuning process rather than training from scratch for self-distillation leads to faster learning time and the state-of-the-art accuracy on data with small quantities of label information. Zhao Ren, Thanh Tam Nguyen, Yi Chang 0004, Björn W. Schuller |
ICASSP | 1 |
| 2023 | Cutting Through the Noise: An Empirical Comparison of Psycho-Acoustic and Envelope-based Features for Machinery Fault DetectionabstractAcoustic-based fault detection has been one of the key instruments to monitor the health condition of mechanical parts. However, the background noise of an industrial environment may negatively influence the performance of fault detection. Limited attention has been paid to improving the robustness of fault detection against industrial environmental noise. Therefore, we present the Lenze production background-noise (LPBN) real-world dataset and an automated and noise-robust auditory inspection (ARAI) system for the end-of-line inspection of geared motors. An acoustic array is used to acquire data from motors with a minor fault, major fault, or which are healthy. A benchmark is provided to compare the psychoacoustic features with different types of envelope features based on expert knowledge of the gearbox. To the best of our knowledge, we are the first to apply time-varying psychoacoustic features for fault detection. We train a state-of-the-art one-class-classifier, on samples from healthy motors and separate the faulty ones for fault detection using a threshold. The best-performing approaches achieve an area under curve of 0.87 (logarithm envelope), 0.86 (time-varying psychoacoustics), and 0.91 (combination of both). Peter Wißbrock, Yvonne Richter, David Pelkmann, Zhao Ren, Gregory Palmer |
ICASSP | 4 |
| 2023 | 10X Faster Subgraph Matching: Dual Matching Networks with Interleaved Diffusion AttentionabstractThe goal of subgraph matching is to determine the presence of a particular query pattern within a large collection of data graphs. Despite being a hard problem, subgraph matching is essential in various disciplines, including bioinformatics, text matching, and graph retrieval. Although traditional approaches could provide exact solutions, their computations are known to be NP-complete, leading to an overwhelmingly querying latency. While recent neural-based approaches have been shown to improve the response time, the oversimplified assumption of the first-order network may neglect the generalisability of fully capturing patterns in varying sizes, causing the performance to drop significantly in datasets in various domains. To overcome these limitations, this paper proposes xDualSM, a dual matching neural network model with interleaved diffusion attention. Specifically, we first embed the structural information of graphs into different adjacency matrices, which explicitly capture the intra-graph and cross-graph structures between the query pattern and the target graph. Then, we introduce a dual matching network with interleaved diffusion attention to carefully capture intra-graph and cross-graph information while reducing computational complexity. Empirically, our proposed framework not only boosted the speed of subgraph matching more than 10x compared to the fastest baseline but also achieved significant improvements of 47.64% in Recall and 34.39% in F1-score compared to the state-of-the-art approximation approach on COX2 dataset. In addition, our results are comparable with exact methods. Zhao Ren, Jun Jo 0001, Nguyen Quoc Viet Hung, Thanh Tam Nguyen |
IJCNN | 3 |
| 2023 | Self-Explaining Neural Networks for Respiratory Sound Classification with Scale-free InterpretabilityabstractAnalysis of respiratory sounds is an area where deep neural networks (DNNs) may benefit clinicians and patients for diagnostic purposes due to their classification power. However, explaining the predictions made by DNNs remains a challenge. Currently, most explanation methods focus on post-hoc explanations, where a separate explanatory model is used to explain a trained DNN. Due to the complex nature of respiratory sound classification pipeline involving signal processing such as frequency analysis and wavelet analysis, post-hoc methods cannot uncover the underlying inference process of DNNs, highlighting the importance of designing DNNs with intrinsic interpretability. In this paper, we propose a self-explaining DNN for respiratory sound classification based on prototype learning. Our model explains its behavior by generating sample prototypes while attaching these prototypes to a layer inside the neural network. Furthermore, we design a scale-free interpretability mechanism, in which the model reaches its final decision by dissecting the input and looking for similarities between several parts of the input and the prototypes. The experimental findings on the largest public respiratory sound database demonstrate that our method achieves comparable, sometimes better, performance with the non-interpretable counterparts while offering state-of-the-art interpretability. The code will be released upon acceptance. Zhao Ren, Thanh Tam Nguyen, Mohammad Mehdi Zahedi, Wolfgang Nejdl |
IJCNN | 1 |
| 2023 | Guest Editorial Trustworthy and Collaborative AI for Personalised Healthcare Through Edge-of-ThingsabstractFrom diagnosis to therapies, the development of artificial intelligence (AI) has facilitated improvements in personalised healthcare applications. The evolution of AI in healthcare is closely related to the changes in the types and volumes of data which we need to deal with. The first generation of healthcare technologies, represented by the highly successful relational databases, are designed to handle structured data involving patient demographics, patient care, treatments, and outcomes of those treatments. Big Data platforms, which are representative of the current mainstream healthcare technologies, are built to process unstructured data from sources like electronic health records, medical imaging, genomic sequencing, and pharmaceutical research. The next generation of healthcare technologies will potentially be Edge-of-Things data, represented by massive amount of streaming data generated from Internet-of-Things frameworks, Cloud systems, and Edge computing platforms. Zhao Ren, Björn W. Schuller, Björn M. Eskofier, Thanh Tam Nguyen, Wolfgang Nejdl |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Prototype Learning for Interpretable Respiratory Sound AnalysisabstractRemote screening of respiratory diseases has been widely studied as a non-invasive and early instrument for diagnosis purposes, especially in the pandemic. The respiratory sound classification task has been realized with numerous deep neural network (DNN) models due to their superior performance. However, in the high-stake medical domain where decisions can have significant consequences, it is desirable to develop interpretable models; thus, providing understandable reasons for physicians and patients. To address the issue, we propose a prototype learning framework, that jointly generates exemplar samples for explanation and integrates these samples into a layer of DNNs. The experimental results indicate that our method outperforms the state-of-the-art approaches on the largest public respiratory sound database. Zhao Ren, Thanh Tam Nguyen, Wolfgang Nejdl |
ICASSP | 1 |
| 2022 | Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough SoundsabstractCovid-19 has caused a huge health crisis worldwide in the past two years. Although an early detection of the virus through nucleic acid screening can considerably reduce its spread, the efficiency of this diagnostic process is limited by its complexity and costs. Hence, an effective and inexpensive way to early detect Covid-19 is still needed. Considering that the cough of an infected person contains a large amount of information, we propose an algorithm for the automatic recognition of Covid-19 from cough signals. Our approach generates static log-Mel spectrograms with deltas and delta-deltas from the cough signal and subsequently extracts feature maps through a Convolutional Neural Network (CNN). Following the advances on transformers in the realm of deep learning, our proposed architecture exploits a novel adaptive position embedding structure which can learn the position information of the features from the CNN output. This make the transformer structure rapidly lock the attention feature location by overlaying with the CNN output, which yields better classification. The efficiency of the proposed architecture is shown by the improvement, w. r. t. the baseline, of our experimental results on the INTERPSEECH 2021 Computational Paralinguistics Challenge CCS (Coughing Sub Challenge) database, which reached 72.6 % UAR (Unweighted Average Recall). Tianhao Yan, Shuo Liu 0012, Emilia Parada-Cabaleiro, Zhao Ren, Björn W. Schuller |
ICASSP | 5 |
| 2022 | Example-based Explanations with Adversarial Attacks for Respiratory Sound AnalysisabstractRespiratory sound classification is an important tool for remote screening of respiratory-related diseases such as pneumonia, asthma, and COVID-19. To facilitate the interpretability of classification results, especially ones based on deep learning, many explanation methods have been proposed using prototypes. However, existing explanation techniques often assume that the data is non-biased and the prediction results can be explained by a set of prototypical examples. In this work, we develop a unified example-based explanation method for selecting both representative data (prototypes) and outliers (criticisms). In particular, we propose a novel application of adversarial attacks to generate an explanation spectrum of data instances via an iterative fast gradient sign method. Such unified explanation can avoid over-generalisation and bias by allowing human experts to assess the model mistakes case by case. We performed a wide range of quantitative and qualitative evaluations to show that our approach generates effective and understandable explanation and is robust with many deep learning models. Yi Chang 0004, Zhao Ren, Thanh Tam Nguyen, Wolfgang Nejdl, Björn W. Schuller |
INTERSPEECH | 2 |
| 2022 | High-dimension to high-dimension screening for detecting genome-wide epigenetic and noncoding RNA regulators of gene expressionabstractMOTIVATION: The advancement of high-throughput technology characterizes a wide variety of epigenetic modifications and noncoding RNAs across the genome involved in disease pathogenesis via regulating gene expression. The high dimensionality of both epigenetic/noncoding RNA and gene expression data make it challenging to identify the important regulators of genes. Conducting univariate test for each possible regulator-gene pair is subject to serious multiple comparison burden, and direct application of regularization methods to select regulator-gene pairs is computationally infeasible. Applying fast screening to reduce dimension first before regularization is more efficient and stable than applying regularization methods alone. RESULTS: We propose a novel screening method based on robust partial correlation to detect epigenetic and noncoding RNA regulators of gene expression over the whole genome, a problem that includes both high-dimensional predictors and high-dimensional responses. Compared to existing screening methods, our method is conceptually innovative that it reduces the dimension of both predictor and response, and screens at both node (regulators or genes) and edge (regulator-gene pairs) levels. We develop data-driven procedures to determine the conditional sets and the optimal screening threshold, and implement a fast iterative algorithm. Simulations and applications to long noncoding RNA and microRNA regulation in Kidney cancer and DNA methylation regulation in Glioblastoma Multiforme illustrate the validity and advantage of our method. AVAILABILITY AND IMPLEMENTATION: The R package, related source codes and real datasets used in this article are provided at https://github.com/kehongjie/rPCor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongjie Ke, Zhao Ren, Jianfei Qi, George C. Tseng, Zhenyao Ye, Tianzhou Ma |
Bioinform. | 2 |
| 2021 | Speaking Corona? Human and Machine Recognition of COVID-19 from VoiceabstractWith the COVID-19 pandemic, several research teams have reported successful advances in automated recognition of COVID-19 by voice. Resulting voice-based screening tools for COVID-19 could support large-scale testing efforts. While capabilities of machines on this task are progressing, we approach the so far unexplored aspect whether human raters can distinguish COVID-19 positive and negative tested speakers from voice samples, and compare their performance to a machine learning baseline. To account for the challenging symptom similarity between COVID-19 and other respiratory diseases, we use a carefully balanced dataset of voice samples, in which COVID-19 positive and negative tested speakers are matched by their symptoms alongside COVID-19 negative speakers without symptoms. Both human raters and the machine struggle to reliably identify COVID-19 positive speakers in our dataset. These results indicate that particular attention should be paid to the distribution of symptoms across all speakers of a dataset when assessing the capabilities of existing systems. The identification of acoustic aspects of COVID-19-related symptom manifestations might be the key for a reliable voice-based COVID-19 detection in the future by both trained human raters and machine learning models. Copyright ©2021 ISCA. Pascal Hecker, Florian B. Pokorny, Katrin D. Bartl-Pokorny, Uwe D. Reichel, Zhao Ren, Simone Hantke, Florian Eyben, Dagmar Schuller, Bert Arnrich, Björn W. Schuller |
Interspeech | 5 |
| 2021 | Representation transfer learning from deep end-to-end speech recognition networks for the classification of health states from speechabstractRepresentation transfer learning has been widely used across a range of machine learning tasks. One such notable approach seen in the speech literature is the use of Convolutional Neural Networks, pre-trained for image classification tasks, to extract features from spectrograms of speech signals. Interestingly, despite the strong performance of such approaches, there have been minimal research efforts exploring the suitability of using speech-specific networks to perform feature extraction. In this regard, a novel feature representation learning framework is presented herein. This approach is comprising the use of Automatic Speech Recognition (ASR) deep neural networks as feature extractors, the fusion of several extracted feature representations using Compact Bilinear Pooling (CBP), and finally inference via a specially optimised Recurrent Neural Network (RNN) classifier. To determine the usefulness of these feature representations, they are comprehensively tested on two representative speech-health classification tasks, namely the food-type being eaten and speaker intoxication. Key results indicate the promise of the extracted features, demonstrating comparable results to other state-of-the-art approaches in the literature. Benjamin Sertolli, Zhao Ren, Björn W. Schuller, Nicholas Cummins |
Comput. Speech Lang. | 2 |
| 2021 | Computer Audition for Fighting the SARS-CoV-2 Corona Crisis - Introducing the Multitask Speech Corpus for COVID-19abstractComputer audition (CA) has experienced a fast development in the past decades by leveraging advanced signal processing and machine learning techniques. In particular, for its noninvasive and ubiquitous character by nature, CA-based applications in healthcare have increasingly attracted attention in recent years. During the tough time of the global crisis caused by the coronavirus disease 2019 (COVID-19), scientists and engineers in data science have collaborated to think of novel ways in prevention, diagnosis, treatment, tracking, and management of this global pandemic. On the one hand, we have witnessed the power of 5G, Internet of Things, big data, computer vision, and artificial intelligence in applications of epidemiology modeling, drug and/or vaccine finding and designing, fast CT screening, and quarantine management. On the other hand, relevant studies in exploring the capacity of CA are extremely lacking and underestimated. To this end, we propose a novel multitask speech corpus for COVID-19 research usage. We collected 51 confirmed COVID-19 patients' in-the-wild speech data in Wuhan city, China. We define three main tasks in this corpus, i.e., three-category classification tasks for evaluating the physical and/or mental status of patients, i.e., sleep quality, fatigue, and anxiety. The benchmarks are given by using both classic machine learning methods and state-of-the-art deep learning techniques. We believe this study and corpus cannot only facilitate the ongoing research on using data science to fight against COVID-19, but also the monitoring of contagious diseases for general purpose. Kun Qian 0003, Maximilian Schmitt, Huaiyuan Zheng, Tomoya Koike, Jing Han 0010, Junjun Duan, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE Internet Things J. | 11 |
| 2021 | EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion EmbeddingsabstractDespite remarkable advances in emotion recognition, they are severely restrained from either the essentially limited property of the employed single modality, or the synchronous presence of all involved multiple modalities. Motivated by this, we propose a novel crossmodal emotion embedding framework called EmoBed, which aims to leverage the knowledge from other auxiliary modalities to improve the performance of an emotion recognition system at hand. The framework generally includes two main learning components, i.e., joint multimodal training and crossmodal training. Both of them tend to explore the underlying semantic emotion information but with a shared recognition network or with a shared emotion embedding space, respectively. In doing this, the enhanced system trained with this approach can efficiently make use of the complementary information from other modalities. Nevertheless, the presence of these auxiliary modalities is not demanded during inference. To empirically investigate the effectiveness and robustness of the proposed framework, we perform extensive experiments on the two benchmark databases RECOLA and OMG-Emotion for the tasks of dimensional emotion regression and categorical emotion classification, respectively. The obtained results show that the proposed framework significantly outperforms related baselines in monomodal inference, and are also competitive or superior to the recently reported systems, which emphasises the importance of the proposed crossmodal learning for emotion recognition. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | CAA-Net: Conditional Atrous CNNs With Attention for Explainable Device-Robust Acoustic Scene ClassificationabstractAcoustic Scene Classification (ASC) aims to classify the environment in which the audio signals are recorded. Recently, Convolutional Neural Networks (CNNs) have been successfully applied to ASC. However, the data distributions of the audio signals recorded with multiple devices are different. There has been little research on the training of robust neural networks on acoustic scene datasets recorded with multiple devices, and on explaining the operation of the internal layers of the neural networks. In this article, we focus on training and explaining device-robust CNNs on multi-device acoustic scene data. We propose conditional atrous CNNs with attention for multi-device ASC. Our proposed system contains an ASC branch and a device classification branch, both modelled by CNNs. We visualise and analyse the intermediate layers of the atrous CNNs. A time-frequency attention mechanism is employed to analyse the contribution of each time-frequency bin of the feature maps in the CNNs. On the Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 ASC dataset, recorded with three devices, our proposed model performs significantly better than CNNs trained on single-device data. Zhao Ren, Qiuqiang Kong, Jing Han 0010, Mark D. Plumbley, Björn W. Schuller |
IEEE Trans. Multim. | 1 |
| 2021 | Frustration recognition from speech during game interaction using wide residual networksabstractAlthough frustration is a common emotional reaction while playing games, an excessive level of frustration can negatively impact a user's experience, discouraging them from further game interactions. The automatic detection of frustration can enable the development of adaptive systems that can adapt a game to a user's specific needs through real-time difficulty adjustment, thereby optimizing the player's experience and guaranteeing game success. To this end, we present a speech-based approach for the automatic detection of frustration during game interactions, a specific task that remains underexplored in research. The experiments were performed on the Multimodal Game Frustration Database (MGFD), an audiovisual dataset—collected within the Wizard-of-Oz framework—that is specially tailored to investigate verbal and facial expressions of frustration during game interactions. We explored the performance of a variety of acoustic feature sets, including Mel-Spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and the low-dimensional knowledge-based acoustic feature set eGeMAPS. Because of the continual improvements in speech recognition tasks achieved by the use of convolutional neural networks (CNNs), unlike the MGFD baseline, which is based on the Long Short-Term Memory (LSTM) architecture and Support Vector Machine (SVM) classifier—in the present work, we consider typical CNNs, including ResNet, VGG, and AlexNet. Furthermore, given the unresolved debate on the suitability of shallow and deep networks, we also examine the performance of two of the latest deep CNNs: WideResNet and EfficientNet. Our best result, achieved with WideResNet and Mel-Spectrogram features, increases the system performance from 58.8% unweighted average recall (UAR) to 93.1% UAR for speech-based automatic frustration recognition. Meishu Song, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Zijiang Yang 0007, Shuo Liu 0012, Zhao Ren, Ziping Zhao 0001, Björn W. Schuller |
Virtual Real. Intell. Hardw. | 6 |
| 2020 | Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition ModelsabstractThe development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this study, we aim to train deep models to defending against non-targeted white-box adversarial attacks. Adversarial data is first generated from the real data using the fast gradient sign method. Then in the research field of speech emotion recognition, adversarial-based training is employed as a method for protecting against adversarial attack. We then train deep convolutional models with both real and adversarial data, and compare the performances of two adversarial training procedures - namely, vanilla adversarial training, and similarity-based adversarial training. In our experiments, through the use of adversarial data augmentation, both of the considered adversarial training procedures can improve the performance when validated on the real data. Additionally, the similarity-based adversarial training learns a more robust model when working with adversarial data. Finally, the considered VGG-16 model performs the best across all models, for both real and generated data. Zhao Ren, Alice Baird, Jing Han 0010, Zixing Zhang 0001, Björn W. Schuller |
ICASSP | 1 |
| 2020 | An Early Study on Intelligent Analysis of Speech Under COVID-19: Severity, Sleep Quality, Fatigue, and AnxietyabstractThe COVID-19 outbreak was announced as a global pandemic by the World Health Organisation in March 2020 and has affected a growing number of people in the past few weeks.In this context, advanced artificial intelligence techniques are brought to the fore in responding to fight against and reduce the impact of this global health crisis.In this study, we focus on developing some potential use-cases of intelligent speech analysis for COVID-19 diagnosed patients.In particular, by analysing speech recordings from these patients, we construct audio-onlybased models to automatically categorise the health state of patients from four aspects, including the severity of illness, sleep quality, fatigue, and anxiety.For this purpose, two established acoustic feature sets and support vector machines are utilised.Our experiments show that an average accuracy of .69obtained estimating the severity of illness, which is derived from the number of days in hospitalisation.We hope that this study can foster an extremely fast, low-cost, and convenient way to automatically detect the COVID-19 disease. Jing Han 0010, Kun Qian 0003, Meishu Song, Zijiang Yang 0007, Zhao Ren, Shuo Liu 0012, Huaiyuan Zheng, Tomoya Koike, Zixing Zhang 0001, Yoshiharu Yamamoto, Björn W. Schuller |
INTERSPEECH | 5 |
| 2020 | Squeeze for Sneeze: Compact Neural Networks for Cold and Flu RecognitionabstractIn digital health applications, speech offers advantages over other physiological signals, in that it can be easily collected, transmitted, and stored using mobile and Internet of Things (IoT) technologies.However, to take full advantage of this positioning, speech-based machine learning models need to be deployed on devices that can have considerable memory and power constraints.These constraints are particularly apparent when attempting to deploy deep learning models, as they require substantial amounts of memory and data movement operations.Herein, we test the suitability of pruning and quantisation as two methods to compress the overall size of neural networks trained for a health-driven speech classification task. Key results presented on the Upper Respiratory Tract InfectionCorpus indicate that pruning, then quantising a network can reduce the number of operational weights by almost 90 %.They also demonstrate the overall size of the network can be reduced by almost 95 %, as measured in MB, without affecting overall recognition performance. Merlin Albes, Zhao Ren, Björn W. Schuller, Nicholas Cummins |
INTERSPEECH | 2 |
| 2020 | A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia RecognitionabstractContains fulltext : 228158.pdf (Publisher’s version ) (Open Access) Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä |
INTERSPEECH | 3 |
| 2020 | Towards Speech Robustness for Acoustic Scene ClassificationabstractThis work discusses the impact of human voice on acoustic scene classification (ASC) systems.Typically, such systems are trained and evaluated on data sets lacking human speech.We show experimentally that the addition of speech can be detrimental to system performance.Furthermore, we propose two alternative solutions to mitigate that effect in the context of deep neural networks (DNNs).We first utilise data augmentation to make the algorithm robust against the presence of human speech in the data.We also introduce a voice-suppression algorithm that removes human speech from audio recordings, and test the DNN classifier on those denoised samples.Experimental results show that both approaches reduce the negative effects of human voice in ASC systems.Compared to using data augmentation, applying voice suppression achieved better classification accuracy and managed to perform more stably for different speech intensity. Shuo Liu 0012, Andreas Triantafyllopoulos, Zhao Ren, Björn W. Schuller |
INTERSPEECH | 3 |
| 2020 | Enhancing Transferability of Black-Box Adversarial Attacks via Lifelong Learning for Speech Emotion Recognition ModelsabstractWell-designed adversarial examples can easily fool deep speech emotion recognition models into misclassifications.The transferability of adversarial attacks is a crucial evaluation indicator when generating adversarial examples to fool a new target model or multiple models.Herein, we propose a method to improve the transferability of black-box adversarial attacks using lifelong learning.First, black-box adversarial examples are generated by an atrous Convolutional Neural Network (CNN) model.This initial model is trained to attack a CNN target model.Then, we adapt the trained atrous CNN attacker to a new CNN target model using lifelong learning.We use this paradigm, as it enables multi-task sequential learning, which saves more memory space than conventional multi-task learning.We verify this property on an emotional speech database, by demonstrating that the updated atrous CNN model can attack all target models which have been learnt, and can better attack a new target model than an attack model trained on one target model only. Zhao Ren, Jing Han 0010, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 1 |
| 2020 | Latent Dynamic Factor Analysis of High-Dimensional Neural RecordingsabstractHigh-dimensional neural recordings across multiple brain regions can be used to establish functional connectivity with good spatial and temporal resolution. We designed and implemented a novel method, Latent Dynamic Factor Analysis of High-dimensional time series (LDFA-H), which combines (a) a new approach to estimating the covariance structure among high-dimensional time series (for the observed variables) and (b) a new extension of probabilistic CCA to dynamic time series (for the latent variables). Our interest is in the cross-correlations among the latent variables which, in neural recordings, may capture the flow of information from one brain region to another. Simulations show that LDFA-H outperforms existing methods in the sense that it captures target factors even when within-region correlation due to noise dominates cross-region correlation. We applied our method to local field potential (LFP) recordings from 192 electrodes in Prefrontal Cortex (PFC) and visual area V4 during a memory-guided saccade task. The results capture time-varying lead-lag dependencies between PFC and V4, and display the associated spatial distribution of the signals. Heejong Bong, Zongge Liu, Zhao Ren, Matthew A. Smith 0001, Valérie Ventura, Robert E. Kass |
NeurIPS | 3 |
| 2020 | Machine Listening for Heart Status Monitoring: Introducing and Benchmarking HSS - The Heart Sounds Shenzhen CorpusabstractAuscultation of the heart is a widely studied technique, which requires precise hearing from practitioners as a means of distinguishing subtle differences in heart-beat rhythm. This technique is popular due to its non-invasive nature, and can be an early diagnosis aid for a range of cardiac conditions. Machine listening approaches can support this process, monitoring continuously and allowing for a representation of both mild and chronic heart conditions. Despite this potential, relevant databases and benchmark studies are scarce. In this paper, we introduce our publicly accessible database, the Heart Sounds Shenzhen Corpus (HSS), which was first released during the recent INTERSPEECH 2018 ComParE Heart Sound sub-challenge. Additionally, we provide a survey of machine learning work in the area of heart sound recognition, as well as a benchmark for HSS utilising standard acoustic features and machine learning models. At best our support vector machine with Log Mel features achieves 49.7% unweighted average recall on a three category task (normal, mild, moderate/severe). Fengquan Dong, Kun Qian 0003, Zhao Ren, Alice Baird, Zhenyu Dai, Florian Metze, Yoshiharu Yamamoto, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono ModalityabstractDespite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the training procedure for either speech or facial emotion recognition. Specifically, the model consists of one modality-specific network per individual modality and one shared network to map both audio and visual cues into final predictions. In the training process, we additionally take the loss from one auxiliary modality into account besides the main modality. To evaluate the effectiveness of the implicit fusion model, we conduct extensive experiments for mono-modal emotion classification and regression, and find that the implicit fusion models outperform the standard mono-modal training process. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Björn W. Schuller |
ICASSP | 3 |
| 2019 | Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic ScenesabstractThe goal of Acoustic Scene Classification (ASC) is to recognise the environment in which an audio waveform has been recorded. Recently, deep neural networks have been applied to ASC and have achieved state-of-the-art performance. However, few works have investigated how to visualise and understand what a neural network has learnt from acoustic scenes. Previous work applied local pooling after each convolutional layer, therefore reduced the size of the feature maps. In this paper, we suggest that local pooling is not necessary, but the size of the receptive field is important. We apply atrous Convolutional Neural Networks (CNNs) with global attention pooling as the classification model. The internal feature maps of the attention model can be visualised and explained. On the Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 dataset, our proposed method achieves an accuracy of 72.7 %, significantly outperforming the CNNs without dilation at 60.4 %. Furthermore, our results demonstrate that the learnt feature maps contain rich information on acoustic scenes in the time-frequency domain. Zhao Ren, Qiuqiang Kong, Jing Han 0010, Mark D. Plumbley, Björn W. Schuller |
ICASSP | 1 |
| 2018 | Towards Conditional Adversarial Training for Predicting Emotions from SpeechabstractMotivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consists of two networks, trained in an adversarial manner: The first network tries to predict emotion from acoustic features, while the second network aims at distinguishing between the predictions provided by the first network and the emotion labels from the database using the acoustic features as conditional information. We evaluate the performance of the proposed conditional adversarial training framework on the widely used emotion database RECOLA. Experimental results show that the proposed training strategy outperforms the conventional training method, and is comparable with, or even superior to other recently reported approaches, including deep and end-to-end learning. Jing Han 0010, Zixing Zhang 0001, Zhao Ren, Fabien Ringeval, Björn W. Schuller |
ICASSP | 3 |
| 2018 | Bags in Bag: Generating Context-Aware Bags for Tracking Emotions from SpeechabstractInternational audience Jing Han 0010, Zixing Zhang 0001, Maximilian Schmitt, Zhao Ren, Fabien Ringeval, Björn W. Schuller |
INTERSPEECH | 4 |
| 2018 | The INTERSPEECH 2018 Computational Paralinguistics Challenge: Atypical & Self-Assessed Affect, Crying & Heart BeatsabstractThe INTERSPEECH 2018 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Atypical Affect Sub-Challenge, four basic emotions annotated in the speech of handicapped subjects have to be classified; in the Self-Assessed Affect Sub-Challenge, valence scores given by the speakers themselves are used for a three-class classification problem; in the Crying Sub-Challenge, three types of infant vocalisations have to be told apart; and in the Heart Beats Sub-Challenge, three different types of heart beats have to be determined.We describe the Sub-Challenges, their conditions, and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by end-to-end learning, the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit for the first time in the challenge series. Björn W. Schuller, Stefan Steidl, Anton Batliner, Peter B. Marschik, Harald Baumeister, Fengquan Dong, Simone Hantke, Florian B. Pokorny, Eva-Maria Rathner, Katrin D. Bartl-Pokorny, Christa Einspieler, Dajie Zhang, Alice Baird, Shahin Amiriparian, Kun Qian 0003, Zhao Ren, Maximilian Schmitt, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 16 |
| 2018 | SILGGM: An extensive R package for efficient statistical inference in large-scale gene networksabstractGene co-expression network analysis is extremely useful in interpreting a complex biological process. The recent droplet-based single-cell technology is able to generate much larger gene expression data routinely with thousands of samples and tens of thousands of genes. To analyze such a large-scale gene-gene network, remarkable progress has been made in rigorous statistical inference of high-dimensional Gaussian graphical model (GGM). These approaches provide a formal confidence interval or a p-value rather than only a single point estimator for conditional dependence of a gene pair and are more desirable for identifying reliable gene networks. To promote their widespread use, we herein introduce an extensive and efficient R package named SILGGM (Statistical Inference of Large-scale Gaussian Graphical Model) that includes four main approaches in statistical inference of high-dimensional GGM. Unlike the existing tools, SILGGM provides statistically efficient inference on both individual gene pair and whole-scale gene pairs. It has a novel and consistent false discovery rate (FDR) procedure in all four methodologies. Based on the user-friendly design, it provides outputs compatible with multiple platforms for interactive network visualization. Furthermore, comparisons in simulation illustrate that SILGGM can accelerate the existing MATLAB implementation to several orders of magnitudes and further improve the speed of the already very efficient R package FastGGM. Testing results from the simulated data confirm the validity of all the approaches in SILGGM even in a very large-scale setting with the number of variables or genes to a ten thousand level. We have also applied our package to a novel single-cell RNA-seq data set with pan T cells. The results show that the approaches in SILGGM significantly outperform the conventional ones in a biological sense. The package is freely available via CRAN at https://cran.r-project.org/package=SILGGM. Zhao Ren, Wei Chen 0074 |
PLoS Comput. Biol. | 2 |
| 2017 | Extending the FOV from disparity and color consistencies in multiview light fieldsabstractLight field, which is captured by a plenoptic camera, is always limited in its narrow field of view (FOV) by the physical size of the aperture. To break through the restriction, we propose to extend the FOV using multiview light fields. A series of light fields are acquired by translating the camera at isometric spatial positions. In contrast to previous methods, our algorithm is the first that achieves light field registration and rendering based on epipolar plane image (EPI) properties, including disparity and color consistencies. Furthermore, the aliasing caused by the under-sampling in the angular space is eliminated by synthesizing novel views in the EPI space. Experimental results on the real scene data have demonstrated the effectiveness of our algorithm. Zhao Ren, Qi Zhang 0029, Hao Zhu 0005, Qing Wang 0006 |
ICIP | 1 |
| 2016 | Sincerity and Deception in Speech: Two Sides of the Same Coin? A Transfer- and Multi-Task Learning PerspectiveabstractIn this work, we investigate the coherence between inferable deception and perceived sincerity in speech, as featured in the Deception and Sincerity tasks of the INTERSPEECH 2016 Computational Paralinguistics ChallengE (ComParE).We demonstrate an effective approach that combines the corpora of both Challenge tasks to achieve higher classification accuracy.We show that the naïve label mapping method based on the assumption that sincerity and deception are just 'two sides of the same coin', i. e., taking deceptive speech as equivalent to non-sincere speech and vice versa, does not yield satisfactory results.However, we can exploit the interplay and synergies between these characteristics.To achieve this, we combine our previously introduced approach for data aggregation by semi-supervised cross-task label completion with multi-task learning, and knowledge-based instance selection.In the result, our approach achieves significant error rate reductions compared to the official Challenge baseline. Yue Zhang 0014, Felix Weninger, Zhao Ren, Björn W. Schuller |
INTERSPEECH | 3 |
| 2016 | FastGGM: An Efficient Algorithm for the Inference of Gaussian Graphical Model in Biological NetworksabstractBiological networks provide additional information for the analysis of human diseases, beyond the traditional analysis that focuses on single variables. Gaussian graphical model (GGM), a probability model that characterizes the conditional dependence structure of a set of random variables by a graph, has wide applications in the analysis of biological networks, such as inferring interaction or comparing differential networks. However, existing approaches are either not statistically rigorous or are inefficient for high-dimensional data that include tens of thousands of variables for making inference. In this study, we propose an efficient algorithm to implement the estimation of GGM and obtain p-value and confidence interval for each edge in the graph, based on a recent proposal by Ren et al., 2015. Through simulation studies, we demonstrate that the algorithm is faster by several orders of magnitude than the current implemented algorithm for Ren et al. without losing any accuracy. Then, we apply our algorithm to two real data sets: transcriptomic data from a study of childhood asthma and proteomic data from a study of Alzheimer's disease. We estimate the global gene or protein interaction networks for the disease and healthy samples. The resulting networks reveal interesting interactions and the differential networks between cases and controls show functional relevance to the diseases. In conclusion, we provide a computationally fast algorithm to implement a statistically sound procedure for constructing Gaussian graphical model and making inference with high-dimensional biological data. The algorithm has been implemented in an R package named "FastGGM". Ting Wang 0003, Zhao Ren, Ying Ding 0003, Zhe Sun 0007, Matthew L. MacDonald, Robert A. Sweet, Jieru Wang, Wei Chen 0074 |
PLoS Comput. Biol. | 2 |