EDBT 2026 Demo / reviewers in the wild / expert
Tiago H. Falk
dblp:27/1611 · also Tiago Henrique Falk
· DBLP profile ↗
109ranked-venue papers
18as first author
30since 2021 · last 2025
0000-0002-5739-2514ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 14 first-author · 12 since 2021Artificial intelligence and machine learning · 37 · 9 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 28 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 8 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Domain Adversarial Neural Networks with Adversarial Robustness Evaluation for Intrusion Detection Systems
Ines Guerziz, Tiago H. Falk, Long Bao Le, Zakaria Abou El Houda |
CRiSIS | 2 |
| 2025 | Towards Lightweight On-Device Audio Deepfake Detection Using Squeezeformers
Ashkan Moradi, Yi Zhu 0011, Tiago H. Falk |
CRiSIS | 3 |
| 2025 | Selective Shift: Towards Personalized Domain Adaptation in Multi-Agent Collaborative PerceptionabstractGiven the scarcity of real data and the time-intensive nature of labeling, current multi-agent perception models often rely on simulated sensor data for training and validation. However, perception performance deteriorates significantly due to domain gap between simulated and real data. Existing adaptation methods focus on domain-generalized feature extraction while neglecting multi-agent shift uncertainty and relational semantic loss. To address this issue, we propose a Selective Shift Domain Adaptation method in multi-agent collaborative perception, called SSDA. SSDA incorporates two essential components: the frequency-decoupled feature shift adjustment (FSA) and the entropy-driven staged adaptive alignment (SAA). To mitigate the relational semantic loss, the FSA is proposed to simplify the representation of correlation features and remove redundant information from the source domain, thereby mitigating interference for domain adversarial scenarios. To tackle the shift uncertainty, the SAA is designed to achieve adaptive alignment from global to local guided by information entropy, which dynamically adjusts weights for samples according to their level of uncertainty. The results demonstrate that the SSDA is significantly superior to the SOTA, achieving up to 7.35% improvements on [email protected]. Hui Zhang 0091, Yiteng Xu, Yonglin Tian, Yidong Li, Tiago H. Falk, Fei-Yue Wang 0001 |
ACM Multimedia | 5 |
| 2025 | Subjective and Objective Evaluation of the Benefits of Multisensory Virtual Nature Immersion for Patients with Post-Traumatic Stress DisorderabstractMultisensory virtual nature immersion has been shown to assist with mental health management. In this paper, we look at the benefits of two multisensory immersion modalities for patients diagnosed with post-traumatic stress disorder (PTSD). The modalities include: (1) a portable solution that integrates scent diffusion within the head-mounted display (HMD), and (2) a pod-based solution, in which a multisensory pod provides olfactory and haptic feedback (i.e., wind and vibrations) alongside audio-visual information via the HMD. We assess the two methods subjectively, via numerous questionnaires, and objectively, via measures computed from electroencephalography signals. The pod-based intervention resulted in greater improvements in PTSD symptoms and in sustained attention, with improvements in several objective measures of user experience, including engagement, valence, arousal, and sense of immersion/presence. Marilia Karla Soares Lopes, Léa Perreault, Belmir Jose de Jesus, Marie-Claude Roberge, Tiago H. Falk |
QoMEX | 5 |
| 2025 | Quantifying the Risk of Private Information Leakage in the Metaverse with EEG-Instrumented Virtual Reality HeadsetsabstractAs virtual reality (VR) technologies become more immersive and metaverse applications burgeon, modern headsets are increasingly becoming equipped with sensors capable of capturing multiple neurophysiological signals. While these signals can be used to measure, in real-time, quality of experience metrics that can be used to enhance interactivity and user engagement, they may also introduce novel privacy risks by unintentionally leaking sensitive personal attributes. In this paper, we explore the extent in which electroencephalography (EEG) signals, recorded during an immersive VR memory task, can be used to infer users’ private information, such as age, biological sex, and identity. We employ both classical machine learning models with hand-crafted features, as well as end-to-end deep learning approaches. Our findings demonstrate that EEG-based features can, indeed, leak information about biological sex, age, and user identity, with end-to-end models obtaining the best performance. Feature importance ranking and deep neural network saliency maps were used to provide explainability on the neural patterns used by the models. We conclude with recommendations on how these findings can also be used to help secure future metaverse applications. Mina Jaberi, Stéphane Bouchard, Tiago H. Falk |
SMC | 3 |
| 2025 | Audio-Visual Cross-Attention for Improved Deepfake Video Detection and Forgery LocalizationabstractWith the emergence of multi-modal generative models, synthesized videos are becoming increasingly realistic, making the detection of deepfakes extremely challenging. While several video deepfake detection models have shown promising performance, their focus has been primarily on the visual modality. To overcome this limitation, we propose a dualstream framework that fuses visual and auditory information via cross-attention computed between embeddings extracted from pre-trained video and audio encoders. Additionally, we design a weakly-supervised forgery localization head that infers frame-level forgery scores from coarse segment-level labels, minimizing the need for fine-grained annotations and allowing for forgery location characterization. In this paper, we describe our preliminary results showing the proposed model outperforming state-of-the-art detectors on both frame-level localization and sequence-level deepfake detection tasks. Ongoing work focuses on investigating the complementarity between the visual and auditory modalities to improve model robustness and explainability. Oussama Jalleli, Yi Zhu 0011, Tiago H. Falk |
SMC | 3 |
| 2025 | Electroencephalography Neuromarkers to Predict the Response of a Multisensory Virtual Reality Nature Immersion Intervention for Patients Diagnosed with Post-Traumatic Stress DisorderabstractImmersive virtual reality (VR) applications rapidly expand across domains, including training, gaming, and healthcare. More recently, multisensory immersive experiences, including olfactory and haptic stimulation, have emerged and shown great promise, especially for interventions in well-being and mental health management. Multisensory experiences, however, are very subjective (e.g., one subject may like certain smells, while others do not), and recent results have suggested that some participants may not respond positively to the treatment. As multisensory VR interventions can be costly and time-consuming for both patients and clinicians, being able to find neuromarkers that predict intervention outcomes would be invaluable. Here, we aim to take the first steps in the development of a neuromarker to predict the response to a multisensory nature immersion VR intervention. A pilot experiment was performed with twenty patients diagnosed with post-traumatic stress disorder. Potential neuromarkers are extracted from electroencephalography (EEG) signals measured from an instrumented VR headset. We show that some EEG patterns start to differ between responders and non-responders as early as the fourth session, i.e., one-third of the way into the entire intervention. This suggests that neuromarkers to predict the outcomes of a multisensory VR immersion intervention may exist. These markers could be used not only to save time and resources for clinicians and patients but also to promote precision treatment where interventions are adjusted to each patient, maximizing success rates. Belmir Jose de Jesus, Marilia Karla Soares Lopes, Léa Perreault, Marie-Claude Roberge, Alcyr Alves de Oliveira, Tiago H. Falk |
SMC | 6 |
| 2025 | Enhancing Network Intrusion Detection Systems: A Multi-Layer Ensemble Approach to Mitigate Adversarial AttacksabstractAdversarial examples can represent a serious threat to machine learning (ML) algorithms. If used to manipulate the behaviour of ML-based Network Intrusion Detection Systems (NIDS), they can jeopardize network security. In this work, we aim to mitigate such risks by increasing the robustness of NIDS towards adversarial attacks. To that end, we explore two adversarial methods for generating malicious network traffic. The first method is based on Generative Adversarial Networks (GAN) and the second one is the Fast Gradient Sign Method (FGSM). The adversarial examples generated by these methods are then used to evaluate a novel multilayer defense mechanism, specifically designed to mitigate the vulnerability of ML-based NIDS. Our solution consists of one layer of stacking classifiers and a second layer based on an autoencoder. If the incoming network data are classified as benign by the first layer, the second layer is activated to ensure that the decision made by the stacking classifier is correct. We also incorporated adversarial training to further improve the robustness of our solution. Experiments on two datasets, namely UNSW-NB15 and NSL-KDD, demonstrate that the proposed approach increases resilience to adversarial attacks. Nasim Soltani, Shayan Nejadshamsi, Zakaria Abou El Houda, Raphaël Khoury, Kelton A. P. Costa, Tiago H. Falk, Anderson R. Avila |
SMC | 6 |
| 2025 | DeepSick: Deceiving Voice-Based Diagnostic Models with Synthetic Multilingual Pathological Speech SignalsabstractVoice-based diagnostic systems offer a scalable solution for remote health assessment. However, recent advances in generative voice models may enable malicious manipulation of voice samples to simulate or conceal disease-related speech characteristics, which poses new risks to diagnostic systems. This paper investigates the vulnerability of diagnostic and detection models to such types of "deepfake" attacks. We show that it is possible to train a generative model to convert between healthy voices and pathological ones, which in turn, can successfully deceive existing diagnostic systems. Here, focus is placed on COVID-19 infection and respiratory abnormalities, but the method can be applied across different pathological conditions affecting vocal attributes. We also benchmark four state-of-the-art synthesized voice detection models on both real and generated pathological speech from three datasets. Our results show that current synthetic voice detectors, typically trained on healthy speech data, perform poorly on generated pathological samples. While fine-tuning with real pathological voices improves detection, a substantial performance gap remains. This work provides initial insights on an emerging threat to remote voice diagnostic systems that needs further work. Yi Zhu 0011, Alan Davoust, Tiago H. Falk |
SMC | 3 |
| 2025 | Benchmarking Self-Supervised Audio Representations for IoT-Enabled Acoustic Beehive Monitoring
Heitor R. Guimarães, Mahsa Abdollahi, Yi Zhu 0011, Ségolène Maucourt, Nico Coallier, Pierre Giovenazzo, Tiago H. Falk |
IEEE Internet Things J. | 7 |
| 2025 | WavRx: A Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic ModelabstractSpeech is known to carry health-related attributes, which has emerged as a novel venue for remote and long-term health monitoring. However, existing models are usually tailored for a specific type of disease, and have been shown to lack generalizability across datasets. Furthermore, concerns have been raised recently towards the leakage of speaker identity from health embeddings. To mitigate these limitations, we propose WavRx, a speech health diagnostics model that captures the respiration and articulation related dynamics from a universal speech representation. Our in-domain and cross-domain experiments on six pathological speech datasets demonstrate WavRx as a new state-of-the-art health diagnostic model. Furthermore, we show that the amount of speaker identity entailed in the WavRx health embeddings is significantly reduced without extra guidance during training. An in-depth analysis of the model was performed, thus providing physiological interpretation of its improved generalizability and privacy-preserving ability. Yi Zhu 0011, Tiago H. Falk |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Promoting Engagement in Remote Patient Monitoring Using Asynchronous MessagingabstractRemote patient monitoring is becoming increasingly instrumental to healthcare delivery but can substantially hamper the interpersonal communication that underlies standard clinical practice. In this work, we explore the benefits imparted to patients, clinicians, and researchers by an asynchronous messaging feature within a platform called COVIDFree@Home. We created COVIDFree@Home to assist the healthcare system in a large metropolitan city in North America during the COVID-19 pandemic. Clinicians used COVIDFree@Home to monitor the self-reported symptoms and vital signs of over 350 COVID-19 patients post-infection. Using thematic analysis of user-initiated messages, we found the messaging feature helped maintain protocol adherence while allowing patients to ask questions about their health and clinicians to convey empathetic care. This feedback cycle also led to higher quality data for hospitalization prediction, as the revisions significantly improved the AUROC of a machine learning model trained on demographic variables, vital signs data, and self-reported symptoms from 0.53 to 0.59. Salaar Liaqat, Daniyal Liaqat, Tatiana Son, Tiago H. Falk, Robert Wu 0002, Andrea Gershon, Eyal de Lara, Alexander Mariakakis |
CHI | 4 |
| 2024 | VIC-KD: Variance-Invariance-Covariance Knowledge Distillation to Make Keyword Spotting More Robust Against Adversarial AttacksabstractKeyword spotting (KWS) refers to the task of identifying a set of predefined words in audio streams. With the advances seen recently with deep neural networks, it has become a popular technology to activate and control small devices, such as voice assistants. Relying on such models for edge devices, however, can be challenging due to hardware constraints. Moreover, as adversarial attacks have increased against voice-based technologies, developing solutions robust to such attacks has become crucial. In this work, we propose VIC-KD, a robust distillation recipe for model compression and adversarial robustness. Using self-supervised speech representations, we show that imposing geometric priors to the latent representations of both Teacher and Student models leads to more robust target models. Experiments on the Google Speech Commands datasets show that the proposed methodology improves upon current state-of-the-art robust distillation methods, such as ARD and RSLAD, by 12% and 8% in robust accuracy, respectively. Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila, Tiago H. Falk |
ICASSP | 4 |
| 2024 | Spectral-temporal saliency masks and modulation tensorgrams for generalizable COVID-19 detectionabstractSpeech COVID-19 detection systems have gained popularity as they represent an easy-to-use and low-cost solution that is well suited for at-home long-term monitoring of patients with persistent symptoms. Recently, however, the limited generalization capability of existing deep neural network based systems to unseen datasets has been raised as a serious concern, as has their limited interpretability. In this study, we aim to develop an interpretable and generalizable speech-based COVID-19 detection system. First, we propose the use of a 3-dimensional modulation frequency tensor (called modulation tensorgram representation, MTR) as input to a convolutional recurrent neural network for COVID-19 detection. The MTR representation is known to capture long-term dynamics of speech correlated with articulation and respiration, hence being a potential candidate for characterizing COVID-19 speech. The customized network explores both the spectral and temporal pattern from MTR to learn the underlying COVID-19 speech pattern. Next, we design a spectro-temporal saliency masking to aggregate regions of the MTR related to COVID-19, thus helping further improve the generalizability and interpretability of the model. Experiments are conducted on three public datasets and results show the proposed solution consistently outperforming two benchmark systems in within-, across-, and unseen-dataset tests. The learned salient regions have been shown correlated with whispered speech and vocal hoarseness, which explains the increased generalizability. Furthermore, our model relies on a small amount of parameters, thus offering a promising solution for on-device remote monitoring of COVID-19 infection. Yi Zhu 0011, Tiago H. Falk |
Comput. Speech Lang. | 2 |
| 2024 | On the Impact of Voice Anonymization on Speech Diagnostic Applications: A Case Study on COVID-19 DetectionabstractWith advances seen in deep learning, voice-based applications are burgeoning, ranging from personal assistants, affective computing, to remote disease diagnostics. As the voice contains both linguistic and para-linguistic information (e.g., vocal pitch, intonation, speech rate, loudness), there is growing interest in voice anonymization to preserve speaker privacy and identity. Voice privacy challenges have emerged over the last few years and focus has been placed on removing speaker identity while keeping linguistic content intact. For affective computing and disease monitoring applications, however, the para-linguistic content may be more critical. Unfortunately, the effects that anonymization may have on these systems are still largely unknown. In this paper, we fill this gap and focus on one particular health monitoring application: speech-based COVID-19 diagnosis. We test three anonymization methods and their impact on five different state-of-the-art COVID-19 diagnostic systems using three public datasets. We validate the effectiveness of the anonymization methods, compare their computational complexity, and quantify the impact across different testing scenarios for both within- and across-dataset conditions. Additionally, we provided a comprehensive evaluation of the importance of different speech aspects for diagnostics and showed how they are affected by different types of anonymizers. Lastly, we show the benefits of using anonymized external data as a data augmentation tool to help recover some of the COVID-19 diagnostic accuracy loss seen with anonymization. Yi Zhu 0011, Mohamed Imoussaïne-Aïkous, Carolyn Côté-Lussier, Tiago H. Falk |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Robustdistiller: Compressing Universal Speech Representations for Enhanced Environment RobustnessabstractSelf-supervised speech pre-training enables deep neural network models to capture meaningful and disentangled factors from raw waveform signals. The learned universal speech representations can then be used across numerous down-stream tasks. These representations, however, are sensitive to distribution shifts caused by environmental factors, such as noise and/or room reverberation. Their large sizes, in turn, make them unfeasible for edge applications. In this work, we propose a knowledge distillation methodology termed RobustDistiller which compresses universal representations while making them more robust against environmental artifacts via a multi-task learning objective. The proposed layer-wise distillation recipe is evaluated on top of three well-established universal representations, as well as with three downstream tasks. Experimental results show the proposed methodology applied on top of the WavLM Base+ teacher model outperforming all other benchmarks across noise types and levels, as well as reverberation times. Oftentimes, the obtained results with the student model (24M parameters) achieved results inline with those of the teacher model (95M). Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila, Mehdi Rezagholizadeh, Boxing Chen, Tiago H. Falk |
ICASSP | 6 |
| 2023 | On the Importance of Different Cough Phases for COVID-19 DetectionabstractCough is an important symptom of numerous respiratory diseases, including COVID-19. While different cough phases (i.e., inhalation, compression, and expulsion) have been shown to be related to different pathological origins, existing cough-based COVID-19 detection systems rely on the entire cough recording, thus such phase-related characteristics are overlooked. In this study, our aim is two-fold. First, we have annotated over 1,250 cough recordings from two publicly-available cough sound databases, thus providing the research community with fine-grained cough phase labels. Next, we extract a number of temporal and acoustic features from each cough phase and test their usefulness and complementarity for COVID-19 detection. Experiments show the importance of cough phase segmentation, not only for improved COVID-19 detection, but also for the development of models that are interpretable and can better generalize across datasets. Yi Zhu 0011, Mahil Hussain Shaik, Tiago H. Falk |
ICASSP | 3 |
| 2023 | On the Use of FOOOF for Electroencephalography Quality Measurement and Device AssessmentabstractElectroencephalography (EEG) signals capture the electrical activity of the brain and have traditionally been used in clinical settings. More recently, with the development of mobile, wearable EEG devices their use has been explored for other applications, including emotion recognition, quality of experience monitoring, or fatigue detection, just to name a few. EEG signals are known to have a 1/f-noise like structure and are very low amplitude, thus making them highly susceptible to artefacts, such as power line interference, muscle movement, and eye blinks. Moreover, prototyping and development of new devices to record EEG signals may introduce additional sources of artefacts generated by different instrumentation settings. As such, automated quantification of the quality of the EEG signals has become important and the focus of recent research. Here, we propose a new quality metric based on the 1/f-noise structure of the EEG signal. Experimental results show the proposed metric classifying clean versus noisy EEG segments in a subject-independent setting with an accuracy of 86.0% for the AF7 electrode location and 64.6% for AF8, two electrode locations known to be highly degraded by artefacts. Additionally, the proposed metrics are shown to generalize well to unseen electrode locations. For example, a quality model trained on AF7 noisy EEG data achieved an accuracy of 61.4% when tested on data collected from the AF8 location. EEG, quality, fooof, prototyping Abhishek Tiwari 0003, Gloria Wu, Katrina Innanen, Amin Mahnam, Bastien Moineau, Tiago H. Falk |
SMC | 6 |
| 2023 | Assessing the Vulnerability of Self-Supervised Speech Representations for Keyword Spotting Under White-Box Adversarial AttacksabstractSelf-supervised speech pre-training has emerged as a useful tool to extract representations from speech that can be used across different tasks. While these models are starting to appear in commercial systems, their robustness to so-called adversarial attacks have yet to be fully characterized. This paper evaluates the vulnerability of three self-supervised speech representations (wav2vec 2.0, HuBERT and WavLM) to three white-box adversarial attacks under different signal-to-noise ratios (SNR). The study uses keyword spotting as a downstream task and shows that the models are very vulnerable to attacks, even at high SNRs. The paper also investigates the transferability of attacks between models and analyses the generated noise patterns in order to develop more effective defence mechanisms. The modulation spectrum shows to be a potential tool for detection of adversarial attacks to speech systems. Heitor R. Guimarães, Yi Zhu 0011, Orson Mengara, Anderson R. Avila, Tiago H. Falk |
SMC | 5 |
| 2023 | COVID-19 Detection via Fusion of Modulation Spectrum and Linear Prediction Speech FeaturesabstractThe coronavirus disease 2019 (COVID-19) pandemic has drastically impacted life around the globe. As life returns to pre-pandemic routines, COVID-19 testing has become a key component, assuring that travellers and citizens are free from the disease. Conventional tests can be expensive, time-consuming (results can take up to 48h), and require laboratory testing. Rapid antigen testing, in turn, can generate results within 15-30 minutes and can be done at home, but research shows they achieve very poor sensitivity rates. In this paper, we propose an alternative test based on speech signals recorded at home with a portable device. It has been well-documented that the virus affects many of the speech production systems (e.g., lungs, larynx, and articulators). As such, we propose the use of new modulation spectral features and linear prediction analysis to characterize these changes and design a two-stage COVID-19 prediction system by fusing the proposed features. Experiments with three COVID-19 speech datasets (CSS, DiCOVA2, and Cambridge subset) show that the two-stage feature fusion system outperforms the benchmark systems of CSS and Cambridge datasets while maintaining lower complexity compared to DL-based systems. Furthermore, the two-stage system demonstrates higher generalizability to unseen conditions in a cross-dataset testing evaluation scheme. The generalizability and interpretability of our proposed system demonstrate the potential for accessible, low-cost, at-home COVID-19 testing. Yi Zhu 0011, Abhishek Tiwari 0003, João Monteiro 0002, Shruti Rajendra Kshirsagar, Tiago H. Falk |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Fusion of Modulation Spectral and Spectral Features with Symptom Metadata for Improved Speech-Based Covid-19 DetectionabstractExisting speech-based coronavirus disease 2019 (COVID-19) detection systems provide poor interpretability and limited robustness to unseen data conditions. In this paper, we propose a system to overcome these limitations. In particular, we propose to fuse two different feature modalities with patient metadata in order to capture different properties of the disease. The first feature set is based on modulation spectral properties of speech. The second comprises spectral shape/descriptor features recently used for COVID-19 detection. Lastly, we fuse patient metadata in order to improve robustness and interpretability. Experiments are performed on the 2021 INTERSPEECH COVID Speech Sub-Challenge dataset with several different data partitioning paradigms. Results show the importance of the modulation spectral features. Metadata, in turn, did not perform very well when used alone but provided invaluable insights when fused with the other features. Overall, a system relying on the fusion of all three modalities showed to be robust to unseen conditions and to rely on interpretable features. The simplicity of the model suggests that it can be deployed in portable devices, hence providing accessible COVID-19 diagnostics worldwide. Yi Zhu 0011, Tiago H. Falk |
ICASSP | 2 |
| 2022 | Multisensory Immersive Experiences: A Pilot Study on Subjective and Instrumental Human Influential Factors AssessmentabstractRecent technological advances have allowed for virtual reality applications to burgeon. With virtual reality (VR), so-called human influential factors play a crucial role in the final perceived immersive media experience (IMEx). While two individuals can use the same VR headset, play the same game in the same location, and have the same goals, the two individuals can have very different experiences, with varying perceptions of immersion, presence, realism, engagement, and cybersickness. This can be particularly true in multisensory immersive experiences where, in addition to audio-visual stimuli, olfactory and haptic feedback can be used. In this paper, we describe a pilot study in which a VR game was developed to combine audio-visual, olfactory, and haptic feedback to the user in real-time. After game play, participants were asked about their IMEx using five scales: realism, immersion, presence, engagement, and overall quality of experience (QoE). Moreover, using an instrumented VR headset we measure electroencephalography (EEG), electrocardiography (ECG), and electrooculography (EOG) signals and compute several instrumental measures of human influential factors, including an engagement index, arousal and valence indices, frontal alpha asymmetry, heart rate, several EEG sub-band powers, and eye blink rate. Using the subjective ratings, we measure the contribution that each IMEx subscale has on overall QoE, as a function of the type of sensory stimuli used. Results on 11 participants suggest very different contributions once smells and haptics are incorporated, relative to traditional audio-visual experiences. We also report on several instrumental measures that showed significant correlations with the IMEx subscales, suggesting that, in the future, real-time instrumental QoE measurement of multisensory experiences could be possible. Reza Amini Gougeh, Tiago H. Falk |
QoMEX | 2 |
| 2022 | Human Influential Factors Assessment During At-Home Gaming with an Instrumented VR HeadsetabstractHuman influential factors (HIFs) play a crucial role in a gamer's perceived immersive media experience, especially in virtual reality (VR) applications where cybersickness can be present. Typically, questionnaires have been used to gauge a gamer's experience focusing on factors such as sense of presence and engagement, to name a few. Questionnaires, however, can be time consuming and do not allow real-time assessment. To this end, affective/physiological computing tools have emerged as potential instrumental measures of HIFs. Affective tools, however, typically require the user to be in controlled laboratory settings. Therefore, the results may have limited transferability to everyday “in the wild” settings. This paper aims to fill this gap. First, we developed a plug-and-play instrumented VR headset capable of measuring multiple physiological signals, including electroencephalography (EEG), electrocardiography (ECG), and electro-oculography (EOG). Instrumented headsets were dropped off at the homes of eight novice VR gamers, together with a gaming laptop and a data streaming laptop, and players were asked to play Half-Life Alyx at the comfort of their homes. At the end of the game, users were asked to answer a questionnaire about their experience, including perceived sense of immersion, presence, emotions, realism, and engagement using 10- point scales. Several HIF biometrics were then extracted from the neurophysiological signals and shown to significantly correlate with the questions related to experience features. These findings suggest that real-time VR gamer experience measurement can be possible in highly ecological settings. Marc-Antoine Moinnereau, Alcyr Alves de Oliveira, Tiago H. Falk |
QoMEX | 3 |
| 2022 | Multi-level self-attentive TDNN: A general and efficient approach to summarize speech into discriminative utterance-level representations
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
Speech Commun. | 3 |
| 2022 | Quality-Aware Bag of Modulation Spectrum Features for Robust Speech Emotion RecognitionabstractAutomatic speech emotion recognition (SER) has gained popularity over the last decade and numerous Challenges have emerged. While the latest Challenges have shown that deep neural networks achieve the best results, existing input features are still a bottleneck and cause severe performance degradation in realistic “in-the-wild” scenarios. In this paper, we propose two innovations to tackle this issue. First, we propose to combine the bag-of-audio-words methodology with modulation spectrum features for environmental robustness. Second, we take advantage of the inherent quality-awareness properties of modulation spectrum and propose the use of a quality feature as an additional feature to be used by the speech emotion recognizer. Experiments are conducted with three multi-lingual speech datasets used in recent SER Challenges degraded by different noise sources and levels, and room reverberation. Experimental results show the proposed features i) consistently outperforming benchmark systems, ii) providing complementary information to classical features, hence improving performance with feature fusion, and iii) showing robustness against environment and language mismatch. Moreover, we show that when the proposed system is provided with quality information, further improvements are obtained. Overall, the proposed bag of modulation spectrum features are shown to be a promising candidate for “in-the-wild” SER. Shruti Rajendra Kshirsagar, Tiago H. Falk |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Context-Aware Speech Stress Detection in Hospital Workers Using Bi-LSTM ClassifiersabstractHospital workers are known to work long hours in a highly stressful environment. The COVID-19 pandemic has increased this burden multi-fold. Pre-COVID statistics already showed that one in every three nurses reported burnout, thus affecting patient satisfaction and the quality of their provided service. Real-time monitoring of burnout, and other underlying factors, such as stress, could provide feedback not only to the clinical staff, but also to hospital administrators, thus allowing for supportive measures to be taken early. In this paper, we present a context-aware speech-based system for stress detection. We consider data from 144 hospital workers who were monitored during their daily shifts over a 10-week period; subjective stress readings were collected daily. Wearable devices measured speech features and physiological readings, such as heart rate. Environment sensors, in turn, were used to track staff movement within the hospital. Here, we show the importance of context-awareness for stress level detection based on a bidirectional LSTM deep neural network. In particular, we show the importance of hospital location and circadian rhythm based contextual cues for stress prediction. Overall, we show improvements as high as 14% in F1 scores once context is incorporated, relative to using the speech features alone. Amr Gaballah, Abhishek Tiwari 0003, Shri Narayanan, Tiago H. Falk |
ICASSP | 4 |
| 2021 | Parametric Audio-visual Quality Measurement Based on Interpretable Fuzzy LogicabstractIn this paper, a fuzzy inference based system for the quality of experience assessment is presented. In particular, two models have been proposed to assess the perceived quality of experience. The first one is a global model which estimates the overall audiovisual quality without going through the individual evaluations of visual and auditory modalities while the other model has been created by merging fuzzy inference systems based on separate auditory and visual objective quality scores. Two different sets of parameters have been tested for the second model leading to two different measures. The contribution of this research lies in the application of the fuzzy inference logic to estimate the audiovisual quality of multimedia data. The experimental results on a publicly available quality dataset show competitive predictive performances of the proposed measures when compared to two existing audiovisual metrics based on random forests regression and multilayer perceptron. Fatima Boudjerida, Atidel Lahoulou, Tiago H. Falk, Zahid Akhtar |
QoMEX | 3 |
| 2021 | On the use of blind channel response estimation and a residual neural network to detect physical access attacks to speaker verification systems
Anderson R. Avila, Jahangir Alam 0001, Fabiano O. Costa Prado, Douglas D. O'Shaughnessy, Tiago H. Falk |
Comput. Speech Lang. | 5 |
| 2021 | Automatic speaker verification from affective speech using Gaussian mixture model based estimation of neutral speech characteristics
Anderson R. Avila, Douglas D. O'Shaughnessy, Tiago H. Falk |
Speech Commun. | 3 |
| 2021 | Feature Pooling of Modulation Spectrum Features for Improved Speech Emotion Recognition in the WildabstractInterest in affective computing is burgeoning, in great part due to its role in emerging affective human-computer interfaces (HCI). To date, the majority of existing research on automated emotion analysis has relied on data collected in controlled environments. With the rise of HCI applications on mobile devices, however, so-called “in-the-wild” settings have posed a serious threat for emotion recognition systems, particularly those based on voice. In this case, environmental factors such as ambient noise and reverberation severely hamper system performance. In this paper, we quantify the detrimental effects that the environment has on emotion recognition and explore the benefits achievable with speech enhancement. Moreover, we propose a modulation spectral feature pooling scheme that is shown to outperform a state-of-the-art benchmark system for environment-robust prediction of spontaneous arousal and valence emotional primitives. Experiments on an environment-corrupted version of the RECOLA dataset of spontaneous interactions show the proposed feature pooling scheme, combined with speech enhancement, outperforming the benchmark across different noise-only, reverberation-only and noise-plus-reverberation conditions. Additional tests with the SEWA database show the benefits of the proposed method for in-the-wild applications. Anderson R. Avila, Zahid Akhtar, João Felipe Santos, Douglas D. O'Shaughnessy, Tiago H. Falk |
IEEE Trans. Affect. Comput. | 5 |
| 2020 | An Ensemble Based Approach for Generalized Detection of Spoofing Attacks to Automatic Speaker RecognizersabstractAs automatic speaker recognizer systems become mainstream, voice spoofing attacks are on the rise. Common attack strategies include replay, the use of text-to-speech synthesis, and voice conversion systems. While previouslyproposed end-to-end detection frameworks have shown to be effective in spotting attacks for one particular spoofing strategy, they have relied on different models, architectures, and speech representations, depending on the spoofing strategy. In practice, however, one does not have a priori information regarding the strategy an attacker might employ to fool a speaker recognizer, thus it is necessary to devise approaches which are able to detect attacks regardless of the strategy employed to generate them. In this work, we introduce an end-to-end ensemble based approach such that two models - previously shown to perform well on each considered attack strategy - are trained jointly, while a third model learns how to mix their outputs yielding a single score. Experimental results with replay and text-to-speech/voice conversion attacks show the proposed ensemble method achieving similar or superior performance when compared to systems specialized on each spoofing strategy separately. João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
ICASSP | 3 |
| 2020 | An end-to-end approach for the verification problem: learning the right distanceabstractIn this contribution, we augment the metric learning setting by introducing a parametric pseudo-distance, trained jointly with the encoder. Several interpretations are thus drawn for the learned distance-like model’s output. We first show it approximates a likelihood ratio which can be used for hypothesis tests, and that it further induces a large divergence across the joint distributions of pairs of examples from the same and from different classes. Evaluation is performed under the verification setting consisting of determining whether sets of examples belong to the same class, even if such classes are novel and were never presented to the model during training. Empirical evaluation shows such method defines an end-to-end approach for the verification problem, able to attain better performance than simple scorers such as those based on cosine similarity and further outperforming widely used downstream classifiers. We further observe training is much simplified under the proposed approach compared to metric learning with actual distances, requiring no complex scheme to harvest pairs of examples. João Monteiro 0002, Isabela Albuquerque, Jahangir Alam 0001, R. Devon Hjelm, Tiago H. Falk |
ICML | 5 |
| 2020 | On The Performance of Time-Pooling Strategies for End-to-End Spoken Language IdentificationabstractAutomatic speech processing applications often have to deal with the problem of aggregating local descriptors (i.e., representations of input speech data corresponding to specific portions across the time dimension) and turning them into a single fixed-dimension representation, known as global descriptor, on top of which downstream classification tasks can be performed. In this paper, we provide an empirical assessment of different time pooling strategies when used with state-of-the-art representation learning models. In particular, insights are provided as to when it is suitable to use simple statistics of local descriptors or when more sophisticated approaches are needed. Here, language identification is used as a case study and a database containing ten oriental languages under varying test conditions (short-duration test recordings, confusing languages, unseen languages) is used. Experiments are performed with classifiers trained on top of global descriptors to provide insights on open-set evaluation performance and show that appropriate selection of such pooling strategies yield embeddings able to outperform well-known benchmark systems as well as previously results based on attention only. João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
LREC | 3 |
| 2020 | Initial Investigation into Neurophysiological Correlates of Argentine Tango Flow States: a Case StudyabstractArgentine tango has been shown to help psychological and physical health by reducing perceived levels of depression and stress, similar and at times better than meditation. An often reported experience in Argentine tango is "flow", which is described as total involvement from the dancers. While this state has been self-reported by experienced dancers while dancing, it has yet to be quantified in real-time. However, with the emergence of portable and wearable devices for the acquisition of physiological signals such as electroencephalography (EEG) and electrocardiography (ECG), and recent innovations in EEG artifact removal algorithms, this quantification may now be possible. In this work presents a case study where we aim to first validate the potential of recording usable EEG and ECG data from dancers while dancing, in an unobtrusive manner, as well as investigate the existence of neurophysiological correlates of Argentine tango flow. Raymundo Cassani, Abhishek Tiwari 0003, Ilona Posner, Bruno Afonso, Tiago H. Falk |
SMC | 5 |
| 2020 | Saccadic Eye Movement Classification Using ExG Sensors Embedded into a Virtual Reality HeadsetabstractMeasuring saccadic eye movements when wearing a virtual reality (VR) head-mounted display (HMD) has recently gained a lot of attention, as it allows for enriched user experiences. This has led to an increase in devices showcasing camera-based eye tracking capabilities. Such devices, however, can be orders of magnitude more expensive than conventional off-the-shelf HMDs. In this study, we explore the use of low-cost sensors embedded directly into the faceplate of the HMD to measure electroencephalography (EEG) and electrooculography (EOG) signals. In a "do-it-yourself" manner, we rely on the openBCI biosignal amplifier for data acquisition. A 7-channel system was tested on four participants who attended visually to a moving target in their field-of-view that moved every 10 degrees over a circumference. Time series and handcrafted features were extracted from the measured ExG signals and served as input to two different classifiers: support vector machine (SVM) and a multilayer perceptron (MLP). A hierarchical classification approach was proposed and found to achieve the best results with the fusion of both features sets, resulting in an average accuracy of 76.51% with an SVM. The results are encouraging and suggest that accurate, low-cost classification of saccadic eye movements may be possible. Marc-Antoine Moinnereau, Alcyr Alves de Oliveira, Tiago H. Falk |
SMC | 3 |
| 2020 | Movement Artifact-Robust Mental Workload Assessment During Physical Activity Using Multi-Sensor FusionabstractMental workload assessment is of great importance for safety critical applications, especially in situations that involve physical demands, such as with first responders (e.g., paramedics, firefighters, or police officers). Advancements in physiological signal monitoring with wearable sensors have made way for real-time mental workload assessment using physiological signals. However, these models have typically been conducted in controlled laboratory settings and rely on a single physiological modality. As a result, such models often experience a drop in performance due to movement artifacts introduced in real-life conditions. In this paper, we demonstrate that a multi-modal mental workload model not only improves measurement accuracy, but can also increase robustness against physical activity artifacts. To this end, an experiment was conducted where mental workload and physical activity levels were modulated simultaneously while physiological data was collected from 48 participants using off-the-shelf wearable devices. Results show improved mental workload assessment with multi-modal fusion under varying physical activity conditions. Abhishek Tiwari 0003, Raymundo Cassani, Jean-François Gagnon, Daniel Lafond, Sébastien Tremblay, Tiago H. Falk |
SMC | 6 |
| 2020 | Generalized end-to-end detection of spoofing attacks to automatic speaker recognizers
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
Comput. Speech Lang. | 3 |
| 2020 | Alzheimer's Disease Diagnosis and Severity Level Detection Based on Electroencephalography Modulation Spectral "Patch" FeaturesabstractOver the last two decades, electroencephalography (EEG) has emerged as a reliable tool for the diagnosis of cortical disorders such as Alzheimer's disease (AD). Typically, resting-state EEG (rsEEG) signals have been used, and traditional frequency bands (delta, theta, alpha, beta and gamma) have been explored. Recent studies, however, have suggested that non-conventional bands may lead to improved diagnostic performance. In this work, we propose a new type of features derived from the 2-dimensional modulation spectral domain representation of the rsEEG signal in order to characterize the neuromodulatory deficit emergent with AD. The proposed features are computed as the power in specific "patches" or regions of interest in the power modulation spectrogram, which are shown to be highly discriminant of AD severity levels. The proposed features were compared with traditional features used in the rsEEG AD monitoring literature. Results showed the proposed features not only achieving improved performance at discriminating between healthy normal elderly controls (Nold) and AD patients with varying severity levels, but also at monitoring severity levels (i.e., mild AD versus moderate AD). Moreover, the proposed features were shown to outperform traditional rsEEG features. Finally, we validated the biological origin of the proposed features by using source localization and comparing the obtained results with ones reported in the AD literature. Raymundo Cassani, Tiago H. Falk |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Multi-objective training of Generative Adversarial Networks with multiple discriminatorsabstractRecent literature has demonstrated promising results for training Generative Adversarial Networks by employing a set of discriminators, in contrast to the traditional game involving one generator against a single adversary. Such methods perform single-objective optimization on some simple consolidation of the losses, e.g. an arithmetic average. In this work, we revisit the multiple-discriminator setting by framing the simultaneous minimization of losses provided by different models as a multi-objective optimization problem. Specifically, we evaluate the performance of multiple gradient descent and the hypervolume maximization algorithm on a number of different datasets. Moreover, we argue that the previously proposed methods and hypervolume maximization can all be seen as variations of multiple gradient descent in which the update direction can be computed efficiently. Our results indicate that hypervolume maximization presents a better compromise between sample quality and computational cost than previous methods. Isabela Albuquerque, João Monteiro 0002, Thang Doan, Breandan Considine, Tiago H. Falk, Ioannis Mitliagkas |
ICML | 5 |
| 2019 | Blind Channel Response Estimation for Replay Attack Detection
Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 4 |
| 2019 | Combining Speaker Recognition and Metric Learning for Speaker-Dependent Representation Learning
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
INTERSPEECH | 3 |
| 2019 | Intrusive Quality Measurement of Noisy and Enhanced Speech based on i-Vector SimilarityabstractIn this paper, the i-vector framework is investigated as an intrusive quality measure for noisy and enhanced speech. While widely used across numerous speech applications, the potential of using i-vectors to summarize the quality of a speech recording has been overlooked. This paper aims to fill this gap. We show that the i-vector framework is well-suited for assessing speech signal quality, surpassing well-established instrumental measures such as the Perceptual Evaluation of Speech Quality (PESQ) and Perceptual Objective Listening Quality Analysis (POLQA). Three datasets are used in our experiments. First, the TIMIT database is used to train the i-vector extractor on clean speech. To evaluate the proposed method, the noisy speech corpus (NOIZEUS) and the evaluation set of the 2014 IEEE REVERB challenge are used with both containing subjective ratings of perceived quality. Correlations with mean opinion scores (MOS) as high as 0.90 are achieved. Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
QoMEX | 4 |
| 2019 | Towards the development of a non-intrusive objective quality measure for DNN-enhanced speechabstractRecently, several works have focused on leveraging the advances of deep neural networks (DNN) to a variety of domains, including speech enhancement. While advances in instrumental quality metrics have been made, particularly for enhanced speech, there is still relatively little research assessing how useful such metrics are for DNN-enhanced speech. This work aims to fill this gap. We performed online listening tests using the outputs of three different DNN-based speech enhancement models for both denoising and dereverberation. When assessing the predictive power of several objective metrics, we found that existing non-intrusive methods fail at monitoring signal quality. To overcome this limitation, we propose a new metric based on a combination of a handful of relevant acoustic features. Results inline with those obtained with intrusive measures are then attained. In a leave-one-model-out test, the proposed non-intrusive metric is also shown to outperform two non-intrusive benchmarks for all three DNN enhancement methods, showing the proposed method is capable of generalizing to unseen models. João Felipe Santos, Tiago H. Falk |
QoMEX | 2 |
| 2019 | Cross-Subject Statistical Shift Estimation for Generalized Electroencephalography-based Mental Workload AssessmentabstractAssessment of mental workload in real world conditions is key to ensure the performance of workers executing tasks which demand sustained attention. Previous literature has employed electroencephalography (EEG) to this end. However, EEG correlates of mental workload vary across subjects and physical strain, thus making it difficult to devise models capable of simultaneously presenting reliable performance across users. The field of domain adaptation (DA) aims at developing methods that allow for generalization across different domains by learning domain-invariant representations. Such DA methods, however, rely on the so-called covariate shift assumption, which typically does not hold for EEG-based applications. As such, in this paper we propose a way to measure the statistical (marginal and conditional) shift observed on data obtained from different users and use this measure to quantitatively assess the effectiveness of different adaptation strategies. In particular, we use EEG data collected from individuals performing a mental task while running in a treadmill and explore the effects of different normalization strategies commonly used to mitigate cross-subject variability. We show the effects that different normalization schemes have on statistical shifts and their relationship with the accuracy of mental workload prediction as assessed on unseen participants at train time. Isabela Albuquerque, João Monteiro 0002, Olivier Rosanne, Abhishek Tiwari 0003, Jean-François Gagnon, Tiago H. Falk |
SMC | 6 |
| 2019 | Automated Alzheimer's Disease Diagnosis using a Low-Density EEG Layout and New Features based on the Power of Modulation Spectral "Patches"abstractOver the last decade, biomarkers to detect Alzheimer's disease (AD) have been proposed based on electroencephalography (EEG), in particular resting-state EEG (rsEEG). Typically, medium-or high-density research-grade EEG layouts with 16+ electrodes have been utilized and features extracted from traditional frequency subbands. As the quality of EEG sensors have gone up in recent years and low-density devices are emerging, it is not clear if accurate diagnosis can still be achieved with such solutions. Moreover, recent research has suggested that decomposing EEG into the five traditional frequency subbands may not be optimal for AD. In this paper, we describe a first attempt at developing an automated AD diagnostic system based on a low-density EEG layout (7 channels) directly from raw EEG signals. Such an approach would make the final solution “light”, both in terms of cost and computational complexity. To achieve this goal, new features are required that are robust to artifacts present in the raw EEG. We present results with a cohort of 54 participants and show that AD diagnostic accuracy as high as 79.6% can be achieved with the proposed solution, thus in line with what can be achieved with a 20-channel layout and manual selection of artifact free EEG segments (81.5%). Further optimization is needed for other tasks, such as measuring AD severity level. Raymundo Cassani, Tiago H. Falk |
SMC | 2 |
| 2019 | Generalizable Adversarial Examples Detection Based on Bi-model Decision MismatchabstractModern applications of artificial neural networks have yielded remarkable performance gains in a wide range of tasks. However, recent studies have discovered that such modelling strategy is vulnerable to Adversarial Examples, i.e. examples with subtle perturbations often too small and imperceptible to humans, but that can easily fool neural networks. Defense techniques against adversarial examples have been proposed, but ensuring robust performance against varying or novel types of attacks remains an open problem. In this work, we focus on the detection setting, in which case attackers become identifiable while models remain vulnerable. Particularly, we employ the decision layer of independently trained models as features for posterior detection. The proposed framework does not require any prior knowledge of adversarial examples generation techniques, and can be directly employed along with unmodified off-the-shelf models. Experiments on the standard MNIST and CIFAR10 datasets deliver empirical evidence that such detection approach generalizes well across not only different adversarial examples generation methods but also quality degradation attacks. Non-linear binary classifiers trained on top of our proposed features can achieve a high detection rate (>90%) in a set of white-box attacks and maintain such performance when tested against unseen attacks. João Monteiro 0002, Isabela Albuquerque, Zahid Akhtar, Tiago H. Falk |
SMC | 4 |
| 2019 | A Multimodal Approach to Improve the Robustness of Physiological Stress Prediction During Physical ActivityabstractStress is well known to have negative effects on health and workplace performance. Physiological sensing using wearables shows in turn great potential for realtime stress monitoring. While some off-the-shelf consumer products (e.g. smartwatches) already feature stress detection, there is still a pressing need to improve the robustness of these models in ecological settings where physical activity can hamper detection accuracy. In this paper, we show that using a multimodal physiological stress model can not only improve model accuracy, but can increase robustness to physical activity inference. To do so, we propose a video game based method to elicit emotional responses. More specifically, 48 participants played video games in which psychological stress and physical activity were jointly modulated. Physiological features showing robustness to stress are analyzed in order to guide further research. Mark Parent, Abhishek Tiwari 0003, Isabela Albuquerque, Jean-François Gagnon, Daniel Lafond, Sébastien Tremblay, Tiago H. Falk |
SMC | 7 |
| 2019 | Mental Workload Assessment During Physical Activity Using Non-linear Movement Artefact Robust Electroencephalography FeaturesabstractAssessment of mental workload is crucial in safety-critical applications. Often, such applications require the user to be ambulant, such as first responders (e.g., paramedics, firefighters, or police officers). Typically, mental workload models have relied on electroencephalography (EEG) signals. EEGs, however, are known to be highly sensitive to movement artefacts, thus limited applications exist for ambulant users and studies have mostly occurred in controlled laboratory settings. In this paper, we explore the robustness of new non-linear features against movement artefacts and test their effectiveness in monitoring mental workload for ambulant users with the end goal of developing mitigation measures based on mental state of operators. To this end, an EEG experiment was conducted where mental workload and physical activity levels were modulated simultaneously and data was collected from 48 participants. Classical EEG features used for workload assessment, such as spectral power and amplitude/phase coherence, were used as benchmarks and compared against the proposed non-linear multi-scale permutation entropy features. Experimental results show the proposed features consistently outperforming the benchmark ones, thus high-lighting their robustness to movement artefacts. Abhishek Tiwari 0003, Isabela Albuquerque, Jean-François Gagnon, Daniel Lafond, Mark Parent, Sébastien Tremblay, Tiago H. Falk |
SMC | 7 |
| 2019 | Residual convolutional neural network with attentive feature pooling for end-to-end language identification from short-duration speech
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
Comput. Speech Lang. | 3 |
| 2019 | Non-Intrusive Speech Quality Prediction Using Modulation Energies and LSTM-NetworkabstractMany signal processing algorithms have been proposed to improve the quality of speech recorded in the presence of noise and reverberation. Perceptual measures, i.e., listening tests, are usually considered the most reliable way to evaluate the quality of speech processed by such algorithms but are costly and time-consuming. Consequently, speech enhancement algorithms are often evaluated using signal-based measures, which can be either intrusive or non-intrusive. As the computation of intrusive measures requires a reference signal, only non-intrusive measures can be used in applications for which the clean speech signal is not available. However, many existing non-intrusive measures correlate poorly with the perceived speech quality, particularly when applied over a wide range of algorithms or acoustic conditions. In this paper, we propose a novel non-intrusive measure of the quality of processed speech that combines modulation energy features and a recurrent neural network using long short-term memory cells. We collected a dataset of perceptually evaluated signals representing several acoustic conditions and algorithms and used this dataset to train and evaluate the proposed measure. Results show that the proposed measure yields higher correlation with perceptual speech quality than that of benchmark intrusive and non-intrusive measures when considering various categories of algorithms. Although the proposed measure is sensitive to mismatch between training and testing, results show that it is a useful approach to evaluate specific algorithms over a wide range of acoustic conditions and may, thus, become particularly useful for real-time selection of speech enhancement algorithm settings. Benjamin Cauchi, Kai Siedenburg, João Felipe Santos, Tiago H. Falk, Simon Doclo, Stefan Goetze |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Improved Audio-Visual Laughter Detection Via Multi-Scale Multi-Resolution Image Texture Features and Classifier FusionabstractEfforts are afoot to design better context-aware human-computer interaction techniques that have knowledge of both their surrounding and the affective state of the user. One of the most important nonverbal behavioural cues for affective human-machine interaction is laughter. Automatic detection of laughter is an interesting, yet challenging problem, which in recent years has gained increased attention from both the academic and industrial communities. The majority of existing laughter detection systems rely on either audio or video modalities. Humans, however, typically rely on audio-visual cues during conversation and/or interaction, thus it is expected that improved results can be achieved if both modalities are used. In this work, we propose a multimodal framework that analyzes audio and video channels separately, then fuses their decisions. Conventional speech spectral and prosodic features are used, whereas new multi -scale multiresolution binarized statistical image features are proposed due to their improved expressive power. Experiments with the publicly available MAHNOB Laughter database show that decision level fusion based on support vector machine classifiers leads to improved performance over single modality approaches, as well as over previously-proposed methods, all whilst requiring just a fraction of the computational power. Zahid Akhtar, Stefany Bedoya, Tiago H. Falk |
ICASSP | 3 |
| 2018 | Dual-Channel Modulation Energy Metric for Direct-to-Reverberation Ratio EstimationabstractNon-intrusive estimators for acoustic parameters like the direct-to-reverberation ratio (DRR) are useful tools but still perform weakly as shown in the acoustic characterization of environments (ACE) challenge. In this paper, we develop a novel dual-channel metric based on the modulation energy domain for DRR estimation. In contrast to established modulation based single-channel metrics like the speech-to-reverberation modulation energy ratio (SRMR), we exploit the spatial information from two microphones as well as the temporal dynamics in the modulation energy domain. The developed metric shows a strong linear correlation to the DRR, which allows a simple mapping. It is shown that the metric is robust against the microphone array configuration, room characteristics and the speech signal. The proposed metric is compared to a reference method based on the spectral variance of the room transfer functions, and both metrics are evaluated using simulated and measured data. In our experiments, the proposed metric achieved a higher correlation and lower RMSE compared the reference method, and outperforms existing SRMR based DRR estimators. Sebastian Braun, João Felipe Santos, Emanuël A. P. Habets, Tiago H. Falk |
ICASSP | 4 |
| 2018 | Investigating Speech Enhancement and Perceptual Quality for Speech Emotion Recognition
Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 4 |
| 2018 | A Neurophysiological Sensor-Equipped Head-Mounted Display for Instrumental QoE Assessment of Immersive MultimediaabstractThe last few years have seen a drastic increase in the consumption of virtual- and augmented-reality (VR and AR, respectively) applications. Ultimately, the success of any emerging technology will rely on the experience it provides the end user, and not on the technology itself. Subjective methods for quality-of-experience (QoE) assessment have as main disadvantage that converting such human factors into a quality rating is difficult, particularly for everyday users. To overcome this limitation, recent research has explored the use of objective methods to monitor neurophysiological correlates of relevant perception processes. In this paper, we describe the development of a neurophysiological sensor-equipped head-mounted display that combines a consumer off-the-shelf VR headset, a modified low-cost portable device for electroencephalogram (EEG) acquisition repurposed to simultaneously acquire EEG, electrocardiogram (ECG), and electrooculogram (EOG) signals with high-quality dry electrodes. The device was evaluated under three different scenarios, each one designed to test the different ExG modalities. Initial tests showed promising results and allowed for (1) steady-state visually evoked potentials to be accurately measured from EEG, (2) heart rate variability measurements to discriminate between different affective videos, and (3) EOG measurements to monitor gaze direction and eye blinks, all while users were mobile. Being able to accurately monitor signals from the autonomic and central nervous systems in an unobtrusive and portable manner is an important step for instrumental QoE assessment of emerging VR/AR applications. Raymundo Cassani, Marc-Antoine Moinnereau, Tiago H. Falk |
QoMEX | 3 |
| 2018 | Towards a Neuro-Inspired No-Reference Instrumental Quality Measure for Text-to-Speech SystemsabstractSubjective evaluation of synthesized speech is not an easy task as various quality dimensions can be affected, including naturalness, prosody, pronunciation, and continuity, to name a few. Evaluations typically rely on naive listeners, thus more closely representing the consumers of commercial products. As such, while the results of these costly and time consuming tests may provide text-to-speech (TTS) system developers with feedback on the perceived quality and acceptability of their devices, it provides little information on what the source of the problems are and what can be done about it. In this paper, we propose the use of neuroimaging to probe the unconscious cognitive processing of naive listeners as they listen to synthesized speech generated by different systems of varying quality. The obtained neural insights have allowed us to extract a small subset of very relevant features from the speech signals and to use these features to build a simple, no-reference instrumental quality metric specifically tailored to TTS speech. The metric is tested on an unseen dataset and shown to significantly outperform a benchmark algorithm. Anderson R. Avila, Tiago H. Falk |
QoMEX | 3 |
| 2018 | On the Analysis of EEG Features for Mental Workload Assessment During Physical ActivityabstractAssessment of mental workload is crucial for applications which require constant attention and where conditions such as mental fatigue and drowsiness must be avoided. As such, electroencephalography (EEG) based mental workload models have been developed in the past. The majority of these models, however, have assumed individuals are not ambulant, thus bypassing the issue of movement-related EEG artefacts. While such models may be useful for a number of applications (e.g., operators are sitting), they may not apply in situations in which operators are performing their task under different physical activity levels. Representative examples can include first responders, such as paramedics, firefighters, or police officers. In this work, we take the first steps towards overcoming this limitation and present results of an experiment simultaneously eliciting increasing mental workload states at varying physical activity levels. EEG data from forty-seven participants was collected while they performed the NASA Revised Multi-Attribute Task Battery II (MATB-II) under three different activity level conditions (no, medium, high). In this study, we report the effects of activity on the noise-robustness and distribution of several spectral, amplitude/phase coherence, and amplitude modulation features, with the ultimate goal of deriving a feature set tailored towards automated workload assessment during physical activity. Preliminary results show spectral features acquired from the frontal area of the cortex as the most promising and that activity aware mental workload models should be developed. Isabela Albuquerque, Abhishek Tiwari 0003, Jean-François Gagnon, Daniel Lafond, Mark Parent, Sébastien Tremblay, Tiago H. Falk |
SMC | 7 |
| 2018 | EEG Artifact Removal for Improved Automated Lane Change Detection while DrivingabstractIn this study, we are interested in brain-machine interfaces that extract movement related cortical potentials (MRCP) from an electroencephalogram (EEG) recorded while driving and use this information to classify/detect when a driver intends to do left or right lane change. Collecting EEGs while driving, however, is a challenging task as it introduces numerous artifacts, including head movements (e.g., to look at side mirrors, etc), eye movements, and whole-body movements. Such artifacts can dominate the signal, thus hampering MRCP and lane change intent detection. Here, we explore three EEG artifact removal algorithms tailored for MRCP detection while driving, namely: headset accelerometer-based independent component analysis (acc-ICA), constrained ICA (cICA) and empirical mode decomposition (EMD). Next, we propose to use a recurrent neural network (RNN) classifier for lane change intent detection and compare results with a conventional support vector machine (SVM) based classifiers. Lastly, we explore the effect of EEG window analysis size on artifact removal and classification performance, thus gauging how close to real-time enhancement can be achieved and with what delay can lane change intent be detected. When looking at averaged trial performance, all system combinations achieved reliable accuracy. On the other hand, for single-trial accuracy, only acc-ICA and EMD methods performed well. Overall, average classification accuracy of 54.25% was achieved with SVMs, whereas accuracies of 82.88%, 82.76% and 82.69% were achieved with the RNN using acc-ICA, cICA and EMD artifact removal algorithms, respectively. Window sizes of 4 seconds prior to lane changes achieved the best result and only six EEG electrodes were required. Marc-Antoine Moinnereau, Sam Karimian-Azari, Tsuyoshi Sakuma, Hidenori Boutani, Lucian Andrei Gheorghe, Tiago H. Falk |
SMC | 6 |
| 2018 | Resting-Awake EEG Amplitude Modulation can Predict Performance of an fNIRS-Based Neurofeedback TaskabstractAffective neurofeedback emerges as an innovative approach to relieve symptoms of psychiatric disorders. However, some users are unable to control these systems and are named neurofeedback illiterates. In this paper, we propose a cross-modality approach to predict affective neurofeedback illiteracy using electroencephalography (EEG) amplitude modulation measures. Thirty-one subjects were submitted to five-minute resting state protocol with EEG records, followed by functional near-infrared spectroscopy (fNIRS)-based affective neurofeedback task. EEG amplitude modulation measures were extracted from the resting block and correlated with the task performance. The gamma-m-alpha modulation from one channel presented a negative correlation with performance (r=-0.639, p=0.006). When used as input for a support vector regression model, this feature predicted performance with a mean absolute error of 13.78%. Our findings suggest this cross-modality feature is related to the affective lateralization effect, as well as with the EEG-hemodynamics coupling during resting-state. This feature might be a promising tool to predict performance and choose the best strategy for future therapeutics using affective neurofeedback. Lucas Trambaiolli, Raymundo Cassani, Claudinei E. Biazoli Jr., Andre Mascioli Cravo, Joo Sato, Tiago H. Falk |
SMC | 6 |
| 2018 | Fusion of bottleneck, spectral and modulation spectral features for improved speaker verification of neutral and whispered speech
Milton Orlando Sarria-Paja, Tiago H. Falk |
Speech Commun. | 2 |
| 2018 | Speech Dereverberation With Context-Aware Recurrent Neural NetworksabstractIn this paper, we propose a model to perform speech dereverberation by estimating its spectral magnitude from the reverberant counterpart. Our models are capable of extracting features that take into account both short- and long-term dependencies in the signal through a convolutional encoder (which extracts features from a short, bounded context of frames) and a recurrent neural network for extracting long-term information. Our model outperforms a recently proposed model that uses different context information depending on the reverberation time, without requiring any sort of additional input, yielding improvements of up to 0.4 on perceptual evaluation of speech quality, 0.3 on short-time objective intelligibility, and 1.0 on perceptual objective listening quality assessment relative to reverberant speech. We also show our model is able to generalize to real room impulse responses even when only trained with simulated room impulse responses, different speakers, and high reverberation times. Finally, listening tests show the proposed method outperforming benchmark models in reduction of perceived reverberation. João Felipe Santos, Tiago H. Falk |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Adaptive Spectro-Temporal Filtering for Electrocardiogram Signal EnhancementabstractAdvances in low-cost portable electrocardiogram (ECG) devices have opened doors for numerous new applications, including fitness tracking, remote health, and peak athletic performance monitoring, to name a few. Many such devices, however, have been shown to be highly contaminated by movement and/or muscle contraction artifacts, which, in turn, can lead to erroneous heart rate and heart rate variability (HRV) analyses. Here, we propose a new denoising method based on adaptive spectro-temporal filtering for ECG enhancement. The algorithm relies on the so-called modulation spectral signal representation, which is shown to accurately separate ECG and noise components. The proposed method was tested on synthetic ECG signals corrupted with varying levels of recorded noise and on long-term bedside noisy ECG recordings. Gains over a state-of-the-art wavelet-based denoising algorithm were achieved, particularly for very noisy scenarios. Overall, the proposed algorithm achieved a 61.8% gain in signal-to-noise ratio improvement, a three times reduction in average heart rate measurement error, and a 15% reduction in HRV measurement error relative to the benchmark, thus suggesting that it is an ideal candidate for ECG-based fitness/athletic monitoring applications. Diana P. Tobón, Tiago H. Falk |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Improved heart rate variability measurement based on modulation spectral processing of noisy electrocardiogram signalsabstractWearable device usage is burgeoning, with representative applications ranging from patient/athelete monitoring to stress/fatigue identification to the so-called quantified self movement. Typically, cardiac information is monitored via electrocardiograms (ECG) and information such as heart rate (HR) and heart rate variability (HRV) are used as key health-related metrics. With many wearable devices, however, lower quality sensors are used, thus resulting in devices that are highly sensible to artifacts due to e.g., user's movement. The introduced artifacts hamper HR/HRV analyses, thus ECG enhancement has been the focus of recent research. Existing enhancement algorithms, however, do not perform well in very noisy conditions, as well as add additional computational processing to already battery-hungry wearable applications. Here, we propose to overcome these limitations by describing a new ECG signal representation called the modulation spectrum. By quantifying the rate-of-change of ECG spectral components, signal and artifactual components become separable, thus allowing for accurate HR and HRV measurement from the noisy signal, even in very extreme conditions typically seen in athletic performance training. The proposed MD-HRV (modulation-domain HRV) metric is tested with noise-corrupted synthetic ECG signals and is compared to `true' HRV values obtained from the clean signals. Experimental results show the proposed metric significantly outperforming conventional HRV indices computed on both the noisy, as well as enhanced ECG signals processed by a state-of-the-art wavelet-based algorithm. The obtained findings suggest that the proposed metric is well suited for wearable applications, particularly those involved with intense movement (e.g., in elite athletic training). Diana P. Tobón, Srinivasan Jayaraman, Tiago H. Falk |
BSN | 3 |
| 2017 | Event-related synchronisation responses to N-back memory tasks discriminate between healthy ageing, mild cognitive impairment, and mild Alzheimer's diseaseabstractIn this study we investigate whether or not event-related (de)synchronisation (ERD/ERS) can be used to differentiate between 27 healthy elderly, 21 subjects diagnosed with amnestic mild cognitive impairment (aMCI) and 16 mild Alzheimer's disease (AD) patients. Using 32-channel EEG recordings, we measured ERD responses to a three-level visual N-back task (N = 0, 1, 2) on the well-known delta, theta, alpha, beta and gamma bands. Our findings revealed that healthy elderly (HE) elicited consistently greater beta and alpha ERD responses than MCI and AD patients at many scalp electrodes, most of them located at fronto-central and temporal-parietal areas. Additionally, significant ERD differences were found on the gamma band in the MCI vs. AD comparison. Based on these findings, we conclude that ERD responses to a working memory (N-back) task could be useful for early MCI diagnosis or for improved AD diagnosis, and also for assessing the likelihood of MCI progression to AD. Francisco J. Fraga, Leonardo A. Ferreira, Tiago H. Falk, Erin Johns, Natalie D. Phillips |
ICASSP | 3 |
| 2017 | Speech temporal dynamics fusion approaches for noise-robust reverberation time estimationabstractReverberation and noise are known to be the two most important culprits for poor performance in far-field speech applications, such as automatic speech recognition. Recent research has suggested that reverberation-aware speech enhancement (or speech technologies, in general) could be used to improve performance. However, recent results also show existing blind room acoustics characterization algorithms are not robust under ambient noise and there is still room for improvement under such settings. In this paper, several fusion approaches are proposed for noise-robust reverberation time estimation. More specifically, feature- and score-level fusion of short- and long-term speech temporal dynamics features are proposed. With noise-aware feature-level fusion, gains of up to 15.4% could be seen in root mean square error. Score-level fusion, in turn, showed further improvements of up to 9.8%. Relative to a recently-proposed noise-robust benchmark algorithm, improvements of 30% could be seen, thus showing the advantages of speech temporal dynamics fusion approaches for noise-robust reverberation time estimation. Mohammed Senoussaoui, João Felipe Santos, Tiago H. Falk |
ICASSP | 3 |
| 2017 | EEG coupling features: Towards mental workload measurement based on wearablesabstractAutomated mental workload measurement is particularly important in safety-critical settings, such as in nuclear plants, aviation, air traffic control, shipping, and transportation, to name a few. As an example, recent statistics have suggested that 90% of the accidents in the transport industry are due to human factors. In this paper, we explore the potential of off-the-shelf wearable technologies in monitoring mental workload in real-time, thus potentially reducing the number of accidents due to human errors. Wearable technologies, while providing the user with ease-of-use, comfort, and portability, have several limitations, such as lower quality signal readings (e.g., due to dry electrodes) and smaller number of recording sites. Such limitations place a burden on the accuracy of existing mental workload models. To overcome this limitation, we propose the use of phase-amplitude and amplitude-amplitude coupling features computed from a portable commercial electroencephalography (EEG) device. Experiments with three different tasks, namely N-back, mental rotation and visual search, show the proposed features being significantly correlated with multiple dimensions of the widely-used NASA task load index test and providing complementary information to other conventional features. Alexandre Drouin-Picaro, Isabela Albuquerque, Jean-François Gagnon, Daniel Lafond, Tiago H. Falk |
SMC | 5 |
| 2017 | Fusion of auditory inspired amplitude modulation spectrum and cepstral features for whispered and normal speech speaker verification
Milton Orlando Sarria-Paja, Tiago H. Falk |
Comput. Speech Lang. | 2 |
| 2016 | Feature mapping, score-, and feature-level fusion for improved normal and whispered speech speaker verificationabstractIn this paper, automatic speaker verification using normal and whispered speech is explored. Typically, for speaker verification systems with varying vocal effort inputs, standard solutions such as feature mapping or addition of data during parameter estimation (training) and enrollment stages result in a trade-off between accuracy gains with whispered test data and accuracy losses (up to 70% in equal error rate, EER) with normal test data. To overcome this shortcoming, this paper proposes two innovations. First, we show the complementarity of features derived from AM-FM models over conventional mel-frequency cepstral coefficients, thus signalling the importance of instantaneous phase information for whispered speech speaker verification. Next, two fusion schemes are explored: score- and feature-level fusion. Overall, we show that gains as high as 30% and 84% in EER can be achieved for normal and whispered speech, respectively, using featurelevel fusion. Milton Orlando Sarria-Paja, Mohammed Senoussaoui, Douglas D. O'Shaughnessy, Tiago H. Falk |
ICASSP | 4 |
| 2016 | A Quality Adaptive Multimodal Affect Recognition System for User-Centric Multimedia IndexingabstractThe recent increase in interest for online multimedia streaming platforms has availed massive amounts of multimedia information that need to be indexed to be searchable and retrievable. User-centric implicit affective indexing employing emotion detection based on psycho-physiological signals, such as electrocardiography (ECG), galvanic skin response (GSR), electroencephalography (EEG) and face tracking, has recently gained attention. However, real world psycho-physiological signals obtained from wearable devices and facial trackers are contaminated by various noise sources that can result in spurious emotion detection. Therefore, in this paper we propose the development of psycho-physiological signal quality estimators for unimodal affect recognition systems. The presented systems perform adequately in classifying users affect however, they resulted in high failure rates due to rejection of bad quality samples. Thus, to reduce the affect recognition failure rate, a quality adaptive multimodal fusion scheme is proposed. The proposed scheme yields no failure, while at the same time classify the users' arousal/valence and liking with significantly above chance weighted F1-scores in a cross-user experiment. Another finding of this study is that head movements encode liking perception of users in response to music snippets. This work also includes the release of the employed dataset including psycho-physiological signals, their quality annotations, and users' affective self-assessments. Mojtaba Khomami Abadi, Jesús Alejandro Cárdenes Cabré, Fabio Morreale, Tiago H. Falk, Nicu Sebe |
ICMR | 5 |
| 2016 | Laughter detection based on the fusion of local binary patterns, spectral and prosodic featuresabstractToday, great focus has been placed on context-aware human-machine interaction, where systems are aware not only of the surrounding environment, but also about the mental/affective state of the user. Such knowledge can allow for the interaction to become more human-like. To this end, automatic discrimination between laughter and speech has emerged as an interesting, yet challenging problem. Typically, audio-or video-based methods have been proposed in the literature; humans, however, are known to integrate both sensory modalities during conversation and/or interaction. As such, this paper explores the fusion of support vector machine classifiers trained on local binary pattern (LBP) video features, as well as speech spectral and prosodic features as a way of improving laughter detection performance. Experimental results on the publicly-available MAHNOB Laughter database show that the proposed audio-visual fusion scheme can achieve a laughter detection accuracy of 93.3%, thus outperforming systems trained on audio or visual features alone. Stefany Bedoya, Tiago H. Falk |
MMSP | 2 |
| 2016 | Physiological quality-of-experience assessment of text-to-speech systemsabstractWith the emergence of various text-to-speech (TTS) systems, developers have to provide superior user experience in order to remain competitive. To this end, quality-of-experience (QoE) perception modelling and measurement has become a key priority. QoE models rely on three influence factors: technological, contextual and human. Existing solutions have typically relied on using individual physiological modalities, such as electroen-cephalography (EEG), to model human influence factors (HIFs). In this paper, we show that fusion of physiological modalities, such as EEG, functional near infrared spectroscopy (fNIRS) and heart rate, provide gains of up to 18.4% relative to utilizing only technological factors and 4% relative to using the best performing individual physiological modality. Tiago H. Falk |
MMSP | 2 |
| 2016 | Relevance vector classifier decision fusion and EEG graph-theoretic features for automatic affective state characterization
Khalil ur Rehman Laghari, Tiago H. Falk |
Neurocomputing | 3 |
| 2015 | On the potential for artificial bandwidth extension of bone and tissue conducted speech: a mutual information studyabstractTo enhance the communication experience of workers equipped with hearing protection devices and radio communication in noisy environments, alternative methods of speech capture have been utilized. One such approach uses speech captured by a microphone in an occluded ear canal. Although high in signal-to-noise ratio, bone and tissue conducted speech has a limited bandwidth with a high frequency roll-off at 2 kHz. In this paper, the potential of using various bandwidth extension techniques is investigated by studying the mutual information between the signals of three uniquely placed microphones: inside an occluded ear, outside the ear and in front of the mouth. Using a Gaussian mixture model approach, the mutual information of the low and high-band frequency ranges of the three microphone signals at varied levels of signal-tonoise ratio is measured. Results show that a speech signal with extended bandwidth and high signal-to-noise ratio may be achieved using the available microphone signals. Rachel E. Bouserhal, Tiago H. Falk, Jérémie Voix |
ICASSP | 2 |
| 2014 | Improving the performance of far-field speaker verification using multi-condition training: the case of GMM-UBM and i-vector systems
Anderson R. Avila, Milton Orlando Sarria-Paja, Francisco J. Fraga, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 5 |
| 2014 | Cochlear Implant Filterbank Design and Optimization: A Simulation StudyabstractCochlear implants (CIs) are devices capable of restoring hearing function in profoundly-deaf patients to an acceptable degree of performance. An essential processing step in any cochlear implant is frequency analysis, which is usually performed via banks of filters. Here, we simulate and test the suitability of different filters and filterbank architectures for CIs with respect to their performance in speech intelligibility. Four different filters were implemented in an established model of CI hearing, the tone-excited vocoder, namely: GTF (Gammatone Filter), DAPGF (Differentiated All-Pole GTF), OZGF (One-Zero GTF) and BUTF (Butterworth). Three filterbank parameters, the filter order ( N), the filter quality factor ( Q) and the number of channels ( Ch), and their combinations were tested using objective and subjective metrics. Simulation results show that all filters tested are suitable for CI implementation, but that the choice of Q and N parameter values is crucial. For most conditions, optimal ( N,Q) combinations were within few units away from the combination (2, 4). Stefano Cosentino, Tiago H. Falk, David McAlpine, Torsten Marquardt |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Updating the SRMR-CI metric for improved intelligibility prediction for cochlear implant usersabstractWhen compared to intrusive speech intelligibility metrics, non-intrusive ones show a stronger dependency on speech content, given the lack of a reference signal for distortion level computation. Reduction of this dependency is an important step needed to develop reliable metrics. In this paper, two different updates to SRMR-CI, a recently-proposed speech intelligibility metric tailored for cochlear implant users, are applied. First, modulation energy thresholding is proposed to reduce the variability caused by the differences in modulation spectral representations for different phonemes and speakers, as well as speech enhancement algorithm artifacts. Second, a narrower range of modulation filters is employed to reduce fundamental frequency effects. Experimental results show that the updated metric outperforms two benchmark metrics, namely ModA and ANIQUE+, by as much as 15% in terms of correlation between objective and subjective ratings, and a relative decrease of 47% in root mean square error compared to the previously-proposed SRMR-CI metric. João Felipe Santos, Tiago H. Falk |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Cognitive, affective, and experience correlates of speech quality perception in complex listening conditionsabstractSubjective speech quality assessment depends on listener “quality” opinions after hearing a particular test speech stimulus. Subjective scores are given based on a perception and quality judgment process that is unique to a particular listener. These processes are postulated to be dependent on the listener's internal reference of what good and bad quality sounds like, as well as their mental and emotional states. To overcome this variability, subjective listening tests often average scores over several listeners. In this paper, we use electroencephalography (EEG) and self-assessment tools to investigate the neural and affective correlates of speech quality perception of reverberant speech, with the goal of obtaining new insights into human speech quality perception in complex listening environments. We show that EEG event related potentials (ERP) are a useful tool to monitor the conscious stages of neural-processing during a speech quality assessment task. Significant correlations were obtained between the so-called P300 ERP component and the reverberation time of the room, as well as between the P300 peak amplitude and emotional self-assessment ratings. These insights could lead to more effective ways of characterizing room acoustics for improved speech quality and intelligibility. Jan-Niklas Voigt-Antons, Khalil ur Rehman Laghari, Sebastian Arndt, Robert Schleicher, Sebastian Möller 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
ICASSP | 7 |
| 2013 | Non-intrusive objective speech quality and intelligibility prediction for hearing instruments in complex listening environmentsabstractA non-intrusive objective speech quality and intelligibility measure tailored to hearing restoration instruments is proposed and evaluated in complex listening environments. The measure builds upon the previously-proposed “speech-to-reverberation modulation energy ratio” (SRMR) by incorporating hearing impairment percepts, such as hearing loss thresholds and altered modulation frequency selectivity. Performance is assessed using speech data corrupted by additive noise, reverberation, and noise-plus-reverberation which were subjectively rated by cochlear implant and hearing aid users. Experimental results show that the developed measures outperform the original SRMR metric for hearing impaired listeners and achieve performance levels inline with existing intrusive quality and intelligibility metrics, but with the advantage of not requiring access to a clean reference signal. As such, the measure may be used to develop qualityor intelligibility-aware speech enhancement algorithms for advanced hearing restoration instruments. Tiago H. Falk, Stefano Cosentino, João Felipe Santos, David Suelzle, Vijay Parsa |
ICASSP | 1 |
| 2013 | Towards an EEG-based biomarker for Alzheimer's disease: Improving amplitude modulation analysis featuresabstractIn this paper, an EEG-based biomarker for automated Alzheimer's disease (AD) diagnosis is described, based on extending a recently-proposed “percentage modulation energy” (PME) metric. More specifically, to improve the signal-to-noise ratio of the EEG signal, PME features were averaged over different durations prior to classification. Additionally, two variants of the PME features were developed: the “percentage raw energy” (PRE) and the “percentage envelope energy” (PEE). Experimental results on a dataset of 88 participants (35 controls, 31 with mild-AD and 22 with moderate AD) show that over 98% accuracy can be achieved with a support vector classifier when discriminating between healthy and mild AD patients, thus significantly outperforming the original PME biomarker. Moreover, the proposed system can achieve over 94% accuracy when discriminating between mild and moderate AD, thus opening doors for very early diagnosis. Francisco J. Fraga, Tiago H. Falk, Lucas Trambaiolli, Eliezyer Fermino de Oliveira, Walter H. L. Pinaya, Paulo A. M. Kanda, Renato Anghinah |
ICASSP | 2 |
| 2013 | Whispered speaker verification and gender detection using weighted instantaneous frequenciesabstractIn this paper, automatic speaker verification and gender detection using whispered speech is explored. Whispered speech, despite its reduced perceptibility, has been shown to convey relevant speaker identity and gender information. This study compares the performance of a GMM-UBM speaker verification system trained with normal and whispered speech under different matched and mismatched conditions, and describes the benefits of adaptation in a speaking-style independent model to handle both vocal efforts. It is shown that performance improvements can be achieved by using speaking-style and gender dependent models, as well as by adding features based on the AM-FM signal representation. Moreover, the AM-FM based features showed to be more discriminative than classical MFCCs for whispered speech gender detection. Experimental results suggest that whispered speech carries sufficient information for reliable automatic speaker identification. Milton Orlando Sarria-Paja, Tiago H. Falk, Douglas D. O'Shaughnessy |
ICASSP | 2 |
| 2013 | Very early detection of Autism Spectrum Disorders based on acoustic analysis of pre-verbal vocalizations of 18-month old toddlersabstractWith the increasing prevalence of Autism Spectrum Disorders (ASD), very early detection has become a key priority research topic, as early interventions can increase the chances of success. Since atypical communication is a hallmark of ASD, automated acoustic-prosodic analyses have received prominent attention. Existing studies, however, have focused on verbal children, typically over the age of three (when many children may be reliably diagnosed) and as high as early teens. Here, an acoustic-prosodic analysis of pre-verbal vocalizations (e.g., babbles, cries) of 18-month old toddlers is performed. Data was obtained from a prospective longitudinal study looking at high-risk siblings of children with ASD who were also diagnosed with ASD, as well as low-risk age-matched typically developing controls. Several acoustic-prosodic features were extracted and used to train support vector machine and probabilistic neural network classifiers; classification accuracy as high as 97% was obtained. Our findings suggest that markers of autism may be present in pre-verbal vocalizations of 18-month old toddlers, thus may be used to assist clinicians with very early detection of ASD. João Felipe Santos, Nirit Brosh, Tiago H. Falk, Lonnie Zwaigenbaum, Susan E. Bryson, Wendy Roberts, Isabel M. Smith, Peter Szatmari, Jessica A. Brian |
ICASSP | 3 |
| 2013 | Neurophysiological experimental facility for Quality of Experience (QoE) assessment
Khalil ur Rehman Laghari, Sebastian Arndt, Jan-Niklas Voigt-Antons, Robert Schleicher, Sebastian Möller 0001, Tiago H. Falk |
IM | 7 |
| 2013 | Predicting the bilateral advantage in cochlear implantees using a non-intrusive speech intelligibility measureabstractA measure to predict speech intelligibility in unilateral and bi-lateral cochlear implant (CI) users is proposed that does not need a priori information (i.e. is non-intrusive), such as the room acoustics. Such measure, termed BiSIMCI, combines an equalization-cancellation stage together with a modulation frequency estimation stage. Simulated and actual subjective data from CI users were used to validate the proposed mea-sure. The actual CI subjective data consisted of speech re-ception thresholds (SRTs) collected in anechoic rooms with a total of 28 target/interferer spatial configurations. The simu-lated CI subjective data were generated by running the intrusive algorithm by Culling et al. [Ear & Hearing 33 (6), 673-682 (2012)] across 109 different target/interferer conditions from two environments, one anechoic and one highly reverberant room (RT60 = 0.89s). The experimental results indicate that the proposed non-intrusive measure provides reliable predic-tions when compared with both actual and simulated SRT; an average correlation of 0.94 was reported for these conditions, and an average correlation of 0.97 was obtained when some in-trusive assumptions were made. 1. Stefano Cosentino, Tiago H. Falk, David McAlpine |
INTERSPEECH | 2 |
| 2013 | Objective speech intelligibility measurement for cochlear implant users in complex listening environments
João Felipe Santos, Stefano Cosentino, Oldooz Hazrati, Philipos C. Loizou, Tiago H. Falk |
Speech Commun. | 5 |
| 2013 | Whispered Speech Detection in Noise Using Auditory-Inspired Modulation Spectrum FeaturesabstractRobustness to ambient noise, varying vocal effort, and availability of only short-duration test utterances represent big challenges for developers of automated speech-enabled applications. Recent studies have proposed the use of vocal effort-matched speaker models as a potential solution to such challenges. However, detecting whispered speech in extremely noisy environments is not a trivial task. This letter proposes the use of auditory-inspired modulation spectral-based features as a method of separating speech from environment-based components, thus resulting in accurate whispered speech detection at signal-to-noise ratios as low as 0 dB. Experimental results show the proposed detection algorithm outperforming two benchmark approaches. Milton Orlando Sarria-Paja, Tiago H. Falk |
IEEE Signal Process. Lett. | 2 |
| 2012 | Automated Dysarthria Severity Classification for Improved Objective Intelligibility Assessment of Spastic Dysarthric SpeechabstractIn this paper, automatic dysarthria severity classifica-tion is explored as a tool to advance objective intelli-gibility prediction of spastic dysarthric speech. A Ma-halanobis distance-based discriminant analysis classifier is developed based on a set of acoustic features for-merly proposed for intelligibility prediction and voice pathology assessment. Feature selection is used to sift salient features for both the disorder severity classifica-tion and intelligibility prediction tasks. Experimental re-sults show that a two-level severity classifier combined with a 9-dimensional intelligibility prediction mapping can achieve 0.92 correlation and 12.52 root-mean-square error with subjective intelligibility ratings. The effects of classification errors on intelligibility accuracy are also explored and shown to be insignificant. Index Terms: Intelligibility, dysarthria, diagnosis. 1. Milton Orlando Sarria-Paja, Tiago H. Falk |
INTERSPEECH | 2 |
| 2012 | Performance Comparison of Intrusive Objective Speech Intelligibility and Quality Metrics for Cochlear Implant UsersabstractIn this paper, we evaluate the performance of six intrusive ob-jective measures as intelligibility predictors of degraded speech for cochlear implant (CI) users. Three practical environmental degradation scenarios are considered: reverberation alone, ad-ditive noise alone, and noise-plus-reverberation. A subjective intelligibility test was performed with eleven cochlear implant users and objective measures were evaluated using four per-formance metrics: Pearson, Spearman rank, and sigmoid-fitted correlation coefficients, and the root mean square error. It was observed that existing metrics performed well in the noise-alone scenarios, but obtained lower performance in the reverberation-alone scenario and in many cases, unacceptable results in the noise-plus-reverberation scenario. It is concluded that further work is still needed in order to accurately predict speech intel-ligibility ratings for CI users, particularly in environments cor-rupted by reverberation. João Felipe Santos, Stefano Cosentino, Oldooz Hazrati, Philipos C. Loizou, Tiago H. Falk |
INTERSPEECH | 5 |
| 2012 | Characterization of atypical vocal source excitation, temporal dynamics and prosody for objective measurement of dysarthric word intelligibility
Tiago H. Falk, Wai-Yip Chan, Fraser Shein |
Speech Commun. | 1 |
| 2011 | Quantifying perturbations in temporal dynamics for automated assessment of spastic dysarthric speech intelligibilityabstractSpastic dysarthric speech is often associated with imprecise placement of articulators which, in turn, cause perturbations in speech temporal dynamics, such as unclear distinctions between adjacent phonemes. While these perturbations can lead to a significant reduction in intelligibility, measures to objectively assess their detrimental effect on intelligibility are lacking. In this paper, short- and long-term temporal dynamics measures are proposed and evaluated as correlates of subjective intelligibility. The former is based on log-energy temporal dynamics information, whereas the latter is based on an auditory-inspired modulation spectral signal representation. A composite measure is also developed based on linearly combining the proposed measures with a tone-unit duration parameter. Experiments with the publicly-available 'Universal Access' database of spastic dysarthric speech show that the proposed composite measure can achieve rank correlations with subjective ratings as high as 0.87, thus providing a tool to automatically diagnose speech disorder severity and to evaluate dysarthria treatment outcomes. Tiago H. Falk, Richard Hummel, Wai-Yip Chan |
ICASSP | 1 |
| 2011 | Spectral Features for Automatic Blind Intelligibility Estimation of Spastic Dysarthric SpeechabstractIn this paper, we explore the use of the standard ITU-T P.563 speech quality estimation algorithm for automatic assessment of dysarthric speech intelligibility. A linear mapping consisting of three salient P.563 internal fea-tures is proposed and shown to accurately estimate spas-tic dysarthric speech intelligibility. Delta-energy features are further proposed in order to characterize the atypi-cal spectral dynamics and limited vowel space observed with spastic dysarthria. Experiments using the publicly-available Universal Access database (10 speaker patients) show that when salient delta-energy and internal P.563 features are used, correlations with subjective intelligi-bility ratings as high as 0.98 can be attained. Index Terms: Dysarthria, intelligibility, P.563, subjec-tive quality, delta-energy Richard Hummel, Wai-Yip Chan, Tiago H. Falk |
INTERSPEECH | 3 |
| 2011 | An Assessment of the Improvement Potential of Time-Frequency Masking for Speech DereverberationabstractThe effect of ideal time-frequency masking (ITFM) on the intel-ligibility of reverberated speech is tested using objective mea-surement, namely STI and PESQ scores. The best choice of ITFM threshold is determined for a range of reverberation times (RTs). Four existing dereverberation algorithms are also as-sessed. Objective test results and informal subjective listen-ing show that IFTM provides great intelligibility improvement for all RTs and outperforms the existing dereverberation algo-rithms, one of which assumes perfect knowledge of the room impulse response. While ITFM provides only a best possible performance bound, our results demonstrate the potential im-provement that could be obtained using time-frequency mask-ing for speech dereverberation. Index Terms: Dereverberation, speech intelligibility, speech transmission index, speech quality, time-frequency masking Chenxi Zheng, Tiago H. Falk, Wai-Yip Chan |
INTERSPEECH | 2 |
| 2011 | Automatic speech emotion recognition using modulation spectral features
Siqing Wu, Tiago H. Falk, Wai-Yip Chan |
Speech Commun. | 2 |
| 2010 | Improving the performance of NIRS-based brain-computer interfaces in the presence of background auditory distractionsabstractIn this paper, the effects of auditory distractions on the performance of brain-computer interfaces (BCI) based on near-infrared spectroscopy (NIRS) are investigated. Experiments show that NIRS-BCI specificity decreases by an average 19% when operated in the presence of continuous background noise (relative to operation in silence) and by 13% when operated in the presence of startle noises. To improve BCI performance in noisy environments, a simple yet effective startle noise compensation strategy is proposed. Acoustic environmental conditions are tracked in realtime and false BCI activations that occur within seconds of detected startle noises are suppressed. Experiments show NIRS-BCI systems equipped with the proposed compensation system attaining performances in noisy conditions comparable to those attained in silent conditions. Tiago H. Falk, Kelly Paton, Sarah D. Power, Tom Chau |
ICASSP | 1 |
| 2010 | Comparison of approaches for instrumentally predicting the quality of text-to-speech systemsabstractIn this paper, we compare and combine different approaches for instrumentally predicting the perceived quality of Text-to-Speech systems. First, a log-likelihood is determined by comparing features extracted from the synthesized speech signal with features trained on natural speech. Second, parameters are extracted which capture quality-relevant degradations of the synthesized speech signal. Both approaches are combined and evaluated on three auditory test databases. The results show that auditory quality judgments can in many cases be predicted with a sufficiently high accuracy and reliability, but that there are considerable differences, mainly between male and female speech samples. Sebastian Möller 0001, Florian Hinterleitner, Tiago H. Falk, Tim Polzehl |
INTERSPEECH | 3 |
| 2010 | Modulation Spectral Features for Robust Far-Field Speaker IdentificationabstractIn this paper, auditory inspired modulation spectral features are used to improve automatic speaker identification (ASI) performance in the presence of room reverberation. The modulation spectral signal representation is obtained by first filtering the speech signal with a 23-channel gammatone filterbank. An eight-channel modulation filterbank is then applied to the temporal envelope of each gammatone filter output. Features are extracted from modulation frequency bands ranging from 3-15 H z and are shown to be robust to mismatch between training and testing conditions and to increasing reverberation levels. To demonstrate the gains obtained with the proposed features, experiments are performed with clean speech, artificially generated reverberant speech, and reverberant speech recorded in a meeting room. Simulation results show that a Gaussian mixture model based ASI system, trained on the proposed features, consistently outperforms a baseline system trained on mel-frequency cepstral coefficients. For multimicrophone ASI applications, three multichannel score combination and adaptive channel selection techniques are investigated and shown to further improve ASI performance. Tiago H. Falk, Wai-Yip Chan |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | A Non-Intrusive Quality and Intelligibility Measure of Reverberant and Dereverberated SpeechabstractA modulation spectral representation is investigated for non-intrusive quality and intelligibility measurement of reverberant and dereverberated speech. The representation is obtained by means of an auditory-inspired filterbank analysis of critical-band temporal envelopes of the speech signal. Modulation spectral insights are used to develop an adaptive measure termed speech to reverberation modulation energy ratio. Experimental results show the proposed measure outperforming three standard algorithms for tasks involving estimation of multiple dimensions of perceived coloration, as well as quality measurement and intelligibility estimation of reverberant and dereverberated speech. Tiago H. Falk, Chenxi Zheng, Wai-Yip Chan |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Spectro-temporal features for robust far-field speaker identificationabstractFeatures derived from an auditory spectro-temporal represen-tation of speech are proposed for robust far-field speaker iden-tification. The auditory representation is obtained by first filtering the speech signal with a gammatone filterbank. A modulation filterbank is then applied to the temporal enve-lope of each gammatone filter output. Compared to com-monly used mel-frequency cepstral coefficients (MFCC), the proposed features are shown to be more robust to mismatched conditions between enrollment and test data and are less sen-sitive to increasing reverberation time (RT). Experiments with simulated and recorded far-field speech show that a Gaus-sian mixture model based identification system, trained on the proposed features, attains an average improvement in identifi-cation accuracy of 15 % relative to a system trained on MFCC. Improvements of up to 85 % are attained for larger RT. Tiago H. Falk, Wai-Yip Chan |
INTERSPEECH | 1 |
| 2008 | Long-term spectro-temporal information for improved automatic speech emotion classificationabstractThis paper investigates the contribution of features which con-vey long-term spectro-temporal (ST) information for the pur-pose of automatic emotional speech classification. The ST rep-resentation is obtained by means of a modulation filterbank de-composition of long-term temporal envelopes of the outputs of a gammatone filterbank. The two-dimensional discrete cosine transform is used to reduce the dimensionality of the represen-tation; candidate features are then derived from statistics com-puted from the DCT coefficients. Sequential forward feature selection is used to select the most salient features. Two types of experiments are described which use the Berlin emotional speech database to test the performance of the ST features alone and in combination with prosodic features. In a multi-class experiment, simulation results with a support vector classifier show that a 44 % reduction in classification error is attained once prosodic features are combined with the proposed ST fea-tures. Additionally, in a one-against-all experiment, an average increase in F-score of 33 % is attained when the proposed ST features are included. Index Terms: speech emotion recognition, spectro-temporal features, modulation spectrum, affective computing. Siqing Wu, Tiago H. Falk, Wai-Yip Chan |
INTERSPEECH | 2 |
| 2008 | Towards Signal-Based Instrumental Quality Diagnosis for Text-to-Speech SystemsabstractIn this letter, the first steps toward the development of a signal-based instrumental quality measure for text-to-speech (TTS) systems are described. Hidden Markov models (HMM), trained on naturally-produced speech, serve as artificial text- and speaker-independent reference models against which synthesized speech signals are assessed. A normalized log-likelihood measure, computed between perceptual features extracted from synthesized speech and a gender-dependent HMM reference model, is proposed and shown to be a reliable parameter for multidimensional TTS quality diagnosis. Experiments with subjectively scored synthesized speech data show that the proposed measure attains promising estimation performance for quality dimensions labeled overall impression, listening effort, naturalness, continuity/fluency, and acceptance. Tiago H. Falk, Sebastian Möller 0001 |
IEEE Signal Process. Lett. | 1 |
| 2008 | Hybrid Signal-and-Link-Parametric Speech Quality Measurement for VoIP CommunicationsabstractA hybrid signal-and-link-parametric approach to speech quality measurement for voice-over-Internet protocol (VoIP) communications is described. Connection parameters are used to determine a base quality representative of the transmission link. Degradation factors, computed from perceptual features extracted from the decoded speech signal, are used to quantify distortions not captured by the connection parameters. The algorithm is tested on speech degraded by acoustic noise, temporal clippings, and noise suppression artifacts, thus simulating degradations present in wireless-VoIP tandem connections. Hybrid measurement is shown to overcome the limitations of pure link parametric and pure signal-based measurement methods, resulting in better measurement accuracy for modern VoIP communications. In addition, the proposed algorithm incurs modest computational overhead relative to pure link parametric measurement and attains up to 88% reduction in processing time relative to the ITU-T standard P.563 signal-based algorithm. Tiago H. Falk, Wai-Yip Chan |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | A Hybrid Signal-and-Link-Parametric Approach to Single-Ended Quality Measurement of Packetized SpeechabstractA hybrid signal-and-link-parametric approach to single-ended quality measurement of packetized speech is proposed. Transmission link parameters are used to determine a base quality for the test signal. The base quality is adjusted by degradation factors calculated from perceptual features extracted from the test signal. The degradation factors are based on Kullback-Leibler distances between a parametric model trained online for the extracted features and reference models of normative speech behavior. The proposed method overcomes the limitations of pure link parametric and pure signal-based methods. Tiago H. Falk, Wai-Yip Chan |
ICASSP (4) | 1 |
| 2007 | Noise suppression based on extending a speech-dominated modulation bandabstractPrevious work on bandpass modulation filtering for noise suppression has resulted in unwanted perceptual artifacts and decreased speech clarity. Artifacts are introduced mainly due to half-wave rectification, which is employed to correct for negative power spectral values resultant from the filtering process. In this paper, modulation frequency estimation (i.e., bandwidth extension) is used to improve perceptual quality. Experiments demonstrate that speech-component lowpass modulation content can be reliably estimated from bandpass modulation content of speech-plus-noise components. Subjective listening tests corroborate that improved quality is attained when the removed speech lowpass modulation content is compensated for by the estimate. Tiago H. Falk, Svante Stadler, W. Bastiaan Kleijn, Wai-Yip Chan |
INTERSPEECH | 1 |
| 2007 | Spectro-temporal processing for blind estimation of reverberation time and single-ended quality measurement of reverberant speechabstractAuditory spectro-temporal representations of reverberant speech are investigated for blind estimation of reverber-ation time (RT) and for single-ended measurement of speech quality. The auditory representations are obtained from an eight-filter filterbank which is used to extract the modulation spectra from temporal envelopes of the speech signal. Gaussian mixture models (GMM), one for each modulation channel and trained on clean speech signals, serve as reference models of normative speech behavior. Consistency measures, computed between re-verberant test signals and each GMM, are mapped to an estimated RT and to an estimated quality score. Experi-ments show that the proposed measures achieve superior performance relative to current “state-of-art ” algorithms. Index Terms: Reverberation time, quality measurement, GMM, modulation spectrum, consistency. Tiago H. Falk, Wai-Yip Chan |
INTERSPEECH | 1 |
| 2007 | Degradation-classification assisted single-ended quality measurement of speechabstractWe propose an algorithm to classify speech degradations at network endpoints and to estimate the speech quality based on the degradation classification decision. Percep-tual features from degraded speech signals are used to form statistical reference models of different degradation classes. Consistency measures, calculated between de-graded speech signals and the reference models, are used to train a degradation classifier and mean opinion score (MOS) mappings. The quality of a received speech signal is estimated based on its degradation class and the MOS mapping associated with the class. Experimental results show that the proposed algorithm achieves high classifi-cation accuracy, and degradation classification improves the accuracy of the quality estimate. Index Terms: speech communication network, speech degradations, speech transmission impairments, degrada-tion classification, speech quality measurement. 1. Tiago H. Falk, Wai-Yip Chan |
INTERSPEECH | 2 |
| 2006 | Enhanced Non-Intrusive Speech Quality Measurement Using Degradation ModelsabstractThe speech quality estimation scheme in [1] is improved with the addition of a reference model of the behavior of speech degraded by different transmission and/or coding schemes. Moreover, via maximization of a mutual information measure, we validate the use of segmental SNR as a measure of the amount of multiplicative noise present in the test signal. These two additions result in an algorithm that is more accurate and more robust to certain distortion conditions. When tested on unseen data, the proposed algorithm outperforms the current "state-of- art" P.563 algorithm while requiring considerably lower computational complexity. Tiago H. Falk, Wai-Yip Chan |
ICASSP (1) | 1 |
| 2006 | Nonintrusive speech quality estimation using Gaussian mixture modelsabstractAn algorithm for nonintrusive speech quality estimation based on Gaussian mixture models (GMMs) is presented. GMMs are used to form an artificial reference model of the behavior of features of undegraded speech. Consistency measures between the degraded speech signal and the reference model serve as indicators of speech quality. Consistency values are mapped to an objective speech quality score using a multivariate adaptive regression splines function. When tested on unseen data, the proposed algorithm generally outperforms ITU-T standard P.563, which is the current "state-of-the-art" algorithm. The algorithm computes objective quality scores roughly twice as fast as P.563. Tiago H. Falk, Wai-Yip Chan |
IEEE Signal Process. Lett. | 1 |
| 2006 | Single-Ended Speech Quality Measurement Using Machine Learning MethodsabstractWe describe a novel single-ended algorithm constructed from models of speech signals, including clean and degraded speech, and speech corrupted by multiplicative noise and temporal discontinuities. Machine learning methods are used to design the models, including Gaussian mixture models, support vector machines, and random forest classifiers. Estimates of the subjective mean opinion score (MOS) generated by the models are combined using hard or soft decisions generated by a classifier which has learned to match the input signal with the models. Test results show the algorithm outperforming ITU-T P.563, the current "state-of-art" standard single-ended algorithm. Employed in a distributed double-ended measurement configuration, the proposed algorithm is found to be more effective than P.563 in assessing the quality of noise reduction systems and can provide a functionality not available with P.862 PESQ, the current double-ended standard algorithm Tiago H. Falk, Wai-Yip Chan |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Non-Intrusive GMM-Based Speech Quality MeasurementabstractWe propose a non-intrusive speech quality measurement algorithm based on using Gaussian-mixture probability models of features of undegraded speech signals as an artificial reference model of "clean" speech behaviour. The consistency between the features of the test speech signal and the reference model serves as an indicator of speech quality. Consistency measures are calculated and mapped to an objective speech quality score using a multivariate adaptive regression splines function. Simulation results show that the proposed method offers accurate and yet low-complexity measurement of speech quality. Tiago H. Falk, Qingfeng Xu, Wai-Yip Chan |
ICASSP (1) | 1 |
| 2005 | An improved GMM-based voice quality predictorabstractA voice quality prediction method based on Gaussian mixture models (GMMs) is improved by constructing a feature selection algorithm to provide the best GMM-based prediction quality. The proposed sequential se-lection algorithm performs N-survivor search, allowing for trading between design complexity and performance. Simulation shows that predictors designed using the pro-posed algorithm outperform two benchmark selection al-gorithms. Performance improvements over the ITU-T P.862 PESQ standard are also attained. 1. Tiago H. Falk, Wai-Yip Chan, Peter Kabal |
INTERSPEECH | 1 |
| 2004 | Speech quality estimation using Gaussian mixture modelsabstractAbstract—An algorithm for nonintrusive speech quality esti-mation based on Gaussian mixture models (GMMs) is presented. GMMs are used to form an artificial reference model of the behavior of features of undegraded speech. Consistency measures between the degraded speech signal and the reference model serve as indicators of speech quality. Consistency values are mapped to an objective speech quality score using a multivariate adaptive regression splines function. When tested on unseen data, the proposed algorithm generally outperforms ITU-T standard P.563, which is the current “state-of-the-art ” algorithm. The algorithm computes objective quality scores roughly twice as fast as P.563. Index Terms—Gaussian mixtures, quality assurance, quality measurement, quality of service, speech coding, speech quality, speech transmission, telephony. I. Tiago H. Falk, Wai-Yip Chan, Peter Kabal |
INTERSPEECH | 1 |