Mathew Magimai-Doss

dblp:76/3728 · also Mathew M. Doss · DBLP profile ↗
← Back
141ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0002-8714-1409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 121 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 79 · 3 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Security and privacy · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cross-modal enhancement of speech representations via textual supervision for paralinguistic analysis
L. Felipe Parra-Gallego, Mathew Magimai-Doss, Juan Rafael Orozco-Arroyave
Speech Commun.2
2025 Emotion information recovery potential of wav2vec2 network fine-tuned for speech recognition task
abstract
Fine-tuning has become a norm to achieve state-of-the-art performance when employing pre-trained networks like foundation models. These models are typically pre-trained on large-scale unannotated data using self-supervised learning (SSL) methods. The SSL-based pre-training on large-scale data enables the network to learn the inherent structure/properties of the data, providing it with capabilities in generalization and knowledge transfer for various downstream tasks. However, when fine-tuned for a specific task, these models become task-specific. Finetuning may cause distortions in the patterns learned by the network during pre-training. In this work, we investigate these distortions by analyzing the network’s information recovery capabilities by designing a study where speech emotion recognition is the target task and automatic speech recognition is an intermediary task. We show that the network recovers the task-specific information but with a shift in the decisions also through attention analysis, we demonstrate some layers do not recover the information fully.
Tilak Purohit, Mathew Magimai-Doss
ICASSP2
2025 Automatic Parkinson's disease detection from speech: Layer selection vs adaptation of foundation models
abstract
In this work, we investigate Speech Foundation Models (SFMs) for Parkinson’s Disease (PD) detection. We explore two main approaches: (1) using SFMs as frozen feature extractors and, (2) fine-tuning/adapting SFMs for PD detection. We propose a cross-validation-based layer selection methodology to identify the layer effective for PD detection. Additionally, we compare the performance of the layer selection scheme with full fine-tuning and, parameter-efficient fine-tuning (PEFT) using Low-Rank Adaptation (LoRA). Our results show that layer selection and LoRA-based fine-tuning can perform on par with full fine-tuning, providing a more parameter-efficient alternative. The highest accuracy was achieved by fine-tuning Whisper using LoRA.
Tilak Purohit, Barbara Ruvolo, Juan Rafael Orozco-Arroyave, Mathew Magimai-Doss
ICASSP4
2025 Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
abstract
Self-supervised learning (SSL) foundation models have emerged as powerful, domain-agnostic, general-purpose feature extractors applicable to a wide range of tasks. Such models pre-trained on human speech have demonstrated high transferability for bioacoustic processing. This paper investigates (i) whether SSL models pre-trained directly on animal vocalizations offer a significant advantage over those pre-trained on speech, and (ii) whether fine-tuning speech-pretrained models on automatic speech recognition (ASR) tasks can enhance bioacous-tic classification. We conduct a comparative analysis using three diverse bioacoustic datasets and two different bioacoustic tasks. Results indicate that pre-training on bioacoustic data provides only marginal improvements over speech-pretrained models, with comparable performance in most scenarios. Fine-tuning on ASR tasks yields mixed outcomes, suggesting that the general-purpose representations learned during SSL pre-training are already well-suited for bioacoustic tasks. These findings highlight the robustness of speech-pretrained SSL models for bioacoustics and imply that extensive fine-tuning may not be necessary for optimal performance.
Eklavya Sarkar, Mathew Magimai-Doss
ICASSP2
2025 Towards Dynamic Skeleton-based Handshape Subunits for Sign Language Assessment
abstract
Sign languages convey information through multiple channels. The handshape channel is an important manual component for conveying the message. In the literature, it is mainly modeled as a sequence of images of discrete postures even in the case of dynamic gestures, leading to blurring problems in detection. Furthermore, to model these discrete postures using deep learning frame-level labeling of the sign language videos is also required, which is time consuming and human intensive. In this paper, as opposed to modeling the handshape information through images of discrete postures, we propose dynamic modeling through skeletal information. More precisely, we develop an approach that combines HamNoSys-based prior knowledge and sign language data to derive dynamic handshape units by modeling skeletal features using hidden Markov models. We demonstrate the effectiveness of the proposed approach through sign language assessment study, sign language recognition, and handshape recognition analysis on the SMILE DSGS corpus.
Sandrine Tornay, Mathew Magimai-Doss
ICASSP2
2025 Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
Karl El Hajal, Enno Hermann, Sevada Hovsepyan, Mathew Magimai-Doss
INTERSPEECH4
2025 Speech power spectra: a window into neural oscillations in Parkinson's disease
Sevada Hovsepyan, Mathew Magimai-Doss
INTERSPEECH2
2025 Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alumäe, Mathew Magimai-Doss
INTERSPEECH4
2025 Children's Voice Privacy: First Steps and Emerging Challenges
abstract
International audience
Ajinkya Kulkarni, Francisco Teixeira, Enno Hermann, Thomas Rolland, Isabel Trancoso, Mathew Magimai-Doss
INTERSPEECH6
2025 Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode Prediction
Bogdan Vlasenko, Mathew Magimai-Doss
INTERSPEECH2
2024 Syllable Level Features for Parkinson's Disease Detection from Speech
abstract
Early detection of Parkinson’s disease (PD), one of the most common neurodegenerative diseases, is crucial for successful treatment and symptom management. In this study, we propose a novel approach inspired by neurocomputational models of speech perception, for PD detection from speech samples. Our proposal emphasises the importance of acoustic/linguistic markers to extract features at the syllable level, in contrast to conventional methods that extract features at the frame or state level. Through the use of syllable-level features (SLF), we successfully identify PD in recorded speech samples. Remarkably, the results not only match but potentially exceed the effectiveness of traditional feature sets used for this purpose. We hope that the proposed approach will provide a new basis for integrating linguistic insights into the identification of speech-related diseases.
Sevada Hovsepyan, Mathew Magimai-Doss
ICASSP2
2024 Content-Based Objective Evaluation of Artificially Generated Sign Language Videos
abstract
Sign language is vital for communication within the deaf and hard-of-hearing community. Avatar-based methods and deep learning techniques like Generative Adversarial Networks have shown promise in generating sign language video content. One of the challenges in sign language generation is the evaluation of the generated video content. One possible solution is to subjectively evaluate using human raters. This is time-consuming and costly. The other possible solution is objective evaluation. In the literature, video quality metrics such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) and skeleton-based measures such as MSE have been proposed. A limitation of these approaches is that they do not provide information about the generated video content. In this paper, we propose a novel phonology-based approach that evaluates the generated video along different channels, namely, hand movement and handshape, which convey the linguistic information in sign language. More precisely, in this approach an objective score is obtained by extracting sequences of hand movement sub-units and handshape sub-units class conditional probabilities (posterior features) from the source and generated videos and comparing them using dynamic time warping. Our experimental studies demonstrate that the proposed objective scoring method yields a better correlation to subjective human ratings than PSNR, SSIM, and MSE-based metrics.
Neha Tarigopula, Preyas Garg, Skanda Muralidhar, Sandrine Tornay, Dinesh Babu Jayagopi, Mathew Magimai-Doss
ICASSP6
2024 Comparing data-Driven and Handcrafted Features for Dimensional Emotion Recognition
abstract
Speech Emotion Recognition (SER) has garnered significant attention over the past two decades. In the early stages of SER technology, ’brute force’-based techniques led to a significant expansion in knowledge-based acoustic feature representation (FR) for modeling sparse emotional data. However, as deep learning techniques have become more powerful, their direct application has been limited by the scarcity of well-annotated emotional data. As a result, pretrained neural embeddings on large speech corpora have gained popularity for SER tasks. These embeddings leverage existing transfer learning methods suitable for general-purpose self-supervised learning (SSL) representations. Recent studies on downstream SSL techniques for dimensional SER have shown promising results. In this research, we aim to evaluate the emotion-discriminative characteristics of neural embeddings in general cases (out-of-domain) and when fine-tuned for SER (in-domain). Given that most SSL techniques are pre-trained primarily on English speech, we plan to use speech emotion corpora in both language-matched and mismatched conditions. We will assess the discriminative characteristics of both handcrafted and standalone neural embeddings as FRs.
Bogdan Vlasenko, Sargam Vyas, Mathew Magimai-Doss
ICASSP3
2024 Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
Gasser Elbanna, Zohreh Mostaani, Mathew Magimai-Doss
INTERSPEECH3
2024 Neurocomputational model of speech recognition for pathological speech detection: a case study on Parkinson's disease speech detection
Sevada Hovsepyan, Mathew Magimai-Doss
INTERSPEECH2
2024 Towards interfacing large language models with ASR systems using confidence measures and prompting
Maryam Naderi, Enno Hermann, Alexandre Nanchen, Sevada Hovsepyan, Mathew Magimai-Doss
INTERSPEECH5
2024 Cross-transfer Knowledge between Speech and Text Encoders to Evaluate Customer Satisfaction
L. Felipe Parra-Gallego, Tilak Purohit, Bogdan Vlasenko, Juan Rafael Orozco-Arroyave, Mathew Magimai-Doss
INTERSPEECH5
2024 On the Quantization of Neural Models for Speaker Verification
abstract
This paper addresses the sub-optimality of current post-training quantization (PTQ) and quantization-aware training (QAT) methods for state-of-the-art speaker verification (SV) models featuring intricate architectural elements such as channel aggregation and squeeze excitation modules. To address these limitations, we propose 1) a data-independent PTQ technique employing iterative low-precision calibration on pre-trained models; and 2) a data-dependent QAT method designed to reduce the performance gap between full-precision and integer models. Our QAT involves two progressive stages where FP-32 weights are initially transformed into FP-8, adapting precision based on the gradient norm, followed by the learning of quantizer parameters (scale and zero-point) for INT8 conversion. Experimental validation underscores the ingenuity of our method in model quantization, demonstrating reduced floating-point operations and INT8 inference time, all while maintaining performance on par with full-precision models.
Vinayak Abrol, Mathew Magimai-Doss
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Towards Learning Emotion Information from Short Segments of Speech
abstract
Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this approach has been successful, there is a growing interest in modelling speech emotion information at the short segment level, at around 250ms-500ms (e.g. the 2021-22 MuSe Challenges). This paper investigates both hand-crafted feature-based and end-to-end raw waveform DNN approaches for modelling speech emotion information in such short segments. Through experimental studies on IEMOCAP corpus, we demonstrate that the end-to-end raw waveform modelling approach is more effective than using hand-crafted features for short-segment level modelling. Furthermore, through relevance signal-based analysis of the trained neural networks, we observe that the top performing end-to-end approach tends to emphasize cepstral information instead of spectral information (such as flux and harmonicity).
Tilak Purohit, Sarthak Yadav, Bogdan Vlasenko, S. Pavankumar Dubagunta, Mathew Magimai-Doss
ICASSP5
2023 Few-shot Dysarthric Speech Recognition with Text-to-Speech Data Augmentation
Enno Hermann, Mathew Magimai-Doss
INTERSPEECH2
2023 Using Commercial ASR Solutions to Assess Reading Skills in Children: A Case Report
abstract
Reading is an acquired skill that is essential for integrating and participating in today's society.Yet, becoming literate can be particularly laborious for some children.Identifying reading difficulties early enough is the first, necessary step toward remediation.Here we investigate the opportunities and limitations of integrating commercial, off-the-shelf automatic speech recognition (ASR) services from IBM Watson to ease the administration and evaluation of children's reading assessment tests in French and Italian.
Timothy Piton, Enno Hermann, Angela Pasqualotto, Marjolaine Cohen, Mathew Magimai-Doss, Daphne Bavelier
INTERSPEECH5
2023 Implicit phonetic information modeling for speech emotion recognition
Tilak Purohit, Bogdan Vlasenko, Mathew Magimai-Doss
INTERSPEECH3
2023 Can Self-Supervised Neural Representations Pre-Trained on Human Speech distinguish Animal Callers?
abstract
Self-supervised learning (SSL) models use only the intrinsic structure of a given signal, independent of its acoustic domain, to extract essential information from the input to an embedding space.This implies that the utility of such representations is not limited to modeling human speech alone.Building on this understanding, this paper explores the cross-transferability of SSL neural representations learned from human speech to analyze bio-acoustic signals.We conduct a caller discrimination analysis and a caller detection study on Marmoset vocalizations using eleven SSL models pre-trained with various pretext tasks.The results show that the embedding spaces carry meaningful caller information and can successfully distinguish the individual identities of Marmoset callers without fine-tuning.This demonstrates that representations pre-trained on human speech can be effectively applied to the bio-acoustics domain, providing valuable insights for future investigations in this field.
Eklavya Sarkar, Mathew Magimai-Doss
INTERSPEECH2
2022 Modeling of Pre-Trained Neural Network Embeddings Learned From Raw Waveform for COVID-19 Infection Detection
abstract
COVID-19 is a respiratory system disorder that can disrupt the function of lungs. Effects of dysfunctional respiratory mechanism can reflect upon other modalities which function in close coupling. Audio signals result from modulation of respiration through speech production system, and hence acoustic information can be modeled for detection of COVID-19. In that direction, this paper is addressing the second DiCOVA challenge that deals with COVID-19 detection based on speech, cough and breathing. We investigate modeling of (a) ComParE LLD representations derived at frame- and turn-level resolutions and (b) neural representations obtained from pre-trained neural networks trained to recognize phones and estimate breathing patterns. On Track 1, the ComParE LLD representations yield a best performance of 78.05% area under the curve (AUC). Experimental studies on Track 2 and Track 3 demonstrate that neural representations tend to yield better detection than ComParE LLD representations. Late fusion of different utterance level representations of neural embeddings yielded a best performance of 80.64% AUC.
Zohreh Mostaani, RaviShankar Prasad, Bogdan Vlasenko, Mathew Magimai-Doss
ICASSP4
2022 Towards Accessible Sign Language Assessment and Learning
abstract
Recently, a phonology-based sign language assessment approach has been proposed using sign language production acquired in 3D space using Kinect sensor. In order to scale the sign language assessment system to realistic application, there is need to reduce the dependency on Kinect, which is not accessible to wider community, and develop solutions that can potentially work with web-cameras. This paper takes a step in that direction by investigating sign language recognition and sign language assessment in 2D space either by dropping the depth coordinate in Kinect or using methods for skeleton estimation from videos. Experimental studies on Swiss German Sign Language corpus SMILE show that, while loss of depth information leads to considerable drop in sign language recognition performance, high level of sign language assessment performance can still be obtained.
Neha Tarigopula, Sandrine Tornay, Skanda Muralidhar, Mathew Magimai-Doss
ICMI4
2022 On Breathing Pattern Information in Synthetic Speech
abstract
The respiratory system is an integral part of human speech production. As a consequence, there is a close relation between respiration and speech signal, and the produced speech signal carries breathing pattern related information. Speech can also be generated using speech synthesis systems. In this paper, we investigate whether synthetic speech carries breathing pattern related information in the same way as natural human speech. We address this research question in the framework of logical-access presentation attack detection using embeddings extracted from neural networks pre-trained for speech breathing pattern estimation. Our studies on ASVSpoof 2019 challenge data show that there is a clear distinction between the extracted breathing pattern embedding of natural human speech and synthesized speech, indicating that speech synthesis systems tend to not carry breathing pattern related information in the same way as human speech. Whilst, this is not the case with voice conversion of natural human speech.
Zohreh Mostaani, Mathew Magimai-Doss
INTERSPEECH2
2022 Unsupervised Voice Activity Detection by Modeling Source and System Information using Zero Frequency Filtering
abstract
Voice activity detection (VAD) is an important pre-processing step for speech technology applications. The task consists of deriving segment boundaries of audio signals which contain voicing information. In recent years, it has been shown that voice source and vocal tract system information can be extracted using zero-frequency filtering (ZFF) without making any explicit model assumptions about the speech signal. This paper investigates the potential of zero-frequency filtering for jointly modeling voice source and vocal tract system information, and proposes two approaches for VAD. The first approach demarcates voiced regions using a composite signal composed of different zero-frequency filtered signals. The second approach feeds the composite signal as input to the rVAD algorithm. These approaches are compared with other supervised and unsupervised VAD methods in the literature, and are evaluated on the Aurora-2 database, across a range of SNRs (20 to -5 dB). Our studies show that the proposed ZFF-based methods perform comparable to state-of-art VAD methods and are more invariant to added degradation and different channel characteristics.
Eklavya Sarkar, RaviShankar Prasad, Mathew Magimai-Doss
INTERSPEECH3
2022 Adjustable deterministic pseudonymization of speech
abstract
While public speech resources become increasingly available, there is a growing interest to preserve the privacy of the speakers, through methods that anonymize the speaker information from speech while preserving the spoken linguistic content. In this paper, a method for pseudonymization (reversible anonymization) of speech is presented, that allows to obfuscate the speaker identity in untranscribed running speech. The approach manipulates the spectro-temporal structure of the speech to simulate a different length and structure of the vocal tract by modifying the formant locations, as well as by altering the pitch and speaking rate. The method is deterministic and partially reversible, and the changes are adjustable on a continuous scale. The method has been evaluated in terms of (i) ABX listening experiments, and (ii) automatic speaker verification and speech recognition. ABX experimental results indicate that the speaker identifiability among forced choice pairs reduced from over 90% to less than 70% through pseudonymization, and that de-pseudonymization was partially effective. An evaluation on the VoicePrivacy 2020 challenge data showed that the proposed approach performs better than the signal processing based baseline method that uses McAdams coefficient and performs slightly worse than the neural source filtering based baseline method. Further analysis showed that the proposed approach: (i) is comparable to the neural source filtering baseline based method in terms of phone posterior feature based objective intelligibility measure, (ii) preserves formant tracks better than the McAdams based method, and (iii) preserves paralinguistic aspects such as dysarthria in several speakers.
S. Pavankumar Dubagunta, R. J. J. H. van Son, Mathew Magimai-Doss
Comput. Speech Lang.3
2021 On The Relationship Between Speech-Based Breathing Signal Prediction Evaluation Measures and Breathing Parameters Estimation
abstract
The respiratory system is one of the major components of the speech production system. Any alteration in breathing can result in changes in speech. Specific breathing characteristics, such as breathing rate and tidal volume, can indicate a person’s pathological condition. More recently, neural network-based methods have started emerging for predicting the breathing signal from the speech signal. The neural networks are trained and evaluated with different objective measures, such as mean squared error (MSE) and Pearson’s correlation. This paper investigates whether there is a systematic relationship between the different objective measures used for training and evaluating the neural network models and the end-goal, i.e. estimation of breathing parameters such as, breathing rate and tidal volume. Our investigations on two different data sets with two different neural network-based approaches show that there is no clear systematic relationship. In other words, obtaining a high Pearson’s correlation on the evaluation set does not necessarily mean better breathing parameter estimation. Thus, indicating the need for developing other objective evaluation measures.
Zohreh Mostaani, Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik, Mathew Magimai-Doss
ICASSP5
2021 Approximating the Mental Lexicon from Clinical Interviews as a Support Tool for Depression Detection
abstract
Depression disorder is one of the major causes of disability in the world that can lead to tragic outcomes. In this paper, we propose a method for using an approximation to a mental lexicon to model the communication process of depressed and non-depressed participants in spontaneous North American English clinical interviews. Our approach, inspired by the Lexical Availability theory, identifies the most relevant vocabulary of the interviewed participant, and use it as features in a classification process. We performed an in-depth evaluation on the DAIC-WOZ [20] and the E-DAIC [11] clinical datasets. Obtained results indicate that our approach can compete against recent contextual embeddings when modeling and identifying depression. We show the generalization capabilities of our algorithm using outside data, reaching a macro F1 = 0.83 and F1 = 0.80 in the DAIC-WOZ and E-DAIC datasets respectively. An analysis of our method’s interpretability allows understanding how the classifier is making its decisions. During this process, we observed strong connections between our obtained results and previous research from the psychological field.
Esaú Villatoro-Tello, Gabriela Ramírez-de-la-Rosa, Daniel Gatica-Perez, Mathew Magimai-Doss, Héctor Jiménez-Salazar
ICMI4
2021 Handling Acoustic Variation in Dysarthric Speech Recognition Systems Through Model Combination
abstract
Developing automatic speech recognition (ASR) systems that recognise dysarthric speech as well as control speech from unimpaired speakers remains challenging. Including more highly variable dysarthric speech during training can also negatively affect the performance on control speakers, which is not desirable when developing speech recognisers for a wider audience. In this work, we analyse how the acoustic variability of dysarthric speech affects ASR systems and propose the combination of multiple acoustic models trained on different subsets of speakers to mitigate this effect. This approach shows improvements for both dysarthric and control speakers on the Torgo and UA-Speech corpora.
Enno Hermann, Mathew Magimai-Doss
Interspeech2
2021 Identification of F1 and F2 in Speech Using Modified Zero Frequency Filtering
RaviShankar Prasad, Mathew Magimai-Doss
Interspeech2
2021 On Modeling Glottal Source Information for Phonation Assessment in Parkinson's Disease
abstract
Parkinson's disease produces several motor symptoms, including different speech impairments that are known as hypokinetic dysarthria. Symptoms associated to dysarthria affect different dimensions of speech such as phonation, articulation, prosody, and intelligibility. Studies in the literature have mainly focused on the analysis of articulation and prosody because they seem to be the most prominent symptoms associated to dysarthria severity. However, phonation impairments also play a significant role to evaluate the global speech severity of Parkinson's patients. This paper proposes an extensive comparison of different methods to automatically evaluate the severity of specific phonation impairments in Parkinson's patients. The considered models include the computation of perturbation and glottal-based features, in addition to features extracted from a zero frequency filtered signals. We consider as well end-to-end models based on 1D CNNs, which are trained to learn features from the raw speech waveform, reconstructed glottal signals, and zero-frequency filtered signals. The results indicate that it is possible to automatically classify between speakers with low versus high phonation severity due to the presence of dysarthria and at the same time to evaluate the severity of the phonation impairments on a continuous scale, posed as a regression problem.
Juan Camilo Vásquez-Correa, Julian Fritsch, Juan Rafael Orozco-Arroyave, Elmar Nöth, Mathew Magimai-Doss
Interspeech5
2021 Late Fusion of the Available Lexicon and Raw Waveform-Based Acoustic Modeling for Depression and Dementia Recognition
abstract
Mental disorders, e.g. depression and dementia, are categorized as priority conditions according to the World Health Organization (WHO). When diagnosing, psychologists employ structured questionnaires/interviews, and different cognitive tests. Although accurate, there is an increasing necessity of developing digital mental health support technologies to alleviate the burden faced by professionals. In this paper, we propose a multi-modal approach for modeling the communication process employed by patients being part of a clinical interview or a cognitive test. The language-based modality, inspired by the Lexical Availability (LA) theory from psycho-linguistics, identifies the most accessible vocabulary of the interviewed subject and use it as features in a classification process. The acoustic-based modality is processed by a Convolutional Neural Network (CNN) trained on signals of speech that predominantly contained voice source characteristics. In the end, a late fusion technique, based on majority voting, assigns the final classification. Results show the complementarity of both modalities, reaching an overall Macro-F1 of 84% and 90% for Depression and Alzheimer's dementia respectively.
Esaú Villatoro-Tello, S. Pavankumar Dubagunta, Julian Fritsch, Gabriela Ramírez-de-la-Rosa, Petr Motlícek, Mathew Magimai-Doss
Interspeech6
2021 Deep learning architectures for estimating breathing signal and respiratory parameters from speech recordings
abstract
Respiration is an essential and primary mechanism for speech production. We first inhale and then produce speech while exhaling. When we run out of breath, we stop speaking and inhale. Though this process is involuntary, speech production involves a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors of the utterance. Thus speech and respiration are closely related, and modeling this relationship makes sensing respiratory dynamics directly from the speech plausible, however is not well explored. In this article, we conduct a comprehensive study to explore techniques for sensing breathing signal and breathing parameters from speech using deep learning architectures and address the challenges involved in establishing the practical purpose of this technology. Estimating the breathing pattern from the speech would give us information about the respiratory parameters, thus enabling us to understand the respiratory health using one's speech.
Venkata Srikanth Nallanthighal, Zohreh Mostaani, Aki Härmä, Helmer Strik, Mathew Magimai-Doss
Neural Networks5
2021 Signal-to-signal neural networks for improved spike estimation from calcium imaging data
abstract
Spiking information of individual neurons is essential for functional and behavioral analysis in neuroscience research. Calcium imaging techniques are generally employed to obtain activities of neuronal populations. However, these techniques result in slowly-varying fluorescence signals with low temporal resolution. Estimating the temporal positions of the neuronal action potentials from these signals is a challenging problem. In the literature, several generative model-based and data-driven algorithms have been studied with varied levels of success. This article proposes a neural network-based signal-to-signal conversion approach, where it takes as input raw-fluorescence signal and learns to estimate the spike information in an end-to-end fashion. Theoretically, the proposed approach formulates the spike estimation as a single channel source separation problem with unknown mixing conditions. The source corresponding to the action potentials at a lower resolution is estimated at the output. Experimental studies on the spikefinder challenge dataset show that the proposed signal-to-signal conversion approach significantly outperforms state-of-the-art-methods in terms of Pearson's correlation coefficient, Spearman's rank correlation coefficient and yields comparable performance for the area under the receiver operating characteristics measure. We also show that the resulting system: (a) has low complexity with respect to existing supervised approaches and is reproducible; (b) is layer-wise interpretable, and (c) has the capability to generalize across different calcium indicators.
Jilt Sebastian, Mriganka Sur, Hema A. Murthy, Mathew Magimai-Doss
PLoS Comput. Biol.4
2021 Utterance Verification-Based Dysarthric Speech Intelligibility Assessment Using Phonetic Posterior Features
abstract
In the literature, the task of dysarthric speech intelligibility assessment has been approached through development of different low-level feature representations, subspace modeling, phone confidence estimation or measurement of automatic speech recognition system accuracy. This paper proposes a novel approach where the intelligibility is estimated as the percentage of correct words uttered by a speaker with dysarthria by matching and verifying utterances of the speaker with dysarthria against control speakers' utterances in phone posterior feature space and broad phonetic posterior feature space. Experimental validation of the proposed approach on the UA-Speech database, with posterior feature estimators trained on the data from auxiliary domain and language, obtained a best Pearson's correlation coefficient (r) of 0.950 and Spearman's correlation coefficient (ρ) of 0.957. Furthermore, replacing control speakers' speech with speech synthesized by a neural text-to-speech system obtained a best r of 0.931 and p of 0.961.
Julian Fritsch, Mathew Magimai-Doss
IEEE Signal Process. Lett.2
2021 On Joint Optimization of Automatic Speaker Verification and Anti-Spoofing in the Embedding Space
abstract
Biometric systems are exposed to spoofing attacks which may compromise their security, and voice biometrics based on automatic speaker verification (ASV), is no exception. To increase the robustness against such attacks, anti-spoofing systems have been proposed for the detection of replay, synthesis and voice conversion-based attacks. However, most proposed anti-spoofing techniques are loosely integrated with the ASV system. In this work, we develop a new integration neural network which jointly processes the embeddings extracted from ASV and anti-spoofing systems in order to detect both zero-effort impostors and spoofing attacks. Moreover, we propose a new loss function based on the minimization of the area under the expected (AUE) performance and spoofability curve (EPSC), which allows us to optimize the integration neural network on the desired operating range in which the biometric system is expected to work. To evaluate our proposals, experiments were carried out on the recent ASVspoof 2019 corpus, including both logical access (LA) and physical access (PA) scenarios. The experimental results show that our proposal clearly outperforms some well-known techniques based on the integration at the score- and embedding-level. Specifically, our proposal achieves up to 23.62% and 22.03% relative equal error rate (EER) improvement over the best performing baseline in the LA and PA scenarios, respectively, as well as relative gains of 27.62% and 29.15% on the AUE metric.
Alejandro Gómez Alanís, José A. González 0001, S. Pavankumar Dubagunta, Antonio M. Peinado, Mathew Magimai-Doss
IEEE Trans. Inf. Forensics Secur.5
2020 Estimating the Degree of Sleepiness by Integrating Articulatory Feature Knowledge in Raw Waveform Based CNNS
abstract
Speech-based degree of sleepiness estimation is an emerging research problem. This paper investigates an end-to-end approach, where given raw waveform as input, a convolutional neural network (CNN) estimates at its output the degree of sleepiness. Within this approach, we investigate constraining the first layer processing and integration of speech production knowledge through transfer learning. We evaluate these methods on the continuous sleepiness corpus of the Interspeech 2019 Computational Paralinguistics (ComParE) Challenge and demonstrate that the proposed approach consistently yields competitive systems. In particular, we observe that integration of speech production knowledge aids in improving the performance and yields systems that are complementary.
Julian Fritsch, S. Pavankumar Dubagunta, Mathew Magimai-Doss
ICASSP3
2020 Dysarthric Speech Recognition with Lattice-Free MMI
abstract
Recognising dysarthric speech is a challenging problem as it differs in many aspects from typical speech, such as speaking rate and pronunciation. In the literature the focus so far has largely been on handling these variabilities in the framework of HMM/GMM and cross-entropy based HMM/DNN systems. This paper focuses on the use of state-of-the-art sequence-discriminative training, in particular lattice-free maximum mutual information (LF-MMI), for improving dysarthric speech recognition. Through a systematic investigation on the Torgo corpus we demonstrate that LF-MMI performs well on such atypical data and compensates much better for the low speaking rates of dysarthric speakers than conventionally trained systems. This can be attributed to inherent aspects of current speech recognition training regimes, like frame subsampling and speed perturbation, which obviate the need for some techniques previously adopted specifically for dysarthric speech.
Enno Hermann, Mathew Magimai-Doss
ICASSP2
2020 Detection Of S1 And S2 Locations In Phonocardiogram Signals Using Zero Frequency Filter
abstract
Heart auscultation is a widely used technique for diagnosing cardiac abnormalities. In that context, capturing of phonocardiogram (PCG) signals and automatically monitoring of the heart by identifying S1 and S2 complexes is an emerging field. One of the first steps involved for identifying S1-S2 complexes is detection of the locations of these events in the PCG signals. Methods proposed in literature, to detect these events in the PCG signal, have largely focused on exploiting the dominant low frequency characteristics of the S1-S2 complexes through frequency-domain processing. In this paper, we propose a purely time-domain processing based method that employs a heavily decaying low pass filter (referred to as zero frequency filter) to suppress extraneous factors and detect S1-S2 locations. We demonstrate the potential of the proposed approach through investigations on two publicly available data sets, namely the PASCAL heart sounds challenge 2011 (PHSC-2011) and Phys- ioNet CinC. The method is also evaluated through an analysis with wearable sensors in the presence and absence of speech activity.
RaviShankar Prasad, Gürkan Yilmaz, Olivier Chételat, Mathew Magimai-Doss
ICASSP4
2020 Towards Multilingual Sign Language Recognition
abstract
Sign language recognition involves modeling of multichannel information such as, hand shapes, hand movements. This requires also sufficient sign language specific data. This is a challenge as sign languages are inherently under-resourced. In the literature, it has been shown that hand shape information can be estimated by pooling resources from multiple sign languages. Such a capability does not exist yet for modeling hand movement information. In this paper, we develop a multilingual sign language approach, where hand movement modeling is also done with target sign language independent data by derivation of hand movement subunits. We validate the proposed approach through an investigation on Swiss German Sign Language, German Sign Language and Turkish Sign Language, and demonstrate that sign language recognition systems can be effectively developed by using multilingual sign language resources.
Sandrine Tornay, Marzieh Razavi, Mathew Magimai-Doss
ICASSP3
2020 A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia Recognition
abstract
Contains fulltext : 228158.pdf (Publisher’s version ) (Open Access)
Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä
INTERSPEECH9
2020 An HMM Approach with Inherent Model Selection for Sign Language and Gesture Recognition
abstract
HMMs have been the one of the first models to be applied for sign recognition and have become the baseline models due to their success in modeling sequential and multivariate data. Despite the extensive use of HMMs for sign recognition, determining the HMM structure has still remained as a challenge, especially when the number of signs to be modeled is high. In this work, we present a continuous HMM framework for modeling and recognizing isolated signs, which inherently performs model selection to optimize the number of states for each sign separately during recognition. Our experiments on three different datasets, namely, German sign language DGS dataset, Turkish sign language HospiSign dataset and Chalearn14 dataset show that the proposed approach achieves better sign language or gesture recognition systems in comparison to the approach of selecting or presetting the number of HMM states based on k-means, and yields systems that perform competitive to the case where the number of states are determined based on the test set performance.
Sandrine Tornay, Oya Aran, Mathew Magimai-Doss
LREC3
2019 Improving Children Speech Recognition through Feature Learning from Raw Speech Signal
abstract
Children speech recognition based on short-term spectral features is a challenging task. One of the reasons is that children speech has high fundamental frequency that is comparable to formant frequency values. Furthermore, as children grow, their vocal apparatus also undergoes changes. This presents difficulties in extracting standard short-term spectral-based features reliably for speech recognition. In recent years, novel acoustic modeling methods have emerged that learn both the feature and phone classifier in an end-to-end manner from the raw speech signal. Through an investigation on PF-STAR corpus we show that children speech recognition can be improved using end-to-end acoustic modeling methods.
S. Pavankumar Dubagunta, Selen Hande Kabil, Mathew Magimai-Doss
ICASSP3
2019 Segment-level Training of ANNs Based on Acoustic Confidence Measures for Hybrid HMM/ANN Speech Recognition
abstract
We show that confidence measures estimated from local posterior probabilities can serve as objective functions for training ANNs in hybrid HMM based speech recognition systems. This leads to a segment-level training paradigm that overcomes the limitation of frame-level updates ignoring the sequence structure in speech. We propose measures that train at the state and phone segment levels, while still decoding in the conventional framework. Experimental results on multiple corpora show that such trainings not only yield better systems in terms of performance, but also give additional improvements with sequence discriminative training. These techniques generalise across front-ends and model architectures, and efficiently handle the effect of segment duration variations on the ANN training.
S. Pavankumar Dubagunta, Mathew Magimai-Doss
ICASSP2
2019 Learning Voice Source Related Information for Depression Detection
abstract
During depression neurophysiological changes can occur, which may affect laryngeal control i.e. behaviour of the vocal folds. Characterising these changes in a precise manner from speech signals is a non trivial task, as this typically involves reliable separation of the voice source information from them. In this paper, by exploiting the abilities of CNNs to learn task-relevant information from the input raw signals, we investigate several methods to model voice source related information for depression detection. Specifically, we investigate modelling of low pass filtered speech signals, linear prediction residual signals, homomorphically filtered voice source signals and zero frequency filtered signals to learn voice source related information for depression detection. Our investigations show that subsegmental level modelling of linear prediction residual signals or zero frequency filtered signals leads to systems better than the state-of-the-art low level descriptor based systems and deep learning based systems modelling the vocal tract system information.
S. Pavankumar Dubagunta, Bogdan Vlasenko, Mathew Magimai-Doss
ICASSP3
2019 HMM-based Approaches to Model Multichannel Information in Sign Language Inspired from Articulatory Features-based Speech Processing
abstract
Sign language conveys information through multiple channels, such as hand shape, hand movement, and mouthing. Modeling this multichannel information is a highly challenging problem. In this paper, we elucidate the link between spoken language and sign language in terms of production phenomenon and perception phenomenon. Through this link we show that hidden Markov model-based approaches developed to model "articulatory" features for spoken language processing can be exploited to model the multichannel information inherent in sign language for sign language processing.
Sandrine Tornay, Marzieh Razavi, Necati Cihan Camgöz, Richard Bowden, Mathew Magimai-Doss
ICASSP5
2019 Using Speech Production Knowledge for Raw Waveform Modelling Based Styrian Dialect Identification
abstract
This paper addresses the Styrian Dialect sub-challenge of the INTERSPEECH 2019 Computational Paralinguistics Challenge.We treat this challenge as dialect identification with no linguistic resources/knowledge and with limited acoustic resources, and develop end-to-end raw waveform modelling based methods that incorporate knowledge related to speech production.In this direction, we investigate two methods: (a) modelling the signals after source system decomposition and (b) transferring knowledge from articulatory feature models trained on English language.Our investigations show that the proposed approaches on the ComParE 2019 Styrian dialect data yield systems that perform better than low level descriptorbased and bag-of-audio-word representation based approaches and comparable to sequence-to-sequence auto-encoder based approach.
S. Pavankumar Dubagunta, Mathew Magimai-Doss
INTERSPEECH2
2019 Understanding and Visualizing Raw Waveform-Based CNNs
abstract
Modeling directly raw waveforms through neural networks for speech processing is gaining more and more attention. Despite its varied success, a question that remains is: what kind of information are such neural networks capturing or learning for different tasks from the speech signal? Such an insight is not only interesting for advancing those techniques but also for understanding better speech signal characteristics. This paper takes a step in that direction, where we develop a gradient based approach to estimate the relevance of each speech sample input on the output score. We show that analysis of the resulting ``relevance signal" through conventional speech signal processing techniques can reveal the information modeled by the whole network. We demonstrate the potential of the proposed approach by analyzing raw waveform CNN-based phone recognition and speaker identification systems.
Hannah Muckenhirn, Vinayak Abrol, Mathew Magimai-Doss, Sébastien Marcel
INTERSPEECH3
2019 End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition
Dimitri Palaz, Mathew Magimai-Doss, Ronan Collobert
Speech Commun.2
2018 Towards Directly Modeling Raw Speech Signal for Speaker Verification Using CNNS
abstract
Speaker verification systems traditionally extract and model cepstral features or filter bank energies from the speech signal. In this paper, inspired by the success of neural network-based approaches to model directly raw speech signal for applications such as speech recognition, emotion recognition and anti-spoofing, we propose a speaker verification approach where speaker discriminative information is directly learned from the speech signal by: (a) first training a CNN-based speaker identification system that takes as input raw speech signal and learns to classify on speakers (unknown to the speaker verification system); and then (b) building a speaker detector for each speaker in the speaker verification system by replacing the output layer of the speaker identification system by two outputs (genuine, impostor), and adapting the system in a discriminative manner with enrollment speech of the speaker and impostor speech data. Our investigations on the Voxforge database shows that this approach can yield systems competitive to state-of-the-art systems. An analysis of the filters in the first convolution layer shows that the filters give emphasis to information in low frequency regions (below 1000 Hz) and implicitly learn to model fundamental frequency information in the speech signal for speaker discrimination.
Hannah Muckenhirn, Mathew Magimai-Doss, Sébastien Marcel
ICASSP2
2018 On Learning to Identify Genders from Raw Speech Signal Using CNNs
abstract
Automatic Gender Recognition (AGR) is the task of identifying the gender of a speaker given a speech signal. Standard approaches extract features like fundamental frequency and cepstral features from the speech signal and train a binary classifier. Inspired from recent works in the area of automatic speech recognition (ASR), speaker recognition and presentation attack detection, we present a novel approach where relevant features and classifier are jointly learned from the raw speech signal in end-to-end manner. We propose a convolutional neural networks (CNN) based gender classifier that consists of: (1) convolution layers, which can be interpreted as a feature learning stage and (2) a multilayer perceptron (MLP), which can be interpreted as a classification stage. The system takes raw speech signal as input, and outputs gender posterior probabilities. Experimental studies conducted on two datasets, namely AVspoof and ASVspoof 2015, with different architectures show that with simple architectures the proposed approach yields better system than standard acoustic features based approach. Further analysis of the CNNs show that the CNNs learn formant and fundamental frequency information for gender identification.
Selen Hande Kabil, Hannah Muckenhirn, Mathew Magimai-Doss
INTERSPEECH3
2018 On Learning Vocal Tract System Related Speaker Discriminative Information from Raw Signal Using CNNs
abstract
In a recent work, we have shown that speaker verification systems can be built where both features and classifiers are directly learned from the raw speech signal with convolutional neural networks (CNNs). In this framework, the training phase also decides the block processing through cross validation. It was found that the first convolution layer, which processes about 20 ms speech, learns to model fundamental frequency information. In the present paper, inspired from speech recognition studies, we build further on that framework to design a CNN-based system, which models sub-segmental speech (about 2ms speech) in the first convolution layer, with an hypothesis that such a system should learn vocal tract system related speaker discriminative information. Through experimental studies on Voxforge corpus and analysis on American vowel dataset, we show that the proposed system (a) indeed focuses on formant regions, (b) yields competitive speaker verification system and (c) is complementary to the CNN-based system that models fundamental frequency information.
Hannah Muckenhirn, Mathew Magimai-Doss, Sébastien Marcel
INTERSPEECH2
2018 Denoising and Raw-waveform Networks for Weakly-Supervised Gender Identification on Noisy Speech
abstract
This paper presents a raw-waveform neural network and uses it along with a denoising network for clustering in weakly supervised learning scenarios under extreme noise conditions. Specifically, we consider language independent Automatic Gender Recognition (AGR) on a set of varied noise conditions and Signal to Noise Ratios (SNRs). We formulate the denoising problem as a source separation task and train the system using a discriminative criterion in order to enhance output SNRs. A denoising Recurrent Neural Network (RNN) is first trained on a small subset (roughly one-fifth) of the data for learning a speech specific mask. The denoised speech signal is then directly fed as input to a raw-waveform convolutional neural network (CNN) trained with denoised speech. We evaluate the standalone performance of denoiser in terms of various signal-to-noise measures and discuss its contribution towards robust AGR. An absolute improvement of 11.06% and 13.33% is achieved by the combined pipeline over the i-vector SVM baseline system for 0 dB and -5 dB SNR conditions, respectively. We further analyse the information captured by the first CNN layer in both noisy and denoised speech.
Jilt Sebastian, Manoj Kumar 0007, Pavan Kumar D. S., Mathew Magimai-Doss, Hema A. Murthy, Shri Narayanan
INTERSPEECH4
2018 Implementing Fusion Techniques for the Classification of Paralinguistic Information
abstract
This work tests several classification techniques and acoustic features and further combines them using late fusion to classify paralinguistic information for the ComParE 2018 challenge. We use Multiple Linear Regression (MLR) with Ordinary Least Squares (OLS) analysis to select the most informative features for Self-Assessed Affect (SSA) sub-Challenge. We also propose to use raw-waveform convolutional neural networks (CNN) in the context of three paralinguistic sub-challenges. By using combined evaluation split for estimating codebook, we obtain better representation for Bag-of-Audio-Words approach. We preprocess the speech to vocalized segments to improve classification performance. For fusion of our leading classification techniques, we use weighted late fusion approach applied for confidence scores. We use two mismatched evaluation phases by exchanging the training and development sets, and this estimates the optimal fusion weight. Weighted late fusion provides better performance on development sets in comparison with baseline techniques. Raw-waveform techniques perform comparable to the baseline.
Bogdan Vlasenko, Jilt Sebastian, Pavan Kumar D. S., Mathew Magimai-Doss
INTERSPEECH4
2018 SMILE Swiss German Sign Language Dataset
Sarah Ebling, Necati Cihan Camgöz, Penny Boyes Braem, Katja Tissi, Sandra Sidler-Miserez, Stephanie Stoll, Simon Hadfield, Tobias Haug, Richard Bowden, Sandrine Tornay, Marzieh Razavi, Mathew Magimai-Doss
LREC12
2018 Towards weakly supervised acoustic subword unit discovery and lexicon development using hidden Markov models
Marzieh Razavi, Ramya Rasipuram, Mathew Magimai-Doss
Speech Commun.3
2017 End-to-End convolutional neural network-based voice presentation attack detection
abstract
Development of countermeasures to detect attacks performed on speaker verification systems through presentation of forged or altered speech samples is a challenging and open research problem. Typically, this problem is approached by extracting features through conventional short-term speech processing and feeding them to a binary classifier. In this article, we develop a convolutional neural network-based approach that learns in an end-to-end manner both the features and the binary classifier from the raw signal. Through investigations on two publicly available databases, namely, ASVspoof and AVspoof, we show that it yields systems comparable to or better than the state-of-the-art approaches for both physical access attacks and logical access attacks. Furthermore, the approach is shown to be complementary to a spectral statistics-based approach, which, similarly to the proposed approach, does not use prior assumptions related to speech signals.
Hannah Muckenhirn, Mathew Magimai-Doss, Sébastien Marcel
IJCB2
2017 A Posterior-Based Multistream Formulation for G2P Conversion
abstract
In the literature, a number of approaches have been proposed for learning grapheme-to-phoneme (G2P) relationship and inferring pronunciations. In this letter, we present a novel multistream framework for G2P conversion, where various machine learning techniques providing different estimates of probability of phonemes given graphemes can be effectively combined during pronunciation inference. More precisely, analogous to multistream automatic speech recognition, the framework involves obtaining different streams of estimates of probability of phonemes given graphemes, combining them based on probability combination rules, and inferring pronunciations by decoding the probabilities resulting after combination. We demonstrate the potential of the proposed approach by combining probabilities estimated by the state-of-the-art conditional random field-based G2P conversion approach and acoustic data-driven G2P conversion approach in the Kullback-Leibler-divergence-based hidden Markov model framework on the PhoneBook 600-word task.
Marzieh Razavi, Mathew Magimai-Doss
IEEE Signal Process. Lett.2
2017 Long-Term Spectral Statistics for Voice Presentation Attack Detection
abstract
Automatic speaker verification systems can be spoofed through recorded, synthetic, or voice converted speech of target speakers. To make these systems practically viable, the detection of such attacks, referred to as presentation attacks, is of paramount interest. In that direction, this paper investigates two aspects: 1) a novel approach to detect presentation attacks where, unlike conventional approaches, no speech signal modeling related assumptions are made, rather the attacks are detected by computing first-order and second-order spectral statistics and feeding them to a classifier, and 2) generalization of the presentation attack detection systems across databases. Our investigations on ASVspoof 2015 challenge database and AVspoof database show that, when compared to the approaches based on conventional short-term spectral features, the proposed approach with a linear discriminative classifier yields a better system, irrespective of whether the spoofed signal is replayed to the microphone or is directly injected into the system software process. Cross-database investigations show that neither the short-term spectral processing-based approaches nor the proposed approach yield systems which are able to generalize across databases or methods of attack. Thus, revealing the difficulty of the problem and the need for further resources and research.
Hannah Muckenhirn, Pavel Korshunov, Mathew Magimai-Doss, Sébastien Marcel
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 HMM-Based Non-Native Accent Assessment Using Posterior Features
abstract
Automatic non-native accent assessment has potential benefits in language learning and speech technologies. The three fundamental challenges in automatic accent assessment are to characterize, model and assess individual variation in speech of the non-native speaker. In our recent work, accentedness score was automatically obtained by comparing two phone probability sequences obtained through instances of non-native and native speech. Although automatic accentedness ratings of the approach correlated well with human accent ratings, the approach is critically constrained because of the requirement of native speech instance. In this paper, we build on the previous work and obtain the native latent symbol probability sequence through the word hypothesis modeled as a hidden Markov model (HMM). The latent symbols are either context-independent phonemes or clustered context-dependent phonemes. The advantage of the proposed approach is that it requires just reference text transcription instead of native speech recordings. Using the HMMs trained on an auxiliary native speech corpus, the proposed approach achieves a correlation of 0.68 with human accent ratings on the ISLE corpus. This is further interesting considering that the approach does not use any non-native data and human accent ratings at any stage of the system development.
Ramya Rasipuram, Milos Cernak, Mathew Magimai-Doss
INTERSPEECH3
2016 Improving Under-Resourced Language ASR Through Latent Subword Unit Space Discovery
abstract
LIDIAP
Marzieh Razavi, Mathew Magimai-Doss
INTERSPEECH2
2016 Articulatory feature based continuous speech recognition using probabilistic lexical modeling
Ramya Rasipuram, Mathew Magimai-Doss
Comput. Speech Lang.2
2016 Acoustic data-driven grapheme-to-phoneme conversion in the probabilistic lexical modeling framework
Marzieh Razavi, Ramya Rasipuram, Mathew Magimai-Doss
Speech Commun.3
2015 Convolutional Neural Networks-based continuous speech recognition using raw speech signal
abstract
State-of-the-art automatic speech recognition systems model the relationship between acoustic speech signal and phone classes in two stages, namely, extraction of spectral-based features based on prior knowledge followed by training of acoustic model, typically an artificial neural network (ANN). In our recent work, it was shown that Convolutional Neural Networks (CNNs) can model phone classes from raw acoustic speech signal, reaching performance on par with other existing feature-based approaches. This paper extends the CNN-based approach to large vocabulary speech recognition task. More precisely, we compare the CNN-based approach against the conventional ANN-based approach on Wall Street Journal corpus. Our studies show that the CNN-based approach achieves better performance than the conventional ANN-based approach with as many parameters. We also show that the features learned from raw speech by the CNN-based approach could generalize across different databases.
Dimitri Palaz, Mathew Magimai-Doss, Ronan Collobert
ICASSP2
2015 Integrated pronunciation learning for automatic speech recognition using probabilistic lexical modeling
abstract
Standard automatic speech recognition (ASR) systems use phoneme-based pronunciation lexicon prepared by linguistic experts. When the hand crafted pronunciations fail to cover the vocabulary of a new domain, a grapheme-to-phoneme (G2P) converter is used to extract pronunciations for new words and then a phoneme based ASR system is trained. G2P converters are typically trained only on the existing lexicons. In this paper, we propose a grapheme based ASR approach in the framework of probabilistic lexical modeling that integrates pronunciation learning as a stage in ASR system training, and exploits both acoustic and lexical resources (not necessarily from the domain or language of interest). The proposed approach is evaluated on four lexical resource constrained ASR tasks and compared with the conventional two stage approach where G2P training is followed by ASR system development.
Ramya Rasipuram, Marzieh Razavi, Mathew Magimai-Doss
ICASSP3
2015 An HMM-based formalism for automatic subword unit derivation and pronunciation generation
abstract
We propose a novel hidden Markov model (HMM) formalism for automatic derivation of subword units and pronunciation generation using only transcribed speech data. In this approach, the subword units are derived from the clustered context-dependent units in a grapheme based system using maximum-likelihood criterion. The subword unit based pronunciations are then learned in the framework of Kullback-Leibler divergence based HMM. The automatic speech recognition (ASR) experiments on WSJ0 English corpus show that the approach leads to 12.7% relative reduction in word error rate compared to grapheme-based system. Our approach can be beneficial in reducing the need for expert knowledge in development of ASR as well as text-to-speech systems.
Marzieh Razavi, Mathew Magimai-Doss
ICASSP2
2015 Objective speech intelligibility assessment through comparison of phoneme class conditional probability sequences
abstract
Assessment of speech intelligibility is important for the development of speech systems, such as telephony systems and text-to-speech (TTS) systems. Existing approaches to the automatic assessment of intelligibility in telephony typically compare a reference speech signal to a degraded copy, which requires that both signals be from the same speaker. In this paper, we propose a novel approach that does not have such a requirement, making it possible to also evaluate TTS systems and recent very low bit rate codecs that may modify speaker characteristics. More specifically, our approach is based on comparing sequences of phoneme class conditional probabilities. We show the potential of our approach on low bit rate telephony conditions, and compare it against subjective TTS intelligibility scores from the 2011 Blizzard Challenge.
Raphael Ullmann, Mathew Magimai-Doss, Hervé Bourlard
ICASSP2
2015 Analysis of CNN-based speech recognition system using raw speech as input
abstract
Automatic speech recognition systems typically model the relationship between the acoustic speech signal and the phones in two separate steps: feature extraction and classifier training.In our recent works, we have shown that, in the framework of convolutional neural networks (CNN), the relationship between the raw speech signal and the phones can be directly modeled and ASR systems competitive to standard approach can be built.In this paper, we first analyze and show that, between the first two convolutional layers, the CNN learns (in parts) and models the phone-specific spectral envelope information of 2-4 ms speech.Given that we show that the CNN-based approach yields ASR trends similar to standard short-term spectral based ASR system under mismatched (noisy) conditions, with the CNN-based approach being more robust.
Dimitri Palaz, Mathew Magimai-Doss, Ronan Collobert
INTERSPEECH2
2015 Automatic accentedness evaluation of non-native speech using phonetic and sub-phonetic posterior probabilities
abstract
Automatic evaluation of non-native speech accentedness has potential implications for not only language learning and accent identification systems but also for speaker and speech recognition systems. From the perspective of speech production, the two primary factors influencing the accentedness are the phonetic and prosodic structure. In this paper, we propose an approach for automatic accentedness evaluation based on comparison of instances of native and non-native speakers at the acoustic-phonetic level. Specifically, the proposed approach measures accentedness by comparing phone class conditional probability sequences corresponding to the instances of native and non-native speakers, respectively. We evaluate the proposed approach on the EMIME bilingual and EMIME Mandarin bilingual corpora, which contains English speech from native English speakers and various non-native English speakers, namely Finnish, German and Mandarin. We also investigate the influence of the granularity of the phonetic unit representation on the performance of the proposed accentedness measure. Our results indicate that the accentedness ratings by the proposed approach correlate consistently with the human ratings of accentedness. In addition, our studies show that the granularity of the phonetic unit representation that yields the best correlation with the human accentedness ratings varies with respect to the native language of the non-native speakers.
Ramya Rasipuram, Milos Cernak, Alexandre Nanchen, Mathew Magimai-Doss
INTERSPEECH4
2015 Objective intelligibility assessment of text-to-speech systems through utterance verification
abstract
Objective assessment of synthetic speech intelligibility can be a useful tool for the development of text-to-speech (TTS) systems, as it provides a reproducible and inexpensive alternative to subjective listening tests. In a recent work, it was shown that the intelligibility of synthetic speech could be assessed objectively by comparing two sequences of phoneme class conditional probabilities, corresponding to instances of synthetic and human reference speech, respectively. In this paper, we build on those findings to propose a novel approach that formulates objective intelligibility assessment as an utterance verification problem using hidden Markov models, thereby alleviating the need for human reference speech. Specifically, given each text input to the TTS system, the proposed approach automatically verifies the words in the output synthetic speech signal and estimates an intelligibility score based on word recall statistics. We evaluate the proposed approach on the 2011 Blizzard Challenge data, and show that the estimated scores and the subjective intelligibility scores are highly correlated (Pearson’s |R| = 0.94).
Raphael Ullmann, Ramya Rasipuram, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2015 Acoustic and lexical resource constrained ASR using language-independent acoustic model and language-dependent probabilistic lexical model
Ramya Rasipuram, Mathew Magimai-Doss
Speech Commun.2
2014 On modeling context-dependent clustered states: Comparing HMM/GMM, hybrid HMM/ANN and KL-HMM approaches
abstract
Deep architectures have recently been explored in hybrid hidden Markov model/artificial neural network (HMM/ANN) framework where the ANN outputs are usually the clustered states of context-dependent phones derived from the best performing HMM/Gaussian mixture model (GMM) system. We can view a hybrid HMM/ANN system as a special case of recently proposed Kullback-Leibler divergence based hidden Markov model (KL-HMM) approach. In KL-HMM approach a probabilistic relationship between the ANN outputs and the context-dependent HMM states is modeled. In this paper, we show that in KL-HMM framework we may not require as many clustered states as the best HMM/GMM system in the ANN output layer. Our experimental results on German part of Media-Parl database show that KL-HMM system achieves better performance compared to hybrid HMM/ANN and HMM/GMM systems with much fewer number of clustered states than is required for HMM/GMM system. The reduction in number of clustered states has broader implications on model complexity and data sparsity issues.
Marzieh Razavi, Ramya Rasipuram, Mathew Magimai-Doss
ICASSP3
2014 On recognition of non-native speech using probabilistic lexical model
abstract
Despite various advances in automatic speech recognition (ASR) technology, recognition of speech uttered by non-native speakers is still a challenging problem.In this paper, we investigate the role of different factors such as type of lexical model and choice of acoustic units in recognition of speech uttered by non-native speakers.More precisely, we investigate the influence of the probabilistic lexical model in the framework of Kullback-Leibler divergence based hidden Markov model (KL-HMM) approach in handling pronunciation variabilities by comparing it against hybrid HMM/artificial neural network (ANN) approach where the lexical model is deterministic.Moreover, we study the effect of acoustic units (being context-independent or clustered context-dependent phones) on ASR performance in both KL-HMM and hybrid HMM/ANN frameworks.Our experimental studies on French part of Me-diaParl as a bilingual corpus indicate that the probabilistic lexical modeling approach in the KL-HMM framework can capture the pronunciation variations present in non-native speech effectively.More precisely, the experimental results show that the KL-HMM system using context-dependent acoustic units and trained solely on native speech data can lead to better ASR performance than adaptation techniques such as maximum likelihood linear regression.
Marzieh Razavi, Mathew Magimai-Doss
INTERSPEECH2
2014 Feature mapping of multiple beamformed sources for robust overlapping speech recognition using a microphone array
abstract
This paper introduces a nonlinear vector-based feature mapping approach to extract robust features for automatic speech recognition (ASR) of overlapping speech using a microphone array. We explore different configurations and additional sources of information to improve the effectiveness of the feature mapping. First, we investigate the full-vector based mapping of different sources in a log mel-filterbank energy (log MFBE) domain, and demonstrate that retraining the acoustic model using the generated training data can help improve the recognition performance. Then we investigate the feature mapping between different domains. Finally in order to improve the qualities of the mapping inputs we propose a nonlinear mapping of the features from multiple beamformed sources, which are directed at the target and interfering speakers, respectively. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, Longbiao Wang, Yicong Zhou, John Dines, Mathew Magimai-Doss, Hervé Bourlard, Qingmin Liao
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 Probabilistic lexical modeling and unsupervised training for zero-resourced ASR
abstract
Standard automatic speech recognition (ASR) systems rely on transcribed speech, language models, and pronunciation dictionaries to achieve state-of-the-art performance. The unavailability of these resources constrains the ASR technology to be available for many languages. In this paper, we propose a novel zero-resourced ASR approach to train acoustic models that only uses list of probable words from the language of interest. The proposed approach is based on Kullback-Leibler divergence based hidden Markov model (KL-HMM), grapheme subword units, knowledge of grapheme-to-phoneme mapping, and graphemic constraints derived from the word list. The approach also exploits existing acoustic and lexical resources available in other resource rich languages. Furthermore, we propose unsupervised adaptation of KL-HMM acoustic model parameters if untranscribed speech data in the target language is available. We demonstrate the potential of the proposed approach through a simulated study on Greek language.
Ramya Rasipuram, Marzieh Razavi, Mathew Magimai-Doss
ASRU3
2013 A probabilistic framework for multiple speaker localization
abstract
This paper presents a novel probabilistic framework for localizing multiple speakers with a microphone array. In this framework, the generalized cross correlation function (GCC) of each microphone pair is interpreted as a probability distribution of the time difference of arrival (TDOA) and subsequently approximated as a Gaussian mixture. The distribution parameters are estimated with a weighted expectation maximization algorithm. Then, the joint distribution of the TDOA Gaussian mixtures is mapped to a multimodal distribution in the location space, where each mode represents a potential source location. The approach taken here performs the localization by 1) reducing the search space to some regions that are likely to contain a source and then 2) extracting the actual speaker locations with a numerical optimization algorithm. The effectiveness of the proposed approach is shown using the AV16.3 corpus.
Youssef Oualil, Mathew Magimai-Doss, Friedrich Faubel, Dietrich Klakow
ICASSP2
2013 Grapheme and multilingual posterior features for under-resourced speech recognition: A study on Scottish Gaelic
abstract
Standard automatic speech recognition (ASR) systems use phonemes as subword units. Thus, one of the primary resource required to build a good ASR system is a well developed phoneme pronunciation lexicon. However, under-resourced languages typically lack such lexical resources. In this paper, we investigate recently proposed grapheme-based ASR in the framework of Kullback-Leibler divergence based hidden Markov model (KL-HMM) for underresourced languages, particularly Scottish Gaelic which has no lexical resources. More specifically, we study the use of grapheme and multilingual phoneme class conditional probabilities (posterior features) as feature observations in KL-HMM. ASR studies conducted show that the proposed approach yields better system compared to the conventional HMM/GMM approach using cepstral features. Furthermore, grapheme posterior features estimated using both auxiliary data and Gaelic data yield the best system.
Ramya Rasipuram, Peter Bell 0001, Mathew Magimai-Doss
ICASSP3
2013 Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks
abstract
In hybrid hidden Markov model/artificial neural networks (HMM/ANN) automatic speech recognition (ASR) system, the phoneme class conditional probabilities are estimated by first extracting acoustic features from the speech signal based on prior knowledge such as, speech perception or/and speech production knowledge, and, then modeling the acoustic features with an ANN. Recent advances in machine learning techniques, more specifically in the field of image processing and text processing, have shown that such divide and conquer strategy (i.e., separating feature extraction and modeling steps) may not be necessary. Motivated from these studies, in the framework of convolutional neural networks (CNNs), this paper investigates a novel approach, where the input to the ANN is raw speech signal and the output is phoneme class conditional probability estimates. On TIMIT phoneme recognition task, we study different ANN architectures to show the benefit of CNNs and compare the proposed approach against conventional approach where, spectral-based feature MFCC is extracted and modeled by a multilayer perceptron. Our studies show that the proposed approach can yield comparable or better phoneme recognition performance when compared to the conventional approach. It indicates that CNNs can learn features relevant for phoneme classification automatically from the raw speech signal.
Dimitri Palaz, Ronan Collobert, Mathew Magimai-Doss
INTERSPEECH3
2013 Improving grapheme-based ASR by probabilistic lexical modeling approach
abstract
There is growing interest in using graphemes as subword units, especially in the context of the rapid development of hidden Markov model (HMM) based automatic speech recognition (ASR) system, as it eliminates the need to build a phoneme pronunciation lexicon. However, directly modeling the relationship between acoustic feature observations and grapheme states may not be always trivial. It usually depends upon the grapheme-to-phoneme relationship within the language. This paper builds upon our recent interpretation of Kullback-Leibler divergence based HMM (KL-HMM) as a probabilistic lexical modeling approach to propose a novel grapheme-based ASR approach where, first a set of acoustic units are derived by modeling context-dependent graphemes in the framework of conventional HMM/Gaussian mixture model (HMM/GMM) system, and then the probabilistic relationship between the derived acoustic units and the lexical units representing graphemes is modeled in the framework of KL-HMM. Through experimental studies on English, where the grapheme-to-phoneme relationship is irregular, we show that the proposed grapheme-based ASR approach (without using any phoneme information) can achieve performance comparable to standard phoneme-based ASR approach.
Ramya Rasipuram, Mathew Magimai-Doss
INTERSPEECH2
2013 A Savitzky-Golay Filtering Perspective of Dynamic Feature Computation
abstract
We address the classical problem of delta feature computation, and interpret the operation involved in terms of Savitzky-Golay (SG) filtering. Features such as the mel-frequency cepstral coefficients (MFCCs), obtained based on short-time spectra of the speech signal, are commonly used in speech recognition tasks. In order to incorporate the dynamics of speech, auxiliary delta and delta-delta features, which are computed as temporal derivatives of the original features, are used. Typically, the delta features are computed in a smooth fashion using local least-squares (LS) polynomial fitting on each feature vector component trajectory. In the light of the original work of Savitzky and Golay, and a recent article by Schafer in IEEE Signal Processing Magazine, we interpret the dynamic feature vector computation for arbitrary derivative orders as SG filtering with a fixed impulse response. This filtering equivalence brings in significantly lower latency with no loss in accuracy, as validated by results on a TIMIT phoneme recognition task. The SG filters involved in dynamic parameter computation can be viewed as modulation filters, proposed by Hermansky.
Sunder Ram Krishnan, Mathew Magimai-Doss, Chandra Sekhar Seelamantula
IEEE Signal Process. Lett.2
2013 Applying Multi- and Cross-Lingual Stochastic Phone Space Transformations to Non-Native Speech Recognition
abstract
In the context of hybrid HMM/MLP Automatic Speech Recognition (ASR), this paper describes an investigation into a new type of stochastic phone space transformation, which maps “source” phone (or phone HMM state) posterior probabilities (as obtained at the output of a Multilayer Perceptron/MLP) into “destination” phone (HMM phone state) posterior probabilities. The resulting stochastic matrix transformation can be used within the same language to automatically adapt to different phone formats (e.g., IPA) or across languages. Additionally, as shown here, it can also be applied successfully to non-native speech recognition. In the same spirit as MLLR adaptation, or MLP adaptation, the approach proposed here is directly mapping posterior distributions, and is trained by optimizing on a small amount of adaptation data a Kullback-Leibler based cost function, along a modified version of an iterative EM algorithm. On a non-native English database (HIWIRE), and comparing with multiple setups (monophone and triphone mapping, MLLR adaptation) we show that the resulting posterior mapping yields state-of-the-art results using very limited amounts of adaptation data in mono-, cross- and multi-lingual setups. We also show that “universal” phone posteriors, trained on a large amount of multilingual data, can be transformed to English phone posteriors, resulting in an ASR system that significantly outperforms a system trained on English data only. Finally, we demonstrate that the proposed approach outperforms alternative data-driven, as well as a knowledge-based, mapping techniques.
David Imseng, Hervé Bourlard, John Dines, Philip N. Garner, Mathew Magimai-Doss
IEEE Trans. Speech Audio Process.5
2012 Acoustic data-driven grapheme-to-phoneme conversion using KL-HMM
abstract
This paper proposes a novel grapheme-to-phoneme (G2P) conversion approach where first the probabilistic relation between graphemes and phonemes is captured from acoustic data using Kullback-Leibler divergence based hidden Markov model (KL-HMM) system. Then, through a simple decoding framework the information in this probabilistic relation is integrated with the sequence information in the orthographic transcription of the word to infer the phoneme sequence. One of the main application of the proposed G2P approach is in the area of low linguistic resource based automatic speech recognition or text-to-speech systems. We demonstrate this potential through a simulation study where linguistic resources from one domain is used to create linguistic resources for a different domain.
Ramya Rasipuram, Mathew Magimai-Doss
ICASSP2
2012 Combining Acoustic Data Driven G2P and Letter-to-Sound Rules for Under Resource Lexicon Generation
abstract
In a recent work, we proposed an acoustic data-driven grapheme-to-phoneme (G2P) conversion approach, where the probabilistic relationship between graphemes and phonemes learned through acoustic data is used along with the orthographic transcription of words to infer the phoneme sequence. In this paper, we extend our studies to under-resourced lexicon development problem. More precisely, given a small amount of transcribed speech data consisting of few words along with its pronunciation lexicon, the goal is to build a pronunciation lexicon for unseen words. In this framework, we compare our G2P approach with standard letter-to-sound (L2S) rule based conversion approach. We evaluated the generated lexicons on PhoneBook 600 words task in terms of pronunciation errors and ASR performance. The G2P approach yields a best ASR performance of 14.0% word error rate (WER), while L2S approach yields a best ASR performance of 13.7% WER. A combination of G2P approach and L2S approach yields a best ASR performance of 9.3% WER.
Ramya Rasipuram, Mathew Magimai-Doss
INTERSPEECH2
2012 Synthetic References for Template-based ASR using posterior features
abstract
Recently, the use of phoneme class-conditional probabilities as features (posterior features) for template-based ASR has been proposed. These features have been found to generalize well to unseen data and yield better systems than standard spectral-based features. In this paper, motivated by the high quality of current text-to-speech systems and the robustness of posterior features toward undesired variability, we investigate the use of synthetic speech to generate reference templates. The use of synthetic speech in template-based ASR not only allows to address the issue of in-domain data collection but also expansion of vocabulary. Using 75- and 600-word task-independent and speaker-independent setup on Phonebook database, we investigate different synthetic voices produced by the Festival HTS-based synthesizer trained on CMU ARCTIC databases. Our study shows that synthetic speech templates can yield performance comparable to the natural speech templates, especially with synthetic voices that have high intelligibility.
Serena Soldo, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH2
2012 Using Sparse Classification Outputs as Feature Observations for Noise-robust ASR
abstract
Contains fulltext : 102118.pdf (author's version ) (Open Access) Contains fulltext : 102118.pdf (Publisher’s version ) (Open Access)
Bert Cranen, Jort F. Gemmeke, Louis ten Bosch, Lou Boves, Mathew Magimai-Doss
INTERSPEECH6
2012 Combination of Sparse Classification and Multilayer Perceptron for Noise-robust ASR
abstract
Contains fulltext : 101559.pdf (Publisher’s version ) (Open Access)
Mathew Magimai-Doss, Jort F. Gemmeke, Bert Cranen, Louis ten Bosch, Lou Boves
INTERSPEECH2
2012 Phase AutoCorrelation (PAC) features for noise robust speech recognition
Shajith Ikbal, Hemant Misra, Hynek Hermansky, Mathew Magimai-Doss
Speech Commun.4
2012 A Fast Parts-Based Approach to Speaker Verification Using Boosted Slice Classifiers
abstract
Speaker verification (SV) on portable devices like smartphones is gradually becoming popular. In this context, two issues need to be considered: 1) such devices have relatively limited computation resources, and 2) they are liable to be used everywhere, possibly in very noisy, uncontrolled environments. This work aims to address both these issues by proposing a computationally efficient yet robust SV system. This novel parts-based system draws inspiration from face and object detection systems in the computer vision domain. The system involves boosted ensembles of simple threshold-based classifiers. It uses a novel set of features extracted from speech spectra, called "slice features." The performance of the proposed system was evaluated through extensive studies involving a wide range of experimental conditions using the TIMIT, HTIMIT, and MOBIO corpus, against standard cepstral features and Gaussian Mixture Model-based SV systems.
Anindya Roy, Mathew Magimai-Doss, Sébastien Marcel
IEEE Trans. Inf. Forensics Secur.2
2011 Fast and flexible Kullback-Leibler divergence based acoustic modeling for non-native speech recognition
abstract
One of the main challenge in non-native speech recognition is how to handle acoustic variability present in multi-accented non-native speech with limited amount of training data. In this paper, we investigate an approach that addresses this challenge by using Kullback-Leibler divergence based hidden Markov models (KL-HMM). More precisely, the acoustic variability in the multi-accented speech is handled by using multilingual phoneme posterior probabilities, estimated by a multilayer perceptron trained on auxiliary data, as input feature for the KL-HMM system. With limited training data, we then build better acoustic models by exploiting the advantage that the KL-HMM system has fewer number of parameters. On HIWIRE corpus, the proposed approach yields a performance of 1.9% word error rate (WER) with 149 minutes of training data and a performance of 5.5% WER with 2 minutes of training data.
David Imseng, Ramya Rasipuram, Mathew Magimai-Doss
ASRU3
2011 Improving Articulatory Feature and Phoneme Recognition Using Multitask Learning
Ramya Rasipuram, Mathew Magimai-Doss
ICANN (1)2
2011 Language dependent universal phoneme posterior estimation for mixed language speech recognition
abstract
This paper presents a new approach to estimate "universal" phoneme posterior probabilities for mixed language speech recognition. More specifically, we propose a new theoretical framework to combine phoneme class posterior probabilities in a principled way by using (statistical) evidence about the language identity. We investigate the proposed approach in a mixed language environment (Speech-Dat(II)) consisting of five European languages. Our studies show that the proposed approach can yield significant improvements on a mixed language task, while maintaining the performance on monolingual tasks. Additionally, through a case study, we also demonstrate the potential benefits of the proposed approach for non-native speech recognition.
David Imseng, Hervé Bourlard, Mathew Magimai-Doss, John Dines
ICASSP3
2011 Integrating articulatory features using Kullback-Leibler divergence based acoustic model for phoneme recognition
abstract
In this paper, we propose a novel framework to integrate articulatory features (AFs) into HMM- based ASR system. This is achieved by using posterior probabilities of different AFs (estimated by multilayer perceptrons) directly as observation features in Kullback-Leibler divergence based HMM (KL-HMM) system. On the TIMIT phoneme recognition task, the proposed framework yields a phoneme recognition accuracy of 72.4% which is comparable to KL-HMM system using posterior probabilities of phonemes as features (72.7%). Furthermore, a best performance of 73.5% phoneme recognition accuracy is achieved by jointly modeling AF probabilities and phoneme probabilities as features. This shows the efficacy and flexibility of the proposed approach.
Ramya Rasipuram, Mathew Magimai-Doss
ICASSP2
2011 Phoneme recognition using Boosted Binary Features
abstract
In this paper, we propose a novel parts-based binary-valued feature for ASR. This feature is extracted using boosted ensembles of simple threshold-based classifiers. Each such classifier looks at a specific pair of time-frequency bins located on the spectro-temporal plane. These features termed as Boosted Binary Features (BBF) are integrated into standard HMM-based system by using multilayer perceptron (MLP) and single layer perceptron (SLP). Preliminary studies on TIMIT phoneme recognition task show that BBF yields similar or better performance compared to MFCC (67.8% accuracy for BBF vs. 66.3% accuracy for MFCC) using MLP, while it yields significantly better performance than MFCC (62.8% accuracy for BBF vs. 45.9% for MFCC) using SLP. This demonstrates the potential of the proposed feature for speech recognition.
Anindya Roy, Mathew Magimai-Doss, Sébastien Marcel
ICASSP2
2011 Posterior features for template-based ASR
abstract
This paper investigates the use of phoneme class conditional probabilities as features (posterior features) for template-based ASR. Using 75 words and 600 words task-independent and speaker-independent setup on Phonebook database, we investigate the use of different posterior distribution estimators, different distance measures that are better suited for posterior distributions, and different training data. The reported experiments clearly demonstrate that posterior features are always superior, and generalize better than other classical acoustic features (at the cost of training a posterior distribution estimator).
Serena Soldo, Mathew Magimai-Doss, Joel Pinto, Hervé Bourlard
ICASSP2
2011 Fast speaker verification on mobile phone data using boosted slice classifiers
abstract
In this work, we Investigate a novel computationally efficient speaker verification (SV) system involving boosted ensembles of simple threshold-based classifiers. The system is based on a novel set of features called "slice features". Both the system and the features were inspired by the recent success of pixel comparison-based ensemble approaches in the computer vision domain. The performance of the proposed system was evaluated through speaker verification experiments on the MOBIO corpus containing mo- bile phone speech, according to a challenging protocol. The system was found to perform reasonably well, compared to multiple state-of-the-art SV systems, with the benefit of significantly lower computational complexity. Its dual characteristics of good performance and computational efficiency could be important factors in the context of SV system implementation on portable devices like mobile phones.
Anindya Roy, Mathew Magimai-Doss, Sébastien Marcel
IJCB2
2011 Improving Non-Native ASR Through Stochastic Multilingual Phoneme Space Transformations
abstract
We propose a stochastic phoneme space transformation technique that allows the conversion of conditional source phoneme posterior probabilities (conditioned on the acoustics) into target phoneme posterior probabilities. The source and target phonemes can be in any language and phoneme format such as the International Phonetic Alphabet. The novel technique makes use of a Kullback-Leibler divergence based hidden Markov model and can be applied to non-native and accented speech recognition or used to adapt systems to under-resourced languages. In this paper, and in the context of hybrid HMM/MLP recognizers, we successfully apply the proposed approach to non-native English speech recognition on the HIWIRE dataset.
David Imseng, Hervé Bourlard, John Dines, Philip N. Garner, Mathew Magimai-Doss
INTERSPEECH5
2011 Grapheme-Based Automatic Speech Recognition Using KL-HMM
abstract
The state-of-the-art automatic speech recognition (ASR) systems typically use phonemes as subword units. In this work, we present a novel grapheme-based ASR system that jointly models phoneme and grapheme information using Kullback-Leibler divergence-based HMM system (KL-HMM). More specifically, the underlying subword unit models are grapheme units and the phonetic information is captured through phoneme posterior probabilities (referred as posterior features) estimated using a multilayer perceptron (MLP). We investigate the proposed approach for ASR on English language, where the correspondence between phoneme and grapheme is weak. In particular, we investigate the effect of contextual modeling on grapheme-based KL-HMM system and the use of MLP trained on auxiliary data. Experiments on DARPA Resource Management corpus have shown that the grapheme-based ASR system modeling longer subword unit context can achieve same performance as phoneme-based ASR system, irrespective of the data on which MLP is trained.
Mathew Magimai-Doss, Ramya Rasipuram, Guillermo Aradilla, Hervé Bourlard
INTERSPEECH1
2011 Hierarchical Tandem Features for ASR in Mandarin
abstract
We apply multilayer perceptron (MLP) based hierarchical Tandem features to large vocabulary continuous speech recognition in Mandarin.Hierarchical Tandem features are estimated using a cascade of two MLP classifiers which are trained independently.The first classifier is trained on perceptual linear predictive coefficients with a 90 ms temporal context.The second classifier is trained using the phonetic class conditional probabilities estimated by the first MLP, but with a relatively longer temporal context of about 150 ms.Experiments on the Mandarin DARPA GALE eval06 dataset show significant reduction (about 7.6% relative) in character error rates by using hierarchical Tandem features over conventional Tandem features.
Joel Pinto, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH2
2011 Analysis and Comparison of Recent MLP Features for LVCSR Systems
abstract
MLP based front-ends have evolved in different ways in re-cent years beyond the seminal TANDEM-PLP features. This paper aims at providing a fair comparison of these recent pro-gresses including the use of different long/short temporal in-puts (PLP,MRASTA,wLP-TRAPS,DCT-TRAPS) and the use of complex architectures (bottleneck, hierarchy, multistream) that go beyond the conventional three layer MLP. Furthermore, the paper identifies which of these actually provide advantages over the conventional TANDEM-PLP. The investigation is car-ried on an LVCSR task for recognition of Mandarin Broadcast speech and results are analyzed in terms of Character Error Rate and phonetic confusions. Results reveal that as stand alone features, multistream front-ends can outperform by 10 % con-ventional MFCC while TANDEM-PLP only improve by 1 %. On the other hand, when used in concatenation with MFCC features, hierarchical/bottleneck front-ends reduce the character error rate by +18 % relative compared to +14 % relative from TANDEM-PLP. The various input long-term representations re-cently developed provide comparable performances.
Fabio Valente, Mathew Magimai-Doss
INTERSPEECH2
2011 Privacy-Sensitive Audio Features for Speech/Nonspeech Detection
abstract
The goal of this paper is to investigate features for speech/nonspeech detection (SND) having low linguistic information from the speech signal. Towards this, we present a comprehensive study of privacy-sensitive features for SND in multiparty conversations. Our study investigates three different approaches to privacy-sensitive features. These approaches are based on: 1) simple, instantaneous feature extraction methods; 2) excitation source information based methods; and 3) feature obfuscation methods such as local (within 130 ms) temporal averaging and randomization applied on excitation source information. To evaluate these approaches for SND, we use multiparty conversational meeting data of nearly 450 hours. On this dataset, we evaluate these features and benchmark them against standard spectral shape based features such as Mel frequency perceptual linear prediction (MFPLP). Fusion strategies combining excitation source with simple features show that comparable performance can be obtained in both close-talking and far-field microphone scenarios. As one way to objectively evaluate the notion of privacy, we conduct phoneme recognition studies on TIMIT. While excitation source features yield phoneme recognition accuracies in between the simple features and the MFPLP features, obfuscation methods applied on the excitation features yield low phoneme accuracies in conjunction with SND performance comparable to that of MFPLP features.
Sree Hari Krishnan Parthasarathi, Daniel Gatica-Perez, Hervé Bourlard, Mathew Magimai-Doss
IEEE ACM Trans. Audio Speech Lang. Process.4
2011 Analysis of MLP-Based Hierarchical Phoneme Posterior Probability Estimator
abstract
We analyze a simple hierarchical architecture consisting of two multilayer perceptron (MLP) classifiers in tandem to estimate the phonetic class conditional probabilities. In this hierarchical setup, the first MLP classifier is trained using standard acoustic features. The second MLP is trained using the posterior probabilities of phonemes estimated by the first, but with a long temporal context of around 150-230 ms. Through extensive phoneme recognition experiments, and the analysis of the trained second MLP using Volterra series, we show that 1) the hierarchical system yields higher phoneme recognition accuracies-an absolute improvement of 3.5% and 9.3% on TIMIT and CTS respectively-over the conventional single MLP-based system, 2) there exists useful information in the temporal trajectories of the posterior feature space, spanning around 230 ms of context, 3) the second MLP learns the phonetic temporal patterns in the posterior features, which include the phonetic confusions at the output of the first MLP as well as the phonotactics of the language as observed in the training data, and 4) the second MLP classifier requires fewer number of parameters and can be trained using lesser amount of training data.
Joel Pinto, Garimella S. V. S. Sivaram, Mathew Magimai-Doss, Hynek Hermansky, Hervé Bourlard
IEEE Trans. Speech Audio Process.3
2011 Transcribing Mandarin Broadcast Speech Using Multi-Layer Perceptron Acoustic Features
abstract
Recently, several multi-layer perceptron (MLP)-based front-ends have been developed and used for Mandarin speech recognition, often showing significant complementary properties to conventional spectral features. Although widely used in multiple Mandarin systems, no systematic comparison of all the different approaches as well as their scalability has been proposed. The novelty of this correspondence is mainly experimental. In this work, all the MLP front-ends recently developed at multiple sites are described and compared in a systematic manner on a 100 hours setup. The study covers the two main directions along which the MLP features have evolved: the use of different input representations to the MLP and the use of more complex MLP architectures beyond the three-layer perceptron. The results are analyzed in terms of confusion matrices and the paper discusses a number of novel findings that the comparison reveals. Furthermore, the two best front-ends used in the GALE 2008 evaluation, referred as MLP1 and MLP2, are studied in a more complex LVCSR system in order to investigate their scalability in terms of the amount of training data (from 100 hours to 1600 hours) and the parametric system complexity (maximum likelihood versus discriminative training, speaker adaptative training, lattice level combination). Results on 5 hours of evaluation data from the GALE project reveal that the MLP features consistently produce improvements in the range of 15%-23% relative at the different steps of a multipass system when compared to mel-frequency cepstral coefficient (MFCC) and PLP features, suggesting that the improvements scale with the amount of data and with the complexity of the system. The integration of those features into the GALE 2008 evaluation system provide very competitive performances compared to other Mandarin systems.
Fabio Valente, Mathew Magimai-Doss, Christian Plahl, Suman V. Ravuri
IEEE ACM Trans. Audio Speech Lang. Process.2
2010 Evaluating the robustness of privacy-sensitive audio features for speech detection in personal audio log scenarios
abstract
Personal audio logs are often recorded in multiple environments. This poses challenges for robust front-end processing, including speech/nonspeech detection (SND). Motivated by this, we investigate the robustness of four different privacy-sensitive features for SND, namely energy, zero crossing rate, spectral flatness, and kurtosis. We study early and late fusion of these features in conjunction with modeling temporal context. These combinations are evaluated in mismatched conditions on a dataset of nearly 450 hours. While both combinations yield improvements over individual features, generally feature combinations perform better. Comparisons with a state-of-the-art spectral based and a privacy-sensitive feature set are also provided.
Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Hervé Bourlard, Daniel Gatica-Perez
ICASSP2
2010 Boosted binary features for noise-robust speaker verification
abstract
The standard approach to speaker verification is to extract cepstral features from the speech spectrum and model them by generative or discriminative techniques. We propose a novel approach where a set of client-specific binary features carrying maximal discriminative information specific to the individual client are estimated from an ensemble of pair-wise comparisons of frequency components in magnitude spectra, using Adaboost algorithm. The final classifier is a simple linear combination of these selected features. Experiments on the XM2VTS database strictly according to a standard evaluation protocol have shown that although the proposed framework yields comparatively lower performance on clean speech, it significantly outperforms the state-of-the-art MFCC-GMM system in mismatched conditions with training on clean speech and testing on speech corrupted by four types of additive noise from the standard Noisex-92 database.
Anindya Roy, Mathew Magimai-Doss, Sébastien Marcel
ICASSP2
2010 Towards mixed language speech recognition systems
abstract
Multilingual speech recognition obviously involves numerous research challenges, including common phoneme sets, adaptation on limited amount of training data, as well as mixed language recognition (common in many countries, like Switzerland). In this latter case, it is not even possible to assume that one knows in advance the language being spoken. This is the context and motivation of the present work. We indeed investigate how current state-of-the-art speech recognition systems can be exploited in multilingual environments, where the language (from an assumed set of five possible languages, in our case) is not a priori known during recognition. We combine monolingual systems and extensively develop and compare different features and acoustic models. On Speech-Dat(II) datasets, and in the context of isolated words, we show that it is actually possible to approach the performances of monolingual systems even if the identity of the spoken language is not a priori known. Index Terms: speech recognition, multilingual speech recognition, combination of mono-lingual speech recognition systems, mixed language recognition. 1
David Imseng, Hervé Bourlard, Mathew Magimai-Doss
INTERSPEECH3
2010 Hierarchical multilayer perceptron based language identification
abstract
Automatic language identification (LID) systems generally exploit acoustic knowledge, possibly enriched by explicit language specific phonotactic or lexical constraints. This paper investigates a new LID approach based on hierarchical multilayer perceptron (MLP) classifiers, where the first layer is a “universal phoneme set MLP classifier”. The resulting (multilingual) phoneme posterior sequence is fed into a second MLP taking a larger temporal context into account. The second MLP can learn/exploit implicitly different types of patterns/information such as confusion between phonemes and/or phonotactics for LID. We investigate the viability of the proposed approach by comparing it against two standard approaches which use phonotactic and lexical constraints with the universal phoneme set MLP classifier as emission probability estimator. On Speech-Dat(II) datasets of five European languages, the proposed approach yields significantly better performance compared to the two standard approaches. Index Terms: Language identification, multilingual processing, hierarchical MLP.
David Imseng, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH2
2010 A comparative large scale study of MLP features for Mandarin ASR
abstract
MLP based front-ends have shown significant complementary properties to conventional spectral features.As part of the DARPA GALE program, different MLP features were developed for Mandarin ASR.In this paper, all the proposed frontends are compared in systematic manner and we extensively investigate the scalability of these features in terms of the amount of training data (from 100 hours to 1600 hours) and system complexity (maximum likelihood training, SAT, lattice level combination, and discriminative training).Results on 5 hours of evaluation data from the GALE project reveal that the MLP features consistently produce relative improvements in the range of 15% -23% at the different steps of a multipass system when compared to the conventional short-term spectral based features like MFCC and PLP.The largest improvement is obtained using a hierarchical MLP approach.
Fabio Valente, Mathew Magimai-Doss, Christian Plahl, Suman V. Ravuri
INTERSPEECH2
2009 MLP based hierarchical system for task adaptation in ASR
abstract
We investigate a multilayer perceptron (MLP) based hierarchical approach for task adaptation in automatic speech recognition. The system consists of two MLP classifiers in tandem. A well-trained MLP available off-the-shelf is used at the first stage of the hierarchy. A second MLP is trained on the posterior features estimated by the first, but with a long temporal context of around 130 ms. By using an MLP trained on 232 hours of conversational telephone speech, the hierarchical adaptation approach yields a word error rate of 1.8% on the 600-word Phonebook isolated word recognition task. This compares favorably to the error rate of 4% obtained by the conventional single MLP based system trained with the same amount of Phonebook data that is used for adaptation. The proposed adaptation scheme also benefits from the ability of the second MLP to model the temporal information in the posterior features.
Joel Pinto, Mathew Magimai-Doss, Hervé Bourlard
ASRU2
2009 Posterior features applied to speech recognition tasks with user-defined vocabulary
abstract
This paper presents a novel approach for those applications where vocabulary is defined by a set of acoustic samples. In this approach, the acoustic samples are used as reference templates in a template matching framework. The features used to describe the reference templates and the test utterances are estimates of phoneme posterior probabilities. These posteriors are obtained from a MLP trained on an auxiliary database. Thus, the speech variability present in the features is reduced by applying the speech knowledge captured by the MLP on the auxiliary database. Moreover, information theoretic dissimilarity measures can be used as local distances between features. When compared to state-of-the-art systems, this approach outperforms acoustic-based techniques and obtains comparable results to orthography-based methods. The proposed method can also be directly combined with other posterior-based HMM systems. This combination successfully exploits the complementarity between templates and parametric models.
Guillermo Aradilla, Hervé Bourlard, Mathew Magimai-Doss
ICASSP3
2009 Non-linear mapping for multi-channel speech separation and robust overlapping spech recognition
abstract
This paper investigates a non-linear mapping approach to extract robust features for ASR and speech separation of overlapping speech. Based on our previous studies, we continue to use two additional sound sources, namely from the target and interfering speakers. The focuses of this work are: 1) We investigate the feature mapping between different domains with the consideration of MMSE criterion and regression optimizations, demonstrating the mapping of log melfilterbank energies to MFCC can be exploited to improve the effectiveness of the regression; 2) We investigate the data-driven filtering for the speech separation by using the mapping method, which can be viewed as a generalized log spectral subtraction and results in better separation performance. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes both non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, John Dines, Mathew Magimai-Doss, Hervé Bourlard
ICASSP3
2009 Volterra series for analyzing MLP based phoneme posterior estimator
abstract
We present a framework to apply Volterra series to analyze multi-layered perceptrons trained to estimate the posterior probabilities of phonemes in automatic speech recognition. The identified Volterra kernels reveal the spectro-temporal patterns that are learned by the trained system for each phoneme. To demonstrate the applicability of Volterra series, we analyze a multilayered perceptron trained using Mel filter bank energy features and analyze its first order Volterra kernels.
Joel Pinto, Garimella S. V. S. Sivaram, Hynek Hermansky, Mathew Magimai-Doss
ICASSP4
2009 Speaker change detection with privacy-preserving audio cues
abstract
In this paper we investigate a set of privacy-sensitive audio features for speaker change detection (SCD) in multiparty conversations. These features are based on three different principles: characterizing the excitation source information using linear prediction residual, characterizing subband spectral information shown to contain speaker information, and characterizing the general shape of the spectrum. Experiments show that the performance of the privacy-sensitive features is comparable or better than that of the state-of-the-art full-band spectral-based features, namely, mel frequency cepstral coefficients, which suggests that socially acceptable ways of recording conversations in real-life is feasible.
Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Daniel Gatica-Perez, Hervé Bourlard
ICMI2
2009 Investigating privacy-sensitive features for speech detection in multiparty conversations
abstract
We investigate four different privacy-sensitive features, namely energy, zero crossing rate, spectral flatness, and kurtosis, for speech detection in multiparty conversations. We liken this scenario to a meeting room and define our datasets and annotations accordingly. The temporal context of these features is modeled. With no temporal context, energy is the best performing single feature. But by modeling temporal context, kurtosis emerges as the most effective feature. Also, we combine the features. Besides yielding a gain in performance, certain combinations of features also reveal that a shorter temporal context is sufficient. We then benchmark other privacy-sensitive features utilized in previous studies. Our experiments show that the performance of all the privacy-sensitive features modeled with context is close to that of state-of-the-art spectral-based features, without extracting and using any features that can be used to reconstruct the speech signal.
Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Hervé Bourlard, Daniel Gatica-Perez
INTERSPEECH2
2009 Hierarchical processing of the modulation spectrum for GALE Mandarin LVCSR system
abstract
This paper aims at investigating the use of TANDEM features based on hierarchical processing of the modulation spectrum.The study is done in the framework of the GALE project for recognition of Mandarin Broadcast data.We describe the improvements obtained using the hierarchical processing and the addition of features like pitch and short-term critical band energy.Results are consistent with previous findings on a different LVCSR task suggesting that the proposed technique is effective and robust across several conditions.Furthermore we describe integration into RWTH GALE LVCSR system trained on 1600 hours of Mandarin data and present progress across the GALE 2007 and GALE 2008 RWTH systems resulting in approximatively 20% CER reduction on several data set.
Fabio Valente, Mathew Magimai-Doss, Christian Plahl, Suman V. Ravuri
INTERSPEECH2
2008 Exploiting contextual information for improved phoneme recognition
abstract
In this paper, we investigate the significance of contextual information in a phoneme recognition system using the hidden Markov model - artificial neural network paradigm. Contextual information is probed at the feature level as well as at the output of the multilayered perceptron. At the feature level, we analyze and compare different methods to model sub-phonemic classes. To exploit the contextual information at the output of the multilayered perceptron, we propose the hierarchical estimation of phoneme posterior probabilities. The best phoneme (excluding silence) recognition accuracy of 73.4% on the TIMIT database is comparable to that of the state-of- the-art systems, but more emphasis is on analysis of the contextual information.
Joel Pinto, Bayya Yegnanarayana, Hynek Hermansky, Mathew Magimai-Doss
ICASSP4
2008 Using KL-based acoustic models in a large vocabulary recognition task
abstract
Posterior probabilities of sub-word units have been shown to be an effective front-end for ASR. However, attempts to model this type of features either do not benefit from modeling context-dependent phonemes, or use an inefficient distribution to estimate the state likelihood. This paper presents a novel acoustic model for posterior features that overcomes these limitations. The proposed model can be seen as a HMM where the score associated with each state is the KL divergence between a distribution characterizing the state and the posterior features from the test utterance. This KL-based acoustic model establishes a framework where other models for posterior features such as hybrid HMM/MLP and discrete HMM can be seen as particular cases. Experiments on the WSJ database show that the KL-based acoustic model can significantly outperform these latter approaches. Moreover, the proposed model can obtain comparable results to complex systems, such as HMM/GMM, using significantly fewer parameters.
Guillermo Aradilla, Hervé Bourlard, Mathew Magimai-Doss
INTERSPEECH3
2008 Neural network based regression for robust overlapping speech recognition using microphone arrays
abstract
This paper investigates a neural network based acoustic feature mapping to extract robust features for automatic speech recognition (ASR) of overlapping speech. In our preliminary studies, we trained neural networks to learn the mapping from log mel filter bank energies (MFBEs) extracted from the distant microphone recordings, including multiple overlapping speakers, to log MFBEs extracted from the clean speech signal. In this paper, we explore the mapping of higher order mel-filterbank cepstral coefficients (MFCC) to lower order coefficients. We also investigate the mapping of features from both target and interfering distant sound sources to the clean target features. This is achieved by using the microphone array to extract features from both the direction of the target and interfering sound sources. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes both non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, John Dines, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2007 Monolingual and crosslingual comparison of tandem features derived from articulatory and phone MLPS
abstract
The features derived from posteriors of a multilayer perceptron (MLP), known as tandem features, have proven to be very effective for automatic speech recognition. Most tandem features to date have relied on MLPs trained for phone classification. We recently showed on a relatively small data set that MLPs trained for articulatory feature classification can be equally effective. In this paper, we provide a similar comparison using MLPs trained on a much larger data set -2000 hours of English conversational telephone speech. We also explore how portable phone-and articulatory feature-based tandem features are in an entirely different language - Mandarin - without any retraining. We find that while the phone-based features perform slightly better than AF-based features in the matched-language condition, they perform significantly better in the cross-language condition. However, in the cross-language condition, neither approach is as effective as the tandem features extracted from an MLP trained on a relatively small amount of in-domain data. Beyond feature concatenation, we also explore novel factored observation modeling schemes that allow for greater flexibility in combining the tandem and standard features.
Özgür Çetin, Mathew Magimai-Doss, Karen Livescu, Arthur Kantor, Simon King 0001, Chris D. Bartels, Joe Frankel
ASRU2
2007 An Articulatory Feature-Based Tandem Approach and Factored Observation Modeling
abstract
The so-called tandem approach, where the posteriors of a multilayer perceptron (MLP) classifier are used as features in an automatic speech recognition (ASR) system has proven to be a very effective method. Most tandem approaches up to date have relied on MLPs trained for phone classification, and appended the posterior features to some standard feature hidden Markov model (HMM). In this paper, we develop an alternative tandem approach based on MLPs trained for articulatory feature (AF) classification. We also develop a factored observation model for characterizing the posterior and standard features at the HMM outputs, allowing for separate hidden mixture and state-tying structures for each factor. In experiments on a subset of Switchboard, we show that the AF-based tandem approach is as effective as the phone-based approach, and that the factored observation model significantly outperforms the simple feature concatenation approach while using fewer parameters.
Özgür Çetin, Arthur Kantor, Simon King 0001, Chris D. Bartels, Mathew Magimai-Doss, Joe Frankel, Karen Livescu
ICASSP (4)5
2007 A Generalized Dynamic Composition Algorithm of Weighted Finite State Transducers for Large Vocabulary Speech Recognition
abstract
We propose a generalized dynamic composition algorithm of weighted finite state transducers (WFST), which avoids the creation of noncoaccessible paths, performs weight look-ahead and does not impose any constraints to the topology of the WFSTs. Experimental results on Wall Street Journal (WSJ1) 20k-word trigram task show that at 17% WER (moderately-wide beam width), the decoding time of the proposed approach is about 48% and 65% of the other two dynamic composition approaches. In comparison with static composition, at the same level of 17% WER, we observe a reduction of about 60% in memory requirement, with an increase of about 60% in decoding time due to extra overheads for dynamic composition.
Octavian Cheng, John Dines, Mathew Magimai-Doss
ICASSP (4)3
2007 Manual Transcription of Conversational Speech at the Articulatory Feature Level
abstract
We present an approach for the manual labeling of speech at the articulatory feature level, and a new set of labeled conversational speech collected using this approach. A detailed transcription, including overlapping or reduced gestures, is useful for studying the great pronunciation variability in conversational speech. It also facilitates the testing of feature classifiers, such as those used in articulatory approaches to automatic speech recognition. We describe an effort to transcribe a small set of utterances drawn from the Switchboard database using eight articulatory tiers. Two transcribers have labeled these utterances in a multi-pass strategy, allowing for correction of errors. We describe the data collection methods and analyze the data to determine how quickly and reliably this type of transcription can be done. Finally, we demonstrate one use of the new data set by testing a set of multilayer perceptron feature classifiers against both the manual labels and forced alignments.
Karen Livescu, Ari Bezman, Nash M. Borges, Lisa Yung, Özgür Çetin, Joe Frankel, Simon King 0001, Mathew Magimai-Doss, Xuemin Chi, Lisa Lavoie
ICASSP (4)8
2007 Articulatory Feature-Based Methods for Acoustic and Audio-Visual Speech Recognition: Summary from the 2006 JHU Summer workshop
abstract
We report on investigations, conducted at the 2006 Johns Hopkins Workshop, into the use of articulatory features (AFs) for observation and pronunciation models in speech recognition. In the area of observation modeling, we use the outputs of AF classifiers both directly, in an extension of hybrid HMM/neural network models, and as part of the observation vector, an extension of the "tandem" approach. In the area of pronunciation modeling, we investigate a model having multiple streams of AF states with soft synchrony constraints, for both audio-only and audio-visual recognition. The models are implemented as dynamic Bayesian networks, and tested on tasks from the small-vocabulary switchboard (SVitchboard) corpus and the CUAVE audio-visual digits corpus. Finally, we analyze AF classification and forced alignment using a newly collected set of feature-level manual transcriptions.
Karen Livescu, Özgür Çetin, Mark Hasegawa-Johnson, Simon King 0001, Chris D. Bartels, Nash M. Borges, Arthur Kantor, Partha Lal, Lisa Yung, Ari Bezman, Stephen Dawson-Haggerty, Bronwyn Woods, Joe Frankel, Mathew Magimai-Doss, Kate Saenko
ICASSP (4)14
2007 Entropy Based Classifier Combination for Sentence Segmentation
abstract
We describe recent extensions to our previous work, where we explored the use of individual classifiers, namely, boosting and maximum entropy models for sentence segmentation. In this paper we extend the set of classification methods with support vector machine (SVM). We propose a new dynamic entropy-based classifier combination approach to combine these classifiers, and compare it with the traditional classifier combination techniques, namely, voting, linear regression and logistic regression. Furthermore, we also investigate the combination of hidden event language models with the output of the proposed classifier combination, and the output of individual classifiers. Experimental studies conducted on the Mandarin TDT4 broadcast news database shows that the SVM classifier as an individual classifier improves over our previous best system. However, the proposed entropy-based classifier combination approach shows the best improvement in F-measure of 1% absolute, and the voting approach shows the best reduction in NIST error rate of 2.7% absolute when compared to the previous best system.
Mathew Magimai-Doss, Dilek Hakkani-Tür, Özgür Çetin, Elizabeth Shriberg, James G. Fung, Nikki Mirghafori
ICASSP (4)1
2007 Articulatory feature classifiers trained on 2000 hours of telephone speech
abstract
The so-called tandem approach, where the posteriors of a multilayer perceptron (MLP) classifier are used as features in an automatic speech recognition (ASR) system has proven to be a very effective method. Most tandem approaches up to date have relied on MLPs trained for phone classification, and appended the posterior features to some standard feature hidden Markov model (HMM). In this paper, we develop an alternative tandem approach based on MLPs trained for articulatory feature (AF) classification. We also develop a factored observation model for characterizing the posterior and standard features at the HMM outputs, allowing for separate hidden mixture and state-tying structures for each factor. In experiments on a subset of Switchboard, we show that the AFbased tandem approach is as effective as the phone-based approach, and that the factored observation model significantly outperforms the simple feature concatenation approach while using fewer parameters.
Joe Frankel, Mathew Magimai-Doss, Simon King 0001, Karen Livescu, Özgür Çetin
INTERSPEECH2
2007 Cross-linguistic analysis of prosodic features for sentence segmentation
abstract
In this paper, we perform a cross-linguistic study of prosodic features in sentence segmentation by using two different feature selection approaches: a forward search wrapper and feature filtering. Experiments in Arabic, English, and Mandarin show that prosodic features make significant contributions in all three languages. Feature selection results indicate that feature relevancy can vary greatly depending on the target language, and therefore the optimal feature subset varies considerably between languages. We observe patterns in the feature selection and the affinity of the different languages toward certain feature types, which gives us insight into future feature selection and feature design. Index Terms: prosodic features, cross lingual, feature selection, sentence segmentation
James G. Fung, Dilek Hakkani-Tür, Mathew Magimai-Doss, Elizabeth Shriberg, Sébastien Cuendet, Nikki Mirghafori
INTERSPEECH3
2007 Improving speech translation with automatic boundary prediction
abstract
This paper investigates the influence of automatic sentence boundary and sub-sentence punctuation prediction on machine translation (MT) of automatically recognized speech.We use prosodic and lexical cues to determine sentence boundaries, and successfully combine two complementary approaches to sentence boundary prediction.We also introduce a new feature for segmentation prediction that directly considers the assumptions of the phrase translation model.In addition, we show how automatically predicted commas can be used to constrain reordering in MT search.We evaluate the presented methods using a state-of-the-art phrase-based statistical MT system on two large vocabulary tasks.We find that careful optimization of the segmentation parameters directly for translation quality improves the translation results in comparison to independent optimization for segmentation quality of the predicted source language sentence boundaries.
Evgeny Matusov, Dustin Hillard, Mathew Magimai-Doss, Dilek Hakkani-Tür, Mari Ostendorf, Hermann Ney
INTERSPEECH3
2006 Threshold Selection for Unsupervised Detection, With an Application to Microphone Arrays
abstract
Detection is usually done by comparing some criterion to a threshold. It is often desirable to keep a performance metric such as false alarm rate constant across conditions. Using training data to select the threshold may lead to suboptimal results on test data recorded in different conditions. This paper investigates unsupervised approaches, where no training data is used. A probabilistic model is fitted on the test data using the EM algorithm, and the threshold value is selected based on the model. The proposed approach (1) does not use training data, (2) uses the test data itself to compensate for simplifications inherent to the model, and (3) permits the use of more complex models in a straightforward manner. On a microphone array speech detection task, the proposed unsupervised approach achieves similar or better results than the "training" approach. The methodology is general and may be applied to other contexts than microphone arrays, and other performance metrics than FAR
Guillaume Lathoud, Mathew Magimai-Doss, Hervé Bourlard
ICASSP (3)2
2005 HMM/ANN Based Spectral Peak Location Estimation for Noise Robust Speech Recognition
abstract
In this paper, we present an HMM/ANN based algorithm to estimate the spectral peak locations. This algorithm makes use of distinct time-frequency (TF) patterns in the spectrogram for estimating the peak locations. Such a use of TF patterns is expected to impose temporal constraints during the peak estimation task, thereby yielding a smoother estimate of the peaks over time. Additionally, the algorithm uses an ergodic topology for the HMM/ANN, thus allowing an estimation of a varying number of peak locations over time. The usefulness of the proposed algorithm is evaluated in the framework of a recently introduced noise robust feature called the spectro-temporal activity pattern (STAP) feature. Interestingly, the recently introduced phase autocorrelation (PAC) spectrum, with enhanced spectral peaks and smoothed spectral valleys, turns out to be more appropriate for this algorithm than the regular spectrum.
Shajith Ikbal, Hervé Bourlard, Mathew Magimai-Doss
ICASSP (1)3
2005 A sector-based, frequency-domain approach to detection and localization of multiple speakers
abstract
Detection and localization of speakers with microphone arrays is a difficult task due to the wideband nature of speech signals, the large amount of overlap between speakers in spontaneous conversations, and the presence of noise sources. Many existing audio multi-source localization methods rely on prior knowledge of the sectors containing active sources and/or the number of active sources. The paper proposes sector-based, frequency-domain approaches that address both detection and localization problems by measuring relative phases between microphones. The first approach is similar to delay-sum beamforming. The second approach is novel: it relies on systematic optimization of a centroid in phase space, for each sector. It provides a major, systematic improvement over the first approach as well as over previous work. Very good results are obtained on more than one hour of recordings in real meeting room conditions, including cases with up to 3 concurrent speakers.
Guillaume Lathoud, Mathew Magimai-Doss
ICASSP (3)2
2005 A spectrogram model for enhanced source localization and noise-robust ASR
abstract
This paper proposes a simple, computationally efficient 2-mixture model approach to discrimination between speech and background noise. It is directly derived from observations on real data, and can be used in a fully unsupervised manner, with the EM algorithm. A first application to sector-based, joint audio source localization and detection, using multiple microphones, confirms that the model can provide major enhancement. A second application to the single channel speech recognition task in a noisy environment yields major improvement on stationary noise and promising results on non-stationary noise. 1.
Guillaume Lathoud, Mathew Magimai-Doss, Bertrand Mesot
INTERSPEECH2
2004 Joint decoding for phoneme-grapheme continuous speech recognition
abstract
Standard ASR systems typically use phonemes as the subword units. Preliminary studies have shown that the performance of ASR systems could be improved by using graphemes as additional subword units. We investigate such a system where the word models are defined in terms of two different subword units, i.e., phoneme and grapheme. During training, models for both the subword units are trained, and then, during recognition, either both or just one subword unit is used. We have studied this system for a continuous speech recognition task in American English. Our studies show that grapheme information used along with phoneme information improves the performance of ASR.
Mathew Magimai-Doss, Samy Bengio, Hervé Bourlard
ICASSP (1)1
2004 Spectro-temporal activity pattern (STAP) features for noise robust ASR
abstract
In this paper, we introduce a new noise robust representation of speech signal obtained by locating points of potential importance in the spectrogram, and parameterizing the activity of time-frequency pattern around those points. These features are referred to as Spectro-Temporal Activity Pattern (STAP) features. The suitability of these features for noise robust speech recognition is examined for a particular parameterization scheme where spectral peaks are chosen as points of potential importance. The activity in the time-frequency patterns around these points are parameterized by measuring the dynamics of the patterns along both time and frequency axes. As the spectral peaks are considered to constitute an important and robust cue for speech recognition, this representation is expected to yield a robust performance. An interesting result of the study is that inspite of using a relatively less amount of information from the speech signal, STAP features are able to achieve a reasonable recognition performance in clean speech, when compared to the state-of-the-art features. In addition, STAP features produce a significantly better performance in high noise conditions. An entropy based combination technique in tandem frame-work to combine STAP features with standard features yields a system which is more robust in all conditions.
Shajith Ikbal, Mathew Magimai-Doss, Hemant Misra, Hervé Bourlard
INTERSPEECH2
2004 Modeling auxiliary features in tandem systems
abstract
Tandem systems transform the cepstral features into posterior probabilities of subword units using artificial neural networks (ANNs), which are processed to form input features for conventional speech recognition systems. They have been shown to perform better than conventional speech recognition systems using cepstral features. Recent studies have shown that modelling cepstral features with auxiliary sources of knowledge leads to improvement in the performance of speech recognition systems. In this paper, we study two approaches to incorporate auxiliary knowledge sources such as pitch frequency, short-term energy, etc. (referred to as auxiliary features), in a tandem-based automatic speech recognition system. In the first approach, we model the auxiliary features in the process of training an ANN, which is later used to extract tandem-features. In the second approach, we extract the tandem-features from an ANN trained with cepstral features only and then model them jointly with auxiliary features. Recognition studies conducted on a connected word recognition task under clean and noisy conditions show that the performance of the tandem system can be improved by incorporating auxiliary features.
Mathew Magimai-Doss, Shajith Ikbal, Todd A. Stephenson, Hervé Bourlard
INTERSPEECH1
2004 Speech recognition with auxiliary information
abstract
State-of-the-art automatic speech recognition (ASR) systems are usually based on hidden Markov models (HMMs) that emit cepstral-based features which are assumed to be piecewise stationary. While not really robust to noise, these features are also known to be very sensitive to "auxiliary" information, such as pitch, energy, rate-of-speech (ROS), etc. Attempts so far to include such auxiliary information in state-of-the-art ASR systems have often been based on simply appending these auxiliary features to the standard acoustic feature vectors. In the present paper, we investigate different approaches to incorporating this auxiliary information using dynamic Bayesian networks (DBNs) or hybrid HMM/ANNs (HMMs with artificial neural networks). These approaches are motivated by the fact that the auxiliary information is not necessarily (directly) emitted by the HMM states but, rather, carries higher-level information (e.g., speaker characteristics) that is correlated with the standard features. As implicitly done for gender modeling elsewhere, this auxiliary information then appears as a conditional variable in the emission distributions and can be hidden (except in the case of some HMM/ANNs) as its estimates become too noisy. Based on recognition experiments carried out on the OGI Numbers database (free format numbers spoken over the telephone), we show that auxiliary information that conditions the distribution of the standard features can, in certain conditions, provide more robust recognition than using auxiliary information that is appended to the standard features; this is most evident in the case of energy as an auxiliary variable in noisy speech.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
IEEE Trans. Speech Audio Process.2
2003 Speech recognition of spontaneous, noisy speech using auxiliary information in Bayesian networks
abstract
Automatic speech recognition (ASR) currently performs well in the case of clean, read speech. It performs worse, however, when the speech is spontaneous and in noisy conditions. In previous work we showed the improvement that using auxiliary information in the framework of Bayesian networks (BNs) can bring to ASR in clean, read speech. Here we show that auxiliary information of pitch or rate-of-speech in the context of BNs also helps performance in spontaneous speech with noise.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
ICASSP (1)2
2003 Using pitch frequency information in speech recognition
abstract
Automatic Speech Recognition systems typically use smoothed spectral features as acoustic observations. In recent studies, it has been shown that complementing these standard features with pitch frequency could improve the system performance of the system. While previously proposed systems have been studied in the framework of HMM/GMMs, in this paper we study and compare different ways to include pitch frequency in state-of-the-art hybrid HMM/ANN system. We have evaluated the proposed system on two different ASR tasks, namely, isolated word recognition and connected word recognition. Our results show that pitch frequency can indeed be used in ASR systems to improve the recognition performance.
Mathew Magimai-Doss, Todd A. Stephenson, Hervé Bourlard
INTERSPEECH1
2003 Enhancement of speech in multispeaker environment
abstract
In this paper a method based on the excitation source information is proposed for enhancement of speech, degraded by speech from other speakers. Speech from multiple speakers is simultaneously collected over two spatially distributed microphones. Time-delay of each speaker with respect to the two microphones is estimated using the excitation source information. A weight function is derived for each speaker using the knowledge of the timedelay and the excitation source information. Linear prediction (LP) residuals of the microphone signals are processed separately using the weight functions. Speech signals are synthesized from the modified residuals. One speech signal per speaker is derived from each microphone signal. The synthesized speech signals of each speaker are combined to produce enhanced speech. Significant enhancement of the speech of one speaker relative to other was observed from the combined signal.
Bayya Yegnanarayana, S. R. Mahadeva Prasanna, Mathew Magimai-Doss
INTERSPEECH3
2002 Auxiliary variables in conditional Gaussian mixtures for automatic speech recognition
abstract
In previous work, we presented a case study using an estimated pitch value as the conditioning variable in conditional Gaussians that showed the utility of hiding the pitch values in certain situations or in modeling it independently of the hidden state in others. Since only single conditional Gaussians were used in that work, we extend that work here to using conditional Gaussian mixtures in the emission distributions to make this work more comparable to state-of-the-art automatic speech recognition. We also introduce a rate-of-speech (ROS) variable within the conditional Gaussian mixtures. We find that, under the current methods, using observed pitch or ROS in the recognition phase does not provide improvement. However, systems trained on pitch or ROS may provide improvement in the recognition phase over the baseline when the pitch or ROS is marginalized out.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH2
2001 Modeling auxiliary information in Bayesian network based ASR
abstract
Automatic speech recognition bases its models on the acoustic features derived from the speech signal. Some have investigated replacing or supplementing these features with information that can not be precisely measured (articulator positions, pitch, gender, etc.) automatically. Consequently, automatic estimations of the desired information would be generated. This data can degrade performance due to its imprecisions. In this paper, we describe a system that treats pitch as an auxiliary information within the framework of Bayesian networks, resulting in improved performance. 1.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH2