Hervé Bourlard

dblp:12/6705 · DBLP profile ↗
← Back
262ranked-venue papers
23as first author
9since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 220 · 16 first-author · 9 since 2021Artificial intelligence and machine learning · 147 · 13 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Comparison of 5 methods for the evaluation of intelligibility in mild to moderate French dysarthric speech
abstract
Altered quality of the phonetic-acoustic information in the speech signal in the case of motor speech disorders may reduce its intelligibility. Monitoring intelligibility is part of the standard clinical assessment of patients. It is also a valuable tool to index the evolution of the speech disorder. However, measuring intelligibility raises methodological debates concerning: the type of linguistic material on which the assessment is based (non-words, words, continuous speech), the evaluation protocol and type of scores (scale-based rating, transcription or recognition tests), and the advantages and disadvantages of listener vs. automatic-based approaches (subjective vs. objective, expertise level, types of models used). In this paper, the intelligibility of the speech of 32 French patients presenting mild to moderate dysarthria and 17 elderly speakers is assessed with five different methods: impressionistic clinician judgment on continuous speech, number of words recognized in an interactive face-to-face setting and in an on-line testing of the same material by 75 judges, automatic feature-based and automatic speech recognition-based methods (both on short sentences). The implications of the different methods for clinical practice are discussed.
Cécile Fougeron, Nicolas Audibert, Ina Kodrasi, Parvaneh Janbakhshi, Michaela Pernon, Nathalie Lévêque, Stephanie Borel, Marina Laganaro, Hervé Bourlard, Frédéric Assal
INTERSPEECH9
2022 From Undercomplete to Sparse Overcomplete Autoencoders to Improve LF-MMI based Speech Recognition
abstract
Starting from a strong Lattice-Free Maximum Mutual Information (LF-MMI) baseline system, we explore different autoencoder configurations to enhance Mel-Frequency Cepstral Coefficients (MFCC) features. Autoencoders are expected to generate new MFCC features that can be used in our LF-MMI based baseline system (with or without retraining) towards speech recognition improvements. Starting from shallow undercomplete autoencoders, and their known equivalence with Principal Component Analysis (PCA), we go to deeper or sparser architectures. In the spirit of kernel-based learning methods, we explore alternatives where the autoencoder first goes overcomplete (i.e., expand the representation space) in a nonlinear way, and then we restrict the autoencoder by means of a sequent bottleneck layer. Finally, as a third solution, we use sparse overcomplete autoencoders where a sparsity constraint is imposed on the higher-dimensional encoding layer. Experimental results are provided on the Augmented Multiparty Interaction (AMI) dataset, where we show that all aforementioned architectures improve speech recognition performance.
Selen Hande Kabil, Hervé Bourlard
INTERSPEECH2
2021 Speech Dereverberation Using Variational Autoencoders
abstract
This paper presents a statistical method for single-channel speech dereverberation using a variational autoencoder (VAE) for modelling the speech spectra. One popular approach for modelling speech spectra is to use non-negative matrix factorization (NMF) where learned clean speech spectral bases are used as a linear generative model for speech spectra. This work replaces this linear model with a powerful nonlinear deep generative model based on VAE. Further, this paper formulates a unified probabilistic generative model of reverberant speech based on Gaussian and Poisson distributions. We develop a Monte Carlo expectation-maximization algorithm for inferring the latent variables in the VAE and estimating the room impulse response for both probabilistic models. Evaluation results show the superiority of the proposed VAE-based models over the NMF-based counterparts.
Deepak Baby, Hervé Bourlard
ICASSP2
2021 Automatic Dysarthric Speech Detection Exploiting Pairwise Distance-Based Convolutional Neural Networks
abstract
Automatic dysarthric speech detection can provide reliable and cost-effective computer-aided tools to assist the clinical diagnosis and management of dysarthria. In this paper we propose a novel automatic dysarthric speech detection approach based on analyses of pairwise distance matrices using convolutional neural networks (CNNs). We represent utterances through articulatory posteriors and consider pairs of phonetically-balanced representations, with one representation from a healthy speaker (i.e., the reference representation) and the other representation from the test speaker (i.e., test representation). Given such pairs of reference and test representations, features are first extracted using a feature extraction front-end, a frame-level distance matrix is computed, and the obtained distance matrix is considered as an image by a CNN-based binary classifier. The feature extraction, distance matrix computation, and CNN-based classifier are jointly optimized in an end-to-end framework. Experimental results on two databases of healthy and dysarthric speakers for different languages and pathologies show that the proposed approach yields a high dysarthric speech detection performance, outperforming other CNN-based baseline approaches.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
ICASSP3
2021 Automatic And Perceptual Discrimination Between Dysarthria, Apraxia of Speech, and Neurotypical Speech
abstract
Automatic techniques in the context of motor speech disorders (MSDs) are typically two-class techniques aiming to discriminate between dysarthria and neurotypical speech or between dysarthria and apraxia of speech (AoS). Further, although such techniques are proposed to support the perceptual assessment of clinicians, the automatic and perceptual classification accuracy has never been compared. In this paper, we investigate a three-class automatic technique and a set of handcrafted features for the discrimination of dysarthria, AoS and neurotypical speech. Instead of following the commonly used One-versus-One or One-versus-Rest approaches for multi-class classification, a hierarchical approach is proposed. Further, a perceptual study is conducted where speech and language pathologists are asked to listen to recordings of dysarthria, AoS, and neurotypical speech and decide which class the recordings belong to. The proposed automatic technique is evaluated on the same recordings and the automatic and perceptual classification performance are compared. The presented results show that the hierarchical classification approach yields a higher classification accuracy than baseline One-versus-One and One-versus-Rest approaches. Further, the presented results show that the automatic approach yields a higher classification accuracy than the perceptual assessment of speech and language pathologists, demonstrating the potential advantages of integrating automatic tools in clinical practice.
Ina Kodrasi, Michaela Pernon, Marina Laganaro, Hervé Bourlard
ICASSP4
2021 Lattice-Free Mmi Adaptation of Self-Supervised Pretrained Acoustic Models
abstract
In this work, we propose lattice-free MMI (LFMMI) for supervised adaptation of self-supervised pretrained acoustic model. We pretrain a Transformer model on thousand hours of untranscribed Librispeech data followed by supervised adaptation with LFMMI on three different datasets. Our results show that fine-tuning with LFMMI, we consistently obtain relative WER improvements of 10% and 35.3% on the clean and other test sets of Librispeech (100h), 10.8% on Switchboard (300h), and 4.3% on Swahili (38h) and 4.4% on Tagalog (84h) compared to the baseline trained only with supervised data.
Apoorv Vyas, Srikanth R. Madikeri, Hervé Bourlard
ICASSP3
2021 Multitask Adaptation with Lattice-Free MMI for Multi-Genre Speech Recognition of Low Resource Languages
abstract
In this paper, we develop Automatic Speech Recognition (ASR) systems for multi-genre speech recognition of low-resource languages where training data is predominantly conversational speech but test data can be in one of the following genres: news broadcast, topical broadcast and conversational speech. ASR for low-resource languages is often developed by adapting a pre-trained model to a target language. When training data is predominantly from one genre and limited, the system's performance for other genres suffer. To handle such out-of-domain scenarios, we employ multitask adaptation by using auxiliary conversational speech data from other languages in addition to the target-language data. We aim to (1) improve adaptation through implicit data augmentation by adding other languages as auxiliary tasks, and (2) prevent the acoustic model from overfitting to the dominant genre in the training set. Pre-trained parameters are obtained from a multilingual model trained with data from 18 languages using the Lattice-Free Maximum Mutual Information (LF-MMI) criterion. The adaptation is performed with the LF-MMI criterion. We present results on MATERIAL datasets for three languages: Kazakh and Farsi and Pashto.
Srikanth R. Madikeri, Petr Motlícek, Hervé Bourlard
Interspeech3
2021 Comparing CTC and LFMMI for Out-of-Domain Adaptation of wav2vec 2.0 Acoustic Model
abstract
In this work, we investigate if the wav2vec 2.0 self-supervised pretraining helps mitigate the overfitting issues with connectionist temporal classification (CTC) training to reduce its performance gap with flat-start lattice-free MMI (E2E-LFMMI) for automatic speech recognition with limited training data. Towards that objective, we use the pretrained wav2vec 2.0 BASE model and fine-tune it on three different datasets including out-of-domain (Switchboard) and cross-lingual (Babel) scenarios. Our results show that for supervised adaptation of the wav2vec 2.0 model, both E2E-LFMMI and CTC achieve similar results; significantly outperforming the baselines trained only with supervised data. Fine-tuning the wav2vec 2.0 model with E2E-LFMMI and CTC we obtain the following relative WER improvements over the supervised baseline trained with E2E-LFMMI. We get relative improvements of 40% and 44% on the clean-set and 64% and 58% on the test set of Librispeech (100h) respectively. On Switchboard (300h) we obtain relative improvements of 33% and 35% respectively. Finally, for Babel languages, we obtain relative improvements of 26% and 23% on Swahili (38h) and 18% and 17% on Tagalog (84h) respectively.
Apoorv Vyas, Srikanth R. Madikeri, Hervé Bourlard
Interspeech3
2021 Subspace-Based Learning for Automatic Dysarthric Speech Detection
abstract
To assist the clinical diagnosis and treatment of speech dysarthria, automatic dysarthric speech detection techniques providing reliable and cost-effective assessment are indispensable. Based on clinical evidence on spectro-temporal distortions associated with dysarthric speech, we propose to automatically discriminate between healthy and dysarthric speakers exploiting spectro-temporal subspaces of speech. Spectro-temporal subspaces are extracted using singular value decomposition, and dysarthric speech detection is achieved by applying a subspace-based discriminant analysis. Experimental results on databases of healthy and dysarthric speakers for different languages and pathologies show that the proposed subspace-based approach using temporal subspaces is more advantageous than using spectral subspaces, also outperforming several state-of-the-art automatic dysarthric speech detection techniques.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
IEEE Signal Process. Lett.3
2020 Synthetic Speech References for Automatic Pathological Speech Intelligibility Assessment
abstract
Automatic pathological speech intelligibility measures are crucial to assist the clinical diagnosis and treatment of speech disorders. The recently proposed pathological short-time objective intelligibility (P-ESTOI) measure was shown to be very advantageous, yielding a high performance for several speech pathologies. However, to assess the intelligibility of an utterance from a patient, P-ESTOI relies on the availability of recordings of the same utterance by several healthy speakers such that an intelligible reference model can be created. Such recordings are not always easily available, limiting the practical applicability of P-ESTOI. To be able to use P-ESTOI in such scenarios, in this paper we propose to use synthetic speech generated by state-of-the-art high-quality text-to-speech systems to create an intelligible reference model. Experimental results on a database of Cerebral Palsy patients show that the performance of P-ESTOI using synthetic speech references is comparable to using natural speech references, making P-ESTOI a flexible measure which does not require healthy speech recordings and which outperforms state-of-the-art pathological speech intelligibility measures.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
ICASSP3
2020 Incremental Semi-Supervised Learning for Multi-Genre Speech Recognition
abstract
In this work, we explore a data scheduling strategy for semi-supervised learning (SSL) for acoustic modeling in automatic speech recognition. The conventional approach uses a seed model trained with supervised data to automatically recognize the entire set of unlabeled (auxiliary) data to generate new labels for subsequent acoustic model training. In this paper, we propose an approach in which the unlabelled set is divided into multiple equal-sized subsets. These subsets are processed in an incremental fashion: for each iteration a new subset is added to the data used for SSL, starting from only one subset in the first iteration. The acoustic model from the previous iteration becomes the seed model for the next one. This scheduling strategy is compared to the approach employing all unlabeled data in one-shot for training. Experiments using lattice-free maximum mutual information based acoustic model training on Fisher English gives 80% word error recovery rate. On the multi-genre evaluation sets on Lithuanian and Bulgarian relative improvements of up to 17.2% in word error rate are observed.
Banriskhem K. Khonglah, Srikanth R. Madikeri, Subhadeep Dey, Hervé Bourlard, Petr Motlícek, Jayadev Billa
ICASSP4
2020 Automatic Discrimination of Apraxia of Speech and Dysarthria Using a Minimalistic Set of Handcrafted Features
abstract
To assist clinicians in the differential diagnosis and treatment of motor speech disorders, it is imperative to establish objective tools which can reliably characterize different subtypes of disorders such as apraxia of speech (AoS) and dysarthria.Objective tools in the context of speech disorders typically rely on thousands of acoustic features, which raises the risk of difficulties in the interpretation of the underlying mechanisms, overadaptation to training data, and weak generalization capabilities to test data.Seeking to use a small number of acoustic features and motivated by the clinical-perceptual signs used for the differential diagnosis of AoS and dysarthria, we propose to characterize differences between AoS and dysarthria using only six handcrafted acoustic features, with three features reflecting segmental distortions, two features reflecting loudness and hypernasality, and one feature reflecting syllabification.These three different sets of features are used to separately train three classifiers.At test time, the decisions of the three classifiers are combined through a simple majority voting scheme.Preliminary results show that the proposed approach achieves a discrimination accuracy of 90%, outperforming using state-of-the-art features such as openSMILE which yield a discrimination accuracy of 65%.
Ina Kodrasi, Michaela Pernon, Marina Laganaro, Hervé Bourlard
INTERSPEECH4
2020 Lattice-Free Maximum Mutual Information Training of Multilingual Speech Recognition Systems
abstract
Multilingual acoustic model training combines data from multiple languages to train an automatic speech recognition system.Such a system is beneficial when training data for a target language is limited.Lattice-Free Maximum Mutual Information (LF-MMI) training performs sequence discrimination by introducing competing hypotheses through a denominator graph in the cost function.The standard approach to train a multilingual model with LF-MMI is to combine the acoustic units from all languages and use a common denominator graph.The resulting model is either used as a feature extractor to train an acoustic model for the target language or directly fine-tuned.In this work, we propose a scalable approach to train the multilingual acoustic model using a typical multitask network for the LF-MMI framework.A set of language-dependent denominator graphs is used to compute the cost function.The proposed approach is evaluated under typical multilingual ASR tasks using GlobalPhone and BABEL datasets.Relative improvements up to 13.2% in WER are obtained when compared to the corresponding monolingual LF-MMI baselines.The implementation is made available as a part of the Kaldi speech recognition toolkit.
Srikanth R. Madikeri, Banriskhem K. Khonglah, Sibo Tong, Petr Motlícek, Hervé Bourlard, Daniel Povey
INTERSPEECH5
2020 On quantifying the quality of acoustic models in hybrid DNN-HMM ASR
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
Speech Commun.3
2020 Automatic Pathological Speech Intelligibility Assessment Exploiting Subspace-Based Analyses
abstract
Competitive state-of-the-art automatic pathological speech intelligibility measures typically rely on regression training on a large number of features, require a large amount of healthy speech training data, or are applicable only to phonetically balanced scenarios where healthy and pathological speakers utter the same utterances. As a result, their performance in unseen data is unsatisfactory, and they cannot be used in low-resource languages or in phonetically unbalanced scenarios. To overcome these drawbacks, we propose a subspace-based intelligibility (SBI) measure. The SBI measure operates based on the hypothesis that dominant spectral patterns of pathological speech differ from intelligible speech (where the pathological and intelligible speech signals do not need to match in phonetic content), with the difference increasing as pathological speech intelligibility decreases. The SBI measure uses a minimal number of speech recordings to compute dominant spectral basis vectors spanning intelligible and pathological speech. The subspaces spanned by the intelligible and pathological spectral basis vectors are compared to each other through a subspace distance measure, which is directly used (i.e., without any training) as the pathological speech intelligibility estimate. Exploiting psychoacoustic evidence on the importance of spectral modulation cues to the perceived speech intelligibility and clinical evidence on the degradation of these cues in pathological speech, we show that the power of the proposed SBI measure lies in capturing the effect of spectral modulation degradation. To be able to additionally track possible degradations in the temporal structure of the pathological speech signal, we also propose two extensions of the SBI measure by incorporating short-time temporal information. Experimental results for different languages and speech pathologies show that the proposed intelligibility measures yield high and significant correlations with subjective intelligibility ratings, while not requiring any regression training or a large number of healthy speech recordings and being applicable to phonetically unbalanced scenarios.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Spectro-Temporal Sparsity Characterization for Dysarthric Speech Detection
abstract
To assist the clinical diagnosis and treatment of neurological diseases that cause speech dysarthria such as Parkinson's disease (PD), it is of paramount importance to craft robust features which can be used to automatically discriminate between healthy and dysarthric speech. Since dysarthric speech of patients suffering from PD is breathy, semi-whispery, and is characterized by abnormal pauses and imprecise articulation, it can be expected that its spectro-temporal sparsity differs from the spectro-temporal sparsity of healthy speech. While we have recently successfully used temporal sparsity characterization for dysarthric speech detection, characterizing spectral sparsity poses the challenge of constructing a valid feature vector from signals with a different number of unaligned time frames. Further, although several non-parametric and parametric measures of sparsity exist, it is unknown which sparsity measure yields the best performance in the context of dysarthric speech detection. The objective of this paper is to demonstrate the advantages of spectro-temporal sparsity characterization for automatic dysarthric speech detection. To this end, we first provide a numerical analysis of the suitability of different non-parametric and parametric measures (i.e., l1-norm, kurtosis, Shannon entropy, Gini index, shape parameter of a Chi distribution, and shape parameter of a Weibull distribution) for sparsity characterization. It is shown that kurtosis, the Gini index, and the parametric sparsity measures are advantageous sparsity measures, whereas the l1-norm and entropy measures fail to robustly characterize the temporal sparsity of signals with a different number of time frames. Second, we propose to characterize the spectral sparsity of an utterance by initially time-aligning it to the same utterance uttered by a (arbitrarily selected) reference speaker using dynamic time warping. Experimental results on a Spanish database of healthy and dysarthric speech show that estimating the spectro-temporal sparsity using the Gini index or the parametric sparsity measures and using it as a feature in a support vector machine results in a high classification accuracy of 83.3%.
Ina Kodrasi, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Neural Network Based End-to-End Query by Example Spoken Term Detection
abstract
This article focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. State-of-the-art approaches primarily rely on dynamic time warping (DTW) based template matching techniques using phone posterior or bottleneck features extracted from a deep neural network (DNN). We use both monolingual and multilingual bottleneck features, and show that multilingual features perform increasingly better with more training languages. Previously, it has been shown that the DTW based matching can be replaced with a CNN based matching while using posterior features. Here, we show that the CNN based matching outperforms DTW based matching using bottleneck features as well. In this case, the feature extraction and pattern matching stages of our QbE-STD system are optimized independently of each other. We propose to integrate these two stages in a fully neural network based end-to-end learning framework to enable joint optimization of those two stages simultaneously. The proposed approaches are evaluated on two challenging multilingual datasets: Spoken Web Search 2013 and Query by Example Search on Speech Task 2014, demonstrating in each case significant improvements.
Dhananjay Ram, Lesly Miculicich, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Multilingual Bottleneck Features for Query by Example Spoken Term Detection
abstract
State of the art solutions to query by example spoken term detection (QbE-STD) rely on bottleneck feature representation of the query and audio document. Here, we present a study on QbE-STD performance using several monolingual as well as multilingual bottleneck features extracted from feed forward networks. In contrast to previous works, we use multitask learning to train the multilingual networks which perform significantly better than the concatenated monolingual features. Additionally, we propose to employ residual networks (ResNet) to estimate the bottleneck features and show significant improvements over the corresponding feed forward network based features. The neural networks are trained on GlobalPhone corpus and QbE-STD experiments are performed on a very challenging QUESST 2014 database
Dhananjay Ram, Lesly Miculicich, Hervé Bourlard
ASRU3
2019 Pathological Speech Intelligibility Assessment Based on the Short-time Objective Intelligibility Measure
abstract
Impaired speech intelligibility in motor speech disorders arising due to neurological diseases negatively affects the communication ability and quality of life of patients. Reliable and cost-effective measures to automatically assess speech intelligibility are necessary for the management of such disorders. In this paper, we propose to automatically assess the intelligibility of pathological speech based on short-time objective intelligibility measures typically used in speech enhancement, which however require a reference signal that is time-aligned to the test signal. We propose a method to create an utterance-dependent reference signal of intelligible speech from multiple healthy speakers. In order to assess intelligibility, the pathological speech signal is aligned to the created reference signal using dynamic time warping and the divergence between the two signals is quantified using either the short-time or the spectral correlation. Experiments on databases of English and French patients suffering from Cerebral Palsy and Amyotrophic Lateral Sclerosis show that the proposed intelligibility measures can obtain a high correlation with subjective intelligibility ratings, outperforming several state-of-the-art pathological speech intelligibility measures.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
ICASSP3
2019 Super-gaussianity of Speech Spectral Coefficients as a Potential Biomarker for Dysarthric Speech Detection
abstract
Parkinson's disease (PD) and Amyotrophic Lateral Sclerosis (ALS) are progressive neurodegenerative diseases which, among other symptoms, cause dysarthria of speech. To assist the clinical diagnosis and treatment of neurological diseases, several studies have addressed the characterization and classification of healthy and dysarthric speech. However, most contributions deal with PD speech, with significantly fewer results presented for ALS speech. The objective of this paper is to show that ALS speech has a similar statistical distribution as PD speech, with the complex spectral coefficients being significantly less super-Gaussian than healthy speech spectral coefficients. In addition, a method to exploit the super-Gaussianity of speech signals as a feature to classify healthy and dysarthric speech is presented and evaluated. The proposed approach is evaluated on a French database of healthy and dysarthric (PD and ALS) speech. Experimental results show that the use of the super-Gaussianity of speech signals yields a significantly higher classification accuracy than state-of-the-art features such as fundamental frequency, jitter, shimmer, harmonics-to-noise ratio, or Mel frequency cepstral coefficients.
Ina Kodrasi, Hervé Bourlard
ICASSP2
2019 An End-to-end Network to Synthesize Intonation Using a Generalized Command Response Model
abstract
The generalized command response (GCR) model represents intonation as a superposition of muscle responses to spike command signals. We have previously shown that the spikes can be predicted by a two-stage system, consisting of a recurrent neural network and a post-processing procedure, but the responses themselves were fixed dictionary atoms. We propose an end-to-end neural architecture that replaces the dictionary atoms with trainable second-order recurrent elements analogous to recursive filters. We demonstrate gradient stability under modest conditions, and show that the system can be trained by imposing temporal sparsity constraints. Subjective listening tests demonstrate that the system can synthesize intonation with high naturalness, comparable to state-of-the-art acoustic models, and retains the physiological plausibility of the GCR model.
François Marelli, Bastian Schnell, Hervé Bourlard, Thierry Dutoit, Philip N. Garner
ICASSP3
2019 An Investigation of Multilingual ASR Using End-to-end LF-MMI
abstract
The end-to-end lattice-free maximum mutual information (LF-MMI) approach has recently been shown to be beneficial for automatic speech recognition (ASR) in general. More specifically, its end-to-end nature and use of context independent phone labels make it attractive for multilingual ASR. We show that end-to-end LF-MMI is indeed competitive on a low-resourced multilingual task, comfortably outperforming a connectionist temporal classification (CTC) baseline. We further investigate the feasibility of biphone contexts, being a candidate compromise between the context independent approach and the triphone contexts that usually perform well. We show that biphones do not initially perform well, but can do so after language adaptive training, concluding that biphones carry language variability but are promising for multilingual ASR.
Sibo Tong, Philip N. Garner, Hervé Bourlard
ICASSP3
2019 Analyzing Uncertainties in Speech Recognition Using Dropout
abstract
The performance of Automatic Speech Recognition (ASR) systems is often measured using Word Error Rates (WER) which requires time-consuming and expensive manually transcribed data. In this paper, we use state-of-the-art ASR systems based on Deep Neural Networks (DNN) and propose a novel framework which uses "Dropout" at the test time to model uncertainty in prediction hypotheses. We systematically exploit this uncertainty to estimate WER without the need for explicit transcriptions. In addition, we show that the predictive uncertainty can also be used to accurately localize the errors made by the ASR system. We study the performance of our approach on Switchboard database where it predicts WER accurately within a range of 2.6% and 5.0% for HMM-DNN and Connectionist Temporal Classification (CTC) ASR systems, respectively.
Apoorv Vyas, Pranay Dighe, Sibo Tong, Hervé Bourlard
ICASSP4
2019 Spectral Subspace Analysis for Automatic Assessment of Pathological Speech Intelligibility
abstract
Speech intelligibility is an important assessment criterion of the communicative performance of pathological speakers. To assist clinicians in their assessment, time- and cost-efficient automatic intelligibility measures offering a repeatable and reliable assessment are desired. In this paper, we propose to automatically assess pathological speech intelligibility based on a distance measure between the subspaces of spectral patterns of the pathological speech signal and of a fully intelligible (healthy) speech signal. To extract the subspace of spectral patterns we investigate two linear decomposition methods, i.e., Principal Component Analysis and Approximate Joint Diagonalization. Pathological speech intelligibility is then derived using a Grassman distance measure which quantifies the difference between the extracted subspaces of pathological and healthy speech. Experiments on an English database of Cerebral Palsy patients show that the proposed intelligibility measure is significantly correlated with subjective intelligibility ratings. In addition, comparisons to state-of-the-art measures show that the proposed subspace-based measure achieves a high performance with a significantly lower computational cost and without imposing any constraints on the speech material of the speakers.
Parvaneh Janbakhshi, Ina Kodrasi, Hervé Bourlard
INTERSPEECH3
2019 Unbiased Semi-Supervised LF-MMI Training Using Dropout
Sibo Tong, Apoorv Vyas, Philip N. Garner, Hervé Bourlard
INTERSPEECH4
2019 Low-rank and sparse subspace modeling of speech for DNN based acoustic modeling
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
Speech Commun.3
2018 Phonological Posterior Hashing for Query by Example Spoken Term Detection
abstract
State of the art query by example spoken term detection (QbE-STD) systems in zero-resource conditions rely on representation of speech in terms of sequences of class-conditional posterior probabilities estimated by deep neural network (DNN). The posteriors are often used for pattern matching or dynamic time warping (DTW). Exploiting posterior probabilities as speech representation propounds diverse advantages in a classification system. One key property of the posterior representations is that they admit a highly effective hashing strategy that enables indexing a large audio archive in divisions for reducing the search complexity. Moreover, posterior indexing leads to a compressed representation and enables pronunciation dewarping and partial detection with no need for DTW. We exploit these characteristics of the posterior space in the context of redundant hash addressing for query-by-example spoken term detection (QbE-STD). We evaluate the QbE-STD system on AMI corpus and demonstrate that tremendous speedup and superior accuracy is achieved compared to the state-of-the-art pattern matching solution based on DTW. The system has the potential to enable massively large scale spoken query detection.
Afsaneh Asaei, Dhananjay Ram, Hervé Bourlard
INTERSPEECH3
2018 Evolution of Neural Network Architectures for Speech Recognition
Hervé Bourlard
INTERSPEECH1
2018 Single-channel Late Reverberation Power Spectral Density Estimation Using Denoising Autoencoders
abstract
In order to suppress the late reverberation in the spectral domain, many single-channel dereverberation techniques rely on an estimate of the late reverberation power spectral density (PSD).In this paper, we propose a novel approach to late reverberation PSD estimation using a denoising autoencoder (DA), which is trained to learn a mapping from the microphone signal PSD to the late reverberation PSD.Simulation results show that the proposed approach yields a high PSD estimation accuracy and generalizes well to unseen data.Furthermore, simulation results show that the proposed DA-based PSD estimate yields a higher PSD estimation accuracy and a similar dereverberation performance than a state-of-the-art statistical PSD estimate, which additionally also requires knowledge of the reverberation time.
Ina Kodrasi, Hervé Bourlard
INTERSPEECH2
2018 CNN Based Query by Example Spoken Term Detection
abstract
In this work, we address the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. State of the art solutions usually rely on dynamic time warping (DTW) based template matching. In contrast, we propose here to tackle the problem as binary classification of images. Similar to the DTW approach, we rely on deep neural network (DNN) based posterior probabilities as feature vectors. The posteriors from a spoken query and a test utterance are used to compute frame-level similarities in a matrix form. This matrix contains somewhere a quasi-diagonal pattern if the query occurs in the test utterance. We propose to use this matrix as an image and train a convolutional neural network (CNN) for identifying the pattern and make a decision about the occurrence of the query. This language independent system is evaluated on SWS 2013 and is shown to give 10% relative improvement over a highly competitive baseline system based on DTW. Experiments on QUESST 2014 database gives similar improvements showing that the approach generalizes to other databases as well.
Dhananjay Ram, Lesly Miculicich, Hervé Bourlard
INTERSPEECH3
2018 Fast Language Adaptation Using Phonological Information
abstract
Phoneme-based multilingual connectionist temporal classification (CTC) model is easily extensible to a new language by concatenating parameters of the new phonemes to the output layer. In the present paper, we improve cross-lingual adaptation in the context of phoneme-based CTC models by using phonological information. A universal (IPA) phoneme classifier is first trained on phonological features generated from a phonological attribute detector. When adapting the multilingual CTC to a new, never seen, language, phonological attributes of the unseen phonemes are derived based on phonology and fed into the phoneme classifier. Posteriors given by the classifier are used to initialize the parameters of the unseen phonemes when extending the multilingual CTC output layer to the target language. Adaptation experiments show that the proposed initialization approaches further improve the cross-lingual adaptation on CTC models and yield significant improvements over Deep Neural Network / Hidden Markov Model (DNN/HMM)-based adaptation using limited data.
Sibo Tong, Philip N. Garner, Hervé Bourlard
INTERSPEECH3
2018 Far-Field ASR Using Low-Rank and Sparse Soft Targets from Parallel Data
abstract
Far-field automatic speech recognition (ASR) of conversational speech is often considered to be a very challenging task due to the poor quality of alignments available for training the DNN acoustic models. A common way to alleviate this problem is to use clean alignments obtained from parallelly recorded close-talk speech data. In this work, we advance the parallel data approach by obtaining enhanced low-rank and sparse soft targets from a close-talk ASR system and using them for training more accurate far-field acoustic models. Specifically, we (i) exploit eigenposteriors and Compressive Sensing dictionaries to learn low-dimensional senone subspaces in DNN posterior space, and (ii) enhance close-talk DNN posteriors to achieve high quality soft targets for training far-field DNN acoustic models. We show that the enhanced soft targets encode the structural and temporal interrelationships among senone classes which are easily accessible in the DNN posterior space of close-talk speech but not in its noisy far-field counterpart. We exploit enhanced soft targets to improve the mapping of far-field acoustics to close-talk senone classes. The experiments are performed on AMI meeting corpus where our approach improves DNN based acoustic modeling by 4.4% absolute (~8% rel.) reduction in WER as compared to a system which doesn't use parallel data. Finally, the approach is also validated on state-of-the-art recurrent and time delay neural network architectures.
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
SLT3
2018 Phonetic subspace features for improved query by example spoken term detection
Dhananjay Ram, Afsaneh Asaei, Hervé Bourlard
Speech Commun.3
2018 Cross-lingual adaptation of a CTC-based multilingual acoustic model
Sibo Tong, Philip N. Garner, Hervé Bourlard
Speech Commun.3
2018 Sparse Subspace Modeling for Query by Example Spoken Term Detection
abstract
This paper focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. Current state-of-the-art approaches to tackle this problem rely on dynamic programming based template matching techniques using phone posterior features extracted at the output of a deep neural network. Previously, it has been shown that the space of phone posteriors is highly structured, as a union of low-dimensional subspaces. To exploit the temporal and sparse structure of the speech data, we investigate here three different QbE-STD systems based on sparse model recovery. More specifically, we use query examples to model the query subspace using dictionary for sparse coding. Reconstruction errors calculated using sparse representation of feature vectors are then used to characterize the underlying subspaces. The first approach uses these reconstruction errors in a dynamic programming framework to detect the spoken query, resulting in a much faster search compared to standard template matching. The other two methods aim at merging template matching and sparsity-based approaches to further improve the performance. The first one proposes to regularize the template matching local distances using sparse reconstruction errors. The second approach aims at using the sparse reconstruction errors to rescore (improve) the template matching likelihood. Experiments on two different databases (AMI and MediaEval) show that the proposed hybrid systems perform better than a highly competitive QbE-STD baseline system.
Dhananjay Ram, Afsaneh Asaei, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Low-rank and sparse soft targets to learn better DNN acoustic models
abstract
Conventional deep neural networks (DNN) for speech acoustic modeling rely on Gaussian mixture models (GMM) and hidden Markov model (HMM) to obtain binary class labels as the targets for DNN training. Subword classes in speech recognition systems correspond to context-dependent tied states or senones. The present work addresses some limitations of GMM-HMM senone alignments for DNN training. We hypothesize that the senone probabilities obtained from a DNN trained with binary labels can provide more accurate targets to learn better acoustic models. However, DNN outputs bear inaccuracies which are exhibited as high dimensional unstructured noise, whereas the informative components are structured and low-dimensional. We exploit principal component analysis (PCA) and sparse coding to characterize the senone subspaces. Enhanced probabilities obtained from low-rank and sparse reconstructions are used as soft-targets for DNN acoustic modeling, that also enables training with untranscribed data. Experiments conducted on AMI corpus shows 4.6% relative reduction in word error rate.
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
ICASSP3
2017 Exploiting Eigenposteriors for Semi-Supervised Training of DNN Acoustic Models with Sequence Discrimination
abstract
LIDIAP
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
INTERSPEECH3
2017 An Investigation of Deep Neural Networks for Multilingual Speech Recognition Training and Adaptation
abstract
Different training and adaptation techniques for multilingual Automatic Speech Recognition (ASR) are explored in the context of hybrid systems, exploiting Deep Neural Networks (DNN) and Hidden Markov Models (HMM). In multilingual DNN training, the hidden layers (possibly extracting bottleneck features) are usually shared across languages, and the output layer can either model multiple sets of language-specific senones or one single universal IPA-based multilingual senone set. Both architectures are investigated, exploiting and comparing different language adaptive training (LAT) techniques originating from successful DNN-based speaker-adaptation. More specifically, speaker adaptive training methods such as Cluster Adaptive Training (CAT) and Learning Hidden Unit Contribution (LHUC) are considered. In addition, a language adaptive output architecture for IPA-based universal DNN is also studied and tested. Experiments show that LAT improves the performance and adaptation on the top layer further improves the accuracy. By combining state-level minimum Bayes risk (sMBR) sequence training with LAT, we show that a language adaptively trained IPA-based universal DNN outperforms a monolingually sequence trained model.
Sibo Tong, Philip N. Garner, Hervé Bourlard
INTERSPEECH3
2017 Perceptual Information Loss due to Impaired Speech Production
abstract
Phonological classes define articulatory-free and articulatory-bound phone attributes. Deep neural network is used to estimate the probability of phonological classes from the speech signal. In theory, a unique combination of phone attributes form a phoneme identity. Probabilistic inference of phonological classes thus enables estimation of their compositional phoneme probabilities. A novel information theoretic framework is devised to quantify the information conveyed by each phone attribute, and assess the speech production quality for perception of phonemes. As a use case, we hypothesize that disruption in speech production leads to information loss in phone attributes, and thus confusion in phoneme identification. We quantify the amount of information loss due to dysarthric articulation recorded in the TORGO database. A novel information measure is formulated to evaluate the deviation from an ideal phone attribute production leading us to distinguish healthy production from pathological speech.
Afsaneh Asaei, Milos Cernak, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Exploiting low-dimensional structures to enhance DNN based acoustic modeling in speech recognition
abstract
We propose to model the acoustic space of deep neural network (DNN) class-conditional posterior probabilities as a union of low-dimensional subspaces. To that end, the training posteriors are used for dictionary learning and sparse coding. Sparse representation of the test posteriors using this dictionary enables projection to the space of training data. Relying on the fact that the intrinsic dimensions of the posterior subspaces are indeed very small and the matrix of all posteriors belonging to a class has a very low rank, we demonstrate how low-dimensional structures enable further enhancement of the posteriors and rectify the spurious errors due to mismatch conditions. The enhanced acoustic modeling method leads to improvements in continuous speech recognition task using hybrid DNN-HMM (hidden Markov model) framework in both clean and noisy conditions, where upto 15.4% relative reduction in word error rate (WER) is achieved.
Pranay Dighe, Gil Luyet, Afsaneh Asaei, Hervé Bourlard
ICASSP4
2016 System fusion and speaker linking for longitudinal diarization of TV shows
abstract
Performing speaker diarization while uniquely identifying the speakers in a collection of audio recordings is a challenging task. Based on our previous work on speaker diarization and linking, we developed a system for diarizing longitudinal TV show data sets based on the fusion of speaker diarization system outputs and speaker linking. Agreement between multiple diarization outputs is found prior to speaker linking, largely reducing the diarization error rate at the expense of keeping some speech data unlabelled. To deal with noisy clusters, a linear prediction based technique was used to label speakers after linking. Considerable gains for both fusion and labelling are reported. Despite the challenges of the longitudinal diarization task, this system obtained similar performance for linked and non-linked tasks under moderate session variability, highlighting the viability of a linking approach to longitudinal diarization of speech in the presence of noise, music and special audio effects.
Marc Ferras, Srikanth R. Madikeri, Petr Motlícek, Hervé Bourlard
ICASSP4
2016 Phonetic and Phonological Posterior Search Space Hashing Exploiting Class-Specific Sparsity Structures
abstract
This paper shows that exemplar-based speech processing using class-conditional posterior probabilities admits a highly effective search strategy relying on posteriors' intrinsic sparsity structures. The posterior probabilities are estimated for phonetic and phonological classes using deep neural network (DNN) computational framework. Exploiting the class-specific sparsity leads to a simple quantized posterior hashing procedure to reduce the search space of posterior exemplars. To that end, small number of quantized posteriors are regarded as representatives of the posterior space and used as hash keys to index subsets of neighboring exemplars. The $k$ nearest neighbor ($k$NN) method is applied for posterior based classification problems. The phonetic posterior probabilities are used as exemplars for phonetic classification whereas the phonological posteriors are used as exemplars for automatic prosodic event detection. Experimental results demonstrate that posterior hashing improves the efficiency of $k$NN classification drastically. This work encourages the use of posteriors as discriminative exemplars appropriate for large scale speech classification tasks.
Afsaneh Asaei, Gil Luyet, Milos Cernak, Hervé Bourlard
INTERSPEECH4
2016 Sound Pattern Matching for Automatic Prosodic Event Detection
abstract
LIDIAP
Milos Cernak, Afsaneh Asaei, Pierre-Edouard Honnet, Philip N. Garner, Hervé Bourlard
INTERSPEECH5
2016 Inter-Task System Fusion for Speaker Recognition
Marc Ferras, Srikanth R. Madikeri, Subhadeep Dey, Petr Motlícek, Hervé Bourlard
INTERSPEECH5
2016 Low-Rank Representation of Nearest Neighbor Posterior Probabilities to Enhance DNN Based Acoustic Modeling
abstract
LIDIAP
Gil Luyet, Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
INTERSPEECH4
2016 Subspace Detection of DNN Posterior Probabilities via Sparse Representation for Query by Example Spoken Term Detection
abstract
We cast the query by example spoken term detection (QbE-STD) problem as subspace detection where query and background subspaces are modeled as union of low-dimensional subspaces. The speech exemplars used for subspace modeling are class-conditional posterior probabilities estimated using deep neural network (DNN). The query and background training exemplars are exploited to model the underlying low-dimensional subspaces through dictionary learning for sparse representation. Given the dictionaries characterizing the query and background subspaces, QbE-STD is performed based on the ratio of the two corresponding sparse representation reconstruction errors. The proposed subspace detection method can be formulated as the generalized likelihood ratio test for composite hypothesis testing. The experimental evaluation demonstrate that the proposed method is able to detect the query given a single example and performs significantly better than a highly competitive QbE-STD baseline system based on template matching.
Dhananjay Ram, Afsaneh Asaei, Hervé Bourlard
INTERSPEECH3
2016 Computational methods for underdetermined convolutive speech localization and separation via model-based sparse component analysis
Afsaneh Asaei, Hervé Bourlard, Mohammad Javad Taghizadeh, Volkan Cevher
Speech Commun.2
2016 On structured sparsity of phonological posteriors for linguistic parsing
Milos Cernak, Afsaneh Asaei, Hervé Bourlard
Speech Commun.3
2016 Sparse modeling of neural network posterior probabilities for exemplar-based speech recognition
Pranay Dighe, Afsaneh Asaei, Hervé Bourlard
Speech Commun.3
2016 Predicting the intrusiveness of noise through sparse coding with auditory kernels
Raphael Ullmann, Hervé Bourlard
Speech Commun.2
2016 A Large-Scale Open-Source Acoustic Simulator for Speaker Recognition
abstract
The state-of-the-art speaker-recognition systems suffer from significant performance loss on degraded speech conditions and acoustic mismatch between enrolment and test phases. Past international evaluation campaigns, such as the NIST speaker recognition evaluation (SRE), have partly addressed these challenges in some evaluation conditions. This work aims at further assessing and compensating for the effect of a wide variety of speech-degradation processes on speaker-recognition performance. We present an open-source simulator generating degraded telephone, VoIP, and interview-speech recordings using a comprehensive list of narrow-band, wide-band, and audio codecs, together with a database of over 60 h of environmental noise recordings and over 100 impulse responses collected from publicly available data. We provide speaker-verification results obtained with an i-vector-based system using either a clean or degraded PLDA back-end on a NIST SRE subset of data corrupted by the proposed simulator. While error rates increase considerably under degraded speech conditions, large relative equal error rate (EER) reductions were observed when using a PLDA model trained with a large number of degraded sessions per speaker.
Marc Ferras, Srikanth R. Madikeri, Petr Motlícek, Subhadeep Dey, Hervé Bourlard
IEEE Signal Process. Lett.5
2016 Speaker Diarization and Linking of Meeting Data
abstract
Finding who spoke when in a collection of recordings, with speakers being uniquely identified across the database, is a challenging task. In this scenario, reasonable computing times and acoustic variation across recordings remain two major concerns to address in state-of-the-art speaker diarization systems. This paper extends prior work on diarizing large speech datasets using algorithms that scale well with increasing amounts of data while compensating for across-recording variability. We follow a two-stage approach performing speaker diarization and speaker linking, the former focusing on local within-recording speaker changes and the latter focusing on global speaker changes across the database. In this study, we explore how these two modules interact with each other, while proposing a diarization fusion approach that prevents diarization errors from propagating to the linking stage. We further explore the diarization fusion for speaker linking using different linking strategies and speaker modeling variants. Evaluation is performed on single distant microphone data from the augmented multiparty interaction corpus show the effectiveness of the fusion approach after speaker linking and intersession variability modeling via joint factor analysis.
Marc Ferras, Srikanth R. Madikeri, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 On application of non-negative matrix factorization for ad hoc microphone array calibration from incomplete noisy distances
abstract
We propose to use non-negative matrix factorization (NMF) to estimate the unknown pairwise distances and reconstruct a distance matrix for microphone array position calibration. We develop new multiplicative update rules for NMF with incomplete input matrix that take into account the symmetry of the distance matrix. Additionally, we develop a convex matrix completion method which is related to an l2-regularized symmetric NMF. Thorough experiments demonstrate that the proposed methods lead to substantial improvement over the state-of-the-art techniques in a wide range of signal-to-noise and unknown-distance ratios. The convex symmetric matrix completion method was found to be the most robust method with less computational cost.
Afsaneh Asaei, Nasser Mohammadiha, Mohammad Javad Taghizadeh, Simon Doclo, Hervé Bourlard
ICASSP5
2015 KL-HMM based speaker diarization system for meetings
abstract
In this paper, the Kullback-Leibler Hidden Markov Model (KL-HMMs) is applied for unsupervised diarization of speech. A general approach to speaker diarization is to split the audio into uniform segments followed by one or more iterations of clustering of the segments and resegmentation of the audio. In the Information Bottlneck (IB) approach to diarization, short uniform segments are clustered using the IB criterion followed by resegmentation with KL-HMM. The KL-HMM approach has been shown to be an effective resegmentation procedure in this respect. Thus, the potential of KL-HMM as an independent diarization system is explored where the uniform segments are clustered and segmented using a sequence of posteriors obtained from the audio with respect to a Gaussian Mixture Model (GMM). The segmentation is performed using KL divergence, while the Jensen Shanon (JS) divergence is used for clustering. The diarization procedure is stopped by applying a Normalized Mutual Information (NMI) based criterion between two consecutive clustering outputs. The proposed method is tested on the NIST RT datasets. A best case relative improvement of 30% is observed in terms of Speaker Error Rate (SER) on the NIST RT 09 dataset when compared with the IB approach.
Srikanth R. Madikeri, Hervé Bourlard
ICASSP2
2015 Combining SGMM speaker vectors and KL-HMM approach for speaker diarization
abstract
In this paper, a method to use SGMM speaker vectors for speaker diarization is introduced. The architecture of the Information Bottleneck (IB) based speaker diarization is utilized for this purpose. The audio for speaker diarization is split into short uniform segments. Speaker vectors are obtained from a Subspace Gaussian Mixture Model (SGMM) system trained on meeting data. The speaker vectors are clustered using the K-means algorithm. Two types of distance measures are explored in the clustering step: cosine distance of the speaker vectors and that of the vectors in a space projected by Probabilistic Linear Discriminant Analysis (PLDA). The clustering output is used as an initialization step for the Kullback Leibler-Hidden Markov Model (KL-HMM) based speech segmentation approach commonly used in the IB system for diarization. The proposed method is compared to clustering the segments using the IB based approach. A relative improvement of approximately 14% is obtained on the diarization performance for the proposed approach using SGMM speaker vectors with PLDA on the NIST RT 09 dataset.
Srikanth R. Madikeri, Petr Motlícek, Hervé Bourlard
ICASSP3
2015 Robust microphone placement for source localization from noisy distance measurements
abstract
We propose a novel algorithm to design an optimum array geometry for source localization inside an enclosure. We assume a square-law decay propagation model for the sound acquisition so that the additive noise on the measured source-microphone distances is proportional to the distances regardless of the noise distribution. We formulate the source localization as an instance of the “Generalized Trust Region Subproblem” (GTRS) whose solution gives the location of the source. We show that by suitable selection of the microphone locations, one can tremendously decrease the noise-sensitivity of the resulting solution. In particular, by minimizing the noise-sensitivity of the source location in terms of sensor positions, we find the optimal noise-robust array geometry for the enclosure. Simulation results are provided to show the efficiency of the proposed algorithm.
Mohammad Javad Taghizadeh, Saeid Haghighatshoar, Afsaneh Asaei, Philip N. Garner, Hervé Bourlard
ICASSP5
2015 Objective speech intelligibility assessment through comparison of phoneme class conditional probability sequences
abstract
Assessment of speech intelligibility is important for the development of speech systems, such as telephony systems and text-to-speech (TTS) systems. Existing approaches to the automatic assessment of intelligibility in telephony typically compare a reference speech signal to a degraded copy, which requires that both signals be from the same speaker. In this paper, we propose a novel approach that does not have such a requirement, making it possible to also evaluate TTS systems and recent very low bit rate codecs that may modify speaker characteristics. More specifically, our approach is based on comparing sequences of phoneme class conditional probabilities. We show the potential of our approach on low bit rate telephony conditions, and compare it against subjective TTS intelligibility scores from the 2011 Blizzard Challenge.
Raphael Ullmann, Mathew Magimai-Doss, Hervé Bourlard
ICASSP3
2015 Novel GCC-PHAT model in diffuse sound field for microphone array pairwise distance based calibration
abstract
We propose a novel formulation of the generalized cross correlation with phase transform (GCC-PHAT) for a pair of microphones in diffuse sound field. This formulation elucidates the links between the microphone distances and the GCC-PHAT output. Hence, it leads to a new model that enables estimation of the pairwise distances by optimizing over the distances best matching the GCC-PHAT observations. Furthermore, the relation of this model to the coherence function is elaborated along with the dependency on the signal bandwidth. The experiments conducted on real data recordings demonstrate the theories and support the effectiveness of the proposed method.
José F. Velasco, Mohammad Javad Taghizadeh, Afsaneh Asaei, Hervé Bourlard, Carlos Julian Martín-Arguedas, Javier Macías Guarasa, Daniel Pizarro-Perez
ICASSP4
2015 On compressibility of neural network phonological features for low bit rate speech coding
abstract
Phonological features extracted by neural network have shown interesting potential for low bit rate speech vocoding. The span of phonological features is wider than the span of phonetic features, and thus fewer frames need to be transmitted. Moreover, the binary nature of phonological features enables a higher compression ratio at minor quality cost. In this paper, we study the compressibility and structured sparsity of the phonological features. We propose a compressive sampling framework for speech coding and sparse reconstruction for decoding prior to synthesis. Compressive sampling is found to be a principled way for compression in contrast to the conventional pruning approach; it leads to $50$\\% reduction in the bit-rate for better or equal quality of the decoded speech. Furthermore, exploiting the structured sparsity and binary characteristic of these features have shown to enable very low bit-rate coding at 700 bps with negligible quality loss; this coding scheme imposes no latency. If we consider a latency of $256$~ms for supra-segmental structures, the rate of $250-350$~bps is achieved.
Afsaneh Asaei, Milos Cernak, Hervé Bourlard
INTERSPEECH3
2015 Sparse modeling of posterior exemplars for keyword detection
abstract
Sparse representation has been shown to be a powerful modeling framework for classification and detection tasks.In this paper, we propose a new keyword detection algorithm based on sparse representation of the posterior exemplars.The posterior exemplars are phone conditional probabilities obtained from a deep neural network.This method relies on the concept that a keyword exemplar lies in a low-dimensional subspace which can be represented as a sparse linear combination of the training exemplars.The training exemplars are used to learn a dictionary for sparse representation of the keywords and background classes.Given this dictionary, the sparse representation of a test exemplar is used to detect the keywords.The experimental results demonstrate the potential of the proposed sparse modeling approach and it compares favorably with the state-of-the-art HMM-based framework on Numbers'95 database.
Dhananjay Ram, Afsaneh Asaei, Pranay Dighe, Hervé Bourlard
INTERSPEECH4
2015 Objective intelligibility assessment of text-to-speech systems through utterance verification
abstract
Objective assessment of synthetic speech intelligibility can be a useful tool for the development of text-to-speech (TTS) systems, as it provides a reproducible and inexpensive alternative to subjective listening tests. In a recent work, it was shown that the intelligibility of synthetic speech could be assessed objectively by comparing two sequences of phoneme class conditional probabilities, corresponding to instances of synthetic and human reference speech, respectively. In this paper, we build on those findings to propose a novel approach that formulates objective intelligibility assessment as an utterance verification problem using hidden Markov models, thereby alleviating the need for human reference speech. Specifically, given each text input to the TTS system, the proposed approach automatically verifies the words in the output synthetic speech signal and estimates an intelligibility score based on word recall statistics. We evaluate the proposed approach on the 2011 Blizzard Challenge data, and show that the estimated scores and the subjective intelligibility scores are highly correlated (Pearson’s |R| = 0.94).
Raphael Ullmann, Ramya Rasipuram, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH4
2015 Ad hoc microphone array calibration: Euclidean distance matrix completion algorithm and theoretical guarantees
Mohammad Javad Taghizadeh, Reza Parhizkar, Philip N. Garner, Hervé Bourlard, Afsaneh Asaei
Signal Process.4
2015 Automatic Recognition of Emergent Social Roles in Small Group Interactions
abstract
This paper investigates the automatic recognition of social roles that emerge naturally in small groups. These roles represent a flexible classification scheme that can generalize across different scenarios of small group interaction. We systematically investigate various verbal and non-verbal cues extracted from turn-taking patterns, vocal expression, and linguistic style to model speakers behavior. The influence of social roles on the behavior cues exhibited by a speaker is modeled using a discriminative approach based on conditional random fields. Experiments performed on several hours of meeting data reveal that social role recognition using conditional random fields achieves an accuracy of 74% in classifying four social roles and outperforms the baseline method on all social role categories . Furthermore , we also demonstrate the effectiveness of our approach by evaluating it on previously unseen scenarios of small group interactions.
Ashtosh Sapru, Hervé Bourlard
IEEE Trans. Multim.2
2014 Model-based sparse component analysis for reverberant speech localization
abstract
In this paper, the problem of multiple speaker localization via speech separation based on model-based sparse recovery is studies. We compare and contrast computational sparse optimization methods incorporating harmonicity and block structures as well as autoregressive dependencies underlying spectrographic representation of speech signals. The results demonstrate the effectiveness of block sparse Bayesian learning framework incorporating autoregressive correlations to achieve a highly accurate localization performance. Furthermore, significant improvement is obtained using ad-hoc microphones for data acquisition set-up compared to the compact microphone array.
Afsaneh Asaei, Hervé Bourlard, Mohammad Javad Taghizadeh, Volkan Cevher
ICASSP2
2014 Exploiting un-transcribed foreign data for speech recognition in well-resourced languages
abstract
Manual transcription of audio databases for automatic speech recognition (ASR) training is a costly and time-consuming process. State-of-the-art hybrid ASR systems that are based on deep neural networks (DNN) can exploit un-transcribed foreign data during unsupervised DNN pre-training or semi-supervised DNN training. We investigate the relevance of foreign data characteristics, in particular domain and language. Using three different datasets of the MediaParl and Ester databases, our experiments suggest that domain and language are equally important. Foreign data recorded under matched conditions (language and domain) yields the most improvement. The resulting ASR system yields about 5% relative improvement compared to the baseline system only trained on transcribed data. Our studies also reveal that the amount of foreign data used for semi-supervised training can be significantly reduced without degrading the ASR performance if confidence measure based data selection is employed.
David Imseng, Blaise Potard, Petr Motlícek, Alexandre Nanchen, Hervé Bourlard
ICASSP5
2014 Filterbank slope based features for speaker diarization
abstract
In this paper, filterbank slope based features are applied to the Information Bottleneck based system for speaker diarization. The filterbank slope based features have shown promise in the context of speaker recognition systems owing to their ability to emphasize formants. Hence, it is proposed to study their use in the context of speaker diarization as well, where speaker discrimination is equally important. The feature is explored using two different filterbank arrangements, linear and Mel, to form the Linear Filterbank Slope (LFS) and Mel Filterbank Slope (MFS), respectively. Both arrangements are shown to be inherently better at speaker discrimination compared with MFCC (Mel Frequency Cepstral Co-efficients). The feature streams are tested on the NIST RT06, 07 and 09 datasets. A best case relative improvement of 22.1% and 37.1% is observed for LFS and MFS, respectively, when compared with the MFCC-based baseline. The combination with time domain features is also studied and further improvements are observed. Finally, results on the fusion of multiple features are presented.
Srikanth R. Madikeri, Hervé Bourlard
ICASSP2
2014 Improving speaker diarization using social role information
abstract
Speaker diarization systems for meetings commonly model acoustic and spatial information, ignoring that meetings are instances of human interactions. Recent studies have shown that social roles influence the interaction patterns of speakers. This paper proposes a novel method to integrate social roles information in the speaker diarization framework. First, we modify the minimum duration constraint in baseline diarization system by using role information to model the expected duration of speaker's turn. Furthermore, we also propose a social role n-gram model as prior information on speaker interaction patterns. The proposed method is integrated in the state-of-the-art diarization system to reduce the speaker error. Experiments are performed on AMI corpus which is annotated in terms of social roles. The proposed method reduces the speaker error by 16% relative to baseline HMM-GMM system. Furthermore, the paper also investigates the performance of the proposed method on other meeting scenarios like those from NIST Rich Transcription campaigns. Experiments on Rich Transcription meetings reveal that speaker error can be reduced by 13% relative to the baseline system, thus demonstrating the potential of the proposed method.
Ashtosh Sapru, Sree Harsha Yella, Hervé Bourlard
ICASSP3
2014 Multilingual deep neural network based acoustic modeling for rapid language adaptation
abstract
This paper presents a study on multilingual deep neural network (DNN) based acoustic modeling and its application to new languages. We investigate the effect of phone merging on multilingual DNN in context of rapid language adaptation. Moreover, the combination of multilingual DNNs with Kullback-Leibler divergence based acoustic modeling (KL-HMM) is explored. Using ten different languages from the Globalphone database, our studies reveal that crosslingual acoustic model transfer through multilingual DNNs is superior to unsupervised RBM pre-training and greedy layer-wise supervised training. We also found that KL-HMM based decoding consistently outperforms conventional hybrid decoding, especially in low-resource scenarios. Furthermore, the experiments indicate that multilingual DNN training equally benefits from simple phoneset concatenation and manually derived universal phonesets.
Ngoc Thang Vu, David Imseng, Daniel Povey, Petr Motlícek, Tanja Schultz, Hervé Bourlard
ICASSP6
2014 Information bottleneck based speaker diarization of meetings using non-speech as side information
abstract
Background noise and errors in speech/non-speech detection cause significant degradation to the output of a speaker diarization system. In a typical speaker diarization system, non-speech segments are excluded prior to unsupervised clustering. In the current study, we exploit the information present in the non-speech segments of a recording to improve the output of the speaker diarization system based on information bottleneck framework. This is achieved by providing information from non-speech segments as side (irrelevant) information to information bottleneck based clustering. Experiments on meeting recordings from RT 06, 07, 09, evaluation sets have shown that the proposed method decreases the diarization error rate by around 18% relative to the baseline speaker diarization system based on information bottleneck framework. Comparison with a state of the art system based on HMM/GMM framework shows that the proposed method significantly decreases the gap in performance between the information bottleneck system and HMM/GMM system.
Sree Harsha Yella, Hervé Bourlard
ICASSP2
2014 Posterior-based sparse representation for automatic speech recognition
abstract
Posterior features have been shown to yield very good performance in multiple contexts including speech recognition, spoken term detection, and template matching. These days, posterior features are usually estimated at the output of a neural network. More recently, sparse representation has also been shown to potentially provide additional advantages to improve discrimination and robustness. One possible instance of this, is referred to as exemplar-based sparse representation. The present work investigates how to exploit sparse modelling together with posterior space properties to further improve speech recognition features. In that context, we leverage exemplar-based sparse representation, and propose a novel approach to project phone posterior features into a new, high-dimensional, sparse feature space. In fact, exploiting the properties of posterior spaces, we generate, new, high-dimensional, linguistically inspired (sub-phone and words), posterior distributions. Validation experiments are performed on the Phonebook (isolated words) and HIWIRE (continuous speech) databases, which support the effectiveness of the proposed approach for speech recognition tasks.
Sara Bahaadini, Afsaneh Asaei, David Imseng, Hervé Bourlard
INTERSPEECH4
2014 Detecting and labeling speakers on overlapping speech using vector taylor series
abstract
Successfully modeling overlapping speech is a crucial step towards improving the performance of current speaker diarization systems. In this direction, we present ongoing work on a novel Multi-Class Vector Taylor Series (MC-VTS) approach that models overlapping speech from knowledge of the individual speaker models and the feature extraction process. We explore several variants of the MC-VTS technique that aim at modeling overlapping speech more precisely. Bootstrapping the algorithm with both oracle and diarization output segmentations, we show the potential of this approach in terms of overlapping speech detection and speaker labeling performances through a set of experiments on far-field microphone meeting data.
Pranay Dighe, Marc Ferras, Hervé Bourlard
INTERSPEECH3
2014 Multi-source posteriors for speech activity detection on public talks
Marc Ferras, Hervé Bourlard
INTERSPEECH2
2014 Diarizing large corpora using multi-modal speaker linking
abstract
602
Marc Ferras, Stefano Masneri, Oliver Schreer, Hervé Bourlard
INTERSPEECH4
2014 Detecting speaker roles and topic changes in multiparty conversations using latent topic models
abstract
Accessing and browsing archives of multiparty conversations can be significantly facilitated by labeling them in terms of high level information. In this paper, we investigate automatic labeling of speaker roles and topic changes in professional meetings. Using the framework of unsupervised topic modeling we express speaker utterances as mixture of latent variables, each of which is governed by a multinomial distribution. The generated latent topic distributions are then used as features for predicting role and topic changes. Experiments performed on several hours of meeting data selected from AMI corpus reveal that latent topic features are effective predictors of speaker roles and topic changes. Furthermore, experiments also reveal an improvement in performance when latent topic information is combined with other multistream features. Index Terms: speaker role labeling, latent topic models, topic boundary detection
Ashtosh Sapru, Hervé Bourlard
INTERSPEECH2
2014 Phoneme background model for information bottleneck based speaker diarization
abstract
Acoustic variability of speakers arises due to differences in their vocal tract characteristics.These individual speaker characteristics are reflected in a speech signal when speakers pronounce a given phoneme.The current work hypothesizes that clusters within a phoneme spoken by multiple speakers roughly correspond to different speakers.Based on this hypothesis, a Gaussian mixture model (GMM) based phoneme background model (PBM) is estimated.The components of such a PBM are used as a set of relevance variables in information bottleneck based speaker diarization system.Experiments are done using phone transcripts obtained from ground-truth and automatic speech recognition (ASR) system to estimate the PBM.The diarization experiments done on meeting recordings from AMI and NIST-RT corpora show that the proposed method achieves significant improvements over the system using a background model which ignores phoneme information.
Sree Harsha Yella, Petr Motlícek, Hervé Bourlard
INTERSPEECH3
2014 Enhanced diffuse field model for ad hoc microphone array calibration
Mohammad Javad Taghizadeh, Philip N. Garner, Hervé Bourlard
Signal Process.3
2014 Using out-of-language data to improve an under-resourced speech recognizer
David Imseng, Petr Motlícek, Hervé Bourlard, Philip N. Garner
Speech Commun.3
2014 Structured Sparsity Models for Reverberant Speech Separation
abstract
We tackle the speech separation problem through modeling the acoustics of the reverberant chambers. Our approach exploits structured sparsity models to perform speech recovery and room acoustic modeling from recordings of concurrent unknown sources. The speakers are assumed to lie on a two-dimensional plane and the multipath channel is characterized using the image model. We propose an algorithm for room geometry estimation relying on localization of the early images of the speakers by sparse approximation of the spatial spectrum of the virtual sources in a free-space model. The images are then clustered exploiting the low-rank structure of the spectro-temporal components belonging to each source. This enables us to identify the early support of the room impulse response function and its unique map to the room geometry. To further tackle the ambiguity of the reflection ratios, we propose a novel formulation of the reverberation model and estimate the absorption coefficients through a convex optimization exploiting joint sparsity model formulated upon spatio-spectral sparsity of concurrent speech representation. The acoustic parameters are then incorporated for separating individual speech signals through either structured sparse recovery or inverse filtering the acoustic channels. The experiments conducted on real data recordings of spatially stationary sources demonstrate the effectiveness of the proposed approach for speech separation and recognition.
Afsaneh Asaei, Mohammad Golbabaee, Hervé Bourlard, Volkan Cevher
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Feature mapping of multiple beamformed sources for robust overlapping speech recognition using a microphone array
abstract
This paper introduces a nonlinear vector-based feature mapping approach to extract robust features for automatic speech recognition (ASR) of overlapping speech using a microphone array. We explore different configurations and additional sources of information to improve the effectiveness of the feature mapping. First, we investigate the full-vector based mapping of different sources in a log mel-filterbank energy (log MFBE) domain, and demonstrate that retraining the acoustic model using the generated training data can help improve the recognition performance. Then we investigate the feature mapping between different domains. Finally in order to improve the qualities of the mapping inputs we propose a nonlinear mapping of the features from multiple beamformed sources, which are directed at the target and interfering speakers, respectively. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, Longbiao Wang, Yicong Zhou, John Dines, Mathew Magimai-Doss, Hervé Bourlard, Qingmin Liao
IEEE ACM Trans. Audio Speech Lang. Process.6
2014 Overlapping speech detection using long-term conversational features for speaker diarization in meeting room conversations
abstract
Overlapping speech has been identified as one of the main sources of errors in diarization of meeting room conversations. Therefore, overlap detection has become an important step prior to speaker diarization. Studies on conversational analysis have shown that overlapping speech is more likely to occur at specific parts of a conversation. They have also shown that overlap occurrence is correlated with various conversational features such as speech, silence patterns and speaker turn changes. We use features capturing this higher level information from structure of a conversation such as silence and speaker change statistics to improve acoustic feature based classifier of overlapping and single-speaker speech classes. The silence and speaker change statistics are computed over a long-term window (around 3-4 seconds) and are used to predict the probability of overlap in the window. These estimates are then incorporated into a acoustic feature based classifier as prior probabilities of the classes. Experiments conducted on three corpora (AMI, NIST-RT and ICSI) have shown that the proposed method improves the performance of acoustic feature-based overlap detector on all the corpora. They also reveal that the model based on long-term conversational features used to estimate probability of overlap which is learned from AMI corpus generalizes to meetings from other corpora (NIST-RT and ICSI). Moreover, experiments on ICSI corpus reveal that the proposed method also improves laughter overlap detection. Consequently, applying overlap handling techniques to speaker diarization using the detected overlap results in reduction of diarization error rate (DER) on all the three corpora.
Sree Harsha Yella, Hervé Bourlard
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Investigating the Impact of Language Style and Vocal Expression on Social Roles of Participants in Professional Meetings
abstract
This paper investigates the influence of social roles on the language style and vocal expression patterns of participants in professional meeting recordings. Language style features are extracted from automatically generated speech transcripts and characterize word usage in terms of psychologically meaningful categories. Vocal expression patterns are generated by applying statistical functionals to low level prosodic and spectral features. The proposed recognition system combines information from both these feature streams to predict participant's social role. Experiments conducted on almost 12.5 hours of meeting data reveal that recognition system trained using language style features and acoustic features can reach a recognition accuracy of 64% and 68% respectively, in classifying four social roles. Moreover, recognition accuracy increases to 69% when information from both feature streams is taken into consideration.
Ashtosh Sapru, Hervé Bourlard
ACII2
2013 Impact of deep MLP architecture on different acoustic modeling techniques for under-resourced speech recognition
abstract
Posterior based acoustic modeling techniques such as Kullback-Leibler divergence based HMM (KL-HMM) and Tandem are able to exploit out-of-language data through posterior features, estimated by a Multi-Layer Perceptron (MLP). In this paper, we investigate the performance of posterior based approaches in the context of under-resourced speech recognition when a standard three-layer MLP is replaced by a deeper five-layer MLP. The deeper MLP architecture yields similar gains of about 15% (relative) for Tandem, KL-HMM as well as for a hybrid HMM/MLP system that directly uses the posterior estimates as emission probabilities. The best performing system, a bilingual KL-HMM based on a deep MLP, jointly trained on Afrikaans and Dutch data, performs 13% better than a hybrid system using the same bilingual MLP and 26% better than a subspace Gaussian mixture system only trained on Afrikaans data.
David Imseng, Petr Motlícek, Philip N. Garner, Hervé Bourlard
ASRU4
2013 MLP-based factor analysis for tandem speech recognition
abstract
In the last years, latent variable models such as factor analysis, probabilistic principal component analysis or subspace Gaussian mixture models have become almost ubiquitous in speech technologies. The key to its success is the joint modeling of multiple effects in the speech signal they address. In this paper, we propose a novel approach to use phone and speaker variabilities together to estimate phone posterior probabilities on a tandem speech recognition system. A Multilayer Perceptron (MLP) with 5 layers and a central bottleneck linear layer is used as a basic processing block that mimics the processing undergone in factor analysis. With multiple factors, phone and a speaker MLP are merged at the bottleneck level to obtain better estimates for the phone posterior probabilities used in the ASR system. Experiments on the WSJ corpus show that the joint phone-speaker modeling can significantly outperform phone modeling alone in terms of Frame Error and Word Error Rates.
Marc Ferras, Hervé Bourlard
ICASSP2
2013 Speaker adaptive Kullback-Leibler divergence based hidden Markov models
abstract
Kullback-Leibler divergence based hidden Markov models (KL-HMM) have recently been introduced as an efficient and principled way to directly model sequences of posterior vectors to perform Automatic Speech Recognition (ASR). Through efficient feature level adaptation and parsimonious use of parameters, KL-HMM was successfully applied to accented and under-resourced speech recognition tasks. In this paper, inspired from Maximum A Posteriori (MAP) adaptation, we further boost KL-HMM performance by applying Bayesian speaker adaptation, directly applied to posterior features. This approach performs a simple, adaptive regression between phone posteriors estimated with a Multilayer Perceptron (MLP) on large amounts of speaker-independent training data, and speaker-specific phone posteriors generated by the speaker-independent MLP on very limited amount of speaker-specific adaptation data. Using Swiss French data (MediaParl), we show that such speaker adaptive KL-HMM can significantly outperform conventional adaptation techniques on non-native speech while yielding similar performance on native data.
David Imseng, Hervé Bourlard
ICASSP2
2013 Improved overlap speech diarization of meeting recordings using long-term conversational features
abstract
Overlapping speech is a source of significant errors in speaker diarization of spontaneous meeting recordings. Recent works on speaker diarization have attempted to solve the problem of overlap detection using classifiers trained on acoustic and spatial features. This paper proposes a method to improve the short-term spectral feature based overlap detector by incorporating information from long-term conversational features in the form of speaker change statistics. The statistics are obtained at segment level(around few seconds) from the output of a diarization system. The approach is motivated by the observation that segments containing more speaker changes are more probable to have more overlaps. Experiments on AMI meeting corpus reveal that the number of overlaps in a segment follows a Poisson distribution whose rate is directly proportional to the number of speaker changes in the segment. When this information is combined with acoustic information in an HMM/GMM overlap detector, improvements are verified in terms of F-measure and consequently, diarization error (DER) is reduced by 5% relative to the baseline overlap detector.
Sree Harsha Yella, Hervé Bourlard
ICASSP2
2013 Automatic social role recognition in professional meetings using conditional random fields
abstract
Social roles characterize relation between participants in a conversation and, in turn, influence their interaction patterns.This paper investigates automatic social role recognition in professional meetings using a completely discriminative framework based on conditional random fields.We present a novel approach which combines information from multiple layers of data.The conversation layer models the influence of social roles on turn taking patterns of participants present in multiparty interactions.A conditional random field augmented with hidden state sequences is used to estimate the posterior distribution of social roles in this layer.The other novelty of our approach consists in modeling statistical dependencies between roles across adjacent segments of meeting.The posterior distribution estimated in conversation layer is combined with role transition information to improve the model.Experiments conducted on more than 40 hours of data reveal that the proposed approach reaches a recognition accuracy of 67% in classifying four social roles using information from conversation layer.Moreover, recognition accuracy increases to 70% when information from multiple layers is taken into consideration.
Ashtosh Sapru, Hervé Bourlard
INTERSPEECH2
2013 Applying Multi- and Cross-Lingual Stochastic Phone Space Transformations to Non-Native Speech Recognition
abstract
In the context of hybrid HMM/MLP Automatic Speech Recognition (ASR), this paper describes an investigation into a new type of stochastic phone space transformation, which maps “source” phone (or phone HMM state) posterior probabilities (as obtained at the output of a Multilayer Perceptron/MLP) into “destination” phone (HMM phone state) posterior probabilities. The resulting stochastic matrix transformation can be used within the same language to automatically adapt to different phone formats (e.g., IPA) or across languages. Additionally, as shown here, it can also be applied successfully to non-native speech recognition. In the same spirit as MLLR adaptation, or MLP adaptation, the approach proposed here is directly mapping posterior distributions, and is trained by optimizing on a small amount of adaptation data a Kullback-Leibler based cost function, along a modified version of an iterative EM algorithm. On a non-native English database (HIWIRE), and comparing with multiple setups (monophone and triphone mapping, MLLR adaptation) we show that the resulting posterior mapping yields state-of-the-art results using very limited amounts of adaptation data in mono-, cross- and multi-lingual setups. We also show that “universal” phone posteriors, trained on a large amount of multilingual data, can be transformed to English phone posteriors, resulting in an ASR system that significantly outperforms a system trained on English data only. Finally, we demonstrate that the proposed approach outperforms alternative data-driven, as well as a knowledge-based, mapping techniques.
David Imseng, Hervé Bourlard, John Dines, Philip N. Garner, Mathew Magimai-Doss
IEEE Trans. Speech Audio Process.2
2013 Robust Log-Energy Estimation and its Dynamic Change Enhancement for In-car Speech Recognition
abstract
The log-energy parameter, typically derived from a full-band spectrum, is a critical feature commonly used in automatic speech recognition (ASR) systems. However, log-energy is difficult to estimate reliably in the presence of background noise. In this paper, we theoretically show that background noise affects the trajectories of not only the “conventional” log-energy, but also its delta parameters. This results in a poor estimation of the actual log-energy and its delta parameters, which no longer describe the speech signal. We thus propose a new method to estimate log-energy from a sub-band spectrum, followed by dynamic change enhancement and mean smoothing. We demonstrate the effectiveness of the proposed log-energy estimation and its post-processing steps through speech recognition experiments conducted on the in-car CENSREC-2 database. The proposed log-energy (together with its corresponding delta parameters) yields an average improvement of 32.8% compared with the baseline front-ends. Moreover, it is also shown that further improvement can be achieved by incorporating the new Mel-Frequency Cepstral Coefficients (MFCCs) obtained by non-linear spectral contrast stretching.
Weifeng Li 0001, Longbiao Wang, Yicong Zhou, Hervé Bourlard, Qingmin Liao
IEEE Trans. Speech Audio Process.4
2013 Wordless Sounds: Robust Speaker Diarization Using Privacy-Preserving Audio Representations
abstract
This paper investigates robust privacy-sensitive audio features for speaker diarization in multiparty conversations: i.e., a set of audio features having low linguistic information for speaker diarization in a single and multiple distant microphone scenarios. We systematically investigate Linear Prediction (LP) residual. Issues such as prediction order and choice of representation of LP residual are studied. Additionally, we explore the combination of LP residual with subband information from 2.5 kHz to 3.5 kHz and spectral slope. Next, we propose a supervised framework using deep neural architecture for deriving privacy-sensitive audio features. We benchmark these approaches against the traditional Mel Frequency Cepstral Coefficients (MFCC) features for speaker diarization in both the microphone scenarios. Experiments on the RT07 evaluation dataset show that the proposed approaches yield diarization performance close to the MFCC features on the single distant microphone dataset. To objectively evaluate the notion of privacy in terms of linguistic information, we perform human and automatic speech recognition tests, showing that the proposed approaches to privacy-sensitive audio features yield much lower recognition accuracies compared to MFCC features.
Sree Hari Krishnan Parthasarathi, Hervé Bourlard, Daniel Gatica-Perez
IEEE Trans. Speech Audio Process.2
2012 Computational methods for structured sparse component analysis of convolutive speech mixtures
abstract
We cast the under-determined convolutive speech separation as sparse approximation of the spatial spectra of the mixing sources. In this framework we compare and contrast the major practical algorithms for structured sparse recovery of speech signal. Specific attention is paid to characterization of the measurement matrix. We first propose how it can be identified using the Image model of multipath effect where the acoustic parameters are estimated by localizing a speaker and its images in a free space model. We further study the circumstances in which the coherence of the projections induced by microphone array design tend to affect the recovery performance.
Afsaneh Asaei, Mike E. Davies 0001, Hervé Bourlard, Volkan Cevher
ICASSP3
2012 Using KL-divergence and multilingual information to improve ASR for under-resourced languages
abstract
Setting out from the point of view that automatic speech recognition (ASR) ought to benefit from data in languages other than the target language, we propose a novel Kullback-Leibler (KL) divergence based method that is able to exploit multilingual information in the form of universal phoneme posterior probabilities conditioned on the acoustics. We formulate a means to train a recognizer on several different languages, and subsequently recognize speech in a target language for which only a small amount of data is available. Taking the Greek SpeechDat(II) data as an example, we show that the proposed formulation is sound, and show that it is able to out-perform a current state-of-the-art HMM/GMM system. We also use a hybrid Tandem-like system to further understand the source of the benefit.
David Imseng, Hervé Bourlard, Philip N. Garner
ICASSP2
2012 Robust triphone mapping for acoustic modeling
abstract
In this paper we revisit the recently proposed triphone mapping as an alternative to decision tree state clustering. We generalize triphone mapping to Kullback-Leibler based hidden Markov models for acoustic modeling and propose a modified training procedure for the Gaussian mixture model based acoustic modeling. We compare the triphone mapping to decision tree state clustering on the Wall Street Journal task as well as in the context of an under-resourced language by using Greek data from the SpeechDat(II) corpus. Experiments reveal that triphone mapping has the best overall performance and is robust against varying the acoustic modeling technique as well as variable amounts of training data.
Milos Cernak, David Imseng, Hervé Bourlard
INTERSPEECH3
2012 Comparing different acoustic modeling techniques for multilingual boosting
abstract
In this paper, we explore how different acoustic modeling techniques can benefit from data in languages other than the target language. We propose an algorithm to perform decision tree state clustering for the recently proposed Kullback-Leibler divergence based hidden Markov models (KL-HMM) and compare it to subspace Gaussian mixture modeling (SGMM). KL-HMM can exploit multilingual information in the form of universal phoneme posterior features and SGMM benefits from a universal background model that can be trained on multilingual data. Taking the Greek SpeechDat(II) data as an example, we show that KL-HMM performs best for small amounts of target language data.
David Imseng, John Dines, Petr Motlícek, Philip N. Garner, Hervé Bourlard
INTERSPEECH5
2012 Sub-band based Log-energy and Its Dynamic Range Stretching for Robust In-car Speech Recognition
abstract
Log energy and its delta parameters, typically derived from full-band spectrum, are commonly used in automatic speech recognition (ASR) systems.In this paper, we address the problem of estimating log energy in the presence of background noise (usually resulting in a reduction in dynamic ranges of spectral energies).We theoretically show that the background noise affects the trajectories of the "conventional" log energy and its delta parameters, resulting in very poor estimation of the actual log energy and its delta parameters, which no longer describe the speech signal.We thus propose to estimate log energy from the sub-band spectrum, followed by a dynamic range stretching.Based on speech recognition experiments conducted on CENSREC-2 in-car database, the proposed log energy (and its corresponding delta parameters) is shown to perform very well, resulting in an average relative improvement of 27.2% compared with the baseline front-ends.Moreover, it is also shown that further improvement can be achieved by incorporating those new MFCCs obtained through non-linear spectral contrast stretching.
Weifeng Li 0001, Hervé Bourlard
INTERSPEECH2
2012 Synthetic References for Template-based ASR using posterior features
abstract
Recently, the use of phoneme class-conditional probabilities as features (posterior features) for template-based ASR has been proposed. These features have been found to generalize well to unseen data and yield better systems than standard spectral-based features. In this paper, motivated by the high quality of current text-to-speech systems and the robustness of posterior features toward undesired variability, we investigate the use of synthetic speech to generate reference templates. The use of synthetic speech in template-based ASR not only allows to address the issue of in-domain data collection but also expansion of vocabulary. Using 75- and 600-word task-independent and speaker-independent setup on Phonebook database, we investigate different synthetic voices produced by the Festival HTS-based synthesizer trained on CMU ARCTIC databases. Our study shows that synthetic speech templates can yield performance comparable to the natural speech templates, especially with synthetic voices that have high intelligibility.
Serena Soldo, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2012 MediaParl: Bilingual mixed language accented speech database
abstract
MediaParl is a Swiss accented bilingual database containing recordings in both French and German as they are spoken in Switzerland. The data were recorded at the Valais Parliament. Valais is a bilingual Swiss canton with many local accents and dialects. Therefore, the database contains data with high variability and is suitable to study multilingual, accented and non-native speech recognition as well as language identification and language switch detection. We also define monolingual and mixed language automatic speech recognition and language identification tasks and evaluate baseline systems. The database is publicly available for download.
David Imseng, Hervé Bourlard, Holger Caesar, Philip N. Garner, Gwénolé Lecorvé, Alexandre Nanchen
SLT2
2012 Multistream speaker diarization of meetings recordings beyond MFCC and TDOA features
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
Speech Commun.3
2011 Model-based compressive sensing for multi-party distant speech recognition
abstract
We leverage the recent algorithmic advances in compressive sensing, and propose a novel source separation algorithm for efficient recovery of convolutive speech mixtures in spectro-temporal domain. Compared to the common sparse component analysis techniques, our approach fully exploits structured sparsity models to obtain substantial improvement over the existing state-of-the-art. We evaluate our method for separation and recognition of a target speaker in a multi-party scenario. Our results provide compelling evidence of the effectiveness of sparse recovery formulations in speech recognition.
Afsaneh Asaei, Hervé Bourlard, Volkan Cevher
ICASSP2
2011 Language dependent universal phoneme posterior estimation for mixed language speech recognition
abstract
This paper presents a new approach to estimate "universal" phoneme posterior probabilities for mixed language speech recognition. More specifically, we propose a new theoretical framework to combine phoneme class posterior probabilities in a principled way by using (statistical) evidence about the language identity. We investigate the proposed approach in a mixed language environment (Speech-Dat(II)) consisting of five European languages. Our studies show that the proposed approach can yield significant improvements on a mixed language task, while maintaining the performance on monolingual tasks. Additionally, through a case study, we also demonstrate the potential benefits of the proposed approach for non-native speech recognition.
David Imseng, Hervé Bourlard, Mathew Magimai-Doss, John Dines
ICASSP2
2011 Posterior features for template-based ASR
abstract
This paper investigates the use of phoneme class conditional probabilities as features (posterior features) for template-based ASR. Using 75 words and 600 words task-independent and speaker-independent setup on Phonebook database, we investigate the use of different posterior distribution estimators, different distance measures that are better suited for posterior distributions, and different training data. The reported experiments clearly demonstrate that posterior features are always superior, and generalize better than other classical acoustic features (at the cost of training a posterior distribution estimator).
Serena Soldo, Mathew Magimai-Doss, Joel Pinto, Hervé Bourlard
ICASSP4
2011 Just-in-time multimodal association and fusion from home entertainment
abstract
In this paper, we describe a real-time multimodal analysis system with just-in-time multimodal association and fusion for a living room environment, where multiple people may enter, interact and leave the observable world with no constraints. It comprises detection and tracking of up to 4 faces, detection and localisation of verbal and paralinguistic events, their association and fusion. The system is designed to be used in open, unconstrained environments like in next generation video conferencing systems that automatically "orchestrate" the transmitted video streams to improve the overall experience of interaction between spatially separated families and friends. Performance levels achieved to date on hand-labelled dataset have shown sufficient reliability at the same time as fulfilling real-time processing requirements.
Danil Korchagin, Petr Motlícek, Stefan Duffner, Hervé Bourlard
ICME4
2011 Multi-Party Speech Recovery Exploiting Structured Sparsity Models
abstract
We study the sparsity of spectro-temporal representation of speech in reverberant acoustic conditions. This study motivates the use of structured sparsity models for efficient speech recov-ery. We formulate the underdetermined convolutive speech sep-aration in spectro-temporal domain as the sparse signal recovery where we leverage model-based recovery algorithms. To tackle the ambiguity of the real acoustics, we exploit the Image Model of the enclosures to estimate the room impulse response func-tion through a structured sparsity constraint optimization. The experiments conducted on real data recordings demonstrate the effectiveness of the proposed approach for multi-party speech applications. Index Terms: speech sparsity, structured sparsity models, un-derdetermined convolutive speech separation, Image Model
Afsaneh Asaei, Mohammad Javad Taghizadeh, Hervé Bourlard, Volkan Cevher
INTERSPEECH3
2011 Improving Non-Native ASR Through Stochastic Multilingual Phoneme Space Transformations
abstract
We propose a stochastic phoneme space transformation technique that allows the conversion of conditional source phoneme posterior probabilities (conditioned on the acoustics) into target phoneme posterior probabilities. The source and target phonemes can be in any language and phoneme format such as the International Phonetic Alphabet. The novel technique makes use of a Kullback-Leibler divergence based hidden Markov model and can be applied to non-native and accented speech recognition or used to adapt systems to under-resourced languages. In this paper, and in the context of hybrid HMM/MLP recognizers, we successfully apply the proposed approach to non-native English speech recognition on the HIWIRE dataset.
David Imseng, Hervé Bourlard, John Dines, Philip N. Garner, Mathew Magimai-Doss
INTERSPEECH2
2011 Grapheme-Based Automatic Speech Recognition Using KL-HMM
abstract
The state-of-the-art automatic speech recognition (ASR) systems typically use phonemes as subword units. In this work, we present a novel grapheme-based ASR system that jointly models phoneme and grapheme information using Kullback-Leibler divergence-based HMM system (KL-HMM). More specifically, the underlying subword unit models are grapheme units and the phonetic information is captured through phoneme posterior probabilities (referred as posterior features) estimated using a multilayer perceptron (MLP). We investigate the proposed approach for ASR on English language, where the correspondence between phoneme and grapheme is weak. In particular, we investigate the effect of contextual modeling on grapheme-based KL-HMM system and the use of MLP trained on auxiliary data. Experiments on DARPA Resource Management corpus have shown that the grapheme-based ASR system modeling longer subword unit context can achieve same performance as phoneme-based ASR system, irrespective of the data on which MLP is trained.
Mathew Magimai-Doss, Ramya Rasipuram, Guillermo Aradilla, Hervé Bourlard
INTERSPEECH4
2011 LP Residual Features for Robust, Privacy-Sensitive Speaker Diarization
abstract
We present a comprehensive study of linear prediction residual for speaker diarization on single and multiple distant microphone conditions in privacy-sensitive settings, a requirement to analyze a wide range of spontaneous conversations. Two representations of the residual are compared, namely real-cepstrum and MFCC, with the latter performing better. Experiments on RT06eval show that residual with subband information from 2.5 kHz to 3.5 kHz and spectral slope yields a performance close to traditional MFCC features. As a way to objectively evaluate privacy in terms of linguistic information, we perform phoneme recognition. Residual features yield low phoneme accuracies compared to traditional MFCC features.
Sree Hari Krishnan Parthasarathi, Hervé Bourlard, Daniel Gatica-Perez
INTERSPEECH2
2011 Hierarchical Tandem Features for ASR in Mandarin
abstract
We apply multilayer perceptron (MLP) based hierarchical Tandem features to large vocabulary continuous speech recognition in Mandarin.Hierarchical Tandem features are estimated using a cascade of two MLP classifiers which are trained independently.The first classifier is trained on perceptual linear predictive coefficients with a 90 ms temporal context.The second classifier is trained using the phonetic class conditional probabilities estimated by the first MLP, but with a relatively longer temporal context of about 150 ms.Experiments on the Mandarin DARPA GALE eval06 dataset show significant reduction (about 7.6% relative) in character error rates by using hierarchical Tandem features over conventional Tandem features.
Joel Pinto, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2011 Privacy-Sensitive Audio Features for Speech/Nonspeech Detection
abstract
The goal of this paper is to investigate features for speech/nonspeech detection (SND) having low linguistic information from the speech signal. Towards this, we present a comprehensive study of privacy-sensitive features for SND in multiparty conversations. Our study investigates three different approaches to privacy-sensitive features. These approaches are based on: 1) simple, instantaneous feature extraction methods; 2) excitation source information based methods; and 3) feature obfuscation methods such as local (within 130 ms) temporal averaging and randomization applied on excitation source information. To evaluate these approaches for SND, we use multiparty conversational meeting data of nearly 450 hours. On this dataset, we evaluate these features and benchmark them against standard spectral shape based features such as Mel frequency perceptual linear prediction (MFPLP). Fusion strategies combining excitation source with simple features show that comparable performance can be obtained in both close-talking and far-field microphone scenarios. As one way to objectively evaluate the notion of privacy, we conduct phoneme recognition studies on TIMIT. While excitation source features yield phoneme recognition accuracies in between the simple features and the MFPLP features, obfuscation methods applied on the excitation features yield low phoneme accuracies in conjunction with SND performance comparable to that of MFPLP features.
Sree Hari Krishnan Parthasarathi, Daniel Gatica-Perez, Hervé Bourlard, Mathew Magimai-Doss
IEEE ACM Trans. Audio Speech Lang. Process.3
2011 Analysis of MLP-Based Hierarchical Phoneme Posterior Probability Estimator
abstract
We analyze a simple hierarchical architecture consisting of two multilayer perceptron (MLP) classifiers in tandem to estimate the phonetic class conditional probabilities. In this hierarchical setup, the first MLP classifier is trained using standard acoustic features. The second MLP is trained using the posterior probabilities of phonemes estimated by the first, but with a long temporal context of around 150-230 ms. Through extensive phoneme recognition experiments, and the analysis of the trained second MLP using Volterra series, we show that 1) the hierarchical system yields higher phoneme recognition accuracies-an absolute improvement of 3.5% and 9.3% on TIMIT and CTS respectively-over the conventional single MLP-based system, 2) there exists useful information in the temporal trajectories of the posterior feature space, spanning around 230 ms of context, 3) the second MLP learns the phonetic temporal patterns in the posterior features, which include the phonetic confusions at the output of the first MLP as well as the phonotactics of the language as observed in the training data, and 4) the second MLP classifier requires fewer number of parameters and can be trained using lesser amount of training data.
Joel Pinto, Garimella S. V. S. Sivaram, Mathew Magimai-Doss, Hynek Hermansky, Hervé Bourlard
IEEE Trans. Speech Audio Process.5
2011 An Information Theoretic Combination of MFCC and TDOA Features for Speaker Diarization
abstract
This correspondence describes a novel system for speaker diarization of meetings recordings based on the combination of acoustic features (MFCC) and time delay of arrivals (TDOAS). The first part of the paper analyzes differences between MFCC and TDOA features which possess completely different statistical properties. When Gaussian mixture models are used, experiments reveal that the diarization system is sensitive to the different recording scenarios (i.e., meeting rooms with varying number of microphones). In the second part, a new multistream diarization system is proposed extending previous work on information theoretic diarization. Both speaker clustering and speaker realignment steps are discussed; in contrary to current systems, the proposed method avoids to perform the feature combination averaging log-likelihood scores. Experiments on meetings data reveal that the proposed approach outperforms the GMM-based system when the recording is done with varying number of microphones.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
IEEE Trans. Speech Audio Process.3
2010 Analysis of phone posterior feature space exploiting class-specific sparsity and MLP-based similarity measure
abstract
Class posterior distributions have recently been used quite successfully in Automatic Speech Recognition (ASR), either for frame or phone level classification or as acoustic features, which can be further exploited (usually after some “ad hoc” transformations) in different classifiers (e.g., in Gaussian Mixture based HMMs). In the present paper, we show preliminary results showing that it may be possible to perform speech recognition without explicit subword unit (phone) classification or likelihood estimation, simply answering the question whether two acoustic (posterior) vectors belong to the same subword unit class or not. In this paper, we first exhibit specific properties of the posterior acoustic space before showing how those properties can be exploited to reach very high performance in deciding (based on an appropriate, trained, distance metric, and hypothesis testing approaches) whether two posterior vectors belong to the same class or not. Performance as high as 90% correct decision rates are reported on the TIMIT database, before reporting kNN phone classification rates.
Afsaneh Asaei, Benjamin Picart, Hervé Bourlard
ICASSP3
2010 Using audio and visual cues for speaker diarisation initialisation
abstract
In this paper we present a novel approach to audio visual speaker diarisation (the task of estimating “who spoke when” using audio and visual cues) in a challenging meeting domain. Our approach is based on the initialisation of the agglomerative speaker clustering using psychology inspired visual features, including Visual Focus of Attention (VFoA) and motion intensities. This method, providing initial speaker clusters of high purity, achieved consistent improvements over the widely adopted linear initialisation method. Moreover, the initialisation using both visual and Time Delay of Arrival (TDoA) cues was also investigated in conjunction with the multi-stream combination of acoustic and visual features (MFCC, TDoA, VFoA, motion intensity, and head pose likelihoods). This speaker diarisation framework allowed to successfully integrate three feature streams, further exploiting the complementarity between multimodal cues.
Giulia Garau, Hervé Bourlard
ICASSP2
2010 Evaluating the robustness of privacy-sensitive audio features for speech detection in personal audio log scenarios
abstract
Personal audio logs are often recorded in multiple environments. This poses challenges for robust front-end processing, including speech/nonspeech detection (SND). Motivated by this, we investigate the robustness of four different privacy-sensitive features for SND, namely energy, zero crossing rate, spectral flatness, and kurtosis. We study early and late fusion of these features in conjunction with modeling temporal context. These combinations are evaluated in mismatched conditions on a dataset of nearly 450 hours. While both combinations yield improvements over individual features, generally feature combinations perform better. Comparisons with a state-of-the-art spectral based and a privacy-sensitive feature set are also provided.
Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Hervé Bourlard, Daniel Gatica-Perez
ICASSP3
2010 Multistream speaker diarization beyond two acoustic feature streams
abstract
Speaker diarization for meetings data are recently converging towards multistream systems. The most common complementary features used in combination with MFCC are Time Delay of Arrival (TDOA). Also other features have been proposed although, there are no reported improvements on top of MFCC+TDOA systems. In this work we investigate the combination of other feature sets along with MFCC+TDOA. We discuss issues and problems related to the weighting of four different streams proposing a solution based on a smoothed version of the speaker error. Experiments are presented on NIST RT06 meeting diarization evaluation. Results reveal that the combination of four acoustic feature streams results in a 30% relative improvement with respect to the MFCC+TDOA feature combination. To the authors' best knowledge, this is the first successful attempt to improve the MFCC+TDOA baseline including other feature streams.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
ICASSP3
2010 Sparse component analysis for speech recognition in multi-speaker environment
abstract
Sparse Component Analysis is a relatively young technique that relies upon a representation of signal occupying only a small part of a larger space. Mixtures of sparse components are disjoint in that space. As a particular application of sparsity of speech signals, we investigate the DUET blind source separation algorithm in the context of speech recognition for multi-party recordings. We show how DUET can be tuned to the particular case of speech recognition with interfering sources, and evaluate the limits of performance as the number of sources increases. We show that the separated speech fits a common metric for sparsity, and conclude that sparsity assumptions lead to good performance in speech separation and hence ought to benefit other aspects of the speech recognition chain.
Afsaneh Asaei, Hervé Bourlard, Philip N. Garner
INTERSPEECH2
2010 Floor holder detection and end of speaker turn prediction in meetings
abstract
We propose a novel fully automatic framework to detect which meeting participant is currently holding the conversational floor and when the current speaker turn is going to finish. Two sets of experiments were conducted on a large collection of multiparty conversations: the AMI meeting corpus. Unsupervised speaker turn detection was performed by post-processing the speaker diarization and the speech activity detection outputs. A supervised end-of-speaker-turn prediction framework, based on Dynamic Bayesian Networks and automatically extracted multimodal features (related to prosody, overlapping speech, and visual motion), was also investigated. These novel approaches resulted in good floor holder detection rates (13:2% Floor Error Rate), attaining state of the art end-of-speaker-turn prediction performances.
Alfred Dielmann, Giulia Garau, Hervé Bourlard
INTERSPEECH3
2010 Audio-visual synchronisation for speaker diarisation
abstract
The role of audio–visual speech synchrony for speaker diarisation is investigated on the multiparty meeting domain. We measured both mutual information and canonical correlation on different sets of audio and video features. As acoustic features we considered energy and MFCCs. As visual features we experimented both with motion intensity features, computed on the whole image, and Kanade Lucas Tomasi motion estimation. Thanks to KLT we decomposed the motion in its horizontal and vertical components. The vertical component was found to be more reliable for speech synchrony estimation. The mutual information between acoustic energy and KLT vertical motion of skin pixels, not only resulted in a 20 % relative improvement over a MFCC only diarisation system, but also outperformed visual features such as motion intensities and head poses. Index Terms: multimodal speaker diarisation, audio–visual speech synchrony, multiparty meetings, mutual information, canonical correlation analysis 1.
Giulia Garau, Alfred Dielmann, Hervé Bourlard
INTERSPEECH3
2010 Towards mixed language speech recognition systems
abstract
Multilingual speech recognition obviously involves numerous research challenges, including common phoneme sets, adaptation on limited amount of training data, as well as mixed language recognition (common in many countries, like Switzerland). In this latter case, it is not even possible to assume that one knows in advance the language being spoken. This is the context and motivation of the present work. We indeed investigate how current state-of-the-art speech recognition systems can be exploited in multilingual environments, where the language (from an assumed set of five possible languages, in our case) is not a priori known during recognition. We combine monolingual systems and extensively develop and compare different features and acoustic models. On Speech-Dat(II) datasets, and in the context of isolated words, we show that it is actually possible to approach the performances of monolingual systems even if the identity of the spoken language is not a priori known. Index Terms: speech recognition, multilingual speech recognition, combination of mono-lingual speech recognition systems, mixed language recognition. 1
David Imseng, Hervé Bourlard, Mathew Magimai-Doss
INTERSPEECH2
2010 Hierarchical multilayer perceptron based language identification
abstract
Automatic language identification (LID) systems generally exploit acoustic knowledge, possibly enriched by explicit language specific phonotactic or lexical constraints. This paper investigates a new LID approach based on hierarchical multilayer perceptron (MLP) classifiers, where the first layer is a “universal phoneme set MLP classifier”. The resulting (multilingual) phoneme posterior sequence is fed into a second MLP taking a larger temporal context into account. The second MLP can learn/exploit implicitly different types of patterns/information such as confusion between phonemes and/or phonotactics for LID. We investigate the viability of the proposed approach by comparing it against two standard approaches which use phonotactic and lexical constraints with the universal phoneme set MLP classifier as emission probability estimator. On Speech-Dat(II) datasets of five European languages, the proposed approach yields significantly better performance compared to the two standard approaches. Index Terms: Language identification, multilingual processing, hierarchical MLP.
David Imseng, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2010 Advances in fast multistream diarization based on the information bottleneck framework
abstract
Multistream diarization is an effective way to improve the diarization performance, MFCC and Time Delay Of Arrivals (TDOA) being the most commonly used features. This paper extends our previous work on information bottleneck diarization aiming to include large number of features besides MFCC and TDOA while keeping computational costs low. At first HMM/GMM and IB systems are compared in case of two and four feature streams and analysis of errors is performed. Results on a dataset of 17 meetings show that, in spite of comparable oracle performances, the IB system is more robust to feature weight variations. Then a sequential optimization is introduced that further improves the speaker error by 5 − 8% relative. In the last part, computational issues are discussed. The proposed approach is significantly faster and its complexity marginally grows with the number of feature streams running in 0.75 realtime even with four streams achieving a speaker error equal to 6%.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
INTERSPEECH3
2010 Mobile social signal processing: vision and research issues
abstract
This paper introduces the First International Workshop on Mobile Social Signal Processing (SSP).The Workshop aims at bringing together the Mobile HCI and Social Signal Processing research communities.The former investigates approaches for effective interaction with mobile and wearable devices, while the latter focuses on modeling, analysis and synthesis of nonverbal behavior in human-human and humanmachine interactions.While dealing with similar problems, the two domains have different goals and methodologies.However, mutual exchange of expertise is likely to raise new research questions as well as to improve approaches in both domains.After providing a brief survey of Mobile HCI and SSP, the paper introduces general aspects of the workshop (including topics, keynote speakers and dissemination means).
Alessandro Vinciarelli, Roderick Murray-Smith, Hervé Bourlard
Mobile HCI3
2010 Enhanced Phone Posteriors for Improving Speech Recognition Systems
abstract
Using phone posterior probabilities has been increasingly explored for improving automatic speech recognition (ASR) systems. In this paper, we propose two approaches for hierarchically enhancing these phone posteriors, by integrating long acoustic context, as well as phonetic and lexical knowledge. In the first approach, phone posteriors estimated with a multilayer perceptron (MLP), are used as emission probabilities in hidden Markov model (HMM) forward-backward recursions. This yields new enhanced posterior estimates integrating HMM topological constraints (encoding specific phonetic and lexical knowledge), and context. In the second approach, temporal contexts of the regular MLP posteriors are postprocessed by a secondary MLP, in order to learn inter- and intra-dependencies between the phone posteriors. These dependencies are phonetic knowledge. The learned knowledge is integrated in the posterior estimation during the inference (forward pass) of the second MLP, resulting in enhanced phone posteriors. We investigate the use of the enhanced posteriors in hybrid HMM/artificial neural network (ANN) and Tandem configurations. We propose using the enhanced posteriors as replacement, or as complementary evidences to the regular MLP posteriors. The proposed methods have been tested on different small and large vocabulary databases, always resulting in consistent improvements in frame, phone, and word recognition rates.
Hamed Ketabdar, Hervé Bourlard
IEEE Trans. Speech Audio Process.2
2009 MLP based hierarchical system for task adaptation in ASR
abstract
We investigate a multilayer perceptron (MLP) based hierarchical approach for task adaptation in automatic speech recognition. The system consists of two MLP classifiers in tandem. A well-trained MLP available off-the-shelf is used at the first stage of the hierarchy. A second MLP is trained on the posterior features estimated by the first, but with a long temporal context of around 130 ms. By using an MLP trained on 232 hours of conversational telephone speech, the hierarchical adaptation approach yields a word error rate of 1.8% on the 600-word Phonebook isolated word recognition task. This compares favorably to the error rate of 4% obtained by the conventional single MLP based system trained with the same amount of Phonebook data that is used for adaptation. The proposed adaptation scheme also benefits from the ability of the second MLP to model the temporal information in the posterior features.
Joel Pinto, Mathew Magimai-Doss, Hervé Bourlard
ASRU3
2009 Posterior features applied to speech recognition tasks with user-defined vocabulary
abstract
This paper presents a novel approach for those applications where vocabulary is defined by a set of acoustic samples. In this approach, the acoustic samples are used as reference templates in a template matching framework. The features used to describe the reference templates and the test utterances are estimates of phoneme posterior probabilities. These posteriors are obtained from a MLP trained on an auxiliary database. Thus, the speech variability present in the features is reduced by applying the speech knowledge captured by the MLP on the auxiliary database. Moreover, information theoretic dissimilarity measures can be used as local distances between features. When compared to state-of-the-art systems, this approach outperforms acoustic-based techniques and obtains comparable results to orthography-based methods. The proposed method can also be directly combined with other posterior-based HMM systems. This combination successfully exploits the complementarity between templates and parametric models.
Guillermo Aradilla, Hervé Bourlard, Mathew Magimai-Doss
ICASSP2
2009 Non-linear mapping for multi-channel speech separation and robust overlapping spech recognition
abstract
This paper investigates a non-linear mapping approach to extract robust features for ASR and speech separation of overlapping speech. Based on our previous studies, we continue to use two additional sound sources, namely from the target and interfering speakers. The focuses of this work are: 1) We investigate the feature mapping between different domains with the consideration of MMSE criterion and regression optimizations, demonstrating the mapping of log melfilterbank energies to MFCC can be exploited to improve the effectiveness of the regression; 2) We investigate the data-driven filtering for the speech separation by using the mapping method, which can be viewed as a generalized log spectral subtraction and results in better separation performance. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes both non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, John Dines, Mathew Magimai-Doss, Hervé Bourlard
ICASSP4
2009 Mutual information based channel selection for speaker diarization of meetings data
abstract
In the meeting case scenario, audio is often recorded using Multiple Distance Microphones (MDM) in a non-intrusive manner. Typically a beamforming is performed in order to obtain a single enhanced signal out of the multiple channels. This paper investigates the use of mutual information for selecting the channel subset that produces the lowest error in a diarization system. Conventional systems perform channel selection on the basis of signal properties such as SNR, cross correlation. In this paper, we propose the use of a mutual information measure that is directly related to the objective function of the diarization system. The proposed algorithms are evaluated on the NIST RT 06 eval dataset. Channel selection improves the speaker error by 1.1% absolute (6.5% relative) w.r.t. the use of all channels.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
ICASSP3
2009 Speaker change detection with privacy-preserving audio cues
abstract
In this paper we investigate a set of privacy-sensitive audio features for speaker change detection (SCD) in multiparty conversations. These features are based on three different principles: characterizing the excitation source information using linear prediction residual, characterizing subband spectral information shown to contain speaker information, and characterizing the general shape of the spectrum. Experiments show that the performance of the privacy-sensitive features is comparable or better than that of the state-of-the-art full-band spectral-based features, namely, mel frequency cepstral coefficients, which suggests that socially acceptable ways of recording conversations in real-life is feasible.
Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Daniel Gatica-Perez, Hervé Bourlard
ICMI4
2009 Investigating privacy-sensitive features for speech detection in multiparty conversations
abstract
We investigate four different privacy-sensitive features, namely energy, zero crossing rate, spectral flatness, and kurtosis, for speech detection in multiparty conversations. We liken this scenario to a meeting room and define our datasets and annotations accordingly. The temporal context of these features is modeled. With no temporal context, energy is the best performing single feature. But by modeling temporal context, kurtosis emerges as the most effective feature. Also, we combine the features. Besides yielding a gain in performance, certain combinations of features also reveal that a shorter temporal context is sufficient. We then benchmark other privacy-sensitive features utilized in previous studies. Our experiments show that the performance of all the privacy-sensitive features modeled with context is close to that of state-of-the-art spectral-based features, without extracting and using any features that can be used to reconstruct the speech signal.
Sree Hari Krishnan Parthasarathi, Mathew Magimai-Doss, Hervé Bourlard, Daniel Gatica-Perez
INTERSPEECH3
2009 KL realignment for speaker diarization with multiple feature streams
abstract
This paper aims at investigating the use of Kullback-Leibler (KL) divergence based realignment with application to speaker diarization. The use of KL divergence based realignment operates directly on the speaker posterior distribution estimates and is compared with traditional realignment performed using HMM/GMM system. We hypothesize that using posterior estimates to re-align speaker boundaries is more robust than gaussian mixture models in case of multiple feature streams with different statistical properties. Experiments are run on the NIST RT06 data. These experiments reveal that in case of conventional MFCC features the two approaches yields the same performance while the KL based system outperforms the HMM/GMM re-alignment in case of combination of multiple feature streams (MFCC and TDOA). Index Terms: speaker diarization, information bottleneck, feature combination
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
INTERSPEECH3
2009 Investigating the use of visual focus of attention for audio-visual speaker diarisation
abstract
Audio-visual speaker diarisation is the task of estimating ``who spoke when'' using audio and visual cues.
Giulia Garau, Sileye O. Ba, Hervé Bourlard, Jean-Marc Odobez
ACM Multimedia3
2009 Social signal processing: Survey of an emerging domain
Alessandro Vinciarelli, Maja Pantic, Hervé Bourlard
Image Vis. Comput.3
2009 An Information Theoretic Approach to Speaker Diarization of Meeting Data
abstract
A speaker diarization system based on an information theoretic framework is described. The problem is formulated according to the information bottleneck (IB) principle. Unlike other approaches where the distance between speaker segments is arbitrarily introduced, the IB method seeks the partition that maximizes the mutual information between observations and variables relevant for the problem while minimizing the distortion between observations. This solves the problem of choosing the distance between speech segments, which becomes the Jensen-Shannon divergence as it arises from the IB objective function optimization. We discuss issues related to speaker diarization using this information theoretic framework such as the criteria for inferring the number of speakers, the tradeoff between quality and compression achieved by the diarization system, and the algorithms for optimizing the objective function. Furthermore, we benchmark the proposed system against a state-of-the-art system on the NIST RT06 (rich transcription) data set for speaker diarization of meetings. The IB-based system achieves a diarization error rate of 23.2% compared to 23.6% for the baseline system. This approach being mainly based on nonparametric clustering, it runs significantly faster than the baseline HMM/GMM based system, resulting in faster-than-real-time diarization.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
IEEE Trans. Speech Audio Process.3
2008 Hierarchical integration of phonetic and lexical knowledge in phone posterior estimation
abstract
Phone posteriors has recently quite often used (as additional features or as local scores) to improve state-of-the-art automatic speech recognition (ASR) systems. Usually, better phone posterior estimates yield better ASR performance. In the present paper we present some initial, yet promising, work towards hierarchically improving these phone posteriors, by implicitly integrating phonetic and lexical knowledge. In the approach investigated here, phone posteriors estimated with a multilayer perceptron (MLP) and short (9 frames) temporal context, are used as input to a second MLP, spanning a longer temporal context (e.g. 19 frames of posteriors) and trained to refine the phone posterior estimates. The rationale behind this is that at the output of every MLP, the information stream is getting simpler (converging to a sequence of binary posterior vectors), and can thus be further processed (using a simpler classifier) by looking at a larger temporal window. Longer term dependencies can be interpreted as phonetic, sub-lexical and lexical knowledge. The resulting enhanced posteriors can then be used for phone and word recognition, in the same way as regular phone posteriors, in hybrid HMM/ANN or Tandem systems. The proposed method has been tested on TIMIT, OGI Numbers and Conversational Telephone Speech (CTS) databases, always resulting in consistent and significant improvements in both phone and word recognition rates.
Hamed Ketabdar, Hervé Bourlard
ICASSP2
2008 Combination of agglomerative and sequential clustering for speaker diarization
abstract
This paper aims at investigating the use of sequential clustering for speaker diarization. Conventional diarization systems are based on parametric models and agglomerative clustering. In our previous work we proposed a non-parametric method based on the agglomerative information bottleneck for very fast diarization. Here we consider the combination of sequential and agglomerative clustering for avoiding local maxima of the objective function and for purification. Experiments are run on the RT06 eval data. Sequential Clustering with oracle model selection can reduce the speaker error by 10% w.r.t. agglomerative clustering. When the model selection is based on Normalized Mutual Information criterion, a relative improvement of 5% is obtained using a combination of agglomerative and sequential clustering.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
ICASSP3
2008 Social signals, their function, and automatic analysis: a survey
abstract
Social Signal Processing (SSP) aims at the analysis of social behaviour in both Human-Human and Human-Computer interactions. SSP revolves around automatic sensing and interpretation of social signals, complex aggregates of nonverbal behaviours through which individuals express their attitudes towards other human (and virtual) participants in the current social context. As such, SSP integrates both engineering (speech analysis, computer vision, etc.) and human sciences (social psychology, anthropology, etc.) as it requires multimodal and multidisciplinary approaches. As of today, SSP is still in its early infancy, but the domain is quickly developing, and a growing number of works is appearing in the literature. This paper provides an introduction to nonverbal behaviour involved in social signals and a survey of the main results obtained so far in SSP. It also outlines possibilities and challenges that SSP is expected to face in the next years if it is to reach its full maturity.
Alessandro Vinciarelli, Maja Pantic, Hervé Bourlard, Alex Pentland
ICMI3
2008 Using KL-based acoustic models in a large vocabulary recognition task
abstract
Posterior probabilities of sub-word units have been shown to be an effective front-end for ASR. However, attempts to model this type of features either do not benefit from modeling context-dependent phonemes, or use an inefficient distribution to estimate the state likelihood. This paper presents a novel acoustic model for posterior features that overcomes these limitations. The proposed model can be seen as a HMM where the score associated with each state is the KL divergence between a distribution characterizing the state and the posterior features from the test utterance. This KL-based acoustic model establishes a framework where other models for posterior features such as hybrid HMM/MLP and discrete HMM can be seen as particular cases. Experiments on the WSJ database show that the KL-based acoustic model can significantly outperform these latter approaches. Moreover, the proposed model can obtain comparable results to complex systems, such as HMM/GMM, using significantly fewer parameters.
Guillermo Aradilla, Hervé Bourlard, Mathew Magimai-Doss
INTERSPEECH2
2008 Neural network based regression for robust overlapping speech recognition using microphone arrays
abstract
This paper investigates a neural network based acoustic feature mapping to extract robust features for automatic speech recognition (ASR) of overlapping speech. In our preliminary studies, we trained neural networks to learn the mapping from log mel filter bank energies (MFBEs) extracted from the distant microphone recordings, including multiple overlapping speakers, to log MFBEs extracted from the clean speech signal. In this paper, we explore the mapping of higher order mel-filterbank cepstral coefficients (MFCC) to lower order coefficients. We also investigate the mapping of features from both target and interfering distant sound sources to the clean target features. This is achieved by using the microphone array to extract features from both the direction of the target and interfering sound sources. We demonstrate the effectiveness of the proposed approach through extensive evaluations on the MONC corpus, which includes both non-overlapping single speaker and overlapping multi-speaker conditions.
Weifeng Li 0001, John Dines, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH4
2008 Integration of TDOA features in information bottleneck framework for fast speaker diarization
abstract
In this paper we address the combination of multiple feature streams in a fast speaker diarization system for meeting recordings. Whenever Multiple Distant Microphones (MDM) are used, it is possible to estimate the Time Delay of Arrival (TDOA) for different channels. In \\cite{xavi_comb}, it is shown that TDOA can be used as additional features together with conventional spectral features for improving speaker diarization. We investigate here the combination of TDOA and spectral features in a fast diarization system based on the Information Bottleneck principle. We evaluate the algorithm on the NIST RT06 diarization task. Adding TDOA features to spectral features reduces the speaker error by 3\\% absolute. Results are comparable to those of conventional HMM/GMM based systems with consistent reduction in computational complexity.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
INTERSPEECH3
2008 Social signal processing: state-of-the-art and future perspectives of an emerging domain
abstract
The ability to understand and manage social signals of a person we are communicating with is the core of social intelligence. Social intelligence is a facet of human intelligence that has been argued to be indispensable and perhaps the most important for success in life. This paper argues that next-generation computing needs to include the essence of social intelligence - the ability to recognize human social signals and social behaviours like politeness, and disagreement - in order to become more effective and more efficient. Although each one of us understands the importance of social signals in everyday life situations, and in spite of recent advances in machine analysis of relevant behavioural cues like blinks, smiles, crossed arms, laughter, and similar, design and development of automated systems for Social Signal Processing (SSP) are rather difficult. This paper surveys the past efforts in solving these problems by a computer, it summarizes the relevant findings in social psychology, and it proposes a set of recommendations for enabling the development of the next generation of socially-aware computing.
Alessandro Vinciarelli, Maja Pantic, Hervé Bourlard, Alex Pentland
ACM Multimedia3
2007 Recognition and understanding of meetings the AMI and AMIDA projects
abstract
The AMI and AMIDA projects are concerned with the recognition and interpretation of multiparty meetings. Within these projects we have: developed an infrastructure for recording meetings using multiple microphones and cameras; released a 100 hour annotated corpus of meetings; developed techniques for the recognition and interpretation of meetings based primarily on speech recognition and computer vision; and developed an evaluation framework at both component and system levels. In this paper we present an overview of these projects, with an emphasis on speech recognition and content extraction.
Steve Renals, Thomas Hain, Hervé Bourlard
ASRU3
2007 Agglomerative information bottleneck for speaker diarization of meetings data
abstract
In this paper, we investigate the use of agglomerative information bottleneck (aIB) clustering for the speaker diarization task of meetings data. In contrary to the state-of-the-art diarization systems that models individual speakers with Gaussian mixture models, the proposed algorithm is completely non parametric . Both clustering and model selection issues of non-parametric models are addressed in this work. The proposed algorithm is evaluated on meeting data on the RT06 evaluation data set. The system is able to achieve diarization error rates comparable to state-of-the-art systems at a much lower computational complexity.
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard
ASRU3
2007 An Acoustic Model Based on Kullback-Leibler Divergence for Posterior Features
abstract
This paper investigates the use of features based on posterior probabilities of subword units such as phonemes. These features are typically transformed when used as inputs for a hidden Markov model with mixture of Gaussians as emission distribution (HMM/GMM). In this work, we introduce a novel acoustic model that avoids the Gaussian assumption and directly uses posterior features without any transformation. This model is described by a finite state machine where each state is characterized by a target distribution and the cost function associated to each state is given by the Kullback-Leibler (KL) divergence between its target distribution and the posterior features. Furthermore, hybrid HMM/ANN system can be seen as a particular case of this KL-based model where state target distributions are predefined. A recursive training algorithm to estimate the state target distributions is also presented.
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
ICASSP (4)3
2007 In-context phone posteriors as complementary features for tandem ASR
abstract
In this paper, we present a method for integrating possible prior knowledge (such as phonetic and lexical knowledge), as well as acoustic context (e.g., the whole utterance) in the phone posterior estimation, and we propose to use the obtained posteriors as complementary posterior features in Tandem ASR configuration. These posteriors are estimated based on HMM state posterior probability definition (typically used in standard HMMs training). In this way, by integrating the appropriate prior knowledge and context, we enhance the estimation of phone posteriors. These new posteriors are called ?in-context? or HMM posteriors. We combine these posteriors as complementary evidences with the posteriors estimated from a Multi Layer Percep- tron (MLP), and use the combined evidence as features for training and inference in Tandem configuration. This approach has improved the performance, as compared to using only MLP estimated posteriors as features in Tandem, on OGI Numbers , Conversational Telephone speech (CTS), and Wall Street Journal (WSJ) databases.
Hamed Ketabdar, Hervé Bourlard
INTERSPEECH2
2007 Non-linear spectral contrast stretching for in-car speech recognition
abstract
In this paper, we present a novel feature normalization method in the log-scaled spectral domain for improving the noise robustness of speech recognition front-ends. In the proposed scheme, a non-linear contrast stretching is added to the outputs of log mel-filterbanks (MFB) to imitate the adaptation of the auditory system under adverse conditions. This is followed by a two-dimensional filter to smooth out the processing artifacts. The proposed MFCC front-ends perform remarkably well on CENSREC-2 in-car database with an average relative improvement of 29.3\\% compared to baseline MFCC system. It is also confirmed that the proposed processing in log MFB domain can be integrated with conventional cepstral post-processing techniques to yield further improvements. The proposed algorithm is simple and requires only a small extra computation load.
Weifeng Li 0001, Hervé Bourlard
INTERSPEECH2
2006 Using Pitch as Prior Knowledge in Template-Based Speech Recognition
abstract
In a previous paper on speech recognition, we showed that templates can better capture the dynamics of speech signal compared to parametric models such as hidden Markov models. The key point in template matching approaches is finding the most similar templates to the test utterance. Traditionally, this selection is given by a distortion measure on the acoustic features. In this work, we propose to improve this template selection with the use of meta-linguistic information as prior knowledge. In this way, similarity is not only based on acoustic features but also on other sources of information that are present in the speech signal. Results on a continuous digit recognition task confirm the statement that similarity between words does not only depend on acoustic features since we obtained 24% relative improvement over the baseline. Interestingly, results are better even when compared to a system with no prior information but a larger number of templates
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
ICASSP (1)3
2006 Using More Informative Posterior Probabilities for Speech Recognition
abstract
In this paper, we present initial investigations towards boosting posterior probability based speech recognition systems by estimating more informative posteriors taking into account acoustic context (e.g., the whole utterance), as well as possible prior information (such as phonetic and lexical knowledge). These posteriors are estimated based on HMM state posterior probability definition (typically used in standard HMMs training). This approach provides a new, principled, theoretical framework for hierarchical estimation/use of more informative posteriors integrating appropriate context and prior knowledge. In the present work, we used the resulting posteriors as local scores for decoding. On the OGI numbers database, this resulted in significant performance improvement, compared to using MLP estimated posteriors for decoding (hybrid HMM/ANN approach) for clean and more specially for noisy speech. The system is also shown to be much less sensitive to tuning factors (such as phone deletion penalty, language model scaling) compared to the standard HMM/ANN and HMM/GMM systems, thus practically it does not need to be tuned to achieve the best possible performance
Hamed Ketabdar, Jithendra Vepa, Samy Bengio, Hervé Bourlard
ICASSP (1)4
2006 Threshold Selection for Unsupervised Detection, With an Application to Microphone Arrays
abstract
Detection is usually done by comparing some criterion to a threshold. It is often desirable to keep a performance metric such as false alarm rate constant across conditions. Using training data to select the threshold may lead to suboptimal results on test data recorded in different conditions. This paper investigates unsupervised approaches, where no training data is used. A probabilistic model is fitted on the test data using the EM algorithm, and the threshold value is selected based on the model. The proposed approach (1) does not use training data, (2) uses the test data itself to compensate for simplifications inherent to the model, and (3) permits the use of more complex models in a straightforward manner. On a microphone array speech detection task, the proposed unsupervised approach achieves similar or better results than the "training" approach. The methodology is general and may be applied to other contexts than microphone arrays, and other performance metrics than FAR
Guillaume Lathoud, Mathew Magimai-Doss, Hervé Bourlard
ICASSP (3)3
2006 Using posterior-based features in template matching for speech recognition
abstract
Given the availability of large speech corpora, as well as the increasing of memory and computational resources, the use of template matching approaches for automatic speech recognition (ASR) have recently attracted new attention. In such template-based approaches, speech is typically represented in terms of acoustic vector sequences, using spectral-based features such as MFCC of PLP, and local distances are usually based on Euclidean or Mahalanobis distances. In the present paper, we further investigate template-based ASR and show (on a continuous digit recognition task) that the use of posterior-based features significantly improves the standard template-based approaches, yielding to systems that are very competitive to state-of-the-art HMMs, even when using a very limited number (e.g., 10) of reference templates. Since those posteriors-based features can also be interpreted as a probability distribution, we also show that using Kullback-Leibler (KL) divergence as a local distance further improves the performance of the template-based approach, now beating state-of-the-art of more complex posterior-based HMMs systems (usually referred to as "Tandem").
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
INTERSPEECH3
2006 Posterior based keyword spotting with a priori thresholds
abstract
In this paper, we propose a new posterior based scoring approach for keyword and non keyword (garbage) elements. The estimation of these scores is based on HMM state posterior probability definition, taking into account long contextual information and the prior knowledge (e.g. keyword model topology). The state posteriors are then integrated into keyword and garbage posteriors for every frame. These posteriors are used to make a decision on detection of the keyword at each frame. The frame level decisions are then accumulated (in this case, by counting) to make a global decision on having the keyword in the utterance. In this way, the contribution of possible outliers are minimized, as opposed to the conventional Viterbi decoding approach which accumulates likelihoods. Experiments on keywords from the Conversational Telephone Speech (CTS) and Numbers'95 databases are reported. Results show that the new scoring approach leads to better trade off between true and false alarms compared to the Viterbi decoding approach, while also providing the possibility to precalculate keyword specific spotting thresholds related to the length of the keywords.
Hamed Ketabdar, Jithendra Vepa, Samy Bengio, Hervé Bourlard
INTERSPEECH4
2006 Multi-stream ASR: an oracle perspective
abstract
Multi-stream based automatic speech recognition (ASR) systems are usually shown to outperform single stream systems, specially in noisy test conditions. And, indeed, there is a trend today in ASR towards using more and more acoustic features combined at the input (early integration, possibly preceded by some linear or nonlinear transformation) or later in the recognition process (e.g., at the level of likelihoods, then referred to as late integration). However, to guarantee optimal exploitation of such multi-stream systems, we need to use features that are as much complementary as possible, while also using the best combination method for those streams. In practice, it is never clear whether we fully exploit the potential of the available streams. This present paper investigates an ‘oracle ’ test to provide some insight in these issues. Although not providing us with an absolute performance upper bound, oracle is shown to indicate the complimentary of the feature streams used, and to provide a reasonable reference target to evaluate combination strategies. The oracle analysis is supported by results obtained on Numbers95 database using different feature streams and entropy based combination method.
Hemant Misra, Jithendra Vepa, Hervé Bourlard
INTERSPEECH3
2006 Understanding and Modeling Communication Scenes
abstract
This paper discusses the technologies required to better understand and model human-human communication, and to use the resulting technologies to build computer-enhanced communication tools. As networks and computers become more pervasive, groups are increasingly using technology to assist communication and collaboration and to reduce travel needs. The addition of new technologies based on advanced signal (audio-visual) processing and multimedia information analysis can have a positive impact on meetings and human communication. However, human communication is complex and is factored across several modalities. To address the problem requires major research efforts in several traditionally separate disciplines including unconstrained speech recognition, visual scene analysis, modeling individuals and groups through the joint processing of multiple information channels, and structuring, indexing and summarizing these multimodal communication scenes. Projects such as AMI/AMIDA (www. amiproiect.org) have made significant progress in these basic areas. AMI/AMIDA research revolves around instrumented meeting rooms and advanced videoconferencing systems which enable the collection, annotation, structuring, and browsing of multimodal meeting recordings.
Hervé Bourlard
SLT1
2006 User-customized password speaker verification using multiple reference and background models
Mohamed Faouzi BenZeghiba, Hervé Bourlard
Speech Commun.2
2006 On variable-scale piecewise stationary spectral analysis of speech signals for ASR
Vivek Tyagi, Hervé Bourlard, Christian Wellekens
Speech Commun.2
2005 HMM/ANN Based Spectral Peak Location Estimation for Noise Robust Speech Recognition
abstract
In this paper, we present an HMM/ANN based algorithm to estimate the spectral peak locations. This algorithm makes use of distinct time-frequency (TF) patterns in the spectrogram for estimating the peak locations. Such a use of TF patterns is expected to impose temporal constraints during the peak estimation task, thereby yielding a smoother estimate of the peaks over time. Additionally, the algorithm uses an ergodic topology for the HMM/ANN, thus allowing an estimation of a varying number of peak locations over time. The usefulness of the proposed algorithm is evaluated in the framework of a recently introduced noise robust feature called the spectro-temporal activity pattern (STAP) feature. Interestingly, the recently introduced phase autocorrelation (PAC) spectrum, with enhanced spectral peaks and smoothed spectral valleys, turns out to be more appropriate for this algorithm than the regular spectrum.
Shajith Ikbal, Hervé Bourlard, Mathew Magimai-Doss
ICASSP (1)2
2005 Multi-resolution Spectral Entropy Feature for Robust ASR
abstract
Recently, entropy measures at different stages of recognition have been used in automatic speech recognition (ASR) tasks. In a recent paper, we proposed that formant positions of a spectrum can be captured by a multi-resolution spectral entropy feature. In this paper, we suggest modifications to the spectral entropy feature extraction approach and compute the entropy contribution from each sub-band to the total entropy of the normalized spectrum. Further, we explore the ideas of overlapping sub-bands and the time derivatives of the spectral entropy feature. The modified feature is robust to additive wide-band noise and performs well at low SNRs. Finally, in the TANDEM framework, we show that the system using combined entropy and PLP (perceptual linear prediction) features works better than the baseline PLP feature for additive wide-band noise at different SNRs.
Hemant Misra, Shajith Ikbal, Sunil Sivadas, Hervé Bourlard
ICASSP (1)4
2005 Improving speech recognition using a data-driven approach
abstract
In this paper, we investigate the possibility of enhancing state-of-the-art HMM-based speech recognition systems using data-driven techniques, where whole set of training utterances is used as reference models and recognition is then performed through the well-known template matching technique, DTW. This approach allows us to better capture the temporal dynamics of the speech signal while avoiding some of the HMM assumptions such as the piecewise stationarity. Potentially, such data-driven techniques also allow us to better exploit meta-data and environmental information, such as speaker, gender, accent and noise conditions. However, we cannot entirely abandon HMMs, which are very powerful and scalable models. Thus, we investigate one way to combine and take advantage of both the approaches, combining scores of HMMs and reference templates. Experiments on the Numbers95 database showed that this combination yields 22\% relative improvement in word error rate over the baseline HMM performance. Applying K-means clustering to the acoustic vectors speeds up the decoding, while still retaining a significant improvement in the recognition accuracy.
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
INTERSPEECH3
2005 Developing and enhancing posterior based speech recognition systems
abstract
Local state or phone posterior probabilities are often investigated as local scores (e.g., hybrid HMM/ANN systems) or as transformed acoustic features (e.g., “Tandem”) to improve speech recognition systems. In this paper, we present initial results towards boosting these approaches by improving posterior estimates, using acoustic context (e.g., as available in the whole utterance), as well as possible prior information (such as topological constraints). In the present work, the enhanced posterior distribution is associated with the “gamma ” distribution typically used in standard HMMs training, and estimated from local likelihoods (GMM) or local posteriors (ANN). This approach results in a family of new HMM based systems, where only posterior probabilities are used, while also providing a new, principled, approach towards a hierarchical use/integration of these posteriors, from the frame level up to the phone and word levels, and integrating the appropriate context and prior knowledge in each level. In the present work, we used the resulting posteriors as local scores in a Viterbi decoder. On the OGI Numbers’95 database, this resulted in improved recognition performance, compared to a state-of-the-art hybrid HMM/ANN system. 1.
Hamed Ketabdar, Jithendra Vepa, Samy Bengio, Hervé Bourlard
INTERSPEECH4
2005 Spectral entropy feature in full-combination multi-stream for robust ASR
abstract
In a recent paper, we reported promising automatic speech recognition results obtained by appending spectral entropy features to PLP features. In the present paper, spectral entropy features are used along with PLP features in the framework of multi-stream combination. In a full-combination multi-stream hidden Markov model/artificial neural network (HMM/ANN) hybrid system, we train a separate multi-layered perceptron (MLP) for PLP features, for spectral entropy features and for both combined by concatenation. The output posteriors from these three MLPs are combined with weights inversely proportional to the entropies of their respective posterior distributions. We show that on the Numbers95 database, this approach yields a significant improvement under both clean and noisy conditions as compared to simply appending the features. Further, in the framework of a Tandem HMM/ANN system, we apply the same inverse entropy weighting to combine the outputs of the MLPs before the softmax non-linearity. Feeding the combined and decorrelated MLP outputs to the HMM gives a 9.2\% relative error reduction as compared to the baseline.
Hemant Misra, Hervé Bourlard
INTERSPEECH2
2005 On variable-scale piecewise stationary spectral analysis of speech signals for ASR
abstract
It is often acknowledged that speech signals contain short-term and long-term temporal properties that are difficult to capture and model by using the usual fixed scale (typically 20ms) short time spectral analysis used in hidden Markov models (HMMs), based on piecewise stationarity and state conditional independence assumptions of acoustic vectors. For example, vowels are typically quasi-stationary over 40-80ms segments, while plosives typically require analysis below 20ms segments. Thus, fixed scale analysis is clearly sub-optimal for ``optimal'' time-frequency resolution and modeling of different stationary phones found in the speech signal. In the present paper, we investigate the potential advantages of using variable size analysis windows towards improving state-of-the-art speech recognition systems. Based on the usual assumption that the speech signal can be modeled by a varying autoregressive (AR) Gaussian process, we estimate the largest piecewise quasi-stationary speech segments, based on the likelihood that a segment was generated by the same AR process. This likelihood is estimated from the Linear Prediction (LP) residual error. Each of these quasi-stationary segments is then used as an analysis window from which spectral features are extracted. Such an approach thus results in variable scale time spectral analysis, adaptively estimating the largest possible analysis window size such that the signal remains quasi-stationary, thus the best temporal/frequency resolution tradeoff. Speech recognition experiments on the OGI Numbers95 database show that the proposed multi-scale piecewise stationary spectral analysis based features indeed yield improved recognition accuracy in clean conditions, compared to features based on minimum cross entropy spectrum as well as those based on fixed scale spectral analysis.
Vivek Tyagi, Christian Wellekens, Hervé Bourlard
INTERSPEECH3
2005 Comparison and combination of features in a hybrid HMM/MLP and a HMM/GMM speech recognition system
abstract
Recently, the advantages of the spectral parameters obtained by frequency filtering (FF) of the logarithmic filter-bank energies (logFBEs) have been reported. These parameters, which are frequency derivatives of the logFBEs, lie in the frequency domain, and have shown good recognition performance with respect to the conventional mel-frequency cepstral coefficients (MFCCs) for hidden Markov models (HMM) based systems. In this paper, the FF features are first compared with the MFCCs and the relative spectral perceptual linear prediction (Rasta-PLP) features using both a hybrid HMM/MLP and a usual HMM/Gaussian mixture models (HMM/GMM) based recognition system, for both clean and noisy speech. Taking advantage of the ability of the hybrid system to deal with correlated features, the inclusion of both the frequency second-derivatives and the raw logFBEs as additional features is proposed and tested. Moreover, the robustness of these features in noisy conditions is enhanced by combining the FF technique with the Rasta temporal filtering approach. Finally, a study of the FF features in the framework of multistream processing is presented. The best recognition results for both clean and noisy speech are obtained from the multistream combination of the J-Rasta-PLP features and the FF features.
Pere Pujol, Susagna Pol, Climent Nadeu, Astrid Hagen, Hervé Bourlard
IEEE Trans. Speech Audio Process.5
2004 Confidence measures in multiple pronunciations modeling for speaker verification
abstract
The paper investigates the use of multiple pronunciations modeling for user-customized password speaker verification (UCP-SV). The main characteristic of UCP-SV is that the system does not have any a priori knowledge about the password used by the speaker. Our aim is to exploit the information about how the speaker pronounces a password in the decision process. This information is extracted automatically using a speaker-independent speech recognizer. We investigate and compare several techniques. Some of them are based on the combination of confidence scores estimated by different models. In this context, we propose a new confidence measure that uses acoustic information extracted during speaker enrollment and based on a log likelihood ratio measure. These techniques show significant improvement (15.7% relative improvement in terms of equal error rate) compared to a UCP-SV baseline system where the speaker is modeled by only one model (corresponding to one utterance).
Mohamed Faouzi BenZeghiba, Hervé Bourlard
ICASSP (1)2
2004 Phase autocorrelation (PAC) features in entropy based multi-stream for robust speech recognition
abstract
Methods to improve noise robustness of speech recognition systems often result in degradation of recognition performance for clean speech. Recently proposed phase autocorrelation (PAC) based features (S. Ikbal et al., Proc. ICASSP-03, p.II-133-6, 2003; Proc. IEEE ASRU 2003 Workshop, 2003), while showing noticeable improvement in noise robustness, also suffer from this drawback. We try to alleviate this problem by using the PAC based features along with regular speech features in a multi-stream framework. The multi-stream system uses the entropy of the posterior probability distribution, computed during recognition, as a confidence measure to combine evidence from different feature streams adaptively (Misra, H. et al., Proc. ICASSP-03, p.II-741-4, 2003). Experimental results obtained on OGI Numbers95 database and Noisex92 noise database show that such a system yields the best possible recognition performance in all conditions. Actually, the combination always performs better than the best performing stream for all the conditions.
Shajith Ikbal, Hemant Misra, Hervé Bourlard, Hynek Hermansky
ICASSP (1)3
2004 Joint decoding for phoneme-grapheme continuous speech recognition
abstract
Standard ASR systems typically use phonemes as the subword units. Preliminary studies have shown that the performance of ASR systems could be improved by using graphemes as additional subword units. We investigate such a system where the word models are defined in terms of two different subword units, i.e., phoneme and grapheme. During training, models for both the subword units are trained, and then, during recognition, either both or just one subword unit is used. We have studied this system for a continuous speech recognition task in American English. Our studies show that grapheme information used along with phoneme information improves the performance of ASR.
Mathew Magimai-Doss, Samy Bengio, Hervé Bourlard
ICASSP (1)3
2004 Spectral entropy based feature for robust ASR
abstract
In general, entropy gives us a measure of the number of bits required to represent some information. When applied to the probability mass function (PMF), entropy can also be used to measure the "peakiness" of a distribution. We propose using the entropy of a short time Fourier transform spectrum, normalised as PMF, as an additional feature for automatic speech recognition (ASR). It is indeed expected that a peaky spectrum, representation of clear formant structure in the case of voiced sounds, will have low entropy, while a flatter spectrum, corresponding to nonspeech or noisy regions, will have higher entropy. Extending this reasoning further, we introduce the idea of a multiband/multiresolution entropy feature where we divide the spectrum into equal size subbands and compute entropy in each subband. The results show that multiband entropy features used in conjunction with normal cepstral features improve the performance of an ASR system.
Hemant Misra, Shajith Ikbal, Hervé Bourlard, Hynek Hermansky
ICASSP (1)3
2004 An online audio indexing system
abstract
This paper presents overview of an online audio indexing system, which creates a searchable index of speech content embedded in digitized audio files. This system is based on our recently proposed offline audio segmentation techniques. As the data arrives continuously, the system first finds boundaries of the acoustically homogenous segments. Next, each of these segments is classified as speech, music or {\\it mixture} classes, where mixtures are defined as regions where speech and other non-speech sounds are present simultaneously and noticeably. The speech segments are then clustered together to provide consistent speaker labels. The speech and mixture segments are converted to text via an ASR system. The resulting words are time-stamped together with other metadata information (speaker identity, speech confidence score) in an XML file to rapidly identify and access target segments. In this paper, we analyze the performance at each stage of this audio indexing system and also compare it with the performance of the corresponding offline modules.
Jitendra Ajmera, Iain McCowan, Hervé Bourlard
INTERSPEECH3
2004 Posteriori probabilities and likelihoods combination for speech and speaker recognition
abstract
This paper investigates a new approach to perform simultaneous speech and speaker recognition. The likelihood estimated by a speaker identification system is combined with the posterior probability estimated by the speech recognizer. So, the joint posterior probability of the pronounced word and the speaker identity is maximized. A comparison study with other standard techniques is carried out in three different applications, (1) closed set speech and speaker identification, (2) open set speech and speaker identification and (3) speaker quantization in speaker-independent speech recognition.
Mohamed Faouzi BenZeghiba, Hervé Bourlard
INTERSPEECH2
2004 Spectro-temporal activity pattern (STAP) features for noise robust ASR
abstract
In this paper, we introduce a new noise robust representation of speech signal obtained by locating points of potential importance in the spectrogram, and parameterizing the activity of time-frequency pattern around those points. These features are referred to as Spectro-Temporal Activity Pattern (STAP) features. The suitability of these features for noise robust speech recognition is examined for a particular parameterization scheme where spectral peaks are chosen as points of potential importance. The activity in the time-frequency patterns around these points are parameterized by measuring the dynamics of the patterns along both time and frequency axes. As the spectral peaks are considered to constitute an important and robust cue for speech recognition, this representation is expected to yield a robust performance. An interesting result of the study is that inspite of using a relatively less amount of information from the speech signal, STAP features are able to achieve a reasonable recognition performance in clean speech, when compared to the state-of-the-art features. In addition, STAP features produce a significantly better performance in high noise conditions. An entropy based combination technique in tandem frame-work to combine STAP features with standard features yields a system which is more robust in all conditions.
Shajith Ikbal, Mathew Magimai-Doss, Hemant Misra, Hervé Bourlard
INTERSPEECH4
2004 Entropy based combination of tandem representations for noise robust ASR
abstract
In this paper, we present an entropy based method to combine tandem representations of the recently proposed Phase AutoCorrelation (PAC) based features and Mel-Frequency Cepstral Coefficients (MFCC) features. PAC based features, derived from a nonlinear transformation of autocorrelation coefficients and shown to be noise robust, improve their robustness to additive noise in their tandem representation. On the other hand, MFCC features in their tandem representation show a significant improvement in recognition performance on clean speech. An entropy based combination method investigated in this paper adaptively gives a higher weighting to the representation of MFCC features in clean speech and to the representation of PAC based features in noisy speech, thus yielding a robust recognition performance in all conditions.
Shajith Ikbal, Hemant Misra, Sunil Sivadas, Hynek Hermansky, Hervé Bourlard
INTERSPEECH5
2004 Modeling auxiliary features in tandem systems
abstract
Tandem systems transform the cepstral features into posterior probabilities of subword units using artificial neural networks (ANNs), which are processed to form input features for conventional speech recognition systems. They have been shown to perform better than conventional speech recognition systems using cepstral features. Recent studies have shown that modelling cepstral features with auxiliary sources of knowledge leads to improvement in the performance of speech recognition systems. In this paper, we study two approaches to incorporate auxiliary knowledge sources such as pitch frequency, short-term energy, etc. (referred to as auxiliary features), in a tandem-based automatic speech recognition system. In the first approach, we model the auxiliary features in the process of training an ANN, which is later used to extract tandem-features. In the second approach, we extract the tandem-features from an ANN trained with cepstral features only and then model them jointly with auxiliary features. Recognition studies conducted on a connected word recognition task under clean and noisy conditions show that the performance of the tandem system can be improved by incorporating auxiliary features.
Mathew Magimai-Doss, Shajith Ikbal, Todd A. Stephenson, Hervé Bourlard
INTERSPEECH4
2004 Text detection, recognition in images and video frames
Datong Chen, Jean-Marc Odobez, Hervé Bourlard
Pattern Recognit.3
2004 Robust speaker change detection
abstract
Most commonly used criteria for speaker change detection like log likelihood ratio (LLR) and Bayesian information criterion (BIC) have an adjustable threshold/penalty parameter to make speaker change decisions. These parameters are not always robust to different acoustic conditions and have to be tuned. In this letter, we present a criterion which can be used to identify speaker changes in an audio stream without such tuning. The criterion consists of calculating the LLR of two models with the same number of parameters. Results on the Hub4 1997 evaluation set indicate that we achieve a performance comparable to using BIC with optimal penalty term.
Jitendra Ajmera, Iain McCowan, Hervé Bourlard
IEEE Signal Process. Lett.3
2004 Speech recognition with auxiliary information
abstract
State-of-the-art automatic speech recognition (ASR) systems are usually based on hidden Markov models (HMMs) that emit cepstral-based features which are assumed to be piecewise stationary. While not really robust to noise, these features are also known to be very sensitive to "auxiliary" information, such as pitch, energy, rate-of-speech (ROS), etc. Attempts so far to include such auxiliary information in state-of-the-art ASR systems have often been based on simply appending these auxiliary features to the standard acoustic feature vectors. In the present paper, we investigate different approaches to incorporating this auxiliary information using dynamic Bayesian networks (DBNs) or hybrid HMM/ANNs (HMMs with artificial neural networks). These approaches are motivated by the fact that the auxiliary information is not necessarily (directly) emitted by the HMM states but, rather, carries higher-level information (e.g., speaker characteristics) that is correlated with the standard features. As implicitly done for gender modeling elsewhere, this auxiliary information then appears as a conditional variable in the emission distributions and can be hidden (except in the case of some HMM/ANNs) as its estimates become too noisy. Based on recognition experiments carried out on the OGI Numbers database (free format numbers spoken over the telephone), we show that auxiliary information that conditions the distribution of the standard features can, in certain conditions, provide more robust recognition than using auxiliary information that is appended to the standard features; this is most evident in the case of energy as an auxiliary variable in noisy speech.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
IEEE Trans. Speech Audio Process.3
2003 Hybrid HMM/ANN and GMM combination for user-customized password speaker verification
abstract
Recently we have proposed an approach for user-customized password speaker verification; in this approach, we combined a hybrid HMM/ANN model (used for utterance verification) and a GMM model (used for speaker verification). In this paper, we extend our investigations. First, we propose a new similarity measure that uses confidence measures developed in the HMM/ANN framework. Secondly, we analyze the contribution of each model using a weighted sum combination technique. Experiments conducted on a subset of the PolyVar database show that for a short password the performance of the combined system did not improve significantly compared to the performance using the GMM model alone, and that the HMM/ANN did not contribute much in the combined system. We discuss possible reasons for this.
Mohamed Faouzi BenZeghiba, Hervé Bourlard
ICASSP (2)2
2003 Phase autocorrelation (PAC) derived robust speech features
abstract
We introduce a new class of noise robust acoustic features derived from a new measure of autocorrelation, and explicitly exploiting the phase variation of the speech signal frame over time. This family of features, referred to as "phase autocorrelation" (PAC) features, include PAC spectrum and PAC MFCC (Mel-frequency cepstral coefficient), among others. In regular autocorrelation based features, the correlation between two signal segments (signal vectors), separated by a particular time interval k, is calculated as a dot product of these two vectors. In our proposed PAC approach, the angle between the two vectors is used as a measure of correlation. Since dot product is usually more affected by noise than the angle, PAC-features are expected to be more robust to noise. This is indeed significantly confirmed by the presented experimental results. The experiments were conducted on the Numbers 95 database, on which "stationary" (car) and "non -stationary" (factory) Noisex 92 noises were added with varying SNR. In most of the cases, without any specific tuning, PAC-MFCC features perform better.
Shajith Ikbal, Hemant Misra, Hervé Bourlard
ICASSP (2)3
2003 Modeling human interaction in meetings
abstract
The paper investigates the recognition of group actions in meetings by modeling the joint behaviour of participants. Many meeting actions, such as presentations, discussions and consensus, are characterised by similar or complementary behaviour across participants. Recognising these meaningful actions is an important step towards the goal of providing effective browsing and summarisation of processed meetings. A corpus of meetings was collected in a room equipped with a number of microphones and cameras. The corpus was labeled in terms of a predefined set of meeting actions characterised by global behaviour. In experiments, audio and visual features for each participant are extracted from the raw data and the interaction of participants is modeled using HMM-based approaches. Initial results on the corpus demonstrate the ability of the system to recognise the set of meeting actions.
Iain McCowan, Samy Bengio, Daniel Gatica-Perez, Guillaume Lathoud, Florent Monay, Darren Moore, Pierre Wellner, Hervé Bourlard
ICASSP (4)8
2003 New entropy based combination rules in HMM/ANN multi-stream ASR
abstract
Classifier performance is often enhanced through combining multiple streams of information. In the context of multi-stream HMM/ANN systems in ASR, a confidence measure widely used in classifier combination is the entropy of the posteriors distribution output from each ANN, which generally increases as classification becomes less reliable. The rule most commonly used is to select the ANN with the minimum entropy. However, this is not necessarily the best way to use entropy in classifier combination. In this article, we test three new entropy based combination rules in a full-combination multi-stream HMM/ANN system for noise robust speech recognition. Best results were obtained by combining all the classifiers having entropy below average using a weighting proportional to their inverse entropy.
Hemant Misra, Hervé Bourlard, Vivek Tyagi
ICASSP (2)2
2003 Speech recognition of spontaneous, noisy speech using auxiliary information in Bayesian networks
abstract
Automatic speech recognition (ASR) currently performs well in the case of clean, read speech. It performs worse, however, when the speech is spontaneous and in noisy conditions. In previous work we showed the improvement that using auxiliary information in the framework of Bayesian networks (BNs) can bring to ASR in clean, read speech. Here we show that auxiliary information of pitch or rate-of-speech in the context of BNs also helps performance in spontaneous speech with noise.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
ICASSP (1)3
2003 On automatic annotation of meeting databases
abstract
In this paper, we discuss meetings as an application domain for multimedia content analysis. Meeting databases are a rich data source suitable for a variety of audio, visual and multi-modal tasks, including speech recognition, people and action recognition, and information retrieval. We specifically focus on the task of semantic annotation of audio-visual (AV) events, where annotation consists of assigning labels (event names) to the data. In order to develop an automatic annotation system in a principled manner, it is essential to have a well-defined task, a standard corpus and an objective performance measure. In this work we address each of these issues to automatically annotate events based on participant interactions.
Mark Barnard, Samy Bengio, Hervé Bourlard, Daniel Gatica-Perez, Iain McCowan
ICIP (3)3
2003 Speech & face based biometric authentication at IDIAP
abstract
We present an overview of research at IDIAP on speech & face based biometric authentication. This paper covers user-customised passwords, adaptation techniques, confidence measures (for use in fusion of audio & visual scores), face verification in difficult image conditions, as well as other related research issues. We also overviewed the open source Torch library, which has aided in the implementation of the above mentioned techniques.
Conrad Sanderson, Samy Bengio, Hervé Bourlard, Johnny Mariéthoz, Ronan Collobert, Mohamed Faouzi BenZeghiba, Fabien Cardinaux, Sébastien Marcel
ICME3
2003 On the combination of speech and speaker recognition
abstract
This paper investigates an approach that maximizes the joint posterior probabil ity of the pronounced word and the speaker identity given the observed data. This probability can be expressed as a product of the posterior probability of the pronounced word estimated through an artificial neural network (ANN), and the likelihood of the data estimated through a Gaussian mixture model (GMM). We show that the posterior probabilities estimated through a speaker-dependent ANN, as usually done in the hybrid HMM/ANN systems, are reliable for speech recognition but they are less reliable for speaker recognition. To alleviate this problem, we thus study how this posterior probability can be combined with the likelihood derived from a speaker-dependent GMM model to improve the speaker recognition performance. We thus end up with a joint model that can be used for text-dependent speaker identification and for speech recognition (and mutually benefiting from each other).
Mohamed Faouzi BenZeghiba, Hervé Bourlard
INTERSPEECH2
2003 Using pitch frequency information in speech recognition
abstract
Automatic Speech Recognition systems typically use smoothed spectral features as acoustic observations. In recent studies, it has been shown that complementing these standard features with pitch frequency could improve the system performance of the system. While previously proposed systems have been studied in the framework of HMM/GMMs, in this paper we study and compare different ways to include pitch frequency in state-of-the-art hybrid HMM/ANN system. We have evaluated the proposed system on two different ASR tasks, namely, isolated word recognition and connected word recognition. Our results show that pitch frequency can indeed be used in ASR systems to improve the recognition performance.
Mathew Magimai-Doss, Todd A. Stephenson, Hervé Bourlard
INTERSPEECH3
2003 On factorizing spectral dynamics for robust speech recognition
abstract
In this paper, we introduce new dynamic speech features based on the modulation spectrum. These features, termed Mel-cepstrum Modulation Spectrum (MCMS), map the time trajectories of the spectral dynamics into a series of slow and fast moving orthogonal components, providing a more general and discriminative range of dynamic features than traditional delta and acceleration features. The features can be seen as the outputs of an array of band-pass filters spread over the cepstral modulation frequency range of interest. In experiments, it is shown that, as well as providing a slight improvement in clean conditions, these new dynamic features yield a significant increase in speech recognition performance in various noise conditions when compared directly to the standard temporal derivative features and RASTA-PLP features.
Vivek Tyagi, Iain McCowan, Hervé Bourlard, Hemant Misra
INTERSPEECH3
2003 Robust speech recognition and feature extraction using HMM2
Katrin Weber, Shajith Ikbal, Samy Bengio, Hervé Bourlard
Comput. Speech Lang.4
2003 Speech/music segmentation using entropy and dynamism features in a HMM classification framework
Jitendra Ajmera, Iain McCowan, Hervé Bourlard
Speech Commun.3
2003 Microphone array post-filter based on noise field coherence
abstract
This paper introduces a novel technique for estimating the signal power spectral density to be used in the transfer function of a microphone array post-filter. The technique is a generalization of the existing Zelinski post-filter, which uses the auto- and cross-spectral densities of the array inputs to estimate the signal and noise spectral densities. The Zelinski technique, however, assumes zero cross-correlation between the noise on different sensors. This assumption is inaccurate, particularly at low frequencies and for arrays with closely spaced sensors, and thus the corresponding post-filter is suboptimal in realistic noise conditions. In this paper, a more general expression of the post-filter estimation is developed based on an assumed knowledge of the complex coherence of the noise field. This general expression can be used to construct a more appropriate post-filter in a variety of different noise fields. In experiments using real noise recordings from a computer office, the modified post-filter results in significant improvement in terms of objective speech quality measures and speech recognition performance using a diffuse noise model.
Iain McCowan, Hervé Bourlard
IEEE Trans. Speech Audio Process.2
2002 Robust HMM-based speech/music segmentation
abstract
In this paper we present a new approach towards high performance speech/music segmentation on realistic tasks related to the automatic transcription of broadcast news. In the approach presented here, the local probability density function (PDF) estimators trained on clean microphone speech are used as a channel model at the output of which the entropy and “dynamism” will be measured and integrated over time through a 2-state (speech and and non-speech) hidden Markov model (HMM) with minimum duration constraints. The parameters of the HMM are trained using the EM algorithm in a completely unsupervised manner. Different experiments, including a variety of speech and music styles, as well as different segment durations of speech and music signals (real data distribution, mostly speech, or mostly music), will illustrate the robustness of the approach, which in each case achieves a frame-level accuracy greater than 94%.
Jitendra Ajmera, Iain McCowan, Hervé Bourlard
ICASSP3
2002 Microphone array post-filter for diffuse noise field
abstract
This paper proposes a novel technique for estimating the signal power spectral density to be used in the transfer function of a microphone array post-filter. The technique is a modification of the existing Zelinski post-filter, which uses the auto- and cross-spectral densities of the array inputs to estimate the signal and noise spectral densities. The Zelinski technique, however, assumes zero cross-correlation between noise on different sensors. This assumption is inaccurate in real conditions, particularly at low frequencies and for arrays with closely spaced sensors. In this paper we replace this with an assumption of a theoretically diffuse noise field, which is more appropriate in a variety of realistic noise environments. In experiments using noise recordings from an office of computer workstations, the modified post-filter results in significant improvement in terms of objective speech quality measures and speech recognition performance.
Iain McCowan, Hervé Bourlard
ICASSP2
2002 Increasing speech recognition robustness with HMM2
abstract
The purpose of this paper is to investigate the behavior of HMM2 models for the recognition of noisy speech. It has previously been shown that HMM2 is able to model dynamically important structural information inherent in the speech signal, often corresponding to formant positions/tracks. As formant regions are known to be robust in adverse conditions, HMM2 seems particularly promising for improving speech recognition robustness. Here, we review different variants of the HMM2 approach with respect to their application to noise-robust automatic speech recognition. It is shown that HMM2 has the potential to tackle the problem of mismatch between training and testing conditions, and that a multi-stream combination of (already noise-robust) cepstral features and formant-like features (extracted by HMM2) improves the noise robustness of a state-of-the-art automatic speech recognition system.
Katrin Weber, Samy Bengio, Hervé Bourlard
ICASSP3
2002 Unknown-multiple speaker clustering using HMM
abstract
An HMM-based speaker clustering framework is presented, where the number of speakers and segmentation boundaries are unknown \\emph{a priori}. Ideally, the system aims to create one pure cluster for each speaker. The HMM is ergodic in nature with a minimum duration topology. The final number of clusters is determined automatically by merging closest clusters and retraining this new cluster, until a decrease in likelihood is observed. In the same framework, we also examine the effect of using only the features from highly voiced frames as a means of improving the robustness and computational complexity of the algorithm. The proposed system is assessed on the 1996 HUB-4 evaluation test set in terms of both cluster and speaker purity. It is shown that the number of clusters found often correspond to the actual number of speakers.
Jitendra Ajmera, Hervé Bourlard, I. Lapidot, Iain McCowan
INTERSPEECH2
2002 User-customized password speaker verification based on HMM/ANN and GMM models
abstract
In this paper, we present a new approach towards user-custom\-ized password speaker verification combining the advantages of hybrid HMM/ANN systems, using Artificial Neural Networks (ANN) to estimate emission probabilities of Hidden Markov Models, and Gaussian Mixture Models. In the approach presented here, we indeed exploit the properties of hybrid HMM/ANN systems, usually resulting in high phonetic recognition rates, to automatically infer the baseline phonetic transcription (HMM topology) associated with the user customized password from a few enrollment utterances and using a large, speaker independent, ANN. The emission probabilities of the resulting HMMs are then modeled in terms of speaker specific/adapted multi-Gaussian HMMs or speaker specific/adapted ANN. In the proposed approach, the hybrid HMM/ANN system is used as a model for utterance (password) verification, while still using a speaker independent GMM for speaker verification. Results (EER) are compared to a state-of-the-art text-dependent approach, using multi-Gaussian HMMs only.
Mohamed Faouzi BenZeghiba, Hervé Bourlard
INTERSPEECH2
2002 Comparison and combination of RASTA-PLP and FF features in a hybrid HMM/MLP speech recognition system
abstract
Recently, the advantages of the spectral parameters obtained by frequency filtering (FF) of the logarithmic filter bank energies (logFBEs) have been reported. These parameters, which are frequency derivatives of the logFBEs, lie in the frequency domain, and have shown good recognition performance with respect to the conventional mel-frequency cepstral coefficients (MFCC) for HMM systems. In this paper, the FF features are compared with the MFCCs and the Rasta-PLP features in the framework of a hybrid HMM/MLP recognition system, for both clean and noisy speech. Taking advantage of the ability of the hybrid system to deal with correlated features, the inclusion of the second derivatives and the FBEs as additional features is proposed. Furthermore, in order to enhance the robustness of these features in noisy conditions, they are combined with the Rasta temporal filtering approach. Finally, a study of the FF in the framework of multistream processing is presented. From the experimental tests, it appears that the new spectral parameters and the tested combinations yield an enhanced recognition performance. 1.
Pere Pujol, Susagna Pol, Astrid Hagen, Hervé Bourlard, Climent Nadeu
INTERSPEECH4
2002 Improving speech recognition performance of small microphone arrays using missing data techniques
abstract
Traditional microphone array speech recognition systems simply recognise the enhanced output of the array. As the level of signal enhancement depends on the number of microphones, such systems do not achieve acceptable speech recognition performance for arrays having only a few microphones. For small microphone arrays, we instead propose using the enhanced output to estimate a reliability mask, which is then used in missing data speech recognition. In missing data speech recognition, the decoded sequence depends on the reliability of each input feature. This reliability is usually based on the signal to noise ratio in each frequency band. In this paper, we use the energy difference between the noisy input and the enhanced output of a small microphone array to determine the frequency band reliability. Recognition experiments with a small array demonstrate the effectiveness of the technique, compared to both traditional microphone array enhancement and a baseline missing data system.
Iain McCowan, Andrew C. Morris, Hervé Bourlard
INTERSPEECH3
2002 Low cost duration modelling for noise robust speech recognition
abstract
State transition matrices as used in standard HMM decoders have two widely perceived limitations. One is that the implicit Geometric state duration distributions which they model do not accurately reflect true duration distributions. The other is that they impose no hard limit on maximum duration with the result that state transition probabilities often have little influence when combined with acoustic probabilities, which are of a different order of magnitude. Explicit duration models were developed in the past to address the first problem. These were not widely taken up because their performance advantage in clean speech recognition was often not sufficiently great to offset the extra complexity which they introduced. However, duration models have much greater potential when applied to noisy speech recognition. In this paper we present a simple and generic form of explicit duration model and show that this leads to strong performance improvements when applied to connected digit recognition in noise.
Andrew C. Morris, Simon Payne, Hervé Bourlard
INTERSPEECH3
2002 Auxiliary variables in conditional Gaussian mixtures for automatic speech recognition
abstract
In previous work, we presented a case study using an estimated pitch value as the conditioning variable in conditional Gaussians that showed the utility of hiding the pitch values in certain situations or in modeling it independently of the hidden state in others. Since only single conditional Gaussians were used in that work, we extend that work here to using conditional Gaussian mixtures in the emission distributions to make this work more comparable to state-of-the-art automatic speech recognition. We also introduce a rate-of-speech (ROS) variable within the conditional Gaussian mixtures. We find that, under the current methods, using observed pitch or ROS in the recognition phase does not provide improvement. However, systems trained on pitch or ROS may provide improvement in the recognition phase over the baseline when the pitch or ROS is marginalized out.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2002 Evaluation of formant-like features for ASR
abstract
This paper investigates possibilities to automatically find a low-dimensional, formant-related physical representation of the speech signal, which is suitable for automatic speech recognition (ASR). This aim is motivated by the fact that formants have been shown to be discriminant features for ASR. Combinations of automatically extracted formant-like features and `conventional', noise-robust, state-of-the-art features (such as MFCCs including spectral subtraction and cepstral mean subtraction) have previously been shown to be more robust in adverse conditions than state-of-the-art features alone. However, it is not clear how these automatically extracted formant-like features behave in comparison with true formants. The purpose of this paper is to investigate two methods to automatically extract formant-like features, and to compare these features to hand-labeled formant tracks as well as to standard MFCCs in terms of their performance on a vowel classification task.
Katrin Weber, Febe de Wet, Bert Cranen, Lou Boves, Samy Bengio, Hervé Bourlard
INTERSPEECH6
2002 Analytic assessment of telephone transmission impact on ASR performance using a simulation model
Sebastian Möller 0001, Hervé Bourlard
Speech Commun.2
2001 Text Identification in Complex Background Using SVM
abstract
The paper presents a fast and robust algorithm to identify text in image or video frames with complex backgrounds and compression effects. The algorithm first extracts the candidate text line on the basis of edge analysis, baseline location and heuristic constraints. Support Vector Machine (SVM) is then used to identify text line from the candidates in edge-based distance map feature space. Experiments based on a large amount of images and video frames from different sources showed the advantages of this algorithm compared to conventional methods in both identification quality and computation time.
Datong Chen, Hervé Bourlard, Jean-Philippe Thiran
CVPR (2)2
2001 Adaptive ML-weighting in multi-band recombination of Gaussian mixture ASR
abstract
Multi-band speech recognition is powerful in band-limited noise, when the recognizer of the noisy band, which is less reliable, can be given less weight in the recombination process. An accurate decision on which bands can be considered as reliable and which bands are less reliable due to corruption by noise is usually hard to take. We investigate a maximum-likelihood (ML) approach to adapting the combination weights of a multi-band system. The Gaussian mixture model parameters are kept constant, while the combination weights are iteratively updated to maximize the data likelihood. Unsupervised offline and online weights adaptation are compared to the use of equal weights, and 'cheating' weights where the noisy band is known, as well as to the fullband system. Initial tests show that both ML-weighting strategies show a robustness gain on band-limited noise.
Astrid Hagen, Hervé Bourlard, Andrew C. Morris
ICASSP2
2001 Error correcting posterior combination for robust multi-band speech recognition
abstract
In human perception, the availability of context enhances recognition and renders it more robust to noise. Even if not all phonemes in a word (or words in a sentence etc.) are correctly perceived, humans can fill in missing parts with the help of cues from the surrounding speech parts. This was proven in studies on human speech perception where recognition of words in sentences under noise was shown to outperform recognition of words in isolation or, even more drastically, of nonsense syllables under noise. A new model for quantifying the influence of contextual information on human recognition performance was recently proposed. Although the authors state that it is not a model for the recognition process itself, we will see how the ideas behind this model can be used in automatic speech recognition to extend our formerly introduced multi-band recognition systems to incorporate frequency contextual information. We will compare the new set-up to our former models such as the full combination subband approach and its approximation.
Astrid Hagen, Hervé Bourlard
INTERSPEECH2
2001 MAP combination of multi-stream HMM or HMM/ANN experts
abstract
Automatic speech recognition (ASR) performance falls dramatically with the level of mismatch between training and test data. The human ability to recognise speech when a large proportion of frequencies are dominated by noise has inspired the "missing data" and "multi-band" approaches to noise robust ASR. "Missing data" ASR identifies low SNR spectral data in each data frame and then ignores it. Multi-band ASR trains a separate model for each position of missing data, estimates a reliability weight for each model, then combines model outputs in a weighted sum. A problem with both approaches is that local data reliability estimation is inherently inaccurate and also assumes that all of the training data was clean. In this article we present a model in which adaptive multi-band expert weighting is incorporated naturally into the maximum a posteriori (MAP) decoding process.
Andrew C. Morris, Astrid Hagen, Hervé Bourlard
INTERSPEECH3
2001 Modeling auxiliary information in Bayesian network based ASR
abstract
Automatic speech recognition bases its models on the acoustic features derived from the speech signal. Some have investigated replacing or supplementing these features with information that can not be precisely measured (articulator positions, pitch, gender, etc.) automatically. Consequently, automatic estimations of the desired information would be generated. This data can degrade performance due to its imprecisions. In this paper, we describe a system that treats pitch as an auxiliary information within the framework of Bayesian networks, resulting in improved performance. 1.
Todd A. Stephenson, Mathew Magimai-Doss, Hervé Bourlard
INTERSPEECH3
2001 HMM2- extraction of formant structures and their use for robust ASR
abstract
As recently introduced, an HMM2 can be considered as a particular case of an HMM mixture in which the HMM emission probabilities (usually estimated through Gaussian mixtures or an artificial neural network) are modeled by state-dependent, feature-based HMM (referred to as frequency HMM). A general EM training algorithm for such a structure has already been developed. Although there are numerous motivations for using such a structure, and many possible ways to exploit it, this paper will mainly focus on one particular instantiation of HMM2 in which the frequency HMM will be used to extract formant structure information, which will then be used as additional acoustic features in a standard Automatic Speech Recognition (ASR) system. While the fact that this architecture is able to automatically extract meaningful formant information is interesting by itself, empirical results will also show the robustness of these features to noise, and their potential to enhance regular HMM-based ASR.
Katrin Weber, Samy Bengio, Hervé Bourlard
INTERSPEECH3
2001 Multi-stream adaptive evidence combination for noise robust ASR
Andrew C. Morris, Astrid Hagen, Hervé Glotin, Hervé Bourlard
Speech Commun.4
2000 A new keyword spotting approach based on iterative dynamic programming
abstract
This paper addresses the problem of detecting keywords in unconstrained speech without explicit modeling of non-keyword segments. The proposed algorithm is based on recent developments in confidence measures using local posterior probabilities, and searches for the segment maximizing the average observation posteriori along the most likely path in the hypothesized keyword model. As known, this approach (sometimes referred to as sliding model method) requires a relaxation of the begin/endpoints of the Viterbi matching, as well as a time normalization of the resulting score, making dynamic programming sub-optimal or more complex (more computation and/or more memory). We present here an alternative (quite simple and efficient) solution to this problem, using an iterative form of Viterbi decoding algorithm, but which does not require scoring for all possible begin/endpoints. Convergence proof of this algorithm is available (Silaghi and Bourlard, 1999). Results obtained with this method on 100 keywords chosen at random from the BREF database are reported.
Marius-Calin Silaghi, Hervé Bourlard
ICASSP2
2000 Using multiple time scales in the framework of multi-stream speech recognition
abstract
In this paper, we present a new approach to incorporating multiple time scale information as independent streams in multi-stream processing. To illustrate the procedure, we take two different sets of multiple time scale features. In the first system, these are features extracted over variable sized windows of three and five times the original window size. In the second system, we take as separate input streams the commonly used difference features, i.e. the first and second order derivatives of the instantaneous features. In the same way, any other kinds of multiple time scale features could be employed. The approach is embedded in the recently introduced ``full combination'' approach to multi-stream processing in which, the phoneme probabilities from all possible combinations of streams are combined in a weighted sum. As an extension of this approach we have found that replacing the sum of probabilities by their product, in the same ``all wise'' context, can result in higher robustness. Capturing different information in each stream, and with the longer time scale features being more robust to noise, the multiple time scale multi-stream system gained a significant performance improvement in both clean speech and in real-environmental noise.
Astrid Hagen, Hervé Bourlard
INTERSPEECH2
2000 Real-time telephone transmission simulation for speech recognizer and dialogue system evaluation and improvement
Sebastian Möller 0001, Hervé Bourlard
INTERSPEECH2
2000 A neural network for classification with incomplete data: application to robust ASR
abstract
If the data vector for input to an automatic classifier is incomplete, the optimal estimate for each class probability must be calculated as the expected value of the classifier output. We identify a form of RBF classifier whose expected outputs can easily be evaluated in terms of the original function parameters. We then describe two ways in which this classifier can be applied to robust automatic speech recognition, depending on whether or not the position of missing data is known
Andrew C. Morris, Ljubomir Josifovski, Hervé Bourlard, Martin Cooke, Phil D. Green
INTERSPEECH3
2000 Automatic speech recognition using dynamic bayesian networks with both acoustic and articulatory variables
abstract
Current technology for automatic speech recognition (ASR) uses hidden Markov models (HMMs) that recognize spoken speech using the acoustic signal. However, no use is made of the causes of the acoustic signal: the articulators. We present here a dynamic Bayesian network (DBN) model that utilizes an additional variable for representing the state of the articulators. A particular strength of the system is that, while it uses measured articulatory data during its training, it does not need to know these values during recognition. As Bayesian networks are not used often in the speech community, we give an introduction to them. After describing how they can be used in ASR, we present a system to do isolated word recognition using articulatory information. Recognition results are given, showing that a system with both acoustics and inferred articulatory positions performs better than a system with only acoustics.
Todd A. Stephenson, Hervé Bourlard, Samy Bengio, Andrew C. Morris
INTERSPEECH2
2000 HMM2- a novel approach to HMM emission probability estimation
abstract
In this paper, we discuss and investigate a new method to estimate local emission probabilities in the framework of hidden Markov models (HMM). Each feature vector is considered to be a sequence and is supposed to be modeled by yet another HMM. Therefore, we call this approach `HMM2'. There is a variety of possible topologies of such HMM2 systems, e.g. incorporating trellis or ergodic HMM structures. Preliminary HMM2 speech recognition experiments on cepstral and spectral features yielded worse results than state-of-the-art systems. However, we believe that HMM2 systems have a lot of potential advantages and are therefore worth investigating further.
Katrin Weber, Samy Bengio, Hervé Bourlard
INTERSPEECH3
2000 Development of Acoustic and Linguistic Resources for Research and Evaluation in Interactive Vocal Information Servers
Giulia Bernardis, Hervé Bourlard, Martin Rajman, Jean-Cédric Chappelier
LREC2
2000 New Approaches Towards Robust, Adaptive Speech Recognition (invited paper)
Hervé Bourlard, Samy Bengio, Katrin Weber
NIPS1
1999 The full combination sub-bands approach to noise robust HMM/ANN based ASR
abstract
The performance of most ASR systems degrades rapidly with data mismatch relative to the data used in training. Under many realistic noise conditions a significant proportion of the spectral representation of a speech signal, which is highly redundant, remains uncorrupted. In the “missing feature ” approach to this problem mismatching data is simply ignored, but the need to base recognition on unorthogonalised spectral features results in reduced performance in clean speech. In multiband ASR the results from independent recognition on a number of within-band orthogonalised sub-bands are combined. This approach more accurately reflects the uncertainty in mismatch detection, but loss of joint information due to independent sub-band processing can also result in reduced performance with clean speech. In this article the “full combination ” approach to noise robust ASR is presented in which multiple data streams are associated not with individual sub-bands but with sub-band combinations. In this way no assumption of sub-band independence is required. Initial tests show some improved robustness to noise with no significant loss of performance with clean speech.
Andrew C. Morris, Astrid Hagen, Hervé Bourlard
EUROSPEECH3
1998 Improving posterior based confidence measures in hybrid HMM/ANN speech recognition systems
abstract
In this paper we define and investigate a set of confidence measures based on hybrid Hidden Markov Model/Artificial Neural Network (HMM/ANN) acoustic models. All these measures are using the neural network to estimate the local phone posterior probabilities, which are then combined and normalized in different ways. Experimental results will indeed show that the use of an appropriate duration normalization is very important to obtain good estimates of the phone and word confidences. The different measures are evaluated at the phone and word levels on both an isolated word task (PHONEBOOK) and a continuous speech recognition task (BREF). It will be shown that one of those confidence measures is well suited for utterance verification, and that (as one could expect) confidence measures at the word level perform better than those at the phone level. Finally, using the resulting approach on PHONEBOOK to rescore the N-best list is shown to yield a 34\\% decrease in word error rate.
Giulia Bernardis, Hervé Bourlard
ICSLP2
1998 Interfacing of CASA and partial recognition based on a multistream technique
abstract
LIDIAP
Frédéric Berthommier, Hervé Glotin, Emmanuel Tessier, Hervé Bourlard
ICSLP4
1997 State-of-the-Art and Recent Progress in Hybrid HMM/ANN Speech Recognition
Hervé Bourlard
ICANN1
1997 Subband-based speech recognition
abstract
In the framework of hidden Markov models (HMM) or hybrid HMM/artificial neural network (ANN) systems, we present a new approach towards automatic speech recognition (ASR). The general idea is to divide up the full frequency band (represented in terms of critical bands) into several subbands, compute phone probabilities for each subband on the basis of subband acoustic features, perform dynamic programming independently for each band, and merge the subband recognizers (recombining the respective, possibly weighted, scores) at some segmental level corresponding to temporal anchor points. The results presented in this paper confirm some preliminary tests reported earlier. On both isolated word and continuous speech tasks, it is indeed shown that even using quite simple recombination strategies, this subband ASR approach can yield at least comparable performance on clean speech while providing better robustness in the case of narrowband noise.
Hervé Bourlard, Stéphane Dupont
ICASSP1
1997 Hybrid HMM/ANN systems for training independent tasks: experiments on Phonebook and related improvements
abstract
In this paper, we evaluate multi-Gaussian HMM systems and hybrid HMM/ANN systems in the framework of task independent training for small size (75 words) and medium size (600 words) vocabularies. To do this, we use the Phonebook database (Pitrelli et al., 1995) which is particularly well suited to this kind of experiment since (1) it is a very large telephone database and (2) the size and content of the test vocabulary is very flexible. For each system, different HMM topologies are compared to test the influence of state tying (with a number of parameters approximately kept constant) on the recognition performance. Two lexica (Phonebook and CMU) are also compared and it is shown that the CMU lexicon leads to significantly better performance. Finally, it is shown that with a quite simple system and a few adaptations to the basic HMM/ANN scheme, recognition performance of 98.5% and 94.7% can easily be achieved, respectively on a lexicon of 75 and 600 words (isolated words, telephone speech and lexicon words not present in the training data).
Stéphane Dupont, Hervé Bourlard, Olivier Deroo, Vincent Fontaine, Jean-Marc Boite
ICASSP2
1997 Speaker-dependent speech recognition based on phone-like units models-application to voice dialling
abstract
This paper presents a speaker dependent speech recognition with application to voice dialling. This work has been developed under the constraints imposed by voice dialling applications, i.e., low memory requirements and limited training material. Two methods for producing speaker dependent word baseforms based on phone-like units (PLU) are presented and compared: (1) a classical vector quantizer is used to divide the space into regions associated with PLUs; (2) a speaker independent hybrid HMM/MLP recognizer is used to generate speaker dependent PLU based models. This work shows that very low error rates can be achieved even with very simple systems, namely a DTW-based recognizer. However, the best results are achieved when using the hybrid HMM/MLP system to generate the word baseforms. Finally, a real-time demonstration simulating voice dialling functions and including keyword spotting and rejection capabilities has been set up and can be tested online.
Vincent Fontaine, Hervé Bourlard
ICASSP2
1997 Using multiple time scales in a multi-stream speech recognition system
abstract
In this paper, we propose and investigate a new approach towards using multiple time scale information in automatic speech recognition (ASR) systems. In this framework, we are using a particular HMM formalism able to process different input streams and to recombine them at some temporal anchor points. While the phonological level of recombination has to be defined a priori, the optimal temporal anchor points are obtained automatically during recognition. In the current approach, those parallel cooperative HMMs will focus on different dynamic properties of the speech signal, defined on differenttime scales. The speech signal is then defined in terms of several information streams, each stream resulting from a particular way of analyzing the speech signal. More specifically, in the currentwork, models aimed at capturing the syllable level temporal structure are used in parallel with classical phoneme-based models. Tests on different continuous speech databases show significant performance improvements, motivating further research to efficiently use large time span information of the order of 200 ms into our standard 10 ms, phone-based ASR systems.
Stéphane Dupont, Hervé Bourlard
EUROSPEECH2
1997 Estimation of global posteriors and forward-backward training of hybrid HMM/ANN systems
abstract
The results of our research presented in this paper is two-fold.First, an estimation of global posteriors is formalized in the framework of hybrid HMM/ANN systems.It is shown that hybrid HMM/ANN systems, in which the ANN part estimates local posteriors, can be used to modelize global model posteriors.This formalization provides us with a clear theory in which both REMAP and \classical" Viterbi trained hybrid systems are unied.Second, a new forward-backward training of hybrid HMM/ANN systems is derived from the previous formulation.Comparisons of performance between Viterbi and forward-backward hybrid systems are presented and discussed.
Jean Hennebert, Christophe Ris, Hervé Bourlard, Steve Renals, Nelson Morgan
EUROSPEECH3
1996 REMAP-experiments with speech recognition
abstract
In this report we present experimental and theoretical results using a framework for training and modeling continuous speech recognition systems based on the theoretically optimal Maximum a Posteriori (MAP) criterion. This is in constrast to most state-of-the-art systems which are trained according to a Maximum Likelihood (ML) criterion. Although the algorithm is quite general, we applied it to a particular form of hybrid system combining Hidden Markov Models (HMMs) and Artificial Neural Networks (ANNs) in which the ANN targets and weights are iteratively re-estimated to guarantee the increase of the posterior probability of the correct model, hence actually minimizing the error rate. More specifically, this training approach is applied to a transition-based model that uses local conditional transition probabilities (i.e., the posterior probability of the current state given the current acoustic vector and the previous state) to estimate the posterior probabilities of sentences. Experi...
Yochai Konig, Hervé Bourlard, Nelson Morgan
ICASSP2
1996 Stochastic perceptual speech models with durational dependence
Jeff A. Bilmes, Nelson Morgan, Su-Lin Wu, Hervé Bourlard
ICSLP4
1996 A new ASR approach based on independent processing and recombination of partial frequency bands
abstract
In the framework of hidden Markov models (HMM) or hybrid HMM/Artificial Neural Network (ANN) systems, we present a new approach towards automatic speech recognition (ASR). The general idea is to split the whole frequency band (represented in terms of critical bands) into a few sub-bands on which different recognizers are independently applied and then recombined at a certain speech unit level to yield global scores and a global recognition decision. The preliminary results presented in this paper show that such an approach, even using quite simple recombination strategies, can yield at least comparable performance on clean speech while providing better robustness in the case of noisy speech.
Hervé Bourlard, Stéphane Dupont
ICSLP1
1996 Towards increasing speech recognition error rates
Hervé Bourlard, Hynek Hermansky, Nelson Morgan
Speech Commun.1
1996 A training algorithm for statistical sequence recognition with applications to transition-based speech recognition
abstract
In this letter, we introduce a discriminant training algorithm for statistical sequence recognition that uses a transition-based stochastic finite state automaton with posterior transition probabilities conditioned on the current input observation and the previous state. This provides a framework for frame-synchronous speech recognition in which posterior probabilities are estimated as the basis for recognition, rather than the state-dependent probability densities that are conventionally used. Preliminary speech recognition experiments support the theory by showing an increase in the estimates of posterior probabilities of the correct sentences and a statistically significant decrease in error rates for independent test sets.
Hervé Bourlard, Yochai Konig, Nelson Morgan
IEEE Signal Process. Lett.1
1995 Stochastic perceptual models of speech
abstract
We have developed a statistical model of speech (based on auditory perceptual criteria) that avoids a number of current constraining assumptions for statistical speech recognition systems, particularly the model of speech as a sequence of stationary segments consisting of uncorrelated acoustic vectors. We further wish to focus statistical modeling power on perceptually-dominant and information-rich portions of the speech signal, which may also be the parts of the speech signal with a better chance to withstand adverse acoustical conditions. We describe some of the theory, along with some preliminary experiments. These experiments suggest that the regions of acoustic signal containing significant spectral change are critical to the recognition of continuous speech.
Nelson Morgan, Hervé Bourlard, Steven Greenberg, Hynek Hermansky, Su-Lin Wu
ICASSP2
1995 Towards increasing speech recognition error rates
Hervé Bourlard
EUROSPEECH1
1995 REMAP: recursive estimation and maximization of a posteriori probabilities in connectionist speech recognition
abstract
In this paper, we briefly describe REMAP, an approach for the training and estimation of posterior probabilities, and report its application to speech recognition. REMAP is a recursive algorithm that is reminiscent of the Expectation Maximization (EM) [5] algorithm for the estimation of data likelihoods. Although very general, the method is developed in the context of a statistical model for transition-based speech recognition using Artificial Neural Networks (ANN) to generate probabilities for Hidden Markov Models (HMMs). In the new approach, we use local conditional posterior probabilities of transitions to estimate global posterior probabilities of word sequences. As with earlier hybrid HMM/ANN systems we have developed, ANNs are used to estimate posterior probabilities. In the new approach, however, the network is trained with targets that are themselves estimates of local posterior probabilities. Initial experimental results support the theory by showing an increase in the estima...
Hervé Bourlard, Yochai Konig, Nelson Morgan
EUROSPEECH1
1995 Digit recognition with stochastic perceptual speech models
abstract
We have recently developed a statistical model of speech that focuses statistical modeling power on phonetic transitions. These are the perceptually-dominant and informationrich portions of the speech signal, which may also be the parts of the speech signal with a better chance to withstand adverse acoustical conditions. We describe here some of the concepts, along with some preliminary experiments on digit recognition. These experiments show that the new models, when used in combination with our more standard models, can significantly improve performance in the presence of noise. 1. BACKGROUND In [5] we reported the development of a statistical model of speech that incorporates some simple temporal properties of speech perception. The primary goal of this theoretical development was to avoid a number of current constraining assumptions for statistical speech recognition systems, particularly the model of speech as a sequence of stationary segments consisting of uncorrelated acoustic...
Nelson Morgan, Su-Lin Wu, Hervé Bourlard
EUROSPEECH3
1995 REMAP: Recursive Estimation and Maximization of A Posteriori Probabilities - Application to Transition-Based Connectionist Speech Recognition
Yochai Konig, Hervé Bourlard, Nelson Morgan
NIPS2
1995 Neural networks for statistical recognition of continuous speech
abstract
In recent years there has been a significant body of work, both theoretical and experimental, that has established the viability of artificial neural networks (ANN's) as a useful technology for speech recognition. It has been shown that neural networks can be used to augment speech recognizers whose underlying structure is essentially that of hidden Markov models (HMM's). In particular, we have demonstrated that fairly simple layered structures, which we lately have termed big dumb neural networks (BDNN's), can be discriminatively trained to estimate emission probabilities for an HMM. Recently simple speech recognition systems (using context-independent phone models) based on this approach have been proved on controlled tests, to be both effective in terms of accuracy (i.e., comparable or better than equivalent state-of-the-art systems) and efficient in terms of CPU and memory run-time requirements. Research is continuing on extending these results to somewhat more complex systems. In this paper, we first give a brief overview of automatic speech recognition (ASR) and statistical pattern recognition in general. We also include a very brief review of HMM's, and then describe the use of ANN's as statistical estimators. We then review the basic principles of our hybrid HMM/ANN approach and describe some experiments. We discuss some current research topics, including new theoretical developments in training ANN's to maximize the posterior probabilities of the correct models for speech utterances. We also discuss some issues of system resources required for training and recognition. Finally, we conclude with some perspectives about fundamental limitations in the current technology and some speculations about where we can go from here.>
Nelson Morgan, Hervé Bourlard
Proc. IEEE2
1995 Comparison of hidden Markov model techniques for automatic speaker verification in real-world conditions
Johan de Veth, Hervé Bourlard
Speech Commun.2
1994 Task independent and dependent training: performance comparison of HMM and hybrid HMM/MLP approaches
abstract
Compares speaker independent isolated word recognition performance obtained with standard phonemic hidden Markov models (HMMs) and hybrid approaches using a multilayer perceptron (MLP) to estimate the HMM emission probabilities. This latter approach has previously been shown particularly effective on a large vocabulary, speaker independent, continuous speech recognition task (i.e., ARPA Resource Management) by using simple context-independent phoneme models and single pronunciation word models. As a consequence, the main goal of the paper is to compare the performance which can be achieved by the different approaches for both task dependent and independent training.>
Jean-Marc Boite, Hervé Bourlard, Bart D'hoore, Sari Accaino, Johan Vantieghem
ICASSP (1)2
1994 Optimizing recognition and rejection performance in wordspotting systems
abstract
Compares the performance which can be achieved by different hidden Markov model (HMM) based wordspotting techniques when their parameters are tuned to optimize recognition and rejection rates. An alternative approach which does not attempt to explicitly model extraneous speech or non-speech noise is also proposed. After optimization of each of these approaches, it appears that the proposed version performs at least as well as the other methods with the advantage of simplicity and possibility to be used in hybrid models using HMMs with a multilayer perceptron (MLP). Test results are reported on a speaker independent telephone database containing 10 keywords as well as on the speaker independent ARPA resource management database in which between 10 and 250 keywords were defined.>
Hervé Bourlard, Bart D'hoore, Jean-Marc Boite
ICASSP (1)1
1994 Comparison of acoustic features and robustness tests of a real-time recogniser using a hardware telephone line simulator
Hugo Van hamme, Guido Gallopyn, Ludwig Weynants, Bart D'hoore, Hervé Bourlard
ICSLP5
1994 Stochastic perceptual auditory-event-based models for speech recognition
Nelson Morgan, Hervé Bourlard, Steven Greenberg, Hynek Hermansky
ICSLP2
1994 Connectionist probability estimators in HMM speech recognition
abstract
The authors are concerned with integrating connectionist networks into a hidden Markov model (HMM) speech recognition system. This is achieved through a statistical interpretation of connectionist networks as probability estimators. They review the basis of HMM speech recognition and point out the possible benefits of incorporating connectionist networks. Issues necessary to the construction of a connectionist HMM recognition system are discussed, including choice of connectionist probability estimator. They describe the performance of such a system using a multilayer perceptron probability estimator evaluated on the speaker-independent DARPA Resource Management database. In conclusion, they show that a connectionist component improves a state-of-the-art HMM system.
Steve Renals, Nelson Morgan, Hervé Bourlard, Horacio Franco
IEEE Trans. Speech Audio Process.3
1993 Limited parameter hidden Markov models for connected digit speaker verification over telephone channels
Johan de Veth, Guido Gallopyn, Hervé Bourlard
ICASSP (2)3
1993 Linear and nonlinear prediction for speech recognition with hidden Markov models
Marco Saerens, Hervé Bourlard
EUROSPEECH2
1993 A new approach towards keyword spotting
Jean-Marc Boite, Hervé Bourlard, Bart D'hoore, Marc Haesen
EUROSPEECH2
1993 Performance comparison of hidden Markov models and neural networks for task dependent and independent isolated word recognition
Hervé Bourlard, Jean-Marc Boite, Bart D'hoore, Marco Saerens
EUROSPEECH1
1993 A neural network based, speaker independent, large vocabulary, continuous speech recognition system: the WERNICKE project
abstract
International Computer Science Institute (ICSI), USA(Author list is alphabetical with the exception of the typist.)ABSTRACTThis paper describes the research underway for the ESPRITWERNICKE project. The project brings together a num-ber of different groups from Europe and the US and focuseson extending the state-of-the-art for hybrid hidden Markovmodel/connectionist approaches to large vocabulary, continu-ous speech recognition. Thispaper describes the specific goalsoftheresearchandpresentstheworkperformedtodate. Resultsare reported for the resource management talker-independentrecognition task. The paper concludes with a discussion of theprojected future work.Keywords: Recognition, Neural Nets, HMM.1. BACKGROUNDW
Tony Robinson, Luís B. Almeida, Jean-Marc Boite, Hervé Bourlard, Frank Fallside, Mike Hochberg, Dan J. Kershaw, Phil Kohn, Yochai Konig, Nelson Morgan, João Paulo da Silva Neto, Steve Renals, Marco Saerens, Chuck Wooters
EUROSPEECH4
1993 Real-time, neural network-based, French alphabet recognition with telephone speech
Philipp Schmid 0001, Ronald A. Cole, Mark A. Fanty, Hervé Bourlard, M. Haessen
EUROSPEECH4
1993 Speaker verification over telephone channels based on concatenated phonemic hidden Markov models
Johan de Veth, Guido Gallopyn, Hervé Bourlard
EUROSPEECH3
1993 Hybrid Neural Network/Hidden Markov Model Systems for Continuous Speech Recognition
abstract
MultiLayer Perceptrons (MLP) are an effective family of algorithms for the smooth estimation of highly-dimensioned probability density functions that are useful in continuous speech recognition. Hidden Markov Models (HMM) provide a structure for the mapping of a temporal sequence of acoustic vectors to a generating sequence of states. For HMMs that are independent of phonetic context, the MLP approaches have consistently provided significant improvements (once we learned how to use them). Recently, these results have been extended to context-dependent models. In this paper, after having reviewed the basic principles of our hybrid HMM/MLP approach, we describe a series of experiments with continuous speech recognition. The hybrid methods directly trade off computational complexity for reduced requirements of memory and memory bandwidth. Results are presented on the widely used Resource Management speech database that is distributed by the National Institute of Standards and Technology. These results demonstrate performance that is at least as good as any other reported continuous speech recognition system (for this task).
Nelson Morgan, Hervé Bourlard, Steve Renals, Horacio Franco
Int. J. Pattern Recognit. Artif. Intell.2
1993 Continuous speech recognition by connectionist statistical methods
abstract
Over the period of 1987-1991, a series of theoretical and experimental results have suggested that multilayer perceptrons (MLP) are an effective family of algorithms for the smooth estimation of high-dimension probability density functions that are useful in continuous speech recognition. The early form of this work has focused on hidden Markov models (HMM) that are independent of phonetic context. More recently, the theory has been extended to context-dependent models. The authors review the basic principles of their hybrid HMM/MLP approach and describe a series of improvements that are analogous to the system modifications instituted for the leading conventional HMM systems over the last few years. Some of these methods directly trade off computational complexity for reduced requirements of memory and memory bandwidth. Results are presented on the widely used Resource Management speech database that has been distributed by the US National Institute of Standards and Technology.
Hervé Bourlard, Nelson Morgan
IEEE Trans. Neural Networks1
1992 CDNN: a context dependent neural network for continuous speech recognition
abstract
A series of theoretical and experimental results have suggested that multilayer perceptrons (MLPs) are an effective family of algorithms for the smooth estimate of highly dimensioned probability density functions that are useful in continuous speech recognition. All of these systems have exclusively used context-independent phonetic models, in the sense that the probabilities or costs are estimated for simple speech units such as phonemes or words, rather than biphones or triphones. Numerous conventional systems based on hidden Markov models (HMMs) have been reported that use triphone or triphone like context-dependent models. In one case the outputs of many context-dependent MLPs (one per context class) were used to help choose the best sentence from the N best sentences as determined by a context-dependent HMM system. It is shown how, without any simplifying assumptions, one can estimate likelihoods for context-dependent phonetic models with nets that are not substantially larger than context-independent MLPs.>
Hervé Bourlard, Nelson Morgan, Chuck Wooters, Steve Renals
ICASSP1
1992 Factoring Networks by a Statistical Method
abstract
We show that it is possible to factor a multilayered classification network with a large output layer into a number of smaller networks, where the product of the sizes of the output layers equals the size of the original output layer. No assumptions of statistical independence are required.
Nelson Morgan, Hervé Bourlard
Neural Comput.2
1992 Neural nets and hidden Markov models: Review and generalizations
Hervé Bourlard, Nelson Morgan, Steve Renals
Speech Commun.1
1991 Continuous speech recognition using PLP analysis with multilayer perceptrons
abstract
The authors investigate the use of continuous features derived by perceptual linear predictive (PLP) analysis, examine the effect of adding temporal features, and compare it to the previously studied use of multiframe input. Comparisons of the MLP (multilayer perceptron) and conventional Gaussian classifiers are also reported. The speaker-dependent portion of the Resource Management database was used for this test. Additionally, some experiments were performed with a perplexity-2200 speaker-independent recognition task on a subset of the TIMIT database. In each case, the PLP features were used as input to the networks. The experiments show the advantage of continuous PLP features and their first and second temporal derivatives.>
Nelson Morgan, Hynek Hermansky, Hervé Bourlard, Phil Kohn, Chuck Wooters
ICASSP3
1991 Neural nets and hidden Markov models: review and generalizations
Hervé Bourlard
EUROSPEECH1
1991 Phonetic context in hybrid HMM/MLP continuous speech recognition
Nelson Morgan, Hervé Bourlard, Chuck Wooters, Phil Kohn
EUROSPEECH2
1991 Connectionist Optimisation of Tied Mixture Hidden Markov Models
Steve Renals, Nelson Morgan, Hervé Bourlard, Horacio Franco
NIPS3
1990 Continuous speech recognition using multilayer perceptrons with hidden Markov models
abstract
A phoneme based, speaker-dependent continuous-speech recognition system embedding a multilayer perceptron (MLP) (i.e. a feedforward artificial neural network) into a hidden Markov model (HMM) approach is described. Contextual information from a sliding window on the input frames is used to improve frame or phoneme classification performance over the corresponding performance for simple maximum-likelihood probabilities, or even maximum a posteriori (MAP) probabilities which are estimated without the benefit of context. Performance for a simple discrete density HMM system appears to be somewhat better when MLP methods are used to estimate the probabilities.>
Nelson Morgan, Hervé Bourlard
ICASSP2
1990 Continuous speech recognition on the resource management database using connectionist probability estimation
Nelson Morgan, Chuck Wooters, Hervé Bourlard
ICSLP3
1990 Connectionist Approaches to the Use of Markov Models for Speech Recognition
Hervé Bourlard, Nelson Morgan, Chuck Wooters
NIPS1
1990 Links Between Markov Models and Multilayer Perceptrons
abstract
The statistical use of a particular classic form of a connectionist system, the multilayer perceptron (MLP), is described in the context of the recognition of continuous speech. A discriminant hidden Markov model (HMM) is defined, and it is shown how a particular MLP with contextual and extra feedback input units can be considered as a general form of such a Markov model. A link between these discriminant HMMs, trained along the Viterbi algorithm, and any other approach based on least mean square minimization of an error function (LMSE) is established. It is shown theoretically and experimentally that the outputs of the MLP (when trained along the LMSE or the entropy criterion) approximate the probability distribution over output classes conditioned on the input, i.e. the maximum a posteriori probabilities. Results of a series of speech recognition experiments are reported. The possibility of embedding MLP into HMM is described. Relations with other recurrent networks are also explained.>
Hervé Bourlard, Christian Wellekens
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Speech dynamics and recurrent neural networks
abstract
Recently, connectionist models have been recognized as an interesting alternative tool to hidden Markov models for speech recognition. Their main property lies in their combination of good discriminating power and the ability to capture input-output relations. They have also been proved useful in dealing with statistical data. However, the serial aspect remains difficult to handle in that kind of model, and several authors have proposed original architectures to deal with this problem. This study establishes links among them and compares their respective advantages. Relations with hidden Markov models are explained.>
Hervé Bourlard, Christian Wellekens
ICASSP1
1989 A Continuous Speech Recognition System Embedding MLP into HMM
Hervé Bourlard, Nelson Morgan
NIPS1
1989 Generalization and Parameter Estimation in Feedforward Netws: Some Experiments
Nelson Morgan, Hervé Bourlard
NIPS2
1988 Links Between Markov Models and Multilayer Perceptrons
Hervé Bourlard, Christian Wellekens
NIPS1
1985 Speaker dependent connected speech recognition via phonetic Markov models
abstract
In this paper, a method for speaker dependent connected speech recognition based on phonemic units is described. In this recognition system, each phoneme is characterized by a very simple 3-state Hidden Markov Model (HMM) which is trained on connected speech by a Viterbi algorithm. Each state has associated with it a continuous (Gaussian) or discrete probability density function (pdf). With the phonemic models so obtained, the recognition is then performed either directly at word level (by the reconstruction of reference words from the models of the constituting phonemes) or via a phonemic labelling. Good results are obtained as well with a German ten digit vocabulary (20 phonemes) as with a French 80 word vocabulary (36 phonemes).
Hervé Bourlard, Yves G. Kamp, Christian Wellekens
ICASSP1
1984 Connected digit recognition using vector quantization
abstract
The principles of classification applied to the representation of the words in a vocabulary lead to the clustering of the acoustic vectors into prototype vectors. For a small number of prototypes, recognition scores comparable to those observed with unclustered vocabularies are obtained with a highly reduced computation time. Two different forms (deterministic and stochastic) of the single-level recognition method for concatenated words are described and the improvements obtained by vector quantization are put into evidence. The use of prototypes in the training phase of the finite stochastic automata representing a vocabulary word is also described.
Hervé Bourlard, Christian Wellekens, Hermann Ney
ICASSP1