Jean-Luc Gauvain

dblp:79/1555 · DBLP profile ↗
← Back
162ranked-venue papers
32as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 142 · 27 first-author · 1 since 2021Artificial intelligence and machine learning · 88 · 16 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Speech recognition and synthesis · 65% Language models and text generation · 17% Information extraction and text analysis · 8%
Computer graphics and multimedia
3 papers
Audio and music processing · 80% Image and video processing · 10% Multimedia analysis and retrieval · 10%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speech processing
0.312018
Optimization of RNN-Based Speech Activity Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Audio and music processing › speech processing
voice activity detection
0.312018
Optimization of RNN-Based Speech Activity Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.222013
Structured Output Layer Neural Network Language Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2013
Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Language models and text generation
neural language model
0.212013
Structured Output Layer Neural Network Language Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
speaker adaptation
0.112010
Comparison of Speaker Adaptation Methods as Feature Extraction for SVM-Based Speaker Recognition · IEEE Trans. Speech Audio Process. 2010
Natural language and speech › Speech recognition and synthesis
speaker recognition
0.112010
Comparison of Speaker Adaptation Methods as Feature Extraction for SVM-Based Speaker Recognition · IEEE Trans. Speech Audio Process. 2010
Natural language and speech › Speech recognition and synthesis
front-end processing
0.112018
Optimization of RNN-Based Speech Activity Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Natural language and speech › Speech recognition and synthesis › speech analysis
voice activity detection
0.112018
Optimization of RNN-Based Speech Activity Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Multimedia analysis and retrieval › multimedia browsing
video browsing
0.112009
VoxaleadNews: robust automatic segmentation of video into browsable content · ACM Multimedia 2009
Image and video processing
video segmentation
0.112009
VoxaleadNews: robust automatic segmentation of video into browsable content · ACM Multimedia 2009
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › continuous speech recognition › large vocabulary continuous speech recognition
broadcast news transcription
0.122006
Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system · IEEE Trans. Speech Audio Process. 2006
Processing Broadcast Audio for Information Access · ACL 2001
Natural language and speech › Language models and text generation › language modeling
continuous space language models
0.112006
Continuous Space Language Models for Statistical Machine Translation · ACL 2006
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › spontaneous speech recognition
conversational speech recognition
0.112006
Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Machine translation
statistical machine translation
0.112006
Continuous Space Language Models for Statistical Machine Translation · ACL 2006
Audio and music processing
speaker diarization
0.112006
Multistage speaker diarization of broadcast news · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Information extraction and text analysis › lexical semantics
word clustering
0.012013
Structured Output Layer Neural Network Language Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2013
Machine learning › Kernel, tree and ensemble methods
support vector machine
0.012010
Comparison of Speaker Adaptation Methods as Feature Extraction for SVM-Based Speaker Recognition · IEEE Trans. Speech Audio Process. 2010
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › continuous speech recognition
large vocabulary continuous speech recognition
0.012000
Large-vocabulary continuous speech recognition: advances and applications · Proc. IEEE 2000
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.012006
Continuous Space Language Models for Statistical Machine Translation · ACL 2006
Audio and music processing › speaker recognition
speaker identification
0.012006
Multistage speaker diarization of broadcast news · IEEE Trans. Speech Audio Process. 2006
Natural language and speech › Speech recognition and synthesis
acoustic modeling
0.011994
Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains · IEEE Trans. Speech Audio Process. 1994
Natural language and speech › Speech recognition and synthesis › acoustic modeling
hidden markov model parameter estimation
0.011994
Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains · IEEE Trans. Speech Audio Process. 1994
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference
0.011994
Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains · IEEE Trans. Speech Audio Process. 1994
Information retrieval
search engines
0.012001
Processing Broadcast Audio for Information Access · ACL 2001
Machine learning › Transfer learning and domain adaptation
model adaptation
0.011994
Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains · IEEE Trans. Speech Audio Process. 1994

Methods — techniques the papers use, named apart from their topics

quantum-behaved particle swarm optimization · 0.7long short-term memory · 0.7coordinated-gate LSTM · 0.7word clustering · 0.2softmax layer · 0.2continuous word representation · 0.2acoustic modeling · 0.1support vector machine · 0.1maximum likelihood linear regression · 0.1maximum a posteriori adaptation · 0.1gaussian mixture model · 0.1bayesian information criterion · 0.1speech recognition · 0.0language modeling · 0.0
YearPublicationVenuePosition
2021 Modeling the Effect of Military Oxygen Masks on Speech Characteristics
abstract
International audience
Benjamin Elie, Jodie Gauvain, Jean-Luc Gauvain, Lori Lamel
Interspeech3
2019 Challenges in Audio Processing of Terrorist-Related Data
Jodie Gauvain, Lori Lamel, Viet Bac Le, Julien Despres, Jean-Luc Gauvain, Abdelkhalek Messaoudi, Bianca Vieru-Dimulescu, Waad Ben Kheder
MMM (2)5
2018 Conversational telephone speech recognition for Lithuanian
Rasa Lileikyte, Lori Lamel, Jean-Luc Gauvain, Arseniy Gorin
Comput. Speech Lang.3
2018 Optimization of RNN-Based Speech Activity Detection
abstract
Speech activity detection (SAD) is an essential component of automatic speech recognition systems impacting the overall system performance. This paper investigates an optimization process for recurrent neural network (RNN) based SAD. This process optimizes all system parameters including those used for feature extraction, the NN weights, and the back-end parameters. Three cost functions are considered for SAD optimization: the frame error rate, the NIST detection cost function, and the word error rate of a downstream speech recognizer. Different types of RNN models and optimization methods are investigated. Three types of RNNs are compared: a basic RNN, long short-term memory (LSTM) network with peepholes, and a coordinated-gate LSTM (CG-LSTM) network introduced by Gelly and Gauvain. Well suited for nondifferentiable optimization problems, quantum-behaved particle swarm optimization is used to optimize feature extraction and posterior smoothing, as well as for the initial training of the neural networks. Experimental SAD results are reported on the NIST 2015 SAD evaluation data as well as REPERE and AMI meeting corpora. Speech recognition results are reported on the OpenKWS'13 test data. For all tasks and conditions, the proposed optimization method significantly improves the SAD performance and among all the tested SAD methods the CG-LSTM model gives the best results.
Gregory Gelly, Jean-Luc Gauvain
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 An investigation into language model data augmentation for low-resourced STT and KWS
abstract
This paper reports on investigations using two techniques for language model text data augmentation for low-resourced automatic speech recognition and keyword search. Lowresourced languages are characterized by limited training materials, which typically results in high out-of-vocabulary (OOV) rates and poor language model estimates. One technique makes use of recurrent neural networks (RNNs) using word or subword units. Word-based RNNs keep the same system vocabulary, so they cannot reduce the OOV, whereas subword units can reduce the OOV but generate many false combinations. A complementary technique is based on automatic machine translation, which requires parallel texts and is able to add words to the vocabulary. These methods were assessed on 10 languages in the context of the Babel program and NIST OpenKWS evaluation. Although improvements vary across languages with both methods, small gains were generally observed in terms of word error rate reduction and improved keyword search performance.
Guangpu Huang, Thiago Fraga-Silva, Lori Lamel, Jean-Luc Gauvain, Arseniy Gorin, Antoine Laurent, Rasa Lileikyte, Abdel Messouadi
ICASSP4
2017 Effective keyword search for low-resourced conversational speech
abstract
In this paper we aim to enhance keyword search for conversational telephone speech under low-resourced conditions. Two techniques to improve the detection of out-of-vocabulary keywords are assessed in this study: using extra text resources to augment the lexicon and language model, and via subword units for keyword search. Two approaches for data augmentation are explored to extend the limited amount of transcribed conversational speech: using conversational-like Web data and texts generated by recurrent neural networks. Contrastive comparisons of subword-based systems are performed to evaluate the benefits of multiple subword decodings and single decoding. Keyword search results are reported for all the techniques, but only some improve performance. Results are reported for the Mongolian and Igbo languages using data from the 2016 Babel program.
Rasa Lileikyte, Thiago Fraga-Silva, Lori Lamel, Jean-Luc Gauvain, Antoine Laurent, Guangpu Huang
ICASSP4
2017 Spoken Language Identification Using LSTM-Based Angular Proximity
Gregory Gelly, Jean-Luc Gauvain
INTERSPEECH2
2016 Machine translation based data augmentation for Cantonese keyword spotting
abstract
This paper presents a method to improve a language model for a limited-resourced language using statistical machine translation from a related language to generate data for the target language. In this work, the machine translation model is trained on a corpus of parallel Mandarin-Cantonese subtitles and used to translate a large set of Mandarin conversational telephone transcripts to Cantonese, which has limited resources. The translated transcripts are used to train a more robust language model for speech recognition and for keyword search in Cantonese conversational telephone speech. This method enables the keyword search system to detect 1.5 times more out-of-vocabulary words, and achieve 1.7% absolute improvement on actual term-weighted value.
Guangpu Huang, Arseniy Gorin, Jean-Luc Gauvain, Lori Lamel
ICASSP3
2016 Investigating techniques for low resource conversational speech recognition
abstract
In this paper we investigate various techniques in order to build effective speech to text (STT) and keyword search (KWS) systems for low resource conversational speech. Subword decoding and graphemic mappings were assessed in order to detect out-of-vocabulary keywords. To deal with the limited amount of transcribed data, semi-supervised training and data selection methods were investigated. Robust acoustic features produced via data augmentation were evaluated for acoustic modeling. For language modeling, automatically retrieved conversational-like Webdata was used, as well as neural network based models. We report STT improvements with all the techniques, but interestingly only some improve KWS performance. Results are reported for the Swahili language in the context of the 2015 OpenKWS Evaluation.
Antoine Laurent, Thiago Fraga-Silva, Lori Lamel, Jean-Luc Gauvain
ICASSP4
2016 A Divide-and-Conquer Approach for Language Identification Based on Recurrent Neural Networks
Gregory Gelly, Jean-Luc Gauvain, Viet Bac Le, Abdelkhalek Messaoudi
INTERSPEECH2
2016 Language Model Data Augmentation for Keyword Spotting in Low-Resourced Training Conditions
abstract
International audience
Arseniy Gorin, Rasa Lileikyte, Guangpu Huang, Lori Lamel, Jean-Luc Gauvain, Antoine Laurent
INTERSPEECH5
2015 Improving data selection for low-resource STT and KWS
abstract
This paper extends recent research on training data selection for speech transcription and keyword spotting system development. Selection techniques were explored in the context of the IARPA-Babel Active Learning (AL) task for 6 languages. Different selection criteria were considered with the goal of improving over a system built using a pre-defined 3-hour training data set. Four variants of the entropy-based criterion were explored: words, triphones, phones as well as the use of HMM-states previously introduced in [4]. The influence of the number of HMM-states was assessed as well as whether automatic or manual reference transcripts were used. The combination of selection criteria was investigated, and a novel multi-stage selection method proposed. This method was also assessed using larger data sets than were permitted in the Babel AL task. Results are reported for the 6 languages. The multi-stage selection was also applied to the surprise language (Swahili) in the NIST OpenKWS 2015 evaluation.
Thiago Fraga-Silva, Antoine Laurent, Jean-Luc Gauvain, Lori Lamel, Viet Bac Le, Abdelkhalek Messaoudi
ASRU3
2015 Active learning based data selection for limited resource STT and KWS
abstract
International audience
Thiago Fraga-Silva, Jean-Luc Gauvain, Lori Lamel, Antoine Laurent, Viet Bac Le, Abdelkhalek Messaoudi
INTERSPEECH2
2015 Minimum word error training of RNN-based voice activity detection
Gregory Gelly, Jean-Luc Gauvain
INTERSPEECH2
2015 Lexical speaker identification in TV shows
Anindya Roy, Hervé Bredin, William Hartmann, Viet Bac Le, Claude Barras, Jean-Luc Gauvain
Multim. Tools Appl.6
2014 Comparing decoding strategies for subword-based keyword spotting in low-resourced languages
abstract
For languages with limited training resources, out-of-vocabulary (OOV) words are a significant problem, both for transcription and keyword spotting. This paper investigates the use of subword lexical units for keyword spotting. Three strate-gies for using the sub-word units are explored: 1) converting word-based lattices to subword lattices after decoding, 2) per-forming a separate decoding for each subword type, and 3) a single decoding using all possible subword units. In these ex-periments, the best performance is achieved by carrying out a separate decoding for each subword type. Further gains are at-tained through system combination. We also find that ignor-
William Hartmann, Viet Bac Le, Abdelkhalek Messaoudi, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH5
2014 Developing STT and KWS systems using limited language resources
abstract
This paper presents recent progress in developing speech-to-text (STT) and keyword spotting (KWS) systems for the 2014 IARPA-Babel evaluation. Systems have been developed for the limited language pack condition for four of the five de-velopment languages in this program phase: Assamese, Ben-gali, Haitian Creole and Zulu. The systems have several novel characteristics that support rapid development of KWS systems. On the STT side different acoustic units are explored based on phonemic or graphemic representations, and system combina-tion is used to improve STT performance. The acoustic models are trained on only 10 hours of speech data with manual tran-scriptions, completed with unsupervised training on additional untranscribed data. Both word and subword units (morphologi-cally decomposed, syllables, phonemes) are used for KWS. The KWS systems are based on the multi-hypotheses produced by a consensus network decoding or searching word lattices. The word error rates of the individual STT systems are on the or-der of 50-60%, and the KWS systems obtain Maximum Term Weighted Values ranging from 30-45 % for all keywords (in-vocabulary and out-of-vocabulary (OOV)). Sub-word units are shown to be successful at locating some of the OOV keywords, and system combination improves system performance. Index Terms: STT, KWS, semi-supervised training, lattice, consensus network, sub-word lexical units, Morfessor,
Viet Bac Le, Lori Lamel, Abdelkhalek Messaoudi, William Hartmann, Jean-Luc Gauvain, Cécile Woehrling, Julien Despres, Anindya Roy
INTERSPEECH5
2013 Acoustic unit discovery and pronunciation generation from a grapheme-based lexicon
abstract
We present a framework for discovering acoustic units and generating an associated pronunciation lexicon from an initial grapheme-based recognition system. Our approach consists of two distinct contributions. First, context-dependent grapheme models are clustered using a spectral clustering approach to create a set of phone-like acoustic units. Next, we transform the pronunciation lexicon using a statistical machine translation-based approach. Pronunciation hypotheses generated from a decoding of the training set are used to create a phrase-based translation table. We propose a novel method for scoring the phrase-based rules that significantly improves the output of the transformation process. Results on an English language dataset demonstrate the combined methods provide a 13% relative reduction in word error rate compared to a baseline grapheme-based system. Our approach could potentially be applied to low-resource languages without existing lexicons, such as in the Babel project.
William Hartmann, Anindya Roy, Lori Lamel, Jean-Luc Gauvain
ASRU4
2013 Rapid development of a Latvian speech-to-text system
abstract
This paper describes the development of a Latvian speech-to-text (STT) system at LIMSI within the Quaero project. One of the aims of the speech processing activities in the Quaero project is to cover all official European languages. However, for some of the languages only very limited, if any, training resources are available via corpora agencies such as LDC and ELRA. The aim of this study was to show the way, taking Latvian as example, an STT system can be rapidly developed without any transcribed training data. Following the scheme proposed in this paper, the Latvian STT system was developed in about a month and obtained a word error rate of 20% on broadcast news and conversation data in the Quaero 2012 evaluation campaign.
Ilya Oparin, Lori Lamel, Jean-Luc Gauvain
ICASSP3
2013 Comparison of feedforward and recurrent neural network language models
abstract
Research on language modeling for speech recognition has increasingly focused on the application of neural networks. Two competing concepts have been developed: On the one hand, feedforward neural networks representing an n-gram approach, on the other hand recurrent neural networks that may learn context dependencies spanning more than a fixed number of predecessor words. To the best of our knowledge, no comparison has been carried out between feedforward and state-of-the-art recurrent networks when applied to speech recognition. This paper analyzes this aspect in detail on a well-tuned French speech recognition task. In addition, we propose a simple and efficient method to normalize language model probabilities across different vocabularies, and we show how to speed up training of recurrent neural networks by parallelization.
Martin Sundermeyer, Ilya Oparin, Jean-Luc Gauvain, B. Freiberg, Ralf Schlüter, Hermann Ney
ICASSP3
2013 Interpolation of acoustic models for speech recognition
abstract
International audience
Thiago Fraga-Silva, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2013 Some issues affecting the transcription of Hungarian broadcast audio
abstract
International audience
Anindya Roy, Lori Lamel, Thiago Fraga-Silva, Jean-Luc Gauvain, Ilya Oparin
INTERSPEECH4
2013 Structured Output Layer Neural Network Language Models for Speech Recognition
abstract
This paper extends a novel neural network language model (NNLM) which relies on word clustering to structure the output vocabulary: Structured OUtput Layer (SOUL) NNLM. This model is able to handle arbitrarily-sized vocabularies, hence dispensing with the need for shortlists that are commonly used in NNLMs. Several softmax layers replace the standard output layer in this model. The output structure depends on the word clustering which is based on the continuous word representation determined by the NNLM. Mandarin and Arabic data are used to evaluate the SOUL NNLM accuracy via speech-to-text experiments. Well tuned speech-to-text systems (with error rates around 10%) serve as the baselines. The SOUL model achieves consistent improvements over a classical shortlist NNLM both in terms of perplexity and recognition accuracy for these two languages that are quite different in terms of their internal structure and recognition vocabulary size. An enhanced training scheme is proposed that allows more data to be used at each training iteration of the neural network.
Hai Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon
IEEE Trans. Speech Audio Process.4
2012 Performance analysis of Neural Networks in combination with n-gram language models
abstract
Neural Network language models (NNLMs) have recently become an important complement to conventional n-gram language models (LMs) in speech-to-text systems. However, little is known about the behavior of NNLMs. The analysis presented in this paper aims to understand which types of events are better modeled by NNLMs as compared to n-gram LMs, in what cases improvements are most substantial and why this is the case. Such an analysis is important to take further benefit from NNLMs used in combination with conventional n-gram models. The analysis is carried out for different types of neural network (feed-forward and recurrent) LMs. The results showing for which type of events NNLMs provide better probability estimates are validated on two setups that are different in their size and the degree of data homogeneity.
Ilya Oparin, Martin Sundermeyer, Hermann Ney, Jean-Luc Gauvain
ICASSP4
2012 Phonotactic Language Recognition Using MLP Features
abstract
International audience
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2011 Lattice-based unsupervised acoustic model training
abstract
Unsupervised acoustic model training has been successfully used to improve the performance of automatic speech recognition systems when only a small amount of manually transcribed data is available for the target domain. The most common approach is use automatic transcriptions to guide acoustic model estimation. However, since the best recognition hypotheses are known to contain errors, we propose to consider multiple transcription hypotheses during training. The idea is that the EM process can benefit from the estimated posterior probabilities of the hypotheses to converge to a better solution. The proposed unsupervised training method is based on lattices. Lattice-based training gives a relative improvement of 2.2% over 1-best training on a Broadcast News transcription task and converges faster with the iterative incremental training.
Thiago Fraga-Silva, Jean-Luc Gauvain, Lori Lamel
ICASSP2
2011 Improved models for Mandarin speech-to-text transcription
abstract
This paper describes recent advances at LIMSI in Mandarin Chinese speech-to-text transcription. A number of novel approaches were introduced in the different system components. The acoustic models are trained on over 1600 hours of audio data from a range of sources, and include pitch and MLP features. N-gram and neural network language models are trained on very large corpora, over 3 billion words of texts; and LM adaptation was explored at different adaptation levels: per show, per snippet, or per speaker cluster. Character-based consensus decoding was found to outperform word-based consensus decoding for Mandarin. The improved system reduces the relative character error rate (CER) by about 10% on previous GALE development and evaluation data sets, obtaining a CER of 9.2% on the P4 broadcast news and broadcast conversation evaluation data.
Lori Lamel, Jean-Luc Gauvain, Viet Bac Le, Ilya Oparin, Sha Meng
ICASSP2
2011 Structured Output Layer neural network language model
abstract
This paper introduces a new neural network language model (NNLM) based on word clustering to structure the output vocabulary: Structured Output Layer NNLM. This model is able to handle vocabularies of arbitrary size, hence dispensing with the design of short-lists that are commonly used in NNLMs. Several softmax layers replace the standard output layer in this model. The output structure depends on the word clustering which uses the continuous word representation induced by a NNLM. The GALE Mandarin data was used to carry out the speech-to-text experiments and evaluate the NNLMs. On this data the well tuned baseline system has a character error rate under 10%. Our model achieves consistent improvements over the combination of an n-gram model and classical short-list NNLMs both in terms of perplexity and recognition accuracy.
Hai Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon
ICASSP4
2011 Large Vocabulary SOUL Neural Network Language Models
abstract
International audience
Hai Son Le, Ilya Oparin, Abdelkhalek Messaoudi, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon
INTERSPEECH5
2011 Genre Categorization and Modeling for Broadcast Speech Transcription
abstract
Broadcast News (BN) speech recognition transcription has attracted research due to the challenges of the task since the mid 1990’s. More recently, research has been moving towards more spontaneous broadcast data, commonly called Broadcast Conversation (BC) speech. Considering the large style difference between BN and BC genres, specific modeling of genres should intuitively result in improved system performance. In this paper BNand BC-style speech recognition has been explored by designing genre-specific systems. In order to separate the training data, an automatic genre categorization with two novel features is proposed. Experiments showed that automatic categorization of genre labels of the training data compared favorably to the original manually specified genre labels provided with corpora. When test data sets were classified into BN or BC genres and tested by the corresponding genre-specific speech recognition systems, modest but consistent error reductions were achieved compared to the baseline genre-independent systems.
Lori Lamel, Jean-Luc Gauvain
INTERSPEECH3
2010 Multi-style MLP features for BN transcription
abstract
It has become common practice to adapt acoustic models to specific-conditions (gender, accent, bandwidth) in order to improve the performance of speech-to-text (STT) transcription systems. With the growing interest in the use of discriminative features produced by a multi layer perceptron (MLP) in such systems, the question arise of whether it is necessary to specialize the MLP to particular conditions, and if so, how to incorporate the condition-specific MLP features in the system. This paper explores three approaches (adaptation, full training, and feature merging) to use condition-specific MLP features in a state-of-the-art BN STT system for French. The third approach without condition-specific adaptation was found to outperform the original models with condition-specific adaptation, and was found to perform almost as well as full training of multiple condition-specific HMMs.
Viet Bac Le, Lori Lamel, Jean-Luc Gauvain
ICASSP3
2010 Improved n-gram phonotactic models for language recognition
abstract
This paper investigates various techniques to improve the estimation of n-gram phonotactic models for language recognition using single-best phone transcriptions and phone lattices. More precisely, we first report on the impact of the so-called acoustic scale factor on the system accuracy when using latticebased training, and then we report on the use of n-gram cutoff and entropy pruning techniques. Several system configurations are explored, such as the use of context-independent and context-dependent phone models, the use of single-best phone hypotheses versus phone lattices, and the use of various n-gram orders. Experiments are conducted using the LRE 2007 evaluation data and the results are reported using the a posteriori EER. The results show that the impact of these techniques on the system accuracy is highly dependent on the training conditions and that careful optimization can lead to performance improvements.
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2010 Automatic speech recognition of multiple accented English data
abstract
Accent variability is an important factor in speech that can sig-nificantly degrade automatic speech recognition performance. We investigate the effect of multiple accents on an English broadcast news recognition system. A multi-accented English corpus is used for the task, including broadcast news segments from 6 different geographic regions: US, Great Britain, Aus-tralia, North Africa, Middle East and India. There is signifi-cant performance degradation of a baseline system trained on only US data when confronted with shows from other regions. The results improve significantly when data from all the regions are included for accent-independent acoustic model training. Further improvements are achieved when MAP-adapted accent-dependent models are used in conjunction with a GMM accent classifier. Index Terms: accented speech recognition, accent adaptation 1.
Dimitra Vergyri, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH3
2010 Comparison of Speaker Adaptation Methods as Feature Extraction for SVM-Based Speaker Recognition
abstract
In the last years the speaker recognition field has made extensive use of speaker adaptation techniques. Adaptation allows speaker model parameters to be estimated using less speech data than needed for maximum-likelihood (ML) training. The maximuma posteriori(MAP) and maximum-likelihood linear regression (MLLR) techniques have typically been used for adaptation. Recently, MAP and MLLR adaptation have been incorporated in the feature extraction stage of support vector machine (SVM)-based speaker recognition systems. Two approaches to feature extraction use a SVM to classify either the MAP-adapted Gaussian mean vector parameters (GSV-SVM) or the MLLR transform coefficients (MLLR-SVM). In this paper, we provide an experimental analysis of the GSV-SVM and MLLR-SVM approaches. We largely focus on the latter by exploring constrained and unconstrained transforms and different choices of the acoustic model. A channel-compensated front-end is used to prevent the MLLR transforms to adapt to channel components in the speech data. Additional acoustic models were trained using speaker adaptive training (SAT) to better estimate the speaker MLLR transforms. We provide results on the NIST 2005 and 2006 Speaker Recognition Evaluation (SRE) data and fusion results on the SRE 2006 data. The results show that using the compensated front-end, SAT models and multiple regression classes bring major performance improvements.
Marc Ferras, Cheung-Chi Leung, Claude Barras, Jean-Luc Gauvain
IEEE Trans. Speech Audio Process.4
2009 Gaussian Backend design for open-set language detection
abstract
This paper proposes a new approach to the challenging open-set language detection task. Most state-of-the-art approaches make use of data sources with several out-of-set languages to model such languages. In the proposed approach, no additional data from out-ofset languages is required, only date from the target languages is used. Experiments are conducted using the LRE-05 and the LRE-07 evaluation data sets with the 30s condition. A Cavgof 4.5% and 3.4% is obtained on these data set, respectively. These results are comparable with other reported results.
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
ICASSP2
2009 Lattice-based MLLR for speaker recognition
abstract
Maximum-Likelihod Linear Regression (MLLR) transform coefficients have shown to be useful features for text-independent speaker recognition systems. These use MLLR coefficients computed on a Large Vocabulary Continuous Speech Recognition System (LVCSR) as features and Support Vector machines(SVM) classification. However, performance is limited by transcripts, which are often erroneous with high word error rates (WER) for spontaneous telephone speech applications. In this paper, we propose using lattice-based MLLR to overcome this issue. Using wordlattices instead of 1-best hypotheses, more hypotheses can be considered for MLLR estimation and, thus, better models are more likely to be used. As opposed to standard MLLR, language model probabilities are taken into account as well. We show how systems using lattice MLLR outperform standard MLLR systems in the Speaker Recognition Evaluation (SRE) 2006. Comparison to other standard acoustic systems is provided as well.
Marc Ferras, Claude Barras, Jean-Luc Gauvain
ICASSP3
2009 Modeling characters versuswords for mandarin speech recognition
abstract
Word based models are widely used in speech recognition since they typically perform well. However, the question of whether it is better to use a word-based or a character-based model warrants being for the Mandarin Chinese language. Since Chinese is written without any spaces or word delimiters, a word segmentation algorithm is applied in a pre-processing step prior to training a word-based language model. Chinese characters carry meaning and speakers are free to combine characters to construct new words. This suggests that character information can also be useful in communication. This paper explores both word-based and character-based models, and their complementarity. Although word-based modeling is found to outperform character-based modeling, increasing the vocabulary size from 56 k to 160 k words did not lead to a gain in performance. Results are reported for the Gale Mandarin speech-to-text task.
Lori Lamel, Jean-Luc Gauvain
ICASSP3
2009 Language score calibration using adapted Gaussian back-end
abstract
Generative Gaussian back-end and discriminative logistic regression are the most used approaches for language score fusion and calibration. Combination of these two approaches can significantly improve the performance. This paper proposes the use of an adapted Gaussian back-end, where the mean of the language-dependent Gaussian is adapted from the mean of a language-specific background Gaussian via maximum a posteriori estimation algorithm. Experiments are conducted using the LRE-07 evaluation data. Compared to the conventional Gaussian back-end approach for a closed set task, relative improvements in the Cavg of 50%, 17% and 4.2% are obtained on the 30s, 10s and 3s conditions, respectively. Besides this, the estimated scores are better calibrated. A combination with logistic regression results in a system with the best calibrated scores. Index Terms: Language recognition, Gaussian back-end, Adaptation
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2009 Modeling northern and southern varieties of dutch for STT
abstract
This paper describes how the Northern (NL) and Southern (VL) varieties of Dutch are modeled in the joint LIMSIVecsys Research speech-to-text transcription systems for broadcast news (BN) and conversational telephone speech (CTS). Using the Spoken Dutch Corpus resources (CGN), systems were developed and evaluated in the 2008 N-Best benchmark. Modeling techniques that are used in our systems for other languages were found to be effective for the Dutch language, however it was also found to be important to have acoustic and language models, and statistical pronunciation generation rules adapted to each variety. This was in particular true for the MLP features which were only effective when trained separately for Dutch and Flemish. The joint submissions obtained the lowest WERs in the benchmark by a significant margin. Index Terms: speech recognition, Dutch, Flemish, CGN, Nbest, broadcast news, conversational telephone speech, MLP.
Julien Despres, Petr Fousek, Jean-Luc Gauvain, Sandrine Gay, Yvan Josse, Lori Lamel, Abdelkhalek Messaoudi
INTERSPEECH3
2009 VoxaleadNews: robust automatic segmentation of video into browsable content
abstract
We present an interface to video and audio podcasts that extracts semantics from the speech content, and packages the extracted information in a variety of navigation tools. The user can jump to the relevant sections and browse from relevant section to relevant section. This interface is related to the Yahoo! Challenge: Robust Automatic Segmentation of Video According to Narrative Themes.
Julien Law-To, Gregory Grefenstette, Jean-Luc Gauvain
ACM Multimedia3
2009 Automatic Speech-to-Text Transcription in Arabic
abstract
The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Being a highly inflected language, the Arabic language has a very large lexical variety and typically with several possible (generally semantically linked) vowelizations for each written form. This article summarizes research carried out over the last few years on speech-to-text transcription of broadcast data in Arabic. The initial research was oriented toward processing of broadcast news data in Modern Standard Arabic, and has since been extended to address a larger variety of broadcast data, which as a consequence results in the need to also be able to handle dialectal speech. While standard techniques in speech recognition have been shown to apply well to the Arabic language, taking into account language specificities help to significantly improve system performance.
Lori Lamel, Abdelkhalek Messaoudi, Jean-Luc Gauvain
ACM Trans. Asian Lang. Inf. Process.3
2008 Context-dependent phone models and models adaptation for phonotactic language recognition
abstract
The performance of a PPRLM language recognition system depends on the quality and the consistency of phone decoders. To improve the performance of the decoders, this paper investigates the use of context-dependent instead of contextindependent phone models, and the use of CMLLR for model adaptation. This paper also discusses several improvements to the LIMSI 2007 NIST LRE system, including the use of a 4gram language model, score calibration and fusion using the FoCalMulti-class toolkit (with large development data) and better decoding parameters such as phone insertion penalty. The improved system is evaluated on the NIST LRE-2005 and the LRE-2007 evaluation data sets. Despite its simplicity, the system achieves for the 30s condition a Cavg of 2.4% and 1.6% on these data sets, respectively.
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2008 Transcribing broadcast data using MLP features
abstract
This paper describes incorporating discriminative features from a multi layer perceptron (MLP) into a state-of-the-art Arabic broadcast data transcription system based on cepstral features. The MLP features are based on a recently proposed Bottle-Neck architecture with long-term warped LPTRAP speech representation at the input. It is shown that the previously reported improvements on a development Arabic transcription system carry through to a full system at a state-ofthe-art level. SAT, CMLLR and MLLR adaptation techniques are shown to be useful for both MLP and combined features, though to a lesser degree than for PLPs. Without adaptation, MLP features obtain superior performance to cepstral features in all test conditions, and with adaptation both feature sets give comparable results. Combining the features, either by feature concatenation or system hypotheses, gives significant gains. Gains from MMI model training seem to be additive to the gain coming from discriminative MLP features.
Petr Fousek, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH3
2008 Investigating morphological decomposition for transcription of Arabic broadcast news and broadcast conversation data
abstract
One of the challenges of Arabic speech recognition is to deal with the huge lexical variety. Morphological decomposition has been proposed to address this problem by increasing lexical coverage, thereby reducing errors that are due to words that are unknown to the system. In our previous attempts to develop an Arabic speech-to-text (STT) transcription system with morphological decomposition, an increase in word error rate of about 2% absolute was observed relative to a comparable word based system. Based on an error analysis and a comparison of our approach with that of other sites, two modifications were made. The first modification was to not decompose the most frequent words; and the second to not decompose the prefix ’Al’ for words starting with a solar consonant since due to assimilation with the following consonant, deletion of the prefix was one of the most frequent errors. Comparable recognition performance was achieved using word-based and morphologically decomposed language models, and since the errors made by the systems are different, combining the two gave a performance gain.
Lori Lamel, Abdelkhalek Messaoudi, Jean-Luc Gauvain
INTERSPEECH3
2008 Comparing prosodic models for speaker recognition
abstract
International audience
Cheung-Chi Leung, Marc Ferras, Claude Barras, Jean-Luc Gauvain
INTERSPEECH4
2008 CallSurf: Automatic Transcription, Indexing and Structuration of Call Center Conversational Speech for Knowledge Extraction and Query by Content
Martine Garnier-Rizet, Gilles Adda, Frédérik Cailliau, Jean-Luc Gauvain, Sylvie Guillemin-Lanne, Lori Lamel, Stephan Vanni, Claire Waast-Richard
LREC4
2007 Constrained MLLR for Speaker Recognition
abstract
One particularly difficult challenge for cross-channel MLLR (CMLLR) are two widely-used techniques for speaker introduced in the 2005 and 2006 NIST Speaker Recognition Evaluations, where training uses telephone speech and verification uses speech from multiple auxiliary comparable to that obtained with cepstral features. This paper describes a new feature extraction technique for speaker recognition based on CMLLR speaker adaptation which session effects through latent factor analysis (LFA) and through support vector machines (SVM). Results on the NIST operates directly on the recorded signal with noise well as in combination with two cepstral approaches such as reduction in the performance gap between telephone and auxiliary microphone data.
Marc Ferras, Cheung-Chi Leung, Claude Barras, Jean-Luc Gauvain
ICASSP (4)4
2007 Speech Recognition System Combination for Machine Translation
abstract
The majority of state-of-the-art speech recognition systems make use of system combination. The combination approaches adopted have traditionally been tuned to minimising word error rates (WERs). In recent years there has been a growing interest in taking the output from speech recognition systems in one language and translating it into another. This paper investigates the use of cross-site combination approaches in terms of both WER and impact on translation performance. In addition, the stages involved in modifying the output from a speech-to-text (STT) system to be suitable for translation are described. Two source languages, Mandarin and Arabic, are recognised and then translated using a phrase-based statistical machine translation system into English. Performance of individual systems and cross-site combination using cross-adaptation and ROVER are given. Results show that the best STT combination scheme in terms of WER is not necessarily the most appropriate when translating speech.
Mark J. F. Gales, Xunying Liu, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004, Spyridon Matsoukas, Tim Ng, Kham Nguyen, Long Nguyen 0001, Jean-Luc Gauvain, Lori Lamel, Abdelkhalek Messaoudi
ICASSP (4)10
2007 Modeling Duration via Lattice Rescoring
abstract
It is often acknowledged that HMMs do not properly model phone and word durations. In this paper phone and word duration models are used to improve the accuracy of state-of-the-art large vocabulary speech recognition systems. The duration information is integrated into the systems in a rescoring of word lattices that include phone-level segmentations. Experimental results are given for a conversational telephone speech (CTS) task in French and for the TC-Star EPPS transcription task in Spanish and English. An absolute word error rate reduction of about 0.5% is observed for the CTS task, and smaller but consistent gains are observed for the EPPS task in both languages.
Nicolas Jennequin, Jean-Luc Gauvain
ICASSP (4)2
2007 The LIMSI 2006 TC-STAR EPPS Transcription Systems
abstract
This paper describes the speech recognizers developed to transcribe European Parliament Plenary Sessions (EPPS) in English and Spanish in the 2nd TC-STAR Evaluation Campaign. The speech recognizers are state-of-the-art systems using multiple decoding passes with models (lexicon, acoustic models, language models) trained for the different transcription tasks. Compared to the LIMSI TC-STAR 2005 EPPS systems, relative word error rate reductions of about 30% have been achieved on the 2006 development data. The word error rates with the LIMSI systems on the 2006 EPPS evaluation data are 8.2% for English and 7.8% for Spanish. Experiments with cross-site adaptation and system combination are also described.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Claude Barras, Eric Bilinski, Olivier Galibert, Agusti Pujol, Holger Schwenk, Xuan Zhu 0001
ICASSP (4)2
2007 Improved machine translation of speech-to-text outputs
abstract
International audience
Daniel Déchelotte, Holger Schwenk, Gilles Adda, Jean-Luc Gauvain
INTERSPEECH4
2007 Improved acoustic modeling for transcribing Arabic broadcast data
abstract
ABSTRACT This paper summarizes our recent progress in improving theautomatic transcription of Arabic broadcast audio data, andsome efforts to address the challenges of the broadcast con-versational speech. Our efforts are aimed at improving theacoustic, pronunciation and language models taking into ac-count specificities of the Arabic language. In previous work wedemonstrated that explicit modeling of short vowels improvedrecognition performance, even when producing non-vocalizedhypotheses. In addition to modeling short vowels, consonantgemination and nunation are now explicitly modeled, alterna-tive pronunciations have been introduced to better represent di-alectical variants, and a duration model has been integrated.In order to facilitate training on Arabic audio data with non-vocalized transcripts a generic vowel model has been intro-duced. Compared with the previous system (used in the 2006GALE evaluation) the relative word error rate has been reducedby over 10%. Index Terms – Speech recognition, Arabic, broadcast news,broadcast conversations
Lori Lamel, Abdelkhalek Messaoudi, Jean-Luc Gauvain
INTERSPEECH3
2006 Continuous Space Language Models for Statistical Machine Translation
Holger Schwenk, Daniel Déchelotte, Jean-Luc Gauvain
ACL3
2006 Discriminant Initialization for Factor Analyzed HMM Training
abstract
Factor analysis has been recently used to model the covariance of the feature vector in speech recognition systems. Maximum likelihood estimation of the parameters of factor analyzed HMMs (FAHMMs) is usually done via the EM algorithm, meaning that initial estimates of the model parameters is a key issue. In this paper we report on experiments showing some evidence that the use of a discriminative criterion to initialize the FAHMM maximum likelihood parameter estimation can be effective. The proposed approach relies on the estimation of a discriminant linear transformation to provide initial values for the factor loading matrices, as well as appropriate initializations for the other model parameters. Speech recognition experiments were carried out on the Wall Street Journal LVCSR task with a 65k vocabulary. Contrastive results are reported with various model sizes using discriminant and non discriminant initialization
Fabrice Lefèvre, Jean-Luc Gauvain
ICASSP (1)2
2006 Arabic Broadcast News Transcription Using a One Million Word Vocalized Vocabulary
abstract
Recently it has been shown that modeling short vowels in Arabic can significantly improve performance even when producing a non-vocalized transcript. Since Arabic texts and audio transcripts are almost exclusively non-vocalized, the training methods have to overcome this missing data problem. For the acoustic models the procedure was bootstrapped with manually vocalized data and extended with semi-automatically vocalized data. In order to also capture the vowel information in the language model, a vocalized 4-gram language model trained on the audio transcripts was interpolated with the original 4-gram model trained on the (non-vocalized) written texts. Another challenge of the Arabic language is its large lexical variety. The out-of-vocabulary rate with a 65k word vocabulary is in the range of 4-8% (compared to under 1% for English). To address this problem a vocalized vocabulary containing over 1 million vocalized words, grouped into 200k word classes is used. This reduces the out-of-vocabulary rate to about 2%. The extended vocabulary and vocalized language model trained on the manually annotated data give a 1.2% absolute word error reduction on the DARPA RT04 development data. However, including the automatically vocalized transcripts in the language model reduces performance indicating that automatic vocalization needs to be improved
Abdelkhalek Messaoudi, Jean-Luc Gauvain, Lori Lamel
ICASSP (1)2
2006 Discriminative Classifiers for Language Recognition
abstract
Most language recognition systems consist of a cascade of three stages: (1) tokenizers that produce parallel phone streams, (2) phonotactic models that score the match between each phone stream and the phonotactic constraints in the target language, and (3) a final stage that combines the scores from the parallel streams appropriately [1]. This paper reports a series of contrastive experiments to assess the impact of replacing the second and third stages with large-margin discriminative classifiers. In addition, it investigates how sounds that are not represented in the tokenizers of the first stage can be approximated with composite units that utilize cross-stream dependencies obtained via multi-string alignments. This leads to a discriminative framework that can potentially incorporate a richer set of features such as prosodic and lexical cues. Experiments are reported on the NIST LRE 1996 and 2003 task and the results show that the new techniques give substantial gains over a competitive PPRLM baseline.
Izhak Shafran, Jean-Luc Gauvain
ICASSP (1)3
2006 Multistage speaker diarization of broadcast news
abstract
This paper describes recent advances in speaker diarization with a multistage segmentation and clustering system, which incorporates a speaker identification step. This system builds upon the baseline audio partitioner used in the LIMSI broadcast news transcription system. The baseline partitioner provides a high cluster purity, but has a tendency to split data from speakers with a large quantity of data into several segment clusters. Several improvements to the baseline system have been made. First, the iterative Gaussian mixture model (GMM) clustering has been replaced by a Bayesian information criterion (BIC) agglomerative clustering. Second, an additional clustering stage has been added, using a GMM-based speaker identification method. Finally, a post-processing stage refines the segment boundaries using the output of a transcription system. On the National Institute of Standards and Technology (NIST) RT-04F and ESTER evaluation data, the multistage system reduces the speaker error by over 70% relative to the baseline system, and gives between 40% and 50% reduction relative to a single-stage BIC clustering system
Claude Barras, Xuan Zhu 0001, Sylvain Meignier, Jean-Luc Gauvain
IEEE Trans. Speech Audio Process.4
2006 Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system
abstract
This paper describes the progress made in the transcription of broadcast news (BN) and conversational telephone speech (CTS) within the combined BBN/LIMSI system from May 2002 to September 2004. During that period, BBN and LIMSI collaborated in an effort to produce significant reductions in the word error rate (WER), as directed by the aggressive goals of the Effective, Affordable, Reusable, Speech-to-text [Defense Advanced Research Projects Agency (DARPA) EARS] program. The paper focuses on general modeling techniques that led to recognition accuracy improvements, as well as engineering approaches that enabled efficient use of large amounts of training data and fast decoding architectures. Special attention is given on efforts to integrate components of the BBN and LIMSI systems, discussing the tradeoff between speed and accuracy for various system combination strategies. Results on the EARS progress test sets show that the combined BBN/LIMSI system achieved relative reductions of 47% and 51% on the BN and CTS domains, respectively.
Spyridon Matsoukas, Jean-Luc Gauvain, Gilles Adda, Thomas Colthurst, Chia-Lin Kao, Owen Kimball, Lori Lamel, Fabrice Lefèvre, Jeff Z. Ma, John Makhoul, Long Nguyen 0001, Rohit Prasad, Richard M. Schwartz, Holger Schwenk, Bing Xiang
IEEE Trans. Speech Audio Process.2
2005 Open Vocabulary ASR for Audiovisual Document Indexation
abstract
The paper reports on an investigation of an open vocabulary recognizer that allows new words to be introduced in the recognition vocabulary, without the need to retrain or adapt the language model. This method uses special word classes, whose n-gram probabilities are estimated during the training process by discounting a mass of probability from the out of vocabulary words. A part-of-speech tagger is used to determine the word classes during language model training and for vocabulary adaptation. Metadata information provided by a French audiovisual archive institute are used to identify important document-specific missing words which are added to appropriate word classes in the system vocabulary. Pronunciations for the new words are derived by grapheme-to-phoneme conversion. On over 3 hours of broadcast news data, this approach leads to a reduction of 0.35% in the OOV rate, of 0.6% of the word error rate, with 80% of the occurrences of the newly introduced words being correctly recognized.
Alexandre Allauzen, Jean-Luc Gauvain
ICASSP (1)2
2005 Alternate Phone Models for Conversational Speech
abstract
This paper investigates the use of alternate phone models for the transcription of conversational telephone speech. The focus of this work is to explore alternative ways of modeling different manners of speaking so as to better cover the observed articulatory styles and pronunciation variants. Four alternate phone sets are compared ranging from 38 to 129 units. Two of the phone sets make use of syllable-position dependent phone models. The acoustic models were trained on 2300 hours of conversational telephone speech data from the Switchboard and Fisher corpora, and experimental results are reported on the EARS Dev04 test set which contains 3 hours of speech from 36 Fisher conversations. While no one particular phone set was found to outperform the others for a majority of speakers, the best overall performance was obtained with the original 48 phone set and a reduced 38 phone set, however combining the hypotheses of the individual models reduces the word error rate from 17.5% (original phone set) to 16.8%.
Lori Lamel, Jean-Luc Gauvain
ICASSP (1)2
2005 Diachronic vocabulary adaptation for broadcast news transcription
Alexandre Allauzen, Jean-Luc Gauvain
INTERSPEECH2
2005 Where are we in transcribing French broadcast news?
abstract
International audience
Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Véronique Gendner, Lori Lamel, Holger Schwenk
INTERSPEECH1
2005 Transcribing lectures and seminars
Lori Lamel, Gilles Adda, Eric Bilinski, Jean-Luc Gauvain
INTERSPEECH4
2005 Modeling vowels for Arabic BN transcription
abstract
This paper describes the LIMSI Arabic Broadcast News system which produces a vowelized word transcription. The under 10x system, evaluated in the NIST RT-04F evaluation, uses a 3 pass decoding strategy with gender- and bandwidth-specific acoustic models, a vowelized 65k word class pronunciation lexicon and a word-class 4-gram language model. In order to explicitly represent the vowelized word forms, each nonvowelized word entry is considered as a word class regrouping all of its associated vowelized forms. Since Arabic texts are almost exclusively written without vowels, an important challenge is to be able to use these efficiently in a system producing a vowelized output. Since a portion of the acoustic training data was manually transcribed with short vowels, enabling an initial set of acoustic models to be estimated in a supervised manner. The remaining audio data, for which vowels are not annotated, were trained in an implicit manner using the recognizer to choose the preferred form. The system was trained on a total of about 150 hours of audio data and almost 600 million words of Arabic texts, and achieved word error rates of 16.0% and 18.5% on the dev04 and eval04 data, respectively.
Abdelkhalek Messaoudi, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH3
2005 The 2004 BBN/LIMSI 20xRT English conversational telephone speech recognition system
abstract
In this paper we describe the English Conversational Telephone Speech (CTS) recognition system jointly developed by BBN and LIMSI under the DARPA EARS program for the 2004 evalua-tion conducted by NIST. The 2004 BBN/LIMSI system achieved a word error rate (WER) of 13.5 % at 18.3xRT (real-time as mea-sured on Pentium 4 Xeon 3.4 GHz Processor) on the EARS progress test set. This translates into a 22.8 % relative improvement in WER over the 2003 BBN/LIMSI EARS evaluation system, which was run without any time constraints. In addition to reporting on the system architecture and the evaluation results, we also highlight the significant improvements made at both sites. 1.
Rohit Prasad, Spyridon Matsoukas, Chia-Lin Kao, Jeff Z. Ma, Dongxin Xu, Thomas Colthurst, Owen Kimball, Richard M. Schwartz, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Fabrice Lefèvre
INTERSPEECH9
2005 Building continuous space language models for transcribing european languages
abstract
International audience
Holger Schwenk, Jean-Luc Gauvain
INTERSPEECH2
2005 Combining speaker identification and BIC for speaker diarization
abstract
International audience
Xuan Zhu 0001, Claude Barras, Sylvain Meignier, Jean-Luc Gauvain
INTERSPEECH4
2005 Genericity and portability for task-independent speech recognition
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
Comput. Speech Lang.2
2004 Lightly supervised acoustic model training using consensus networks
abstract
The paper presents some recent work on using consensus networks to improve lightly supervised acoustic model training for the LIMSI Mandarin BN system. Lightly supervised acoustic model training has been attracting growing interest, since it can help to reduce the development costs for speech recognition systems substantially. Compared to supervised training with accurate transcriptions, the key problem in lightly supervised training is getting the approximate transcripts to be as close as possible to manually produced detailed ones, i.e., finding a proper way to provide the information for supervision. Previous work using a language model to provide supervision has been quite successful. The paper extends the original method by presenting a new way to get the information needed for supervision during training. Studies are carried out using the TDT4 Mandarin audio corpus and associated closed-captions. After automatically recognizing the training data, the closed-captions are aligned with a consensus network derived from the hypothesized lattices. As is the case with closed-caption filtering, this method can remove speech segments whose automatic transcripts contain errors, but it can also recover errors in the hypothesis if the information is present in the lattice. Experimental results show that, compared with simply training on all of the data, consensus network based lightly supervised acoustic model training results in a small reduction in the character error rate on the DARPA/NIST RT'03 development and evaluation data.
Langzhou Chen, Lori Lamel, Jean-Luc Gauvain
ICASSP (1)3
2004 Speech transcription in multiple languages
abstract
The paper summarizes recent work underway at LIMSI on speech-to-text transcription in multiple languages. The research has been oriented towards the processing of broadcast audio and conversational speech for information access. Broadcast news transcription systems have been developed for seven languages, and it is planned to address several other languages in the near term. Research on conversational speech has mainly focused on the English language, with some initial work on French, Arabic and Spanish. Automatic processing must take into account the characteristics of the audio data, such as needing to deal with the continuous data stream, specificities of the language and the use of an imperfect word transcription for accessing the information content. Our experience thus far indicates that at today's word error rates, the techniques used in one language can be successfully ported to other languages, and most of the language specificities concern lexical and pronunciation modeling.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Leonardo Canseco-Rodriguez, Langzhou Chen, Olivier Galibert, Abdelkhalek Messaoudi, Holger Schwenk
ICASSP (3)2
2004 Speech recognition in multiple languages and domains: the 2003 BBN/LIMSI EARS system
abstract
We report on the results of the first evaluations for the BBN/LIMSI system under the new DARPA EARS program. The evaluations were carried out for conversational telephone speech (CTS) and broadcast news (BN) for three languages: English, Mandarin, and Arabic. In addition to providing system descriptions and evaluation results, the paper highlights methods that worked well across the two domains and those few that worked well on one domain but not the other. For the BN evaluations, which had to be run under 10 times real-time, we demonstrated that a joint BBN/LIMSI system with a time constraint achieved better results than either system alone.
Richard M. Schwartz, Thomas Colthurst, Nicolae Duta, Herbert Gish, Rukmini Iyer, Chia-Lin Kao, Daben Liu, Owen Kimball, Jeff Z. Ma, John Makhoul, Spyridon Matsoukas, Long Nguyen 0001, Mohammed Noamany, Rohit Prasad, Bing Xiang, Dongxin Xu, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen
ICASSP (3)17
2004 Dynamic language modeling for broadcast news
abstract
ABSTRACT This paper describes some recent experiments on unsuper-vised language model adaptation for transcription of broadcastnews data. In previous work, a framework for automaticallyselecting adaptation data using information retrieval techniqueswas proposed. This work extends the method and presents ex-perimental results with unsupervised language model adapta-tion. Threeprimaryaspectsareconsidered: (1)theperformanceof5widelyusedLMadaptationmethodsusingthesameadapta-tion data is compared; (2) the influence of the temporal distancebetween the training and test data epoch on the adaptation effi-ciency is assessed; and (3) show-based language model adapta-tion is compared with story-based language model adaptation.Experimentshavebeencarriedoutforbroadcastnewstranscrip-tion in English and Mandarin Chinese. A relative word errorrate reduction of 4.7% was obtained in English and a 5.6% rela-tive character error rate reduction in Mandarin withstory-basedMDI adaptation. 1. INTRODUCTION While n-gram models are successfully used in speech recog-nition, their performance is influenced by any mismatch be-tween thetraining and testdata [7]. Theidea of language model(LM) adaptation is to use a small amount of domain specificdata to adjust the LM to reduce the impact of linguistic differ-ences between the training and testing data. Different schemesforLMadaptation have been proposed, such asthecache modelbased on the observation that a word whichoccurred in a recenttext has a higher probability to be seen again [9]; the triggermodel which uses a trigger word pair to get at semantic infor-mation [10]; and structured LMs [1].Broadcast news (BN) transcription is a complicated task forboth acoustic and language modeling. The linguistic attributesof BN data are complex, arising from the many different speak-ing styles, from spontaneous conversation to prepared speech(close in style to written texts). The content of BN data is openand any given BN show covers multiple topics.As a consequence, it is difficult to predict the topics of a BNshow without looking at the data itself. The only informationthat is available for the show are the hypotheses output from thespeech recognizer. However, for any given broadcast, the num-ber of words in the hypothesized transcript is quite small andcontains recognition errors. Therefore the transcripts are notsufficient for use as an adaptive corpus. Information retrieval(IR) methods provide a means to address this problem. Insteadof directly using the ASR hypotheses for LM adaptation, theycan be used as queries to an IR system in order to select ad-ditional on-topic adaptation data from a large general corpus.This approach reduces the effect of transcription errors in thehypotheses and at the same time provides substantially moretextual data for LM estimation.In this paper, a series of experiments are presented exploringthe general framework of unsupervised LM adaptation using IRmethods [3]. The performances of a variety of popular tech-niques for LM adaptation using automatically selected adapta-tion data are compared. The investigated techniques are linearinterpolation,maximum aposteriori(MAP)adaptation, mixturemodels, dynamic mixture models, and minimum discriminationinformation (MDI) adaptation. The effect of the temporal dis-tance between the epoch of the adaptation corpus and of theepoch of the test data is also assessed. As mentioned above, agiven BN show typically covers several stories, with each storybeing related to a different topic. To address the changing prop-erty of BN data, static and dynamic models for LM adaptationare investigated. In static modeling the LM is updated oncefor the whole show, which means that the LM must be simul-taneously fit to multiple topics. Dynamic modeling updates theLM at each automatically detected story change, which entailsestimating multiplestory-based LMs for each BN show. Exper-iments arecarriedout forBNtranscriptioninAmericanEnglishand Mandarin Chinese.
Langzhou Chen, Lori Lamel, Jean-Luc Gauvain, Gilles Adda
INTERSPEECH3
2004 Language recognition using phone latices
abstract
This paper proposes a new phone lattice based method for automatic language recognition from speech data. By using phone lattices some approximations usually made by language identification (LID) systems relying on phonotactic constraints to simplify the training and decoding processes can be avoided. We demonstrate the use of phone lattices both in training and testing significantly improves the accuracy of a phonotactically based LID system. Performance is further enhanced by using a neural network to combine the results of multiple phone recognizers. Using three phone recognizers with context independent phone models, the system achieves an equal error rate of 2.7% on the Eval03 NIST detection test (30s segment, primary condition) with an overall decoding process that runs faster than real-time (0.5xRT).
Jean-Luc Gauvain, Abdelkhalek Messaoudi, Holger Schwenk
INTERSPEECH1
2004 Speaker diarization from speech transcripts
Lori Lamel, Jean-Luc Gauvain, Leonardo Canseco-Rodriguez
INTERSPEECH2
2004 Transcription of arabic broadcast news
abstract
This paper describes recent research on transcribing Modern Standard Arabic broadcast news data. The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Arabic is a highly inflected language where articles and affixes are added to roots in order to change the word’s meaning. A corpus of 50 hours of audio data from 7 television and radio sources and 200 M words of newspaper texts were used to train the acoustic and language models. The transcription system based on these models and a vowelized dictionary obtains an average word error rate on a test set comprised of 12 hours of test data from 8 sources is about 18%.
Abdelkhalek Messaoudi, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH3
2004 Neural network language models for conversational speech recognition
abstract
International audience
Holger Schwenk, Jean-Luc Gauvain
INTERSPEECH2
2003 Feature and score normalization for speaker verification of cellular data
abstract
This paper presents some experiments with feature and score normalization for text-independent speaker verification of cellular data. The speaker verification system is based on cepstral features and Gaussian mixture models with 1024 components. The following methods, which have been proposed for feature and score normalization, are reviewed and evaluated on cellular data: cepstral mean subtraction (CMS), variance normalization, feature warping, T-norm, Z-norm and the cohort method. We found that the combination of feature warping and T-norm gives the best results on the NIST 2002 test data (for the one-speaker detection task). Compared to a baseline system using both CMS and variance normalization and achieving a 0.410 minimal decision cost function (DCF), feature warping and T-norm respectively bring 8% and 12% relative reductions, whereas the combination of both techniques yields a 22% relative reduction, reaching a DCF of 0.320. This result approaches the state-of-the-art performance level obtained for speaker verification with land-line telephone speech.
Claude Barras, Jean-Luc Gauvain
ICASSP (2)2
2003 Unsupervised language model adaptation for broadcast news
abstract
Unsupervised language model adaptation for speech recognition is challenging, particularly for complicated tasks such the transcription of broadcast news (BN) data. This paper presents an unsupervised adaptation method for language modeling based on information retrieval techniques. The method is designed for the broadcast news transcription task where the topics of the audio data cannot be predicted in advance. Experiments are carried out using the LIMSI American English BN transcription system and the NIST 1999 BN evaluation sets. The unsupervised adaptation method reduces the perplexity by 7% relative to the baseline LM and yields a 2% relative improvement for a 10xRT system.
Langzhou Chen, Jean-Luc Gauvain, Lori Lamel, Gilles Adda
ICASSP (1)2
2003 Conversational telephone speech recognition
abstract
This paper describes the development of a speech recognition system for the processing of telephone conversations, starting with a state-of-the-art broadcast news transcription system. We identify major changes and improvements in acoustic and language modeling, as well as decoding, which are required to achieve state-of-the-art performance on conversational speech. Some major changes on the acoustic side include the use of speaker normalization (VTLN), the need to cope with channel variability, and the need for efficient speaker adaptation and better pronunciation modeling. On the linguistic side the primary challenge is to cope with the limited amount of language model training data. To address this issue we make use of a data selection technique, and a smoothing technique based on a neural network language model. At the decoding level lattice rescoring and minimum word error decoding are applied. On the development data, the improvements yield an overall word error rate of 24.9% whereas the original BN transcription system had a word error rate of about 50% on the same data.
Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen, Fabrice Lefèvre
ICASSP (1)1
2003 Multi-source training and adaptation for generic speech recognition
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2002 Transcribing audio-video archives
abstract
This paper addresses the automatic transcription of audiovideo archives using a state-of-the-art broadcast news speech transcription system. A 9-hour corpus spanning the latter half of the 20th century (1945-1995) has been transcribed and an analysis of the transcription quality carried out. In addition to the challenges of transcribing heterogenous broadcast news data, we are faced with changing properties of the archive over time, such as the audio quality, the speaking style, vocabulary items and manner of expression. After assessing the performance of the transcription system, several paths are explored in an attempt to reduce the mismatch between the acoustic and language models and the archived data.
Claude Barras, Alexandre Allauzen, Lori Lamel, Jean-Luc Gauvain
ICASSP4
2002 Unsupervised acoustic model training
abstract
This paper describes some recent experiments using unsupervised techniques for acoustic model training in order to reduce the system development cost. The approach uses a speech recognizer to transcribe unannotated raw broadcast news data. The hypothesized transcription is used to create labels for the training data. Experiments providing supervision only via the language model training materials show that including texts which are contemporaneous with the audio data is not crucial for success of the approach, and that the acoustic models can be initialized with as little as 10 minutes of manually annotated data. These experiments demonstrate that unsupervised training is a viable training scheme and can dramatically reduce the cost of building acoustic models.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda
ICASSP2
2002 Connectionist language modeling for large vocabulary continuous speech recognition
abstract
This paper describes ongoing work on a new approach for language modeling for large vocabulary continuous speech recognition. Almost all state.. o. f-the-art systems use statistical n-gram language models estimated on text corpora. One principle problem with such language models is the fact that many of the n-grams are never observed even in very large training corpora, and therefore it is common to back-off to a lower-order model. In this paper we propose to address this problem by carrying out the estimation task in a continuous space, enabling a smooth interpolation of the probabilities. A neural network is used to learn the projection of the words onto a continuous space and to estimate the n-gram probabilities. The connectionist language model is being evaluated on the DARPA HUB5 conversational telephone speech recognition task and preliminary results show consistent improvements in both perplexity and word error rate.
Holger Schwenk, Jean-Luc Gauvain
ICASSP2
2002 Advances in Large Vocabulary Speech Recognition
Jean-Luc Gauvain, Renato De Mori, Lori Lamel
Comput. Speech Lang.1
2002 Lightly supervised and unsupervised acoustic model training
Lori Lamel, Jean-Luc Gauvain, Gilles Adda
Comput. Speech Lang.2
2002 The LIMSI Broadcast News transcription system
Jean-Luc Gauvain, Lori Lamel, Gilles Adda
Speech Commun.1
2002 User evaluation of the MASK kiosk
Lori Lamel, Samir Bennacef, Jean-Luc Gauvain, Hervé Dartigues, Jean-Noël Temem
Speech Commun.3
2001 Processing Broadcast Audio for Information Access
abstract
This paper addresses recent progress in speaker-independent, large vocabulary, continuous speech recognition, which has opened up a wide range of near and mid-term applications. One rapidly expanding application area is the processing of broadcast audio for information access. At LIMSI, broadcast news transcription systems have been developed for English, French, German, Mandarin and Portuguese, and systems for other languages are under development. Audio indexation must take into account the specificities of audio data, such as needing to deal with the continuous data stream and an imperfect word transcription. Some near-term applications areas are audio data mining, selective dissemination of information and media monitoring.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker, Claude Barras, Langzhou Chen, Yannick de Kercadio
ACL1
2001 Automatic transcription of compressed broadcast audio
abstract
With increasing volumes of audio and video data broadcast over the Web, it is of interest to assess the performance of state-of-the-art automatic transcription systems on compressed audio data for media indexation applications. In this paper the performance of the LIMSI 10x French broadcast news transcription system is measured on a two-hour audio set for a range of MP3 and RealAudio codecs at various bit rates and the GSM codec used for European cellular phone communications. The word error rates are compared with those obtained on high quality PCM recordings prior to compression. For a 6.5 kbps. audio bit rate (the most commonly used on the Web), word error rates under 40% can be achieved, which makes automatic media monitoring systems over the Web a realistic task.
Claude Barras, Lori Lamel, Jean-Luc Gauvain
ICASSP3
2001 Investigating lightly supervised acoustic model training
abstract
The last decade has witnessed substantial progress in speech recognition technology, with todays state-of-the-art systems being able to transcribe broadcast audio data with a word error of about 20%. However, acoustic model development for the recognizers requires large corpora of manually transcribed training data. Obtaining such data is both time-consuming and expensive, requiring trained human annotators with substantial amounts of supervision. We describe some experiments using different levels of supervision for acoustic model training in order to reduce the system development cost. The experiments have been carried out using the DARPA TDT-2 corpus (also used in the SDR99 and SDR00 evaluations). Our experiments demonstrate that light supervision is sufficient for acoustic model development, drastically reducing the development cost.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda
ICASSP2
2001 Towards task-independent speech recognition
abstract
Despite the considerable progress made in the last decade, speech recognition is far from a solved problem. For instance, porting a recognition system to a new task (or language) still requires substantial investment of time and money, as well as expertise in speech recognition. The paper takes a first step at evaluating to what extent a generic state-of-the-art speech recognizer can reduce the manual effort required for system development. We demonstrate the genericity of wide domain models, such as broadcast news acoustic and language models, and techniques to achieve a higher degree of genericity, such as transparent methods to adapt such models to a specific task. This work targets three tasks using commonly available corpora: small vocabulary recognition (TI-digits), text dictation (WSJ), and goal-oriented spoken dialog (ATIS).
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
ICASSP2
2001 Using information retrieval methods for language model adaptation
abstract
In this paper we report experiments on language model adaptation using information retrieval methods, drawing upon recent developments in information extraction and topic tracking. One of the problems is extracting reliable topic information with high confidence from the audio signal in the presence of recognition errors. The work in the information retrieval domain on information extraction and topic tracking suggested a new way to solve this problem. In this work, we make use of information retrieval methods to extract topic information in the word recognizer hypotheses, which are then used to automatically select adaptation data from a very large general text corpus. Two adaptive language models, a mixture based model and a MAP based model, have been investigated using the adaptation data. Experiments carried out with the LIMSI Mandarin broadcast news transcription system gives a relative character error rate reduction of 4.3% with this adaptation method.
Langzhou Chen, Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
INTERSPEECH2
2001 Improving genericity for task-independent speech recognition
abstract
Although there have been regular improvements in speech recognition technology over the past decade, speech recognition is far from being a solved problem. Recognition systems are usually tuned to a particular task and porting the system to a new task (or language) is both time-consuming and expensive. In this paper, issues in speech recognizer portability are addressed through the development of generic core speech recognition technology. First, the genericity of wide domain models is assessed by evaluating performance on several tasks. Then, the use of transparent methods for adapting generic models to a specific task is explored. Finally, further techniques are evaluated aiming at enhancing the genericity of the wide domain models. We show that unsupervised acoustic model adaptation and multi-source training can reduce the performance gap between task-independent and taskdependent acoustic models, and for some tasks even out-perform task-dependent acoustic models.
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2001 Audio Partitioning and Transcription for Broadcast Data Indexation
Jean-Luc Gauvain, Lori Lamel, Gilles Adda
Multim. Tools Appl.1
2000 Transcription and indexation of broadcast data
abstract
We report on recent research on transcribing and indexing broadcast news data for information retrieval purposes. The system described combines an adapted version of the LIMSI 1998 Hub-4E transcription system for speech recognition with text-based IR methods. Experimental results are reported in terms of recognition word error rate and mean average precision for both the TREC SDR98 (100h) and SDR99 (600h) data sets. With query expansion using commercial transcripts, comparable mean average precisions are obtained on manual reference transcriptions and automatic transcriptions with a word error rate of 21.5% measured on a 10 hour data subset.
Jean-Luc Gauvain, Lori Lamel, Yannick de Kercadio, Gilles Adda
ICASSP1
2000 Broadcast news transcription in Mandarin
abstract
In this paper, our work in developing a Mandarin broadcast news transcription system is described. The main focus of this work is a port of the LIMSI American English broadcast news transcription system to the Chinese Mandarin language. The system consists of an audio partitioner and an HMM-based continuous speech recognizer. The acoustic models were trained on about 24 hours of data from the 1997 Hub4 Mandarin corpus available via LDC. In addition to the transcripts, the language models were trained on Mandarin Chinese News Corpus containing about 186 million characters. We investigate recognition performance as a function of lexical size, with and without tone in the lexicon, and with a topic dependent language model. The transcription character error rate on the DARPA 1997 test set is 18.1% using a lexicon with 3 tone levels and a topic-based language model. 1. INTRODUCTION It is well known that radio and television broadcast shows contain different types of speech from the acoust...
Langzhou Chen, Lori Lamel, Gilles Adda, Jean-Luc Gauvain
INTERSPEECH4
2000 Fast decoding for indexation of broadcast data
abstract
Processing time is an important factor in making a speech transcription system viable for automatic indexation of radio and television broadcasts. When only concerned by the word error rate, it is common to design systems that run in 100 times real-time or more. This paper addresses issues in reducing the speech recognition time for automatic indexation of radio and TV broadcasts with the aim of obtaining reasonable performance for close to real-time operation. We investigated computational resources in the range 1 to 10xRT on commonly available platforms. Constraints on the computational resources led us to reconsider design issues, particularly those concerning the acoustic models and the decoding strategy. A new decoder was implemented which transcribes broadcast data in few times real-time with only a slight increase in word error rate when compared to our best system. Experiments with spoken document retrieval show that comparable IR results are obtained with a 10xRT automatic tra...
Jean-Luc Gauvain, Lori Lamel
INTERSPEECH1
2000 Considerations in the design and evaluation of spoken language dialog systems
abstract
In this paper we summarize our experience at LIMSI in the design, development and evaluation of spoken language dialog systems for information retrieval tasks. This work has been for the most part carried out in the context of several European and international projects. Evaluation plays an integral role in the development of spoken language dialog systems. While there are commonly used measures and methodologies for evaluating speech recognizers, the evaluation of spoken dialog systems is considerably more complicated due to the interactive nature and the human perception of performance. It is therefore important to assess not only the individual system components, but the overall system performance using objective and subjective measures. 1. INTRODUCTION In our view, spokenlanguagesystems should provide a natural, user-friendly interface with the computer, allowing easy access to the stored information. At LIMSI we have experience in developing several spoken language dialog system...
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain
INTERSPEECH3
2000 Combining multiple speech recognizers using voting and language model information
Holger Schwenk, Jean-Luc Gauvain
INTERSPEECH2
2000 Large-vocabulary continuous speech recognition: advances and applications
abstract
The past decade (1990-2000) has witnessed substantial advances in speech recognition technology, which when combined with the increase in computational power and storage capacity has resulted in a variety of commercial products already or soon to be on the market. The authors review the state of the art in core technology, large vocabulary continuous speech recognition, with a view toward highlighting recent advances. We then highlight issues in moving toward applications, discussing system efficiency, portability across languages and tasks, and enhancing the system output by adding tags and nonlinguistic information. Current performance in speech recognition and outstanding challenges for three classes of applications (dictation, audio indexation, and spoken language dialogue systems), are discussed.
Jean-Luc Gauvain, Lori Lamel
Proc. IEEE1
2000 Speaker verification over the telephone
Lori Lamel, Jean-Luc Gauvain
Speech Commun.2
2000 The LIMSI ARISE system
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain, Samir Bennacef, Martine Garnier-Rizet, B. Prouts
Speech Commun.3
1999 Large vocabulary speech recognition in French
abstract
We present some design considerations concerning our large vocabulary continuous speech recognition system in French. The impact of the epoch of the text training material on lexical coverage, language model perplexity and recognition performance on newspaper texts is demonstrated. The effectiveness of larger vocabulary sizes and larger text training corpora for language modeling is investigated. French is a highly inflected language producing large lexical variety and a high homophone rate. About 30% of recognition errors are shown to be due to substitutions between inflected forms of a given root form. When word error rates are analysed as a function of word frequency, a significant increase in the error rate can be measured for frequency ranks above 5000.
Martine Adda-Decker, Gilles Adda, Jean-Luc Gauvain, Lori Lamel
ICASSP3
1999 The LIMSI ARISE system for train travel information
abstract
In the context of the LE-3 ARISE (Automatic Railway Information Systems for Europe) project we have been developing a dialog system for vocal access to rail travel information. The system provides schedule information for the main French intercity connections, as well as, simulated fares and reservations, reductions and services. The goal is to obtain high dialog success rates with a very open dialog structure, where the user is free to ask any question or to provide any information at any point in time. In order to improve the performance with such an open dialog strategy, we make use of implicit confirmation using the callers wording (when possible), and change to a more constrained dialog level, when the dialog is not going well. In addition to own assessment, the prototype system undergoes periodic user evaluations carried out by the our partners at the French Railways.
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain, Samir Bennacef
ICASSP3
1999 Using AR HMM state-dependent filtering for speech enhancement
abstract
In this paper we address the problem of enhancing speech which has been degraded by additive noise. As proposed by Ephraim et al. (1989), autoregressive hidden Markov models (AR-HMM) for the clean speech and an autoregressive Gaussian for the noise are used. The filter applied to a given frame of noisy speech is estimated using the noise model and the autoregressive Gaussian having the highest a posteriori probability given the decoded state sequence. The success of this technique is highly dependent on accurate estimation of the best state sequence. A new strategy combining the use of cepstral-based HMMs, autoregressive HMMs, and a model combination technique, is proposed. The intelligibility of the enhanced speech is indirectly assessed via speech recognition, by comparing performance on noisy speech with compensated models to performance on the enhanced speech with clean-speech models. The results on enhanced speech are as good as our best results obtained with noise compensated models.
Driss Matrouf, Jean-Luc Gauvain
ICASSP2
1999 Language modeling for broadcast news transcription
abstract
The advanced intonation model for speech synthesis described here has a three level architecture. An initial abstract characterisation designed to represent intonation at cognitive percept level is rewritten to an intermediate representation which, though speaker-independent, accurately reflects physical pitch contours. At this stage the contours lack the variability we associate with natural speech. This representation is then further rewritten to provide an actual physical contour (now including variability and other ''natural'' phenomena such as micro-intonation). One or two examples are given for stages one and two of the process, and some indication of how we tackle stage three.
Gilles Adda, Michèle Jardino, Jean-Luc Gauvain
EUROSPEECH3
1999 Recent advances in transcribing television and radio broadcasts
abstract
Transcription of broadcast news shows (radio and television) is a major step in developing automatic tools for indexation and retrieval of the vast amounts of information generated on a daily basis. Broadcast shows are challenging to transcribe as they consist of a continuous data stream with segments of different linguistic and acoustic natures. Transcribing such data requires addressing two main problems: those related to the varied acoustic properties of the signal, and those related to the linguistic properties of the speech. Prior to word transcription, the data is partitioned into homogeneous acoustic segments. Non-speech segments are identified and rejected, and the speech segments are clustered and labeled according to bandwidth and gender. The speaker-independent large vocabulary, continuous speech recognizer makes use of n-gram statistics for language modeling and of continuous density HMMs with Gaussian mixtures for acoustic modeling. The LIMSI system has consistently obtain...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Michèle Jardino
EUROSPEECH1
1999 Comparing different model configurations for language identification using a phonotactic approach
abstract
In this paper different model configurations for language identification using a phonotactic approach are explored. Identification experiments were carried out on the 11-language telephone speech corpus OGI-TS, containing calls in French, English, German, Spanish, Japanese, Korean, Mandarin, Tamil, Farsi, Hindi, and Vietnamese. Phone sequences output by one or multiple phone recognizers are rescored with language-dependent phonotactic models approximated by phone bigrams. The parameters of different sets of acoustic phone models were estimated using the 4-language IDEAL corpus. Sets of language-specific phonotactic models were trained using the training portion of the OGITS CORPUS. Error rates are significantly reduced by combining language-dependent and language-independent acoustic decoders, especially for short segments. A 9.9% LID error rate was obtained on the 11-language task using phonotactic models trained on spontaneous speech data. These results show that the phonotactic approach is relative insensitive to an acoustic mismatch between training and test conditions.
Driss Matrouf, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel
EUROSPEECH3
1998 Multilingual phone recognition of spontaneous telephone speech
abstract
In this paper we report on experiments with phone recognition of spontaneous telephone speech. Phone recognizers were trained and assessed on IDEAL, a multilingual corpus containing telephone speech in French, British English, German and Castillan Spanish. We investigated the influence of the training material composition (size and linguistic content) on the recognition performance using context-independent (CI) hidden Markov models (HMMs) and phonotactic bigram models. We found that when testing on spontaneous speech data, using only spontaneous speech training data gave the highest phone accuracies for the four languages, even though this data comprises only 14% of the available training data. The use of context-dependent (CD) HMMs reduced the phone error across the 4 languages, with the average error reduced to 51.9% from the 57.4% obtained with CI models. We suggest a straightforward way of detecting non speech phenomena. The basic idea is to remove sequences of consonants between two silence labels from the recognized phone strings prior to scoring. This simple technique reduces the relative average phone error rate by 5.4%. The lowest phone error with CD models and filtering was obtained for Spanish (39.1%) with 4 language average being 49.1%.
Cristobal Corredor-Ardoy, Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain
ICASSP4
1998 Partitioning and transcription of broadcast news data
abstract
Radio and television broadcasts consist of a continuous stream of data comprised of segments of different linguistic and acoustic natures, which poses challenges for transcription. In this paper we report on our recent work in transcribing broadcast news data[2, 4], including the problem of partitioning the data into homogeneous segments prior to word recognition. Gaussian mixture models are used to identify speech and non-speech segments. A maximumlikelihood segmentation/clustering process is then applied to the speech segments using GMMs and an agglomerative clustering algorithm. The clustered segments are then labeled according to bandwidth and gender. The recognizer is a continuous mixture density, tied-state cross-word context-dependent HMM system with a 65k trigram language model. Decoding is carried out in three passes, with a final pass incorporating cluster-based test-set MLLR adaptation. The overall word transcription error on the Nov'97 unpartitioned evaluation test data was...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda
ICSLP1
1998 User evaluation of the mask kiosk
abstract
In this paper we report on a series of user trials carried out to assess the performance and usability of the Multimodal Multimedia Service Kiosk (MASK) prototype. The aim of the ESPRIT MASK project was to pave the way for advanced public service applications with user interfaces employing multimodal, multimedia input and output. The prototype kiosk was developed after analyzing the technological requirements in the context of users performing travel enquiry tasks, in close collaboration with the French Railways (SNCF) and the Ergonomics group at the University College of London (UCL). The time to complete the transaction with the MASK kiosk is reduced by about 30% compared to that required for the standard kiosk, and the transaction success rate is 85% for novices and 94% once familiar with the system. In addition to meeting or exceeding the performance goals set at the project onset in terms of success rate, transaction time, and user satisfaction, the MASK kiosk was judged to be user-friendly and simple to use.
Lori Lamel, Samir Bennacef, Jean-Luc Gauvain, Hervé Dartigues, Jean-Noël Temem
ICSLP3
1998 Language identification incorporating lexical information
abstract
In this paper we explore the use of lexical information for language identification (LID). Our reference LID system uses language-dependent acoustic phone models and phone-based bigram language models. For each language, lexical information is introduced by augmenting the phone vocabulary with the N most frequent words in the training data. Combined phone and word bigram models are used to provide linguistic constraints during acoustic decoding. Experiments were carried out on a 4-language telephone speech corpus. Using lexical information achieves a relative error reduction of about 20% on spontaneous and read speech compared to the reference phone-based system. Identification rates of 92%, 96% and 99% are achieved for spontaneous, read and task-specific speech segments respectively, with prior speech detection.
Driss Matrouf, Martine Adda-Decker, Lori Lamel, Jean-Luc Gauvain
ICSLP4
1998 On the use of speech and text corpora for speech recognition in French
Martine Adda-Decker, Gilles Adda, Lori Lamel, Jean-Luc Gauvain
LREC4
1998 A multilingual corpus for language identification
Lori Lamel, Gilles Adda, Martine Adda-Decker, Cristobal Corredor-Ardoy, Jean-Jacques Gangolf, Jean-Luc Gauvain
LREC6
1997 Transcribing broadcast news shows
abstract
While significant improvements have been made in large vocabulary continuous speech recognition of large read-speech corpora such as the ARPA Wall Street Journal-based CSR corpus (WSJ) for American English and the BREF corpus for French, these tasks remain relatively artificial. In this paper we report on our development work in moving from laboratory read speech data to real-world speech data in order to build a system for the new ARPA broadcast news transcription task. The LIMSI Nov96 speech recognizer makes use of continuous density HMMs with Gaussian mixtures for acoustic modeling and n-gram statistics estimated on newspaper texts. The acoustic models are trained on the WSJO/WSJ1, and adapted using MAP estimation with task-specific training data. The overall word error on the Nov96 partitioned evaluation test was 27.1%.
Jean-Luc Gauvain, Gilles Adda, Lori Lamel, Martine Adda-Decker
ICASSP1
1997 Speaker recognition with the Switchboard corpus
abstract
We present our development work carried out in preparation for the March'96 speaker recognition test on the Switchboard corpus organized by NIST. The speaker verification system evaluated was a Gaussian mixture model (GMM). We provide experimental results on the development test and evaluation test data, and some experiments carried out since the evaluation comparing the GMM with a phone-based approach. Better performance is obtained by training on data from multiple sessions, and with different handsets. High error rates are obtained even using a phone-based approach both with and without the use of orthographic transcriptions of the training data. We also describe a human perceptual test carried out on a subset of the development data, which demonstrates the difficulty human listeners had with this task.
Lori Lamel, Jean-Luc Gauvain
ICASSP2
1997 Model compensation for noises in training and test data
abstract
It is well known that the performance of speech recognition systems degrade rapidly as the mismatch between the training and test conditions increases. Approaches to compensate for this mismatch generally assume that the training data is noise-free, and the test data is noisy. In practice, this assumption is seldom correct. We propose an iterative technique to compensate for noise in both the training and test data. The adopted approach compensates the speech model parameters using the noise present in the test data, and compensates the test data frames using the noise present in the training data. The training and test data are assumed to come from different and unknown microphones and acoustic environments. The interest of such a compensation scheme has been assessed on the MASK task using a continuous density HMM-based speech recognizer. Experimental results show the advantage of compensating for both test and training noise.
Driss Matrouf, Jean-Luc Gauvain
ICASSP2
1997 Text normalization and speech recognition in French
abstract
In this paper we present a quantitative investigation into the impact of text normalization on lexica and language models for speech recognition in French. The text normalization process defines what is considered to be a word by the recognition system. Depending on this definition we can measure different lexical coverages and language model perplexities, both of which are closely related to the speech recognition accuracies obtained on read newspaper texts. Different text normalizations of up to 185M words of newspaper texts are presented along with corresponding lexical coverage and perplexity measures. Some normalizations were found to be necessary to achieve good lexical coverage, while others were more or less equivalent in this regard. The choice of normalization to create language models for use in the recognition experiments with read newspaper texts was based on these findings. Our best system configuration obtained a 11.2% word error rate in the AUPELF `French-speaking' spee...
Gilles Adda, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel
EUROSPEECH3
1997 Language identification with language-independent acoustic models
abstract
In this paper we explore the use of languageindependent acoustic models for language identification (LID). The phone sequence output by a single language-independent phone recognizer is rescored with language-dependent phonotactic models approximated by phone bigrams. The language-independent phoneme inventory was obtained by Agglomerative Hierarchical Clustering, using a measure of similarity between phones. This system is compared with a parallel language-dependent phone architecture, which uses optimally the acoustic log likelihood and the phonotactic score for language identification. Experiments were carried out on the 4-language telephone speech corpus IDEAL, containing calls in British English, Spanish, French and German. Results show that the language-independent approach performs as well as the language-dependent one: 9% versus 10% of error rate on 10 second chunks, for the 4-language task. 1. INTRODUCTION This paper presents some of our recent research on automatic language ...
Cristobal Corredor-Ardoy, Jean-Luc Gauvain, Martine Adda-Decker, Lori Lamel
EUROSPEECH2
1997 Transcription of broadcast news
abstract
In this paper we report on our recent work in transcribing broadcast news shows. Radio and television broadcasts contain signal segments of various linguistic and acoustic natures. The shows contain both prepared and spontaneous speech. The signal may be studio quality or have been transmitted over a telephone or other noisy channel (ie., corrupted by additive noise and nonlinear distorsions), or may contain speech over music. Transcription of this type of data poses challenges in dealing with the continuous stream of data under varying conditions. Our approach to this problem is to segment the data into a set of categories, which are then processed with category specific acoustic models. We describe our 65k speech recognizer and experiments using different sets of acoustic models for transcription of broadcast news data. The use of prior knowledge of the segment boundaries and types is shown to not crucially affect the performance. 1. INTRODUCTION The goal of this research is to au...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
EUROSPEECH1
1997 Multilingual large vocabulary speech recognition: the European SQALE project
Steve J. Young, Martine Adda-Decker, Xavier L. Aubert, Christian Dugast, Jean-Luc Gauvain, Dan J. Kershaw, Lori Lamel, David A. van Leeuwen, David Pye, Anthony J. Robinson, Herman J. M. Steeneken, Philip C. Woodland
Comput. Speech Lang.5
1997 The LIMSI RailTel System: Field trial of a telephone service for rail travel information
Lori Lamel, Samir Bennacef, Sophie Rosset, Laurence Devillers, S. Foukia, Jean-Jacques Gangolf, Jean-Luc Gauvain
Speech Commun.7
1996 Developments in large vocabulary, continuous speech recognition of German
abstract
We describe our large vocabulary continuous speech recognition system for the German language, the development of which was partly carried out within the context of the European LRE project 62-058 SQALE. The recognition system is the LIMSI recognizer originally developed for French and American English, which has been adapted to German. Specificities of German, as relevant to the recognition system, are presented. These specificities have been accounted for during the recognizer's adaptation process. We present experimental results on a first test set ger-dev95 to measure progress in system development. Results are given with the final system using different acoustic model sets on two test sets ger-dev95 and ger-eval95. This system achieved a word error rate of 17.3% (official word error rate of 16.1% after SQALE adjudication process) on the ger-eval95 test set.
Martine Adda-Decker, Gilles Adda, Lori Lamel, Jean-Luc Gauvain
ICASSP4
1996 Developments in continuous speech dictation using the 1995 ARPA NAB news task
abstract
We report on the LIMSI recognizer evaluated in the ARPA 1995 North American Business (NAB) news benchmark test. In contrast to previous evaluations, the new Hub 3 test aims at improving basic SI, CSR performance on unlimited-vocabulary read speech recorded under more varied acoustical conditions (background environmental noise and unknown microphones). The LIMSI recognizer is an HMM-based system with a Gaussian mixture. Decoding is carried out in multiple forward acoustic passes, where more refined acoustic and language models are used in successive passes and information is transmitted via word graphs. In order to deal with the varied acoustic conditions, channel compensation is performed iteratively, refining the noise estimates before the first three decoding passes. The final decoding pass is carried out with speaker-adapted models obtained via unsupervised adaptation using the MLLR method. On the Sennheiser microphone (average SNR 29 dB) a word error of 9.1% was obtained, which can be compared to 17.5% on the secondary microphone data (average SNR 15 dB) using the same recognition system.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Driss Matrouf
ICASSP1
1996 Speech recognition for an information kiosk
Jean-Luc Gauvain, Jean-Jacques Gangolf, Lori Lamel
ICSLP1
1996 Spoken language processing in a multilingual context
abstract
In this paper we overview the spoken language processing activities at LIMSI, which are carried out in a multilingual framework.These activities include speech-to-text conversion, spoken language systems for information retrieval, speaker and language recognition, and speech response.The Spoken Language Processing Group has also been actively involved in corpora development and evaluation.The group has regularly participated in evaluations organized by ARPA, in the LE-SQALE project, and in the AUPELF-UREF program for provision of linguistic resources and evaluation tests for French.
Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain, Gilles Adda
ICSLP3
1996 A stochastic case frame approach for natural language understanding
abstract
A stochastically based approach for the semantic analysis component of a natural spoken language system for the ATIS task has been developed.The semantic analyzer of the spoken language system already in use at LIMSI makes use of a rule-based case grammar.In this work, the system of rules for the semantic analysis is replaced with a relatively simple, first order Hidden Markov Model.The performance of the two approaches can be compared because they use identical semantic representations despite their rather different methods for meaning extraction.We use an evaluation methodology that assesses performance at different semantic levels, including the database response comparison used in the ARPA ATIS paradigm.
Wolfgang Minker, Samir Bennacef, Jean-Luc Gauvain
ICSLP3
1996 Comments on "Towards increasing speech recognition error rates" by H. Bourlard, H. Hermansky, and N. Morgan
Joseph Mariani, Jean-Luc Gauvain, Lori Lamel
Speech Commun.2
1995 Developments in continuous speech dictation using the ARPA WSJ task
abstract
We report on our recent development work in large vocabulary, American English continuous speech dictation. We have experimented with (1) alternative analyses for the acoustic front end, (2) the use of an enlarged vocabulary so as to reduce the number of errors due to out-of-vocabulary words, (3) extensions to the lexical representation, (4) the use of additional acoustic training data, and (5) modification of the acoustic models for telephone speech. The recognizer was evaluated on Hubs 1 and 2 of the fall 1994 ARPA NAB CSR Hub and Spoke Benchmark test. Experimental results for development and evaluation test data are given, as well as an analysis of the errors on the development data.
Jean-Luc Gauvain, Lori Lamel, Martine Adda-Decker
ICASSP1
1995 Experiments with speaker verification over the telephone
abstract
In this paper we present a study on speaker verification showing achievable performance levels for both high quality speech and telephone speech and for two operational modes, i.e. textdependent and text-independent speaker verification. A statistical modeling approach is taken, where for text independent verification the talker is viewed as a source of phones, modeled by a fully connected Markov chain, where the lexical and syntactic structures of the language are approximated by local phonotactic constraints. A first series of experiments were carried out on high quality speech from the BREF corpus to validate this approach and resulted in an a posteriori equal error rate of 0.3% in textdependent as well as in text-independent mode. A second series of experiments were carried out on a telephone corpus recorded specifically for speaker verification algorithm development. On this data, the lowest equal error rate is 2.9% for the text-dependent mode when 2 trials are allowed per attempt...
Jean-Luc Gauvain, Lori Lamel, B. Prouts
EUROSPEECH1
1995 Issues in Large Vocabulary, Multilingual Speech Recognition
abstract
In this paper we report on our activities in multilingual, speakerindependent, large vocabulary continuous speech recognition. The multilingual aspect of this work is of particular importance in Europe, where each country has its own national language. Our existing recognizer for American English and French, has been ported to British English and German. It has been assessed in the context of the LRESQALE project whose objective was to experiment with installing in Europe a multilingual evaluation paradigm for the assessment of large vocabulary, continuous speech recognition systems. The recognizer makes use of phone-based continuous density HMM for acoustic modeling and n-gram statistics estimated on newspaper texts for language modeling. The system has been evaluated on a dictation task with read, newspaper-based corpora, the ARPA Wall Street Journal corpus of American English, the WSJCAM0 corpus of British English, the BREF-Le Monde corpus of French and the PHONDAT-Frankfurter Runds...
Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain
EUROSPEECH3
1995 Development of spoken language corpora for travel information
abstract
In this paper we report on our ongoing work in developing spoken language corpora in the context of information access in two travel domain tasks, L'ATIS and MASK. The collection of spoken language corpora remains an important research area and represents a significant portion of work in the development of spoken language systems. The use of additional acoustic and language model training data has been shown to almost systematically improve performance in continuous speech recognition. Similarly, progress in spokenlanguage understanding is closely linked to the availability of spoken language corpora. We record subjects on a regular basis using development versions of the spoken language systems for both tasks, obtaining over 1000 queries/month from 20 subjects. To help assess our progress in system development, each subject since March'95 completes a questionnaire addressing the user-friendliness, reliability, ease-of-use of the MASK data collection system. INTRODUCTION The collecti...
Lori Lamel, Sophie Rosset, Samir Bennacef, Hélène Bonneau-Maynard, Laurence Devillers, Jean-Luc Gauvain
EUROSPEECH6
1995 A phone-based approach to non-linguistic speech feature identification
Lori Lamel, Jean-Luc Gauvain
Comput. Speech Lang.2
1994 The LIMSI continuous speech dictation system: evaluation on the ARPA Wall Street Journal task
abstract
We report progress made at LIMSI in speaker-independent large vocabulary speech dictation using the ARPA Wall Street Journal-based CSR corpus. The recognizer makes use of continuous density HMM with Gaussian mixture for acoustic modeling and n-gram statistics estimated on the newspaper texts for language modeling. The recognizer uses a time-synchronous graph-search strategy which is shown to still be viable with vocabularies of up to 20 K words when used with bigram back-off language models. A second forward pass, which makes use of a word graph generated with the bigram, incorporates a trigram language model. Acoustic modeling uses cepstrum-based features, context-dependent phone models (intra and interword), phone duration models, and sex-dependent models. The recognizer has been evaluated in the Nov92 and Nov93 ARPA tests for vocabularies of up to 20,000 words.>
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
ICASSP (1)1
1994 Language identification using phone-based acoustic likelihoods
abstract
Applies the technique of phone-based acoustic likelihoods to the problem of language identification. The basic idea is to process the unknown speech signal by language-specific phone model sets in parallel, and to hypothesize the language associated with the model set having the highest likelihood. Using laboratory quality speech the language can be identified as French or English with better than 99% accuracy with only as little as 2 seconds of speech. On spontaneous telephone speech from the OGI corpus, the language can be identified as French or English with 82% accuracy with 10 seconds of speech. The 10 language identification rate using the OGI corpus is 59.7% with 10 seconds of signal.>
Lori Lamel, Jean-Luc Gauvain
ICASSP (1)2
1994 A spoken language system for information retrieval
Samir Bennacef, Hélène Bonneau-Maynard, Jean-Luc Gauvain, Lori Lamel, Wolfgang Minker
ICSLP3
1994 Continuous speech dictation in French
abstract
A major research activity at LIMSI is multilingual, speakerindependent, large vocabulary speech dictation. In this paper we report on efforts in large vocabulary, speaker-independent continuous speech recognition of French using the BREF corpus. Recognition experiments were carried out with vocabularies containing up to 20k words. The recognizer makes use of continuous density HMM with Gaussian mixture for acoustic modeling and n-gram statistics estimated on 38 million words of newspaper text from Le Monde for language modeling. The recognizer uses a time-synchronous graph-search strategy. When a bigram language model is used, recognition is carried out in a single forward pass. A second forward pass, which makes use of a word graph generated with the bigram language model, incorporates a trigram language model. Acoustic modeling uses cepstrum-based features, contextdependent phone models and phone duration models. An average phone accuracy of 86% was achieved. A word accuracy of 84% h...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
ICSLP1
1994 Speech-To-Text Conversion in French
abstract
Speech-to-text conversion of French necessitates that both the acoustic level recognition and language modeling be tailored to the French language. Work in this area was initiated at LIMSI over 10 years ago. In this paper a summary of the ongoing research in this direction is presented. Included are studies on distributional properties of French text materials; problems specific to speech-to-text conversion particular of French; studies in phoneme-to-grapheme conversion for continuous, error-free phonemic strings; past work on isolated-word speech-to-text conversion; and more recent work on continuous-speech, speech-to-text conversion. Also demonstrated is the use of phone recognition for both language and speaker identification. The continuous speech-to-text conversion for French is based on a speaker-independent, vocabulary-independent recognizer. In this paper phone recognition and word recognition results are reported evaluating this recognizer on read speech taken from the BREF corpus. The recognizer was trained on over 4 hours of speech from 57 speakers, and tested on sentences from an independent set of 19 speakers. A phone accuracy of 78.7% was obtained using a set of 35 phones. The word accuracy was 88% for a 1139 word lexicon and 86% for a 2716 word lexicon, with a word pair grammar with respective perplexities of 100 and 160. Using a bigram grammar, word accuracies of 85.5% and 81.7% were obtained with 5 K and 20 K word vocabularies, with respective perplexities of 122 and 205.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Joseph Mariani
Int. J. Pattern Recognit. Artif. Intell.1
1994 Speaker-independent continuous speech dictation
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
Speech Communication1
1994 Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains
abstract
In this paper, a framework for maximum a posteriori (MAP) estimation of hidden Markov models (HMM) is presented. Three key issues of MAP estimation, namely, the choice of prior distribution family, the specification of the parameters of prior densities, and the evaluation of the MAP estimates, are addressed. Using HMM's with Gaussian mixture state observation densities as an example, it is assumed that the prior densities for the HMM parameters can be adequately represented as a product of Dirichlet and normal-Wishart densities. The classical maximum likelihood estimation algorithms, namely, the forward-backward algorithm and the segmental k-means algorithm, are expanded, and MAP estimation formulas are developed. Prior density estimation issues are discussed for two classes of applications/spl minus/parameter smoothing and model adaptation/spl minus/and some experimental results are given illustrating the practical interest of this approach. Because of its adaptive nature, Bayesian learning is shown to serve as a unified approach for a wide range of speech recognition applications.>
Jean-Luc Gauvain
IEEE Trans. Speech Audio Process.1
1993 Cross-lingual experiments with phone recognition
Lori Lamel, Jean-Luc Gauvain
ICASSP (2)2
1993 Speaker adaptation based on MAP estimation of HMM parameters
Chin-Hu Lee, Jean-Luc Gauvain
ICASSP (2)2
1993 A French version of the MIT-ATIS system: portability issues
Hélène Bonneau-Maynard, Jean-Luc Gauvain, David Goodine, Lori Lamel, Joseph Polifroni, Stephanie Seneff
EUROSPEECH2
1993 Speaker-independent continuous speech dictation
abstract
Abstract In this paper we report on progress made at LIMSI in speaker-independent large vocabulary speech dictation using newspaper-based speech corpora in English and French. The recognizer makes use of continuous density HMMs with Gaussian mixtures for acoustic modeling and n -gram statistics estimated on newspaper texts for language modeling. Acoustic modeling uses cepstrum-based features, context-dependent phone models (intra and interword), phone duration models, and sex-dependent models. For English the ARPA Wall Street Journal -based CSR corpus is used and for French the BREF corpus containing recordings of texts from the French newspaper Le Monde is used. Experiments were carried out with both these corpora at the phone level and at the word level with vocabularies containing up to 20,000 words. Word recognition experiments are also described for the ARPA RM task which has been widely used to evaluate and compare systems.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
EUROSPEECH1
1993 Identifying non-linguistic speech features
abstract
Over the last decade technological advances have been made which enable us to envision real-world applications of speech technologies. It is possible to foresee applications, for example, information centers in public places such as train stations and airports, where the spoken query is to be recognized without even prior knowledge of the languagebeing spoken. Other applications may require accurate identification of the speaker for security reasons, including control of access to confidential information or for telephone-based transactions.
Lori Lamel, Jean-Luc Gauvain
EUROSPEECH2
1993 High performance speaker-independent phone recognition using CDHMM
abstract
In this paper we report high phone accuracies on three corpora: WSJ0, BREF and TIMIT. The main characteristics of the phone recognizer are: high dimensional feature vector (48), context- and genderdependent phone models with duration distribution, continuous density HMM with Gaussian mixtures, and n-gram probabilities for the phonotatic constraints. These models are trained on speech data that have either phonetic or orthographic transcriptions using maximum likelihood and maximum a posteriori estimation techniques. On the WSJ0 corpus with a 46 phone set we obtain phone accuraciesof 72.4% and 74.4% using 500 and 1600 CD phone units, respectively. Accuracy on BREF with 35 phones is as high as 78.7% with only 428 CD phone units. On TIMIT using the 61 phone symbols and only 500 CD phone units, we obtain a phoneaccuracyof 67.2% which correspond to 73.4% when the recognizer output is mapped to the commonly used 39 phone set. Making reference to our work on large vocabularyCSR, we show that ...
Lori Lamel, Jean-Luc Gauvain
EUROSPEECH2
1993 Large vocabulary speech recognition using subword units
Jean-Luc Gauvain, Roberto Pieraccini, Lawrence R. Rabiner
Speech Commun.2
1992 Improved acoustic modeling with Bayesian learning
abstract
The authors study the use of Bayesian learning for the estimation of the parameters of a multivariate mixture Gaussian density. For speech recognition algorithms based on the continuous density hidden Markov model (CDHMM) framework, Bayesian learning serves as a unified approach for the following four applications: parameter smoothing, speaker adaptation, speaker group modeling, and corrective training. In the approach, the authors use Bayesian learning techniques to incorporate prior knowledge into the CDHMM training process in the form of prior densities of the HMM parameters. The theoretical basis for this procedure is presented. All four applications have been evaluated. Experimental results of the TI connected digit task and the Naval Resource Management task are provided to show the effectiveness of Bayesian adaptation of CDHMM.>
Jean-Luc Gauvain
ICASSP1
1992 Experiments on speaker-independent phone recognition using BREF
abstract
A series of experiments for speaker-independent, continuous speech phone recognition have been carried out using the recently recorded BREF corpus. The authors' experiments were the first to use this database, and are meant to provide a baseline performance evaluation for vocabulary independent phone recognition. The system was trained using hand-verified data from 43 speakers. Using 35 context-dependent phone models, a baseline phone accuracy of 60% (no phone grammar) has been obtained on an independent test set of 7635 phone segments from 19 speakers. Including phone bigram probabilities as phonotactic constraints results in a performance of 63.3%. A phone accuracy of 68.6% (73.3% correct) was obtained with 428 context dependent models.>
Lori Lamel, Jean-Luc Gauvain
ICASSP2
1992 A speech understanding system based on statistical representation of semantics
abstract
An understanding system, designed for both speech and text input, has been implemented based on statistical representation of task specific semantic knowledge. The core of the system is the conceptual decoder, which extracts the words and their association to the conceptual structure of the task directly from the acoustic signal. The conceptual information, which is also used to clarify the English sentences, is encoded following a statistical paradigm. A template generator and an SQL (structured query language) translator process the sentence and produce SQL code for querying a relational database. Results of the system on the official DARPA test are given.>
Roberto Pieraccini, Evelyne Tzoukermann, Zakhar Gorelov, Jean-Luc Gauvain, Esther Levin, Jay G. Wilpon
ICASSP4
1992 Bayesian learning for hidden Markov model with Gaussian mixture state observation densities
Jean-Luc Gauvain
Speech Commun.1
1991 Bayesian learning for hidden Markov model with Gaussian mixture state observation densities
Jean-Luc Gauvain
EUROSPEECH1
1991 BREF, a large vocabulary spoken corpus for French
Lori F. Larnel, Jean-Luc Gauvain, Maxine Eskénazi
EUROSPEECH2
1990 Adapting probability-transitions in DP matching process for an oral task-oriented dialogue
abstract
A system is presented which has been developed to evaluate the introduction of voice technologies in air-traffic controller training. The dialogue system is designed to replace the pseudopilot, a human playing the role of the pilot during training exercises and communicating both with the student controller and with the air-traffic simulator which updates the radar image on a screen. The knowledge representation which renders the cooperation of different knowledge sources possible is described. The approach makes use of pragmatic knowledge to predict a sublanguage which dynamically limits the recognition search space and therefore improves the accuracy. Recognition experiments in a task-simulation environment indicate that the combined use of both dynamic probabilities and dialogue strategies allows the system to obtain performance equal to 96.5% at the sentences level.>
K. Matrouf, Jean-Luc Gauvain, Françoise D. Néel, Joseph Mariani
ICASSP2
1990 Design considerations and text selection for BREF, a large French read-speech corpus
abstract
BREF, a large read-speech corpus in French has been designed with several aims: to provide enough speech data to develop dictation machines, to provide data for evaluation of continuous speech recognition systems (both speaker-dependent and speaker-independent), and to provide a corpus of continuous speech to study phonological variations. This paper presents some of the design considerations of BREF, focusing on the text analysis and the selection of text materials. The texts to be read were selected from 4.6 million words of the French newspaper, Le Monde. In total, 11,000 texts were selected, with an emphasis on maximizing the number of distinct triphones. Separate text materials were selected for training and test corpora. The goal is to obtain about 10,000 words (approximately 60-70 min.) of speech from each of 100 speakers, from different French dialects. INTRODUCTION One of the main obstacles to progress in continuous speech recognition has been the lack of sufficient speech m...
Jean-Luc Gauvain, Lori Lamel, Maxine Eskénazi
ICSLP1
1989 Adaptive syntax representation in an oral task-oriented dialogue for air-traffic controller training
K. Matrouf, Françoise D. Néel, Jean-Luc Gauvain, Joseph Mariani
EUROSPEECH3
1987 Vector quantization for speaker adaptation
abstract
In view of designing a speaker-independent large vocabulary recognition system, we evaluate a vector quantization approach to speaker adaptation. Only one speaker (the reference speaker) pronounces the application vocabulary. He also pronounces a small vocabulary called the adaptation vocabulary. Each new speaker then merely pronounces the adaptation vocabulary. Two adaptation methods are investigated, establishing a correspondence between the codebooks of these two speakers. This allows us to transform the reference utterances of the reference speaker into suitable references for the new speaker. Method I uses a transposed codebook to represent the new speaker during the recognition process whereas Method II uses a codebook which is obtained by clustering on the new speaker's pronunciation of the adaptation vocabulary. Experiments were carried out on a 20-speaker database (10 male, 10 female). The adaptation vocabulary contains 136 words; the application one has 104 words. The mean recognition error rate without adaptation is 22.3% for inter-speaker experiments; after one of the two methods has been implemented the mean recognition error rate is 10.5%. Comparison of performance of the two methods shows that a new speaker's codebook is not necessary to represent the new speaker.
Hélène Bonneau-Maynard, Jean-Luc Gauvain
ICASSP2
1986 A syllable-based isolated word recognition experiment
abstract
In view of the automatic recognition of a very large or, eventually, unlimited vocabulary, it is necessary to choose recognition units that are smaller than word size. A series of experiments has been carried out to evaluate the use of the syllable as a concatenative unit for large vocabularies. For isolated word recognition from syllable units, a reference pattern was created for each word of the lexicon by concatenating isolated syllable templates. The test utterance is then matched to each word template through the use of dynamic programming. To build the word reference pattern, the syllable templates were adjusted by using the VLTS procedure and down sampling the beginning and end parts of the syllable. The syllable approach was compared to the classical whole word one on the same 10,400-word vocabulary. The storage required for the syllable dictionary is one sixth of that necessary for the whole word dictionary. For a trained speaker, the recognition error rate obtained using the syllable approach was 12% compared to 6% using the whole word approach. This difference may be reduced by using syllable templates extracted from words to take coarticulation effects between syllables into account.
Jean-Luc Gauvain
ICASSP1
1986 A dynamic time warp VLSI processor for continuous speech recognition
abstract
We present a VLSI processor designed to compute dynamic time warping algorithms for speech recognition with extreme rapidity. This processor works as a coprocessor in a classical system including a standard microprocessor and a digital signal processor. It uses its own local memory for reference utterances and intermediate results. It has been designed to give maximum efficiency on continuous speech recognition applications with or without syntax constraints. Its flexibility permits software optimisation and its use in a large number of different applications. We use a sequential approach for DTW computations and work along the time axis. All the calculations are carried out on each frame of the unknown utterance as soon as it arrives from the DSP and DTW computations therefore take place in real time. Response time is in hundredth of second; intermediate results are obtained before the end of the sentence. A system using this chip will be able to carry out continuous speech recognition in real timee on a vocabulary of 300 references. Many of those chips can be used in parallel on a single system.
Georges Quénot, Jean-Luc Gauvain, Jean-Jacques Gangolf, Joseph Mariani
ICASSP2
1984 Evaluation of time compression for connected word recognition
abstract
Recently, several studies have shown the interesting aspects of nonuniform sampling of the filtered speech signal in the context of an isolated word recognizer. This paper investigates the effect of nonuniform sampling for connected word recognition. Three nonlinear time compression techniques are evaluated, one that brings all reference utterances down to one same length, and others, for which the utterance length is variable. The nonuniform sampling approach is compared to the uniform one by opposing the three non-linear methods to two linear time compression ones. The results show that the variable length trace segmentation technique gives the best scores under all conditions, and that the uniform sampling approach can therefore be advantageously used in connected word recognition processes.
Jean-Luc Gauvain, Joseph Mariani
ICASSP1
1983 On the use of time compression for word-based recognition
abstract
Dynamic time warping is a very efficient technique in dealing with the problem of time distorsion between different pronunciations of any given worm. However when, in a word-based recognition system(isolated or connected words), time normalisation is solely based on a DR-matching process, much processing time is necessitated. Another characteristic is that the general constraints used to optimise the DP-matching algorithm impose a severe limit on acceptable time distorsions. In this paper, we evaluate the interesting aspects of non-linear time compression methods which carry out a first time normalisation prior to DP-matching in word-based recognition. We describe three non-linear compression techniques, which have come under consideration during the study of recognition systems developed at LIMSI. These non-linear compression methods are compared to linear compressions in an isolated word recognition framework for different vocabularies.
Jean-Luc Gauvain, Joseph Mariani, Jean-Sylvain Liénard
ICASSP1
1982 A method for connected word recognition and word spotting on a microprocessor
abstract
In the last few years microprocessors have been used successfully in the construction of single board isolated word recognition systems. The interest that connected word recognition presently attracts, and the work that has been carried out at LIMSI in isolated word recognition, have brought us to implement a connected word and word spotting algorithm on a microprocessor. Herein we describe a speaker-dependent, limited vocabulary system using an 8088 microprocessor. There is only one training pass for each vocabulary word, except in the case of very short words, where we use two. There is no limit to the number of words in each utterance. We employ a compression method which carries out a preliminary time normalisation and reduces the amount of information used. A concise and efficient dynamic time warping procedure eliminates any remaining time distortion during the recognition phase. Real time processing is made possible by compressing and using the time warping method while acquisition is being carried out.
Jean-Luc Gauvain, Joseph Mariani
ICASSP1