Lori Lamel

dblp:14/2303 · also Lori Faith Lamel · DBLP profile ↗
← Back
188ranked-venue papers
38as first author
15since 2021 · last 2026
0000-0001-7443-9938ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 152 · 31 first-author · 9 since 2021Artificial intelligence and machine learning · 126 · 21 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies
Anisia Popescu, Johanna Cronenberg, Ioana Vasilescu, Ioana Chitoran, Lori Lamel, Martine Adda-Decker
LREC5
2025 Tracking /r/ Deletion: Forced Alignment of Pronunciation Variants and Sociophonetic Insights into Post-Obstruent Final /r/ in French
abstract
International audience
Anisia Popescu, Lori Lamel, Marc Evrard, Ioana Vasilescu
INTERSPEECH2
2024 Cross-linguistic transfer of phonological assimilation in early and late bilinguals
Sharon Peperkamp, Sonya Kaiser, Lori Lamel, Martine Adda-Decker
CogSci3
2024 Using Speech Technology to Test Theories of Phonetic and Phonological Typology
abstract
The present paper uses speech technology derived tools and methodologies to test theories about phonetic typology. We specifically look at how the two-way laryngeal contrast (voiced /b, d, g, v, z/ vs. voiceless /p, t, k, f, s/ obstruents) is implemented in European Portuguese, a language that has been suggested to exhibit a different voicing system than its sister Romance languages, more similar to the one found for Germanic languages. A large European Portuguese corpus was force aligned using (1) different combinations of parallel Portuguese (original), Italian (Romance language) and German (Germanic language) acoustic phone models and letting an ASR system choose the best fitting one, and (2) pronunciation variants (/b, d, g, v, z/ produced as either [b, d, g, v, z] or [p, t, k, f, s]) for obstruent consonants. Results support previous accounts in the literature that European Portuguese is diverging from the traditional voicing system known for Romance language, towards a hybrid system where stops and fricatives are specified for different voicing features.
Anisia Popescu, Lori Lamel, Ioana Vasilescu
LREC/COLING2
2024 Crosslinguistic Comparison of Acoustic Variation in the Vowel Sequences /ia/ and /io/ in Four Romance Languages
abstract
International audience
Johanna Cronenberg, Ioana Chitoran, Lori Lamel, Ioana Vasilescu
INTERSPEECH3
2024 Automatic Speech Recognition with parallel L1 and L2 acoustic phone models to evaluate /l/ allophony in L2 English speech production
Anisia Popescu, Lori Lamel, Ioana Vasilescu, Laurence Devillers
INTERSPEECH2
2023 Exploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus
abstract
The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in this paper use the CEMO, which was collected in a French emergency call center. Two pre-trained models based on speech and text were fine-tuned for speech emotion recognition. Using pre-trained Transformer encoders mitigates our data’s limited and sparse nature. This paper explores the different fusion strategies of these modality-specific models. In particular, fusions with and without cross-attention mechanisms were tested to gather the most relevant information from both the speech and text encoders. We show that multimodal fusion brings an absolute gain of 4-9% with respect to either single modality and that the Symmetric multi-headed cross-attention mechanism performed better than late classical fusion approaches. Our experiments also suggest that for the real-life CEMO corpus, the audio component encodes more emotive information than the textual one.
Théo Deschamps-Berger, Lori Lamel, Laurence Devillers
ICASSP2
2023 Mandarin lexical tone duration: Impact of speech style, word length, syllable position and prosodic position
Yaru Wu, Martine Adda-Decker, Lori Lamel
Speech Commun.3
2022 When Phonetics Meets Morphology: Intervocalic Voicing Within and Across Words in Romance Languages
abstract
International audience
Mathilde Hutin, Martine Adda-Decker, Lori Lamel, Ioana Vasilescu
INTERSPEECH3
2022 Voicing neutralization in Romanian fricatives across different speech styles
abstract
International audience
Laura Spinu, Ioana Vasilescu, Lori Lamel, Jason Lilley
INTERSPEECH3
2022 Extracting Linguistic Knowledge from Speech: A Study of Stop Realization in 5 Romance Languages
abstract
This paper builds upon recent work in leveraging the corpora and tools originally used to develop speech technologies for corpus-based linguistic studies. We address the non-canonical realization of consonants in connected speech and we focus on voicing alternation phenomena of stops in 5 standard varieties of Romance languages (French, Italian, Spanish, Portuguese, Romanian). For these languages, both large scale corpora and speech recognition systems were available for the study. We use forced alignment with pronunciation variants and machine learning techniques to examine to what extent such frequent phenomena characterize languages and what are the most triggering factors. The results confirm that voicing alternations occur in all Romance languages. Automatic classification underlines that surrounding contexts and segment duration are recurring contributing factors for modeling voicing alternation. The results of this study also demonstrate the new role that machine learning techniques such as classification algorithms can play in helping to extract linguistic knowledge from speech and to suggest interesting research directions.
Yaru Wu, Mathilde Hutin, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker
LREC4
2022 Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic Variation
abstract
Speech characteristics vary from speaker to speaker. While some variation phenomena are due to the overall communication setting, others are due to diastratic factors such as gender, provenance, age, and social background. The analysis of these factors, although relevant for both linguistic and speech technology communities, is hampered by the need to annotate existing corpora or to recruit, categorise, and record volunteers as a function of targeted profiles. This paper presents a methodology that uses a knowledge base to provide speaker-specific information. This can facilitate the enrichment of existing corpora with new annotations extracted from the knowledge base. The method also helps the large scale analysis by automatically extracting instances of speech variation to correlate with diastratic features. We apply our method to an over 120-hour corpus of broadcast speech in French and investigate variation patterns linked to reduction phenomena and/or specific to connected speech such as disfluencies. We find significant differences in speech rate, the use of filler words, and the rate of non-canonical realisations of frequent segments as a function of different professional categories and age groups.
Yaru Wu, Fabian M. Suchanek, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker
LREC4
2021 End-to-End Speech Emotion Recognition: Challenges of Real-Life Emergency Call Centers Data Recordings
abstract
Recognizing a speaker’s emotion from their speech can be a key element in emergency call centers. End-to-end deep learning systems for speech emotion recognition now achieve equivalent or even better results than conventional machine learning approaches. In this paper, in order to validate the performance of our neural network architecture for emotion recognition from speech, we first trained and tested it on the widely used corpus accessible by the community, IEMOCAP. We then used the same architecture with the real life corpus, CEMO, comprised of 440 dialogs (2h16m) from 485 speakers. The most frequent emotions expressed by callers in these real-life emergency dialogues are fear, anger and positive emotions such as relief. In the IEMOCAP general topic conversations, the most frequent emotions are sadness, anger and happiness. Using the same end-to-end deep learning architecture, an Unweighted Accuracy Recall (UA) of 63% is obtained on IEMOCAP and a UA of 45.6% on CEMO, each with 4 classes. Using only 2 classes (Anger, Neutral), the results for CEMO are 76.9% UA compared to 81.1% UA for IEMOCAP. We expect that these encouraging results with CEMO can be improved by combining the audio channel with the linguistic channel. Real-life emotions are clearly more complex than acted ones, mainly due to the large diversity of emotional expressions of speakers.
Théo Deschamps-Berger, Lori Lamel, Laurence Devillers
ACII2
2021 Modeling the Effect of Military Oxygen Masks on Speech Characteristics
abstract
International audience
Benjamin Elie, Jodie Gauvain, Jean-Luc Gauvain, Lori Lamel
Interspeech4
2021 Synchronic Fortition in Five Romance Languages? A Large Corpus-Based Study of Word-Initial Devoicing
abstract
International audience
Mathilde Hutin, Yaru Wu, Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker
Interspeech5
2020 Ongoing Phonologization of Word-Final Voicing Alternations in Two Romance Languages: Romanian and French
abstract
International audience
Mathilde Hutin, Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker
INTERSPEECH4
2020 Mandarin Lexical Tones: A Corpus-Based Study of Word Length, Syllable Position and Prosodic Position on Duration
abstract
International audience
Yaru Wu, Martine Adda-Decker, Lori Lamel
INTERSPEECH3
2019 " Gra[f] e!" Word-Final Devoicing of Obstruents in Standard French: An Acoustic Study Based on Large Corpora
abstract
International audience
Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker, Nicolas Audibert
INTERSPEECH3
2019 Challenges in Audio Processing of Terrorist-Related Data
Jodie Gauvain, Lori Lamel, Viet Bac Le, Julien Despres, Jean-Luc Gauvain, Abdelkhalek Messaoudi, Bianca Vieru-Dimulescu, Waad Ben Kheder
MMM (2)2
2018 Exploring Temporal Reduction in Dialectal Spanish: A Large-scale Study of Lenition of Voiced Stops and Coda-s
abstract
International audience
Ioana Vasilescu, Nidia Hernández, Bianca Vieru-Dimulescu, Lori Lamel
INTERSPEECH4
2018 Studying Vowel Variation in French-Algerian Arabic Code-switched Speech
abstract
International audience
Jane Wottawa, Djegdjiga Amazouz, Martine Adda-Decker, Lori Lamel
INTERSPEECH4
2018 The French-Algerian Code-Switching Triggered audio corpus (FACST)
Djegdjiga Amazouz, Martine Adda-Decker, Lori Lamel
LREC3
2018 A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noël Kouarata, Lori Lamel, Hélène Bonneau-Maynard, Markus Müller 0001, Annie Rialland, Sebastian Stüker, François Yvon, Marcely Zanon Boito
LREC8
2018 Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville)
Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt
LREC6
2018 Conversational telephone speech recognition for Lithuanian
Rasa Lileikyte, Lori Lamel, Jean-Luc Gauvain, Arseniy Gorin
Comput. Speech Lang.2
2017 An investigation into language model data augmentation for low-resourced STT and KWS
abstract
This paper reports on investigations using two techniques for language model text data augmentation for low-resourced automatic speech recognition and keyword search. Lowresourced languages are characterized by limited training materials, which typically results in high out-of-vocabulary (OOV) rates and poor language model estimates. One technique makes use of recurrent neural networks (RNNs) using word or subword units. Word-based RNNs keep the same system vocabulary, so they cannot reduce the OOV, whereas subword units can reduce the OOV but generate many false combinations. A complementary technique is based on automatic machine translation, which requires parallel texts and is able to add words to the vocabulary. These methods were assessed on 10 languages in the context of the Babel program and NIST OpenKWS evaluation. Although improvements vary across languages with both methods, small gains were generally observed in terms of word error rate reduction and improved keyword search performance.
Guangpu Huang, Thiago Fraga-Silva, Lori Lamel, Jean-Luc Gauvain, Arseniy Gorin, Antoine Laurent, Rasa Lileikyte, Abdel Messouadi
ICASSP3
2017 Effective keyword search for low-resourced conversational speech
abstract
In this paper we aim to enhance keyword search for conversational telephone speech under low-resourced conditions. Two techniques to improve the detection of out-of-vocabulary keywords are assessed in this study: using extra text resources to augment the lexicon and language model, and via subword units for keyword search. Two approaches for data augmentation are explored to extend the limited amount of transcribed conversational speech: using conversational-like Web data and texts generated by recurrent neural networks. Contrastive comparisons of subword-based systems are performed to evaluate the benefits of multiple subword decodings and single decoding. Keyword search results are reported for all the techniques, but only some improve performance. Results are reported for the Mongolian and Igbo languages using data from the 2016 Babel program.
Rasa Lileikyte, Thiago Fraga-Silva, Lori Lamel, Jean-Luc Gauvain, Antoine Laurent, Guangpu Huang
ICASSP3
2017 Addressing Code-Switching in French/Algerian Arabic Speech
abstract
International audience
Djegdjiga Amazouz, Martine Adda-Decker, Lori Lamel
INTERSPEECH3
2017 Developing an Embosi (Bantu C25) Speech Variant Dictionary to Model Vowel Elision and Morpheme Deletion
abstract
International audience
Jamison Cooper-Leavitt, Lori Lamel, Annie Rialland, Martine Adda-Decker, Gilles Adda
INTERSPEECH2
2017 Schwa Realization in French: Using Automatic Speech Processing to Study Phonological and Socio-Linguistic Factors in Large Corpora
abstract
International audience
Yaru Wu, Martine Adda-Decker, Cécile Fougeron, Lori Lamel
INTERSPEECH4
2016 Machine translation based data augmentation for Cantonese keyword spotting
abstract
This paper presents a method to improve a language model for a limited-resourced language using statistical machine translation from a related language to generate data for the target language. In this work, the machine translation model is trained on a corpus of parallel Mandarin-Cantonese subtitles and used to translate a large set of Mandarin conversational telephone transcripts to Cantonese, which has limited resources. The translated transcripts are used to train a more robust language model for speech recognition and for keyword search in Cantonese conversational telephone speech. This method enables the keyword search system to detect 1.5 times more out-of-vocabulary words, and achieve 1.7% absolute improvement on actual term-weighted value.
Guangpu Huang, Arseniy Gorin, Jean-Luc Gauvain, Lori Lamel
ICASSP4
2016 Investigating techniques for low resource conversational speech recognition
abstract
In this paper we investigate various techniques in order to build effective speech to text (STT) and keyword search (KWS) systems for low resource conversational speech. Subword decoding and graphemic mappings were assessed in order to detect out-of-vocabulary keywords. To deal with the limited amount of transcribed data, semi-supervised training and data selection methods were investigated. Robust acoustic features produced via data augmentation were evaluated for acoustic modeling. For language modeling, automatically retrieved conversational-like Webdata was used, as well as neural network based models. We report STT improvements with all the techniques, but interestingly only some improve KWS performance. Results are reported for the Swahili language in the context of the 2015 OpenKWS Evaluation.
Antoine Laurent, Thiago Fraga-Silva, Lori Lamel, Jean-Luc Gauvain
ICASSP3
2016 Language Model Data Augmentation for Keyword Spotting in Low-Resourced Training Conditions
abstract
International audience
Arseniy Gorin, Rasa Lileikyte, Guangpu Huang, Lori Lamel, Jean-Luc Gauvain, Antoine Laurent
INTERSPEECH4
2016 Marginal Contrast Among Romanian Vowels: Evidence from ASR and Functional Load
abstract
International audience
Margaret E. L. Renwick, Ioana Vasilescu, Camille Dutrey, Lori Lamel, Bianca Vieru-Dimulescu
INTERSPEECH4
2015 Improving data selection for low-resource STT and KWS
abstract
This paper extends recent research on training data selection for speech transcription and keyword spotting system development. Selection techniques were explored in the context of the IARPA-Babel Active Learning (AL) task for 6 languages. Different selection criteria were considered with the goal of improving over a system built using a pre-defined 3-hour training data set. Four variants of the entropy-based criterion were explored: words, triphones, phones as well as the use of HMM-states previously introduced in [4]. The influence of the number of HMM-states was assessed as well as whether automatic or manual reference transcripts were used. The combination of selection criteria was investigated, and a novel multi-stage selection method proposed. This method was also assessed using larger data sets than were permitted in the Babel AL task. Results are reported for the 6 languages. The multi-stage selection was also applied to the surprise language (Swahili) in the NIST OpenKWS 2015 evaluation.
Thiago Fraga-Silva, Antoine Laurent, Jean-Luc Gauvain, Lori Lamel, Viet Bac Le, Abdelkhalek Messaoudi
ASRU4
2015 Active learning based data selection for limited resource STT and KWS
abstract
International audience
Thiago Fraga-Silva, Jean-Luc Gauvain, Lori Lamel, Antoine Laurent, Viet Bac Le, Abdelkhalek Messaoudi
INTERSPEECH3
2015 Analysing rhythm in ritual discourse in yucatec maya using automatic speech alignment
abstract
International audience
Valentina Vapnarsky, Claude Barras, Cédric Becquey, David Doukhan, Martine Adda-Decker, Lori Lamel
INTERSPEECH6
2014 Comparing decoding strategies for subword-based keyword spotting in low-resourced languages
abstract
For languages with limited training resources, out-of-vocabulary (OOV) words are a significant problem, both for transcription and keyword spotting. This paper investigates the use of subword lexical units for keyword spotting. Three strate-gies for using the sub-word units are explored: 1) converting word-based lattices to subword lattices after decoding, 2) per-forming a separate decoding for each subword type, and 3) a single decoding using all possible subword units. In these ex-periments, the best performance is achieved by carrying out a separate decoding for each subword type. Further gains are at-tained through system combination. We also find that ignor-
William Hartmann, Viet Bac Le, Abdelkhalek Messaoudi, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH4
2014 Language diversity: speech processing in a multi-lingual context
Lori Lamel
INTERSPEECH1
2014 Developing STT and KWS systems using limited language resources
abstract
This paper presents recent progress in developing speech-to-text (STT) and keyword spotting (KWS) systems for the 2014 IARPA-Babel evaluation. Systems have been developed for the limited language pack condition for four of the five de-velopment languages in this program phase: Assamese, Ben-gali, Haitian Creole and Zulu. The systems have several novel characteristics that support rapid development of KWS systems. On the STT side different acoustic units are explored based on phonemic or graphemic representations, and system combina-tion is used to improve STT performance. The acoustic models are trained on only 10 hours of speech data with manual tran-scriptions, completed with unsupervised training on additional untranscribed data. Both word and subword units (morphologi-cally decomposed, syllables, phonemes) are used for KWS. The KWS systems are based on the multi-hypotheses produced by a consensus network decoding or searching word lattices. The word error rates of the individual STT systems are on the or-der of 50-60%, and the KWS systems obtain Maximum Term Weighted Values ranging from 30-45 % for all keywords (in-vocabulary and out-of-vocabulary (OOV)). Sub-word units are shown to be successful at locating some of the OOV keywords, and system combination improves system performance. Index Terms: STT, KWS, semi-supervised training, lattice, consensus network, sub-word lexical units, Morfessor,
Viet Bac Le, Lori Lamel, Abdelkhalek Messaoudi, William Hartmann, Jean-Luc Gauvain, Cécile Woehrling, Julien Despres, Anindya Roy
INTERSPEECH2
2014 Automatic language identity tagging on word and sentence-level in multilingual text sources: a case-study on Luxembourgish
Thomas Lavergne, Gilles Adda, Martine Adda-Decker, Lori Lamel
LREC4
2014 Human annotation of ASR error regions: Is "gravity" a sharable concept for human annotators?
Daniel Luzzati, Cyril Grouin, Ioana Vasilescu, Martine Adda-Decker, Eric Bilinski, Nathalie Camelin, Juliette Kahn, Carole Lailler, Lori Lamel, Sophie Rosset
LREC9
2013 Acoustic unit discovery and pronunciation generation from a grapheme-based lexicon
abstract
We present a framework for discovering acoustic units and generating an associated pronunciation lexicon from an initial grapheme-based recognition system. Our approach consists of two distinct contributions. First, context-dependent grapheme models are clustered using a spectral clustering approach to create a set of phone-like acoustic units. Next, we transform the pronunciation lexicon using a statistical machine translation-based approach. Pronunciation hypotheses generated from a decoding of the training set are used to create a phrase-based translation table. We propose a novel method for scoring the phrase-based rules that significantly improves the output of the transformation process. Results on an English language dataset demonstrate the combined methods provide a 13% relative reduction in word error rate compared to a baseline grapheme-based system. Our approach could potentially be applied to low-resource languages without existing lexicons, such as in the Babel project.
William Hartmann, Anindya Roy, Lori Lamel, Jean-Luc Gauvain
ASRU3
2013 Score normalization and system combination for improved keyword spotting
abstract
We present two techniques that are shown to yield improved Keyword Spotting (KWS) performance when using the ATWV/MTWV performance measures: (i) score normalization, where the scores of different keywords become commensurate with each other and they more closely correspond to the probability of being correct than raw posteriors; and (ii) system combination, where the detections of multiple systems are merged together, and their scores are interpolated with weights which are optimized using MTWV as the maximization criterion. Both score normalization and system combination approaches show that significant gains in ATWV/MTWV can be obtained, sometimes on the order of 8-10 points (absolute), in five different languages. A variant of these methods resulted in the highest performance for the official surprise language evaluation for the IARPA-funded Babel project in April 2013.
Damianos Karakos, Richard M. Schwartz, Stavros Tsakalidis, Le Zhang 0002, Shivesh Ranjan, Tim Ng, Roger Hsiao, Guruprasad Saikumar, Ivan Bulyko, Long Nguyen 0001, John Makhoul, Frantisek Grézl, Mirko Hannemann, Martin Karafiát, Igor Szöke, Karel Veselý, Lori Lamel, Viet Bac Le
ASRU17
2013 Rapid development of a Latvian speech-to-text system
abstract
This paper describes the development of a Latvian speech-to-text (STT) system at LIMSI within the Quaero project. One of the aims of the speech processing activities in the Quaero project is to cover all official European languages. However, for some of the languages only very limited, if any, training resources are available via corpora agencies such as LDC and ELRA. The aim of this study was to show the way, taking Latvian as example, an STT system can be rapidly developed without any transcribed training data. Following the scheme proposed in this paper, the Latvian STT system was developed in about a month and obtained a word error rate of 20% on broadcast news and conversation data in the Quaero 2012 evaluation campaign.
Ilya Oparin, Lori Lamel, Jean-Luc Gauvain
ICASSP2
2013 Recent evolution of non-standard consonantal variants in French broadcast news
abstract
This paper investigates sociophonetic questions about global tendencies in contemporaneous European spoken French. The authors argue that automatic alignment allowing targeted variants can provide evidence for current hypotheses about possible ongoing sound changes or about destandardization even in formal contexts as broadcast news. This study focused on the evolution over a decade, in radio or TV news, of three 'non- standard' consonantal variants: consonant cluster reduction, affrication/palatalization of dental stops and voiceless fricative epithesis. Measures obtained by this method showed that the first variant remains almost absent in journalists' speech, exactly as affrication of /d/. In contrast, affrication of /t/ is increasing and the fricative epithesis, partially unpredictable, becomes longer. Our findings support the use of automatic alignment as an aid to validate sociolinguistic hypotheses and to develop pattern-driven studies, gathering more variables.
Maria Candea, Martine Adda-Decker, Lori Lamel
INTERSPEECH3
2013 Interpolation of acoustic models for speech recognition
abstract
International audience
Thiago Fraga-Silva, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2013 Discriminative training of a phoneme confusion model for a dynamic lexicon in ASR
abstract
International audience
Panagiota Karanasou, François Yvon, Thomas Lavergne, Lori Lamel
INTERSPEECH4
2013 Some issues affecting the transcription of Hungarian broadcast audio
abstract
International audience
Anindya Roy, Lori Lamel, Thiago Fraga-Silva, Jean-Luc Gauvain, Ilya Oparin
INTERSPEECH2
2013 Blip10000: a social video dataset containing SPUG content for tagging and retrieval
abstract
The increasing amount of digital multimedia content available is inspiring potential new types of user interaction with video data. Users want to easily find the content by searching and browsing. For this reason, techniques are needed that allow automatic categorisation, searching the content and linking to related information. In this work, we present a dataset that contains comprehensive semi-professional user-generated (SPUG) content, including audiovisual content, user-contributed metadata, automatic speech recognition transcripts, automatic shot boundary files, and social information for multiple 'social levels'. We describe the principal characteristics of this dataset and present results that have been achieved on different tasks.
Sebastian Schmiedeke, Isabelle Ferrané, Maria Eskevich, Christoph Kofler, Martha A. Larson, Yannick Estève, Lori Lamel, Gareth J. F. Jones, Thomas Sikora
MMSys8
2012 Phonotactic Language Recognition Using MLP Features
abstract
International audience
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2012 Development and Evaluation of Automatic Punctuation for French and English Speech-to-Text
abstract
International audience
Jáchym Kolár, Lori Lamel
INTERSPEECH2
2012 Cross-lingual studies of ASR errors: paradigms for perceptual evaluations
Ioana Vasilescu, Martine Adda-Decker, Lori Lamel
LREC3
2011 Automatic Generation of a Pronunciation Dictionary with Rich Variation Coverage Using SMT Methods
Panagiota Karanasou, Lori Lamel
CICLing (2)2
2011 Lattice-based unsupervised acoustic model training
abstract
Unsupervised acoustic model training has been successfully used to improve the performance of automatic speech recognition systems when only a small amount of manually transcribed data is available for the target domain. The most common approach is use automatic transcriptions to guide acoustic model estimation. However, since the best recognition hypotheses are known to contain errors, we propose to consider multiple transcription hypotheses during training. The idea is that the EM process can benefit from the estimated posterior probabilities of the hypotheses to converge to a better solution. The proposed unsupervised training method is based on lattices. Lattice-based training gives a relative improvement of 2.2% over 1-best training on a Broadcast News transcription task and converges faster with the iterative incremental training.
Thiago Fraga-Silva, Jean-Luc Gauvain, Lori Lamel
ICASSP3
2011 Pronunciation variants generation using SMT-inspired approaches
abstract
Enriching a pronunciation dictionary with phonological variation is a challenging task, not yet solved despite several decades of research, in particular for speech-to-text transcription of real world data where it is important to cover different pronunciation variants. This paper proposes two alternative methods, inspired by machine translation, to derive pronunciation variants from an initial lexicon with limited variations. In the first case, an n-best pronunciation list is extracted directly from a machine translation tool, used as a grapheme-to-phoneme (g2p) converter. The second is a novel method based on a pivot approach, previously used for the paraphrase extraction task, and here applied as a post-processing step to the g2p converter. Some preliminary speech recognition experiments with the automatically generated pronunciation variants are reported using Quaero development data.
Panagiota Karanasou, Lori Lamel
ICASSP2
2011 Improved models for Mandarin speech-to-text transcription
abstract
This paper describes recent advances at LIMSI in Mandarin Chinese speech-to-text transcription. A number of novel approaches were introduced in the different system components. The acoustic models are trained on over 1600 hours of audio data from a range of sources, and include pitch and MLP features. N-gram and neural network language models are trained on very large corpora, over 3 billion words of texts; and LM adaptation was explored at different adaptation levels: per show, per snippet, or per speaker cluster. Character-based consensus decoding was found to outperform word-based consensus decoding for Mandarin. The improved system reduces the relative character error rate (CER) by about 10% on previous GALE development and evaluation data sets, obtaining a CER of 9.2% on the P4 broadcast news and broadcast conversation evaluation data.
Lori Lamel, Jean-Luc Gauvain, Viet Bac Le, Ilya Oparin, Sha Meng
ICASSP1
2011 On Development of Consistently Punctuated Speech Corpora
Jáchym Kolár, Lori Lamel
INTERSPEECH2
2011 Comparing Multi-Stage Approaches for Cross-Show Speaker Diarization
abstract
International audience
Viet-Anh Tran, Viet Bac Le, Claude Barras, Lori Lamel
INTERSPEECH4
2011 Cross-Lingual Study of ASR Errors: On the Role of the Context in Human Perception of Near-Homophones
abstract
International audience
Ioana Vasilescu, Dahbia Yahia, Natalie D. Snoeren, Martine Adda-Decker, Lori Lamel
INTERSPEECH5
2011 Genre Categorization and Modeling for Broadcast Speech Transcription
abstract
Broadcast News (BN) speech recognition transcription has attracted research due to the challenges of the task since the mid 1990’s. More recently, research has been moving towards more spontaneous broadcast data, commonly called Broadcast Conversation (BC) speech. Considering the large style difference between BN and BC genres, specific modeling of genres should intuitively result in improved system performance. In this paper BNand BC-style speech recognition has been explored by designing genre-specific systems. In order to separate the training data, an automatic genre categorization with two novel features is proposed. Experiments showed that automatic categorization of genre labels of the training data compared favorably to the original manually specified genre labels provided with corpora. When test data sets were classified into BN or BC genres and tested by the corresponding genre-specific speech recognition systems, modest but consistent error reductions were achieved compared to the baseline genre-independent systems.
Lori Lamel, Jean-Luc Gauvain
INTERSPEECH2
2010 Multi-style MLP features for BN transcription
abstract
It has become common practice to adapt acoustic models to specific-conditions (gender, accent, bandwidth) in order to improve the performance of speech-to-text (STT) transcription systems. With the growing interest in the use of discriminative features produced by a multi layer perceptron (MLP) in such systems, the question arise of whether it is necessary to specialize the MLP to particular conditions, and if so, how to incorporate the condition-specific MLP features in the system. This paper explores three approaches (adaptation, full training, and feature merging) to use condition-specific MLP features in a state-of-the-art BN STT system for French. The third approach without condition-specific adaptation was found to outperform the original models with condition-specific adaptation, and was found to perform almost as well as full training of multiple condition-specific HMMs.
Viet Bac Le, Lori Lamel, Jean-Luc Gauvain
ICASSP2
2010 Comparing mono- & multilingual acoustic seed models for a low e-resourced language: a case-study of luxembourgish
abstract
Luxembourgish is embedded in a multilingual context on the divide between Romance and Germanic cultures and has often been viewed as one of Europe’s under-resourced languages. We focus on the acoustic modeling of Luxembourgish. By taking advantage of monolingual acoustic seeds selected from German, French or English model sets via IPA symbol correspondances, we investigated whether Luxembourgish spoken words were globally better represented by one of these languages. Although speech in Luxembourgish is frequently interspersed with French words, forced alignments on these data showed a clear preference for Germanic acoustic models with only a limited usage of French. German models provided the best match with 54% of the data, 35% for English and only 11% for French models. A set of multilingual acoustic models, estimated the pooled German, French, and English audio data, captured 27% to 48% of the data depending on conditions. Index Terms: multilingual alignment, acoustic seed models, under-resourced languages, Luxembourgish, English, French, German.
Martine Adda-Decker, Lori Lamel, Natalie D. Snoeren
INTERSPEECH2
2010 Improved n-gram phonotactic models for language recognition
abstract
This paper investigates various techniques to improve the estimation of n-gram phonotactic models for language recognition using single-best phone transcriptions and phone lattices. More precisely, we first report on the impact of the so-called acoustic scale factor on the system accuracy when using latticebased training, and then we report on the use of n-gram cutoff and entropy pruning techniques. Several system configurations are explored, such as the use of context-independent and context-dependent phone models, the use of single-best phone hypotheses versus phone lattices, and the use of various n-gram orders. Experiments are conducted using the LRE 2007 evaluation data and the results are reported using the a posteriori EER. The results show that the impact of these techniques on the system accuracy is highly dependent on the training conditions and that careful optimization can lead to performance improvements.
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2010 Automatic speech recognition of multiple accented English data
abstract
Accent variability is an important factor in speech that can sig-nificantly degrade automatic speech recognition performance. We investigate the effect of multiple accents on an English broadcast news recognition system. A multi-accented English corpus is used for the task, including broadcast news segments from 6 different geographic regions: US, Great Britain, Aus-tralia, North Africa, Middle East and India. There is signifi-cant performance degradation of a baseline system trained on only US data when confronted with shows from other regions. The results improve significantly when data from all the regions are included for accent-independent acoustic model training. Further improvements are achieved when MAP-adapted accent-dependent models are used in conjunction with a GMM accent classifier. Index Terms: accented speech recognition, accent adaptation 1.
Dimitra Vergyri, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH2
2010 Evaluation Protocol and Tools for Question-Answering on Speech Transcripts
Nicolas Moreau, Olivier Hamon, Djamel Mostefa, Sophie Rosset, Olivier Galibert, Lori Lamel, Jordi Turmo, Pere Comas, Paolo Rosso, Davide Buscaldi, Khalid Choukri
LREC6
2009 Gaussian Backend design for open-set language detection
abstract
This paper proposes a new approach to the challenging open-set language detection task. Most state-of-the-art approaches make use of data sources with several out-of-set languages to model such languages. In the proposed approach, no additional data from out-ofset languages is required, only date from the target languages is used. Experiments are conducted using the LRE-05 and the LRE-07 evaluation data sets with the 30s condition. A Cavgof 4.5% and 3.4% is obtained on these data set, respectively. These results are comparable with other reported results.
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
ICASSP3
2009 Modeling characters versuswords for mandarin speech recognition
abstract
Word based models are widely used in speech recognition since they typically perform well. However, the question of whether it is better to use a word-based or a character-based model warrants being for the Mandarin Chinese language. Since Chinese is written without any spaces or word delimiters, a word segmentation algorithm is applied in a pre-processing step prior to training a word-based language model. Chinese characters carry meaning and speakers are free to combine characters to construct new words. This suggests that character information can also be useful in communication. This paper explores both word-based and character-based models, and their complementarity. Although word-based modeling is found to outperform character-based modeling, increasing the vocabulary size from 56 k to 160 k words did not lead to a gain in performance. Results are reported for the Gale Mandarin speech-to-text task.
Lori Lamel, Jean-Luc Gauvain
ICASSP2
2009 Language score calibration using adapted Gaussian back-end
abstract
Generative Gaussian back-end and discriminative logistic regression are the most used approaches for language score fusion and calibration. Combination of these two approaches can significantly improve the performance. This paper proposes the use of an adapted Gaussian back-end, where the mean of the language-dependent Gaussian is adapted from the mean of a language-specific background Gaussian via maximum a posteriori estimation algorithm. Experiments are conducted using the LRE-07 evaluation data. Compared to the conventional Gaussian back-end approach for a closed set task, relative improvements in the Cavg of 50%, 17% and 4.2% are obtained on the 30s, 10s and 3s conditions, respectively. Besides this, the estimated scores are better calibrated. A combination with logistic regression results in a system with the best calibrated scores. Index Terms: Language recognition, Gaussian back-end, Adaptation
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2009 Modeling northern and southern varieties of dutch for STT
abstract
This paper describes how the Northern (NL) and Southern (VL) varieties of Dutch are modeled in the joint LIMSIVecsys Research speech-to-text transcription systems for broadcast news (BN) and conversational telephone speech (CTS). Using the Spoken Dutch Corpus resources (CGN), systems were developed and evaluated in the 2008 N-Best benchmark. Modeling techniques that are used in our systems for other languages were found to be effective for the Dutch language, however it was also found to be important to have acoustic and language models, and statistical pronunciation generation rules adapted to each variety. This was in particular true for the MLP features which were only effective when trained separately for Dutch and Flemish. The joint submissions obtained the lowest WERs in the benchmark by a significant margin. Index Terms: speech recognition, Dutch, Flemish, CGN, Nbest, broadcast news, conversational telephone speech, MLP.
Julien Despres, Petr Fousek, Jean-Luc Gauvain, Sandrine Gay, Yvan Josse, Lori Lamel, Abdelkhalek Messaoudi
INTERSPEECH6
2009 A perceptual investigation of speech transcription errors involving frequent near-homophones in French and american English
abstract
This article compares the errors made by automatic speech recognizers to those made by humans for near-homophones in American English and French. This exploratory study focuses on the impact of limited word context and the potential resulting ambiguities for automatic speech recognition (ASR) systems and human listeners. Perceptual experiments using 7-gram chunks centered on incorrect or correct words output by an ASR system, show that humans make significantly more transcription errors on the first type of stimuli, thus highlighting the local ambiguity. The long-term aim of this study is to improve the modeling of such ambiguous items in order to reduce ASR errors.
Ioana Vasilescu, Martine Adda-Decker, Lori Lamel, Pierre A. Hallé
INTERSPEECH3
2009 Automatic Speech-to-Text Transcription in Arabic
abstract
The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Being a highly inflected language, the Arabic language has a very large lexical variety and typically with several possible (generally semantically linked) vowelizations for each written form. This article summarizes research carried out over the last few years on speech-to-text transcription of broadcast data in Arabic. The initial research was oriented toward processing of broadcast news data in Modern Standard Arabic, and has since been extended to address a larger variety of broadcast data, which as a consequence results in the need to also be able to handle dialectal speech. While standard techniques in speech recognition have been shown to apply well to the Arabic language, taking into account language specificities help to significantly improve system performance.
Lori Lamel, Abdelkhalek Messaoudi, Jean-Luc Gauvain
ACM Trans. Asian Lang. Inf. Process.1
2009 Automatic Word Decompounding for ASR in a Morphologically Rich Language: Application to Amharic
abstract
This paper investigates a data-driven word decompounding algorithm for use in automatic speech recognition. An existing algorithm, called ldquoMorfessor,rdquo has been enhanced in order to address the problem of increased phonetic confusability arising from word decompounding by incorporating phonetic properties and some constraints on recognition units derived from forced alignments experiments. Speech recognition experiments have been carried out on a broadcast news task for the Amharic language to validate the approach. The out of vocabulary (OOV) word rates were reduced by 35% to 50% and a small reduction in word error rate (WER) has been achieved. The algorithm is relatively language independent and requires minimal adaptation to be applied to other languages.
Thomas Pellegrini, Lori Lamel
IEEE Trans. Speech Audio Process.2
2008 Context-dependent phone models and models adaptation for phonotactic language recognition
abstract
The performance of a PPRLM language recognition system depends on the quality and the consistency of phone decoders. To improve the performance of the decoders, this paper investigates the use of context-dependent instead of contextindependent phone models, and the use of CMLLR for model adaptation. This paper also discusses several improvements to the LIMSI 2007 NIST LRE system, including the use of a 4gram language model, score calibration and fusion using the FoCalMulti-class toolkit (with large development data) and better decoding parameters such as phone insertion penalty. The improved system is evaluated on the NIST LRE-2005 and the LRE-2007 evaluation data sets. Despite its simplicity, the system achieves for the 30s condition a Cavg of 2.4% and 1.6% on these data sets, respectively.
Mohamed Faouzi BenZeghiba, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2008 Transcribing broadcast data using MLP features
abstract
This paper describes incorporating discriminative features from a multi layer perceptron (MLP) into a state-of-the-art Arabic broadcast data transcription system based on cepstral features. The MLP features are based on a recently proposed Bottle-Neck architecture with long-term warped LPTRAP speech representation at the input. It is shown that the previously reported improvements on a development Arabic transcription system carry through to a full system at a state-ofthe-art level. SAT, CMLLR and MLLR adaptation techniques are shown to be useful for both MLP and combined features, though to a lesser degree than for PLPs. Without adaptation, MLP features obtain superior performance to cepstral features in all test conditions, and with adaptation both feature sets give comparable results. Combining the features, either by feature concatenation or system hypotheses, gives significant gains. Gains from MMI model training seem to be additive to the gain coming from discriminative MLP features.
Petr Fousek, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH2
2008 Investigating morphological decomposition for transcription of Arabic broadcast news and broadcast conversation data
abstract
One of the challenges of Arabic speech recognition is to deal with the huge lexical variety. Morphological decomposition has been proposed to address this problem by increasing lexical coverage, thereby reducing errors that are due to words that are unknown to the system. In our previous attempts to develop an Arabic speech-to-text (STT) transcription system with morphological decomposition, an increase in word error rate of about 2% absolute was observed relative to a comparable word based system. Based on an error analysis and a comparison of our approach with that of other sites, two modifications were made. The first modification was to not decompose the most frequent words; and the second to not decompose the prefix ’Al’ for words starting with a solar consonant since due to assimilation with the following consonant, deletion of the prefix was one of the most frequent errors. Comparable recognition performance was achieved using word-based and morphologically decomposed language models, and since the errors made by the systems are different, combining the two gave a performance gain.
Lori Lamel, Abdelkhalek Messaoudi, Jean-Luc Gauvain
INTERSPEECH1
2008 A corpus-based prosodic study of Alsatian, Belgian and Swiss French
abstract
The object of this paper is a prosodic study of the French language as it is spoken in Alsace, Belgium and Switzerland, also compared with standard French through large corpora (over 100 hours) of scripted and spontaneous speech. The data were segmented into phones by automatic alignment; pitch values were extracted and averaged over segments. Two features are addressed: initial stress (through pitch and duration correlates) and penultimate lengthening. Different patterns enable us to distinguish the three varieties under investigation. Swiss speakers exhibit pitch rise and polysyllabic word onset lengthening in clitic–nonclitic sequences, while Alsatians tend to lengthen the initial vowel of nonclitic words. Belgians show prepausal penultimate lengthening whereas the Swiss tend to lengthen the last two prepausal vowels.
Cécile Woehrling, Philippe Boula de Mareüil, Martine Adda-Decker, Lori Lamel
INTERSPEECH4
2008 CallSurf: Automatic Transcription, Indexing and Structuration of Call Center Conversational Speech for Knowledge Extraction and Query by Content
Martine Garnier-Rizet, Gilles Adda, Frédérik Cailliau, Jean-Luc Gauvain, Sylvie Guillemin-Lanne, Lori Lamel, Stephan Vanni, Claire Waast-Richard
LREC6
2008 Question Answering on Speech Transcriptions: the QAST evaluation in CLEF
Lori Lamel, Sophie Rosset, Christelle Ayache, Djamel Mostefa, Jordi Turmo, Pere Comas
LREC1
2008 Multi-level information and automatic dialog act detection in human-human spoken dialogs
Sophie Rosset, Delphine Tribout, Lori Lamel
Speech Commun.3
2007 Speech Recognition System Combination for Machine Translation
abstract
The majority of state-of-the-art speech recognition systems make use of system combination. The combination approaches adopted have traditionally been tuned to minimising word error rates (WERs). In recent years there has been a growing interest in taking the output from speech recognition systems in one language and translating it into another. This paper investigates the use of cross-site combination approaches in terms of both WER and impact on translation performance. In addition, the stages involved in modifying the output from a speech-to-text (STT) system to be suitable for translation are described. Two source languages, Mandarin and Arabic, are recognised and then translated using a phrase-based statistical machine translation system into English. Performance of individual systems and cross-site combination using cross-adaptation and ROVER are given. Results show that the best STT combination scheme in terms of WER is not necessarily the most appropriate when translating speech.
Mark J. F. Gales, Xunying Liu, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004, Spyridon Matsoukas, Tim Ng, Kham Nguyen, Long Nguyen 0001, Jean-Luc Gauvain, Lori Lamel, Abdelkhalek Messaoudi
ICASSP (4)11
2007 The LIMSI 2006 TC-STAR EPPS Transcription Systems
abstract
This paper describes the speech recognizers developed to transcribe European Parliament Plenary Sessions (EPPS) in English and Spanish in the 2nd TC-STAR Evaluation Campaign. The speech recognizers are state-of-the-art systems using multiple decoding passes with models (lexicon, acoustic models, language models) trained for the different transcription tasks. Compared to the LIMSI TC-STAR 2005 EPPS systems, relative word error rate reductions of about 30% have been achieved on the 2006 development data. The word error rates with the LIMSI systems on the 2006 EPPS evaluation data are 8.2% for English and 7.8% for Spanish. Experiments with cross-site adaptation and system combination are also described.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Claude Barras, Eric Bilinski, Olivier Galibert, Agusti Pujol, Holger Schwenk, Xuan Zhu 0001
ICASSP (4)1
2007 Improved acoustic modeling for transcribing Arabic broadcast data
abstract
ABSTRACT This paper summarizes our recent progress in improving theautomatic transcription of Arabic broadcast audio data, andsome efforts to address the challenges of the broadcast con-versational speech. Our efforts are aimed at improving theacoustic, pronunciation and language models taking into ac-count specificities of the Arabic language. In previous work wedemonstrated that explicit modeling of short vowels improvedrecognition performance, even when producing non-vocalizedhypotheses. In addition to modeling short vowels, consonantgemination and nunation are now explicitly modeled, alterna-tive pronunciations have been introduced to better represent di-alectical variants, and a duration model has been integrated.In order to facilitate training on Arabic audio data with non-vocalized transcripts a generic vowel model has been intro-duced. Compared with the previous system (used in the 2006GALE evaluation) the relative word error rate has been reducedby over 10%. Index Terms – Speech recognition, Arabic, broadcast news,broadcast conversations
Lori Lamel, Abdelkhalek Messaoudi, Jean-Luc Gauvain
INTERSPEECH1
2007 Using phonetic features in unsupervised word decompounding for ASR with application to a less-represented language
abstract
In this paper, a data-driven word decompounding algorithm is described and applied to a broadcast news corpus in Amharic. The baseline algorithm has been enhanced in order to address the problem of increased phonetic confusability arising from word decompounding by incorporating phonetic properties and some constraints on recognition units derived from prior forced alignment experiments. Speech recognition experiments have been carried out to validate the approach. Out of vocabulary (OOV) words rates can be reduced by 30% to 40% and an absolute Word Error Rate (WER) reduction of 0.4% has been achieved. The algorithm is relatively language independent and requires minimal adaptation to be applied to other languages. Index Terms: automatic speech recognition, unsupervised word decompounding, less-represented languages
Thomas Pellegrini, Lori Lamel
INTERSPEECH2
2006 Arabic Broadcast News Transcription Using a One Million Word Vocalized Vocabulary
abstract
Recently it has been shown that modeling short vowels in Arabic can significantly improve performance even when producing a non-vocalized transcript. Since Arabic texts and audio transcripts are almost exclusively non-vocalized, the training methods have to overcome this missing data problem. For the acoustic models the procedure was bootstrapped with manually vocalized data and extended with semi-automatically vocalized data. In order to also capture the vowel information in the language model, a vocalized 4-gram language model trained on the audio transcripts was interpolated with the original 4-gram model trained on the (non-vocalized) written texts. Another challenge of the Arabic language is its large lexical variety. The out-of-vocabulary rate with a 65k word vocabulary is in the range of 4-8% (compared to under 1% for English). To address this problem a vocalized vocabulary containing over 1 million vocalized words, grouped into 200k word classes is used. This reduces the out-of-vocabulary rate to about 2%. The extended vocabulary and vocalized language model trained on the manually annotated data give a 1.2% absolute word error reduction on the DARPA RT04 development data. However, including the automatically vocalized transcripts in the language model reduces performance indicating that automatic vocalization needs to be improved
Abdelkhalek Messaoudi, Jean-Luc Gauvain, Lori Lamel
ICASSP (1)3
2006 Investigating automatic decomposition for ASR in less represented languages
abstract
This paper addresses the use of an automatic decomposition method to reduce lexical variety and thereby improve speech recognition of less well-represented languages. The Amharic language has been selected for these experiments since only a small quantity of resources are available compared to well-covered languages. Inspired by the Harris algorithm, the method automatically generates plausible affixes, that combined with decompounding can reduce the size of the lexicon and the OOV rate. Recognition experiments are carried out for four different configurations (full-word and decompounded) and using supervised training with a corpus containing only two hours of manually transcribed data.
Thomas Pellegrini, Lori Lamel
INTERSPEECH2
2006 Experimental detection of vowel pronunciation variants in Amharic
Thomas Pellegrini, Lori Lamel
LREC2
2006 Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system
abstract
This paper describes the progress made in the transcription of broadcast news (BN) and conversational telephone speech (CTS) within the combined BBN/LIMSI system from May 2002 to September 2004. During that period, BBN and LIMSI collaborated in an effort to produce significant reductions in the word error rate (WER), as directed by the aggressive goals of the Effective, Affordable, Reusable, Speech-to-text [Defense Advanced Research Projects Agency (DARPA) EARS] program. The paper focuses on general modeling techniques that led to recognition accuracy improvements, as well as engineering approaches that enabled efficient use of large amounts of training data and fast decoding architectures. Special attention is given on efforts to integrate components of the BBN and LIMSI systems, discussing the tradeoff between speed and accuracy for various system combination strategies. Results on the EARS progress test sets show that the combined BBN/LIMSI system achieved relative reductions of 47% and 51% on the BN and CTS domains, respectively.
Spyridon Matsoukas, Jean-Luc Gauvain, Gilles Adda, Thomas Colthurst, Chia-Lin Kao, Owen Kimball, Lori Lamel, Fabrice Lefèvre, Jeff Z. Ma, John Makhoul, Long Nguyen 0001, Rohit Prasad, Richard M. Schwartz, Holger Schwenk, Bing Xiang
IEEE Trans. Speech Audio Process.7
2005 Alternate Phone Models for Conversational Speech
abstract
This paper investigates the use of alternate phone models for the transcription of conversational telephone speech. The focus of this work is to explore alternative ways of modeling different manners of speaking so as to better cover the observed articulatory styles and pronunciation variants. Four alternate phone sets are compared ranging from 38 to 129 units. Two of the phone sets make use of syllable-position dependent phone models. The acoustic models were trained on 2300 hours of conversational telephone speech data from the Switchboard and Fisher corpora, and experimental results are reported on the EARS Dev04 test set which contains 3 hours of speech from 36 Fisher conversations. While no one particular phone set was found to outperform the others for a majority of speakers, the best overall performance was obtained with the original 48 phone set and a reduced 38 phone set, however combining the hypotheses of the individual models reduces the word error rate from 17.5% (original phone set) to 16.8%.
Lori Lamel, Jean-Luc Gauvain
ICASSP (1)1
2005 Do speech recognizers prefer female speakers?
Martine Adda-Decker, Lori Lamel
INTERSPEECH2
2005 Where are we in transcribing French broadcast news?
abstract
International audience
Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Véronique Gendner, Lori Lamel, Holger Schwenk
INTERSPEECH6
2005 Transcribing lectures and seminars
Lori Lamel, Gilles Adda, Eric Bilinski, Jean-Luc Gauvain
INTERSPEECH1
2005 Modeling vowels for Arabic BN transcription
abstract
This paper describes the LIMSI Arabic Broadcast News system which produces a vowelized word transcription. The under 10x system, evaluated in the NIST RT-04F evaluation, uses a 3 pass decoding strategy with gender- and bandwidth-specific acoustic models, a vowelized 65k word class pronunciation lexicon and a word-class 4-gram language model. In order to explicitly represent the vowelized word forms, each nonvowelized word entry is considered as a word class regrouping all of its associated vowelized forms. Since Arabic texts are almost exclusively written without vowels, an important challenge is to be able to use these efficiently in a system producing a vowelized output. Since a portion of the acoustic training data was manually transcribed with short vowels, enabling an initial set of acoustic models to be estimated in a supervised manner. The remaining audio data, for which vowels are not annotated, were trained in an implicit manner using the recognizer to choose the preferred form. The system was trained on a total of about 150 hours of audio data and almost 600 million words of Arabic texts, and achieved word error rates of 16.0% and 18.5% on the dev04 and eval04 data, respectively.
Abdelkhalek Messaoudi, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH2
2005 The 2004 BBN/LIMSI 20xRT English conversational telephone speech recognition system
abstract
In this paper we describe the English Conversational Telephone Speech (CTS) recognition system jointly developed by BBN and LIMSI under the DARPA EARS program for the 2004 evalua-tion conducted by NIST. The 2004 BBN/LIMSI system achieved a word error rate (WER) of 13.5 % at 18.3xRT (real-time as mea-sured on Pentium 4 Xeon 3.4 GHz Processor) on the EARS progress test set. This translates into a 22.8 % relative improvement in WER over the 2003 BBN/LIMSI EARS evaluation system, which was run without any time constraints. In addition to reporting on the system architecture and the evaluation results, we also highlight the significant improvements made at both sites. 1.
Rohit Prasad, Spyridon Matsoukas, Chia-Lin Kao, Jeff Z. Ma, Dongxin Xu, Thomas Colthurst, Owen Kimball, Richard M. Schwartz, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Fabrice Lefèvre
INTERSPEECH10
2005 Genericity and portability for task-independent speech recognition
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
Comput. Speech Lang.3
2005 Challenges in real-life emotion annotation and machine learning based detection
Laurence Devillers, Laurence Vidrascu, Lori Lamel
Neural Networks3
2005 Investigating syllabic structures and their variation in spontaneous French
Martine Adda-Decker, Philippe Boula de Mareüil, Gilles Adda, Lori Lamel
Speech Commun.4
2004 Lightly supervised acoustic model training using consensus networks
abstract
The paper presents some recent work on using consensus networks to improve lightly supervised acoustic model training for the LIMSI Mandarin BN system. Lightly supervised acoustic model training has been attracting growing interest, since it can help to reduce the development costs for speech recognition systems substantially. Compared to supervised training with accurate transcriptions, the key problem in lightly supervised training is getting the approximate transcripts to be as close as possible to manually produced detailed ones, i.e., finding a proper way to provide the information for supervision. Previous work using a language model to provide supervision has been quite successful. The paper extends the original method by presenting a new way to get the information needed for supervision during training. Studies are carried out using the TDT4 Mandarin audio corpus and associated closed-captions. After automatically recognizing the training data, the closed-captions are aligned with a consensus network derived from the hypothesized lattices. As is the case with closed-caption filtering, this method can remove speech segments whose automatic transcripts contain errors, but it can also recover errors in the hypothesis if the information is present in the lattice. Experimental results show that, compared with simply training on all of the data, consensus network based lightly supervised acoustic model training results in a small reduction in the character error rate on the DARPA/NIST RT'03 development and evaluation data.
Langzhou Chen, Lori Lamel, Jean-Luc Gauvain
ICASSP (1)2
2004 Speech transcription in multiple languages
abstract
The paper summarizes recent work underway at LIMSI on speech-to-text transcription in multiple languages. The research has been oriented towards the processing of broadcast audio and conversational speech for information access. Broadcast news transcription systems have been developed for seven languages, and it is planned to address several other languages in the near term. Research on conversational speech has mainly focused on the English language, with some initial work on French, Arabic and Spanish. Automatic processing must take into account the characteristics of the audio data, such as needing to deal with the continuous data stream, specificities of the language and the use of an imperfect word transcription for accessing the information content. Our experience thus far indicates that at today's word error rates, the techniques used in one language can be successfully ported to other languages, and most of the language specificities concern lexical and pronunciation modeling.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Leonardo Canseco-Rodriguez, Langzhou Chen, Olivier Galibert, Abdelkhalek Messaoudi, Holger Schwenk
ICASSP (3)1
2004 Speech recognition in multiple languages and domains: the 2003 BBN/LIMSI EARS system
abstract
We report on the results of the first evaluations for the BBN/LIMSI system under the new DARPA EARS program. The evaluations were carried out for conversational telephone speech (CTS) and broadcast news (BN) for three languages: English, Mandarin, and Arabic. In addition to providing system descriptions and evaluation results, the paper highlights methods that worked well across the two domains and those few that worked well on one domain but not the other. For the BN evaluations, which had to be run under 10 times real-time, we demonstrated that a joint BBN/LIMSI system with a time constraint achieved better results than either system alone.
Richard M. Schwartz, Thomas Colthurst, Nicolae Duta, Herbert Gish, Rukmini Iyer, Chia-Lin Kao, Daben Liu, Owen Kimball, Jeff Z. Ma, John Makhoul, Spyridon Matsoukas, Long Nguyen 0001, Mohammed Noamany, Rohit Prasad, Bing Xiang, Dongxin Xu, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen
ICASSP (3)18
2004 Dynamic language modeling for broadcast news
abstract
ABSTRACT This paper describes some recent experiments on unsuper-vised language model adaptation for transcription of broadcastnews data. In previous work, a framework for automaticallyselecting adaptation data using information retrieval techniqueswas proposed. This work extends the method and presents ex-perimental results with unsupervised language model adapta-tion. Threeprimaryaspectsareconsidered: (1)theperformanceof5widelyusedLMadaptationmethodsusingthesameadapta-tion data is compared; (2) the influence of the temporal distancebetween the training and test data epoch on the adaptation effi-ciency is assessed; and (3) show-based language model adapta-tion is compared with story-based language model adaptation.Experimentshavebeencarriedoutforbroadcastnewstranscrip-tion in English and Mandarin Chinese. A relative word errorrate reduction of 4.7% was obtained in English and a 5.6% rela-tive character error rate reduction in Mandarin withstory-basedMDI adaptation. 1. INTRODUCTION While n-gram models are successfully used in speech recog-nition, their performance is influenced by any mismatch be-tween thetraining and testdata [7]. Theidea of language model(LM) adaptation is to use a small amount of domain specificdata to adjust the LM to reduce the impact of linguistic differ-ences between the training and testing data. Different schemesforLMadaptation have been proposed, such asthecache modelbased on the observation that a word whichoccurred in a recenttext has a higher probability to be seen again [9]; the triggermodel which uses a trigger word pair to get at semantic infor-mation [10]; and structured LMs [1].Broadcast news (BN) transcription is a complicated task forboth acoustic and language modeling. The linguistic attributesof BN data are complex, arising from the many different speak-ing styles, from spontaneous conversation to prepared speech(close in style to written texts). The content of BN data is openand any given BN show covers multiple topics.As a consequence, it is difficult to predict the topics of a BNshow without looking at the data itself. The only informationthat is available for the show are the hypotheses output from thespeech recognizer. However, for any given broadcast, the num-ber of words in the hypothesized transcript is quite small andcontains recognition errors. Therefore the transcripts are notsufficient for use as an adaptive corpus. Information retrieval(IR) methods provide a means to address this problem. Insteadof directly using the ASR hypotheses for LM adaptation, theycan be used as queries to an IR system in order to select ad-ditional on-topic adaptation data from a large general corpus.This approach reduces the effect of transcription errors in thehypotheses and at the same time provides substantially moretextual data for LM estimation.In this paper, a series of experiments are presented exploringthe general framework of unsupervised LM adaptation using IRmethods [3]. The performances of a variety of popular tech-niques for LM adaptation using automatically selected adapta-tion data are compared. The investigated techniques are linearinterpolation,maximum aposteriori(MAP)adaptation, mixturemodels, dynamic mixture models, and minimum discriminationinformation (MDI) adaptation. The effect of the temporal dis-tance between the epoch of the adaptation corpus and of theepoch of the test data is also assessed. As mentioned above, agiven BN show typically covers several stories, with each storybeing related to a different topic. To address the changing prop-erty of BN data, static and dynamic models for LM adaptationare investigated. In static modeling the LM is updated oncefor the whole show, which means that the LM must be simul-taneously fit to multiple topics. Dynamic modeling updates theLM at each automatically detected story change, which entailsestimating multiplestory-based LMs for each BN show. Exper-iments arecarriedout forBNtranscriptioninAmericanEnglishand Mandarin Chinese.
Langzhou Chen, Lori Lamel, Jean-Luc Gauvain, Gilles Adda
INTERSPEECH2
2004 Speaker diarization from speech transcripts
Lori Lamel, Jean-Luc Gauvain, Leonardo Canseco-Rodriguez
INTERSPEECH1
2004 Transcription of arabic broadcast news
abstract
This paper describes recent research on transcribing Modern Standard Arabic broadcast news data. The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Arabic is a highly inflected language where articles and affixes are added to roots in order to change the word’s meaning. A corpus of 50 hours of audio data from 7 television and radio sources and 200 M words of newspaper texts were used to train the acoustic and language models. The transcription system based on these models and a vowelized dictionary obtains an average word error rate on a test set comprised of 12 hours of test data from 8 sources is about 18%.
Abdelkhalek Messaoudi, Lori Lamel, Jean-Luc Gauvain
INTERSPEECH2
2004 Automatic detection of dialog acts based on multilevel information
abstract
Recently there has been growing interest in using dialog acts to characterize human-human and human-machine dialogs. This paper reports on our experience in the annotation and the automatic detection of dialog acts in human-human spoken dialog corpora. Our work is based on two hypotheses: first, word position is more important than the exact word in identifying the dialog act; and second, there is a strong grammar constraining the sequence of dialog acts. A memory based learning approach has been used to detect dialog acts. In a first set of experiments the number of utterances per turn is known, and in a second set, the number of utterances is hypothesized using a language model for utterance boundary detection. In order to verify our first hypothesis, the model trained on a French corpus was tested on a corpus for a similar task in English and for a second French corpus from a different domain. A correct dialog act detection rate of about 84 % is obtained for the same domain and language condition and about 75 % for the cross-language and cross-domain conditions. 1.
Sophie Rosset, Lori Lamel
INTERSPEECH2
2003 Unsupervised language model adaptation for broadcast news
abstract
Unsupervised language model adaptation for speech recognition is challenging, particularly for complicated tasks such the transcription of broadcast news (BN) data. This paper presents an unsupervised adaptation method for language modeling based on information retrieval techniques. The method is designed for the broadcast news transcription task where the topics of the audio data cannot be predicted in advance. Experiments are carried out using the LIMSI American English BN transcription system and the NIST 1999 BN evaluation sets. The unsupervised adaptation method reduces the perplexity by 7% relative to the baseline LM and yields a 2% relative improvement for a 10xRT system.
Langzhou Chen, Jean-Luc Gauvain, Lori Lamel, Gilles Adda
ICASSP (1)3
2003 Conversational telephone speech recognition
abstract
This paper describes the development of a speech recognition system for the processing of telephone conversations, starting with a state-of-the-art broadcast news transcription system. We identify major changes and improvements in acoustic and language modeling, as well as decoding, which are required to achieve state-of-the-art performance on conversational speech. Some major changes on the acoustic side include the use of speaker normalization (VTLN), the need to cope with channel variability, and the need for efficient speaker adaptation and better pronunciation modeling. On the linguistic side the primary challenge is to cope with the limited amount of language model training data. To address this issue we make use of a data selection technique, and a smoothing technique based on a neural network language model. At the decoding level lattice rescoring and minimum word error decoding are applied. On the development data, the improvements yield an overall word error rate of 24.9% whereas the original BN transcription system had a word error rate of about 50% on the same data.
Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen, Fabrice Lefèvre
ICASSP (1)2
2003 Emotion detection in task-oriented spoken dialogues
abstract
Detecting emotions in the context of automated call center services can be helpful for following the evolution of the human-computer dialogues, enabling dynamic modification of the dialogue strategies and influencing the final outcome. The emotion detection work reported here is a part of larger study aiming to model user behavior in real interactions. We make use of a corpus of real agent-client spoken dialogues in which the manifestation of emotion is quite complex, and it is common to have shaded emotions since the interlocutors attempt to control the expression of their internal attitude. Our aims are to define appropriate emotions for call center services, to annotate the dialogues and to validate the presence of emotions via perceptual tests and to find robust cues for emotion detection. In contrast to research carried out with artificial data with simulated emotions, for real-life corpora the set of appropriate emotion labels must be determined. Two studies are reported: the first investigates automatic emotion detection using linguistic information, whereas the second concerns perceptual tests for identifying emotions as well as the prosodic and textual cues which signal them. About 11% of the utterances are annotated with non-neutral emotion labels. Preliminary experiments using lexical cues detect about 70% of these labels.
Laurence Devillers, Lori Lamel, Ioana Vasilescu
ICME2
2003 Multi-source training and adaptation for generic speech recognition
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2002 Transcribing audio-video archives
abstract
This paper addresses the automatic transcription of audiovideo archives using a state-of-the-art broadcast news speech transcription system. A 9-hour corpus spanning the latter half of the 20th century (1945-1995) has been transcribed and an analysis of the transcription quality carried out. In addition to the challenges of transcribing heterogenous broadcast news data, we are faced with changing properties of the archive over time, such as the audio quality, the speaking style, vocabulary items and manner of expression. After assessing the performance of the transcription system, several paths are explored in an attempt to reduce the mismatch between the acoustic and language models and the archived data.
Claude Barras, Alexandre Allauzen, Lori Lamel, Jean-Luc Gauvain
ICASSP3
2002 Unsupervised acoustic model training
abstract
This paper describes some recent experiments using unsupervised techniques for acoustic model training in order to reduce the system development cost. The approach uses a speech recognizer to transcribe unannotated raw broadcast news data. The hypothesized transcription is used to create labels for the training data. Experiments providing supervision only via the language model training materials show that including texts which are contemporaneous with the audio data is not crucial for success of the approach, and that the acoustic models can be initialized with as little as 10 minutes of manually annotated data. These experiments demonstrate that unsupervised training is a viable training scheme and can dramatically reduce the cost of building acoustic models.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda
ICASSP1
2002 Annotations for Dynamic Diagnosis of the Dialog State
Laurence Devillers, Sophie Rosset, Hélène Bonneau-Maynard, Lori Lamel
LREC4
2002 Advances in Large Vocabulary Speech Recognition
Jean-Luc Gauvain, Renato De Mori, Lori Lamel
Comput. Speech Lang.3
2002 Lightly supervised and unsupervised acoustic model training
Lori Lamel, Jean-Luc Gauvain, Gilles Adda
Comput. Speech Lang.1
2002 The LIMSI Broadcast News transcription system
Jean-Luc Gauvain, Lori Lamel, Gilles Adda
Speech Commun.2
2002 User evaluation of the MASK kiosk
Lori Lamel, Samir Bennacef, Jean-Luc Gauvain, Hervé Dartigues, Jean-Noël Temem
Speech Commun.1
2002 Automatic transcription of Broadcast News data
David S. Pallett, Lori Lamel
Speech Commun.2
2002 Corrigendum to "Editorial" [Speech Communication 37 (2002) (1-2)]
David S. Pallett, Lori Lamel
Speech Commun.2
2001 Processing Broadcast Audio for Information Access
abstract
This paper addresses recent progress in speaker-independent, large vocabulary, continuous speech recognition, which has opened up a wide range of near and mid-term applications. One rapidly expanding application area is the processing of broadcast audio for information access. At LIMSI, broadcast news transcription systems have been developed for English, French, German, Mandarin and Portuguese, and systems for other languages are under development. Audio indexation must take into account the specificities of audio data, such as needing to deal with the continuous data stream and an imperfect word transcription. Some near-term applications areas are audio data mining, selective dissemination of information and media monitoring.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker, Claude Barras, Langzhou Chen, Yannick de Kercadio
ACL2
2001 Automatic transcription of compressed broadcast audio
abstract
With increasing volumes of audio and video data broadcast over the Web, it is of interest to assess the performance of state-of-the-art automatic transcription systems on compressed audio data for media indexation applications. In this paper the performance of the LIMSI 10x French broadcast news transcription system is measured on a two-hour audio set for a range of MP3 and RealAudio codecs at various bit rates and the GSM codec used for European cellular phone communications. The word error rates are compared with those obtained on high quality PCM recordings prior to compression. For a 6.5 kbps. audio bit rate (the most commonly used on the Web), word error rates under 40% can be achieved, which makes automatic media monitoring systems over the Web a realistic task.
Claude Barras, Lori Lamel, Jean-Luc Gauvain
ICASSP2
2001 Investigating lightly supervised acoustic model training
abstract
The last decade has witnessed substantial progress in speech recognition technology, with todays state-of-the-art systems being able to transcribe broadcast audio data with a word error of about 20%. However, acoustic model development for the recognizers requires large corpora of manually transcribed training data. Obtaining such data is both time-consuming and expensive, requiring trained human annotators with substantial amounts of supervision. We describe some experiments using different levels of supervision for acoustic model training in order to reduce the system development cost. The experiments have been carried out using the DARPA TDT-2 corpus (also used in the SDR99 and SDR00 evaluations). Our experiments demonstrate that light supervision is sufficient for acoustic model development, drastically reducing the development cost.
Lori Lamel, Jean-Luc Gauvain, Gilles Adda
ICASSP1
2001 Towards task-independent speech recognition
abstract
Despite the considerable progress made in the last decade, speech recognition is far from a solved problem. For instance, porting a recognition system to a new task (or language) still requires substantial investment of time and money, as well as expertise in speech recognition. The paper takes a first step at evaluating to what extent a generic state-of-the-art speech recognizer can reduce the manual effort required for system development. We demonstrate the genericity of wide domain models, such as broadcast news acoustic and language models, and techniques to achieve a higher degree of genericity, such as transparent methods to adapt such models to a specific task. This work targets three tasks using commonly available corpora: small vocabulary recognition (TI-digits), text dictation (WSJ), and goal-oriented spoken dialog (ATIS).
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
ICASSP3
2001 Using information retrieval methods for language model adaptation
abstract
In this paper we report experiments on language model adaptation using information retrieval methods, drawing upon recent developments in information extraction and topic tracking. One of the problems is extracting reliable topic information with high confidence from the audio signal in the presence of recognition errors. The work in the information retrieval domain on information extraction and topic tracking suggested a new way to solve this problem. In this work, we make use of information retrieval methods to extract topic information in the word recognizer hypotheses, which are then used to automatically select adaptation data from a very large general text corpus. Two adaptive language models, a mixture based model and a MAP based model, have been investigated using the adaptation data. Experiments carried out with the LIMSI Mandarin broadcast news transcription system gives a relative character error rate reduction of 4.3% with this adaptation method.
Langzhou Chen, Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
INTERSPEECH3
2001 Improving genericity for task-independent speech recognition
abstract
Although there have been regular improvements in speech recognition technology over the past decade, speech recognition is far from being a solved problem. Recognition systems are usually tuned to a particular task and porting the system to a new task (or language) is both time-consuming and expensive. In this paper, issues in speech recognizer portability are addressed through the development of generic core speech recognition technology. First, the genericity of wide domain models is assessed by evaluating performance on several tasks. Then, the use of transparent methods for adapting generic models to a specific task is explored. Finally, further techniques are evaluated aiming at enhancing the genericity of the wide domain models. We show that unsupervised acoustic model adaptation and multi-source training can reduce the performance gap between task-independent and taskdependent acoustic models, and for some tasks even out-perform task-dependent acoustic models.
Fabrice Lefèvre, Jean-Luc Gauvain, Lori Lamel
INTERSPEECH3
2001 Audio Partitioning and Transcription for Broadcast Data Indexation
Jean-Luc Gauvain, Lori Lamel, Gilles Adda
Multim. Tools Appl.2
2000 Transcription and indexation of broadcast data
abstract
We report on recent research on transcribing and indexing broadcast news data for information retrieval purposes. The system described combines an adapted version of the LIMSI 1998 Hub-4E transcription system for speech recognition with text-based IR methods. Experimental results are reported in terms of recognition word error rate and mean average precision for both the TREC SDR98 (100h) and SDR99 (600h) data sets. With query expansion using commercial transcripts, comparable mean average precisions are obtained on manual reference transcriptions and automatic transcriptions with a word error rate of 21.5% measured on a 10 hour data subset.
Jean-Luc Gauvain, Lori Lamel, Yannick de Kercadio, Gilles Adda
ICASSP2
2000 Investigating text normalization and pronunciation variants for German broadcast transcription
abstract
In this paper we describe our ongoing work concerning lexical modeling in the LIMSI broadcast transcription system for German. Lexical decomposition is investigated with a twofold goal: lexical coverage optimization and improved letter-to-sound conversion. A set of about 450 decompounding rules, developed using statistics from a 300M word corpus, reduces the OOV rate from 4.5% to 4.0% on a 30k development text set. Adding partial inflection stripping, the OOV rate drops to 2.9%. For letterto -sound conversion, decompounding reduces cross-lexeme ambiguities and thus contributes to more consistent pronunciation dictionaries. Another point of interest concerns reduced pronunciation modeling. Word error rates, measured on 1.3 hours of ARTE TV broadcast, vary between 18 and 24% depending on the show and the system configuration. Our experiments indicate that using reduced pronunciations slightly decreases word error rates. 1. INTRODUCTION The German language, more than other major western ...
Martine Adda-Decker, Gilles Adda, Lori Lamel
INTERSPEECH3
2000 Broadcast news transcription in Mandarin
abstract
In this paper, our work in developing a Mandarin broadcast news transcription system is described. The main focus of this work is a port of the LIMSI American English broadcast news transcription system to the Chinese Mandarin language. The system consists of an audio partitioner and an HMM-based continuous speech recognizer. The acoustic models were trained on about 24 hours of data from the 1997 Hub4 Mandarin corpus available via LDC. In addition to the transcripts, the language models were trained on Mandarin Chinese News Corpus containing about 186 million characters. We investigate recognition performance as a function of lexical size, with and without tone in the lexicon, and with a topic dependent language model. The transcription character error rate on the DARPA 1997 test set is 18.1% using a lexicon with 3 tone levels and a topic-based language model. 1. INTRODUCTION It is well known that radio and television broadcast shows contain different types of speech from the acoust...
Langzhou Chen, Lori Lamel, Gilles Adda, Jean-Luc Gauvain
INTERSPEECH2
2000 Fast decoding for indexation of broadcast data
abstract
Processing time is an important factor in making a speech transcription system viable for automatic indexation of radio and television broadcasts. When only concerned by the word error rate, it is common to design systems that run in 100 times real-time or more. This paper addresses issues in reducing the speech recognition time for automatic indexation of radio and TV broadcasts with the aim of obtaining reasonable performance for close to real-time operation. We investigated computational resources in the range 1 to 10xRT on commonly available platforms. Constraints on the computational resources led us to reconsider design issues, particularly those concerning the acoustic models and the decoding strategy. A new decoder was implemented which transcribes broadcast data in few times real-time with only a slight increase in word error rate when compared to our best system. Experiments with spoken document retrieval show that comparable IR results are obtained with a 10xRT automatic tra...
Jean-Luc Gauvain, Lori Lamel
INTERSPEECH2
2000 Considerations in the design and evaluation of spoken language dialog systems
abstract
In this paper we summarize our experience at LIMSI in the design, development and evaluation of spoken language dialog systems for information retrieval tasks. This work has been for the most part carried out in the context of several European and international projects. Evaluation plays an integral role in the development of spoken language dialog systems. While there are commonly used measures and methodologies for evaluating speech recognizers, the evaluation of spoken dialog systems is considerably more complicated due to the interactive nature and the human perception of performance. It is therefore important to assess not only the individual system components, but the overall system performance using objective and subjective measures. 1. INTRODUCTION In our view, spokenlanguagesystems should provide a natural, user-friendly interface with the computer, allowing easy access to the stored information. At LIMSI we have experience in developing several spoken language dialog system...
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain
INTERSPEECH1
2000 Towards best practice in the development and evaluation of speech recognition components of a spoken language dialog system
abstract
This article provides a global overview of the main aspects of current practice in the design, implementation and evaluation of speech recognition components for Spoken Language Dialog Systems (SLDSs), and presents the results of the DISC European project related to speech recognition. DISC and its successor DISC-2 are efforts towards the definition of best practice guidelines for SLDS development and evaluation. SLDSs aim at using natural spoken input for performing an information processing task such as automated standards, call routing or travel planning and reservations. The main functionality of an SLDS are speech recognition, natural language understanding, dialog management, database access and interpretation, response generation and speech synthesis. Speech recognition, which transforms the acoustic signal into a string of words, is a key technology in any SLDS.
Lori Lamel, Wolfgang Minker, Patrick Paroubek
Nat. Lang. Eng.1
2000 Large-vocabulary continuous speech recognition: advances and applications
abstract
The past decade (1990-2000) has witnessed substantial advances in speech recognition technology, which when combined with the increase in computational power and storage capacity has resulted in a variety of commercial products already or soon to be on the market. The authors review the state of the art in core technology, large vocabulary continuous speech recognition, with a view toward highlighting recent advances. We then highlight issues in moving toward applications, discussing system efficiency, portability across languages and tasks, and enhancing the system output by adding tags and nonlinguistic information. Current performance in speech recognition and outstanding challenges for three classes of applications (dictation, audio indexation, and spoken language dialogue systems), are discussed.
Jean-Luc Gauvain, Lori Lamel
Proc. IEEE2
2000 Speaker verification over the telephone
Lori Lamel, Jean-Luc Gauvain
Speech Commun.1
2000 The LIMSI ARISE system
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain, Samir Bennacef, Martine Garnier-Rizet, B. Prouts
Speech Commun.1
1999 Large vocabulary speech recognition in French
abstract
We present some design considerations concerning our large vocabulary continuous speech recognition system in French. The impact of the epoch of the text training material on lexical coverage, language model perplexity and recognition performance on newspaper texts is demonstrated. The effectiveness of larger vocabulary sizes and larger text training corpora for language modeling is investigated. French is a highly inflected language producing large lexical variety and a high homophone rate. About 30% of recognition errors are shown to be due to substitutions between inflected forms of a given root form. When word error rates are analysed as a function of word frequency, a significant increase in the error rate can be measured for frequency ranks above 5000.
Martine Adda-Decker, Gilles Adda, Jean-Luc Gauvain, Lori Lamel
ICASSP4
1999 The LIMSI ARISE system for train travel information
abstract
In the context of the LE-3 ARISE (Automatic Railway Information Systems for Europe) project we have been developing a dialog system for vocal access to rail travel information. The system provides schedule information for the main French intercity connections, as well as, simulated fares and reservations, reductions and services. The goal is to obtain high dialog success rates with a very open dialog structure, where the user is free to ask any question or to provide any information at any point in time. In order to improve the performance with such an open dialog strategy, we make use of implicit confirmation using the callers wording (when possible), and change to a more constrained dialog level, when the dialog is not going well. In addition to own assessment, the prototype system undergoes periodic user evaluations carried out by the our partners at the French Railways.
Lori Lamel, Sophie Rosset, Jean-Luc Gauvain, Samir Bennacef
ICASSP1
1999 Recent advances in transcribing television and radio broadcasts
abstract
Transcription of broadcast news shows (radio and television) is a major step in developing automatic tools for indexation and retrieval of the vast amounts of information generated on a daily basis. Broadcast shows are challenging to transcribe as they consist of a continuous data stream with segments of different linguistic and acoustic natures. Transcribing such data requires addressing two main problems: those related to the varied acoustic properties of the signal, and those related to the linguistic properties of the speech. Prior to word transcription, the data is partitioned into homogeneous acoustic segments. Non-speech segments are identified and rejected, and the speech segments are clustered and labeled according to bandwidth and gender. The speaker-independent large vocabulary, continuous speech recognizer makes use of n-gram statistics for language modeling and of continuous density HMMs with Gaussian mixtures for acoustic modeling. The LIMSI system has consistently obtain...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Michèle Jardino
EUROSPEECH2
1999 Comparing different model configurations for language identification using a phonotactic approach
abstract
In this paper different model configurations for language identification using a phonotactic approach are explored. Identification experiments were carried out on the 11-language telephone speech corpus OGI-TS, containing calls in French, English, German, Spanish, Japanese, Korean, Mandarin, Tamil, Farsi, Hindi, and Vietnamese. Phone sequences output by one or multiple phone recognizers are rescored with language-dependent phonotactic models approximated by phone bigrams. The parameters of different sets of acoustic phone models were estimated using the 4-language IDEAL corpus. Sets of language-specific phonotactic models were trained using the training portion of the OGITS CORPUS. Error rates are significantly reduced by combining language-dependent and language-independent acoustic decoders, especially for short segments. A 9.9% LID error rate was obtained on the 11-language task using phonotactic models trained on spontaneous speech data. These results show that the phonotactic approach is relative insensitive to an acoustic mismatch between training and test conditions.
Driss Matrouf, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel
EUROSPEECH4
1999 Overview of the ARISE project
Els den Os, Lou Boves, Lori Lamel, Paolo Baggia
EUROSPEECH3
1999 Design strategies for spoken language dialog systems
Sophie Rosset, Samir Bennacef, Lori Lamel
EUROSPEECH3
1999 Pronunciation variants across system configuration, language and speaking style
Martine Adda-Decker, Lori Lamel
Speech Commun.2
1998 Multilingual phone recognition of spontaneous telephone speech
abstract
In this paper we report on experiments with phone recognition of spontaneous telephone speech. Phone recognizers were trained and assessed on IDEAL, a multilingual corpus containing telephone speech in French, British English, German and Castillan Spanish. We investigated the influence of the training material composition (size and linguistic content) on the recognition performance using context-independent (CI) hidden Markov models (HMMs) and phonotactic bigram models. We found that when testing on spontaneous speech data, using only spontaneous speech training data gave the highest phone accuracies for the four languages, even though this data comprises only 14% of the available training data. The use of context-dependent (CD) HMMs reduced the phone error across the 4 languages, with the average error reduced to 51.9% from the 57.4% obtained with CI models. We suggest a straightforward way of detecting non speech phenomena. The basic idea is to remove sequences of consonants between two silence labels from the recognized phone strings prior to scoring. This simple technique reduces the relative average phone error rate by 5.4%. The lowest phone error with CD models and filtering was obtained for Spanish (39.1%) with 4 language average being 49.1%.
Cristobal Corredor-Ardoy, Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain
ICASSP2
1998 Partitioning and transcription of broadcast news data
abstract
Radio and television broadcasts consist of a continuous stream of data comprised of segments of different linguistic and acoustic natures, which poses challenges for transcription. In this paper we report on our recent work in transcribing broadcast news data[2, 4], including the problem of partitioning the data into homogeneous segments prior to word recognition. Gaussian mixture models are used to identify speech and non-speech segments. A maximumlikelihood segmentation/clustering process is then applied to the speech segments using GMMs and an agglomerative clustering algorithm. The clustered segments are then labeled according to bandwidth and gender. The recognizer is a continuous mixture density, tied-state cross-word context-dependent HMM system with a 65k trigram language model. Decoding is carried out in three passes, with a final pass incorporating cluster-based test-set MLLR adaptation. The overall word transcription error on the Nov'97 unpartitioned evaluation test data was...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda
ICSLP2
1998 User evaluation of the mask kiosk
abstract
In this paper we report on a series of user trials carried out to assess the performance and usability of the Multimodal Multimedia Service Kiosk (MASK) prototype. The aim of the ESPRIT MASK project was to pave the way for advanced public service applications with user interfaces employing multimodal, multimedia input and output. The prototype kiosk was developed after analyzing the technological requirements in the context of users performing travel enquiry tasks, in close collaboration with the French Railways (SNCF) and the Ergonomics group at the University College of London (UCL). The time to complete the transaction with the MASK kiosk is reduced by about 30% compared to that required for the standard kiosk, and the transaction success rate is 85% for novices and 94% once familiar with the system. In addition to meeting or exceeding the performance goals set at the project onset in terms of success rate, transaction time, and user satisfaction, the MASK kiosk was judged to be user-friendly and simple to use.
Lori Lamel, Samir Bennacef, Jean-Luc Gauvain, Hervé Dartigues, Jean-Noël Temem
ICSLP1
1998 Language identification incorporating lexical information
abstract
In this paper we explore the use of lexical information for language identification (LID). Our reference LID system uses language-dependent acoustic phone models and phone-based bigram language models. For each language, lexical information is introduced by augmenting the phone vocabulary with the N most frequent words in the training data. Combined phone and word bigram models are used to provide linguistic constraints during acoustic decoding. Experiments were carried out on a 4-language telephone speech corpus. Using lexical information achieves a relative error reduction of about 20% on spontaneous and read speech compared to the reference phone-based system. Identification rates of 92%, 96% and 99% are achieved for spontaneous, read and task-specific speech segments respectively, with prior speech detection.
Driss Matrouf, Martine Adda-Decker, Lori Lamel, Jean-Luc Gauvain
ICSLP3
1998 On the use of speech and text corpora for speech recognition in French
Martine Adda-Decker, Gilles Adda, Lori Lamel, Jean-Luc Gauvain
LREC3
1998 The disc approach to spoken language systems development and evaluation
Laila Dybkjær, Niels Ole Bernsen, Rolf Carlson, Lin Chase, Niels Dahlbäck, Klaus Failenschmid, Ulrich Heid, Paul Kleinheisterkamp, Arne Jönsson, Hans Kamp, Inger Karlson, Jan van Kuppevelt, Lori Lamel, Patrick Paroubek
LREC13
1998 A multilingual corpus for language identification
Lori Lamel, Gilles Adda, Martine Adda-Decker, Cristobal Corredor-Ardoy, Jean-Jacques Gangolf, Jean-Luc Gauvain
LREC1
1997 Transcribing broadcast news shows
abstract
While significant improvements have been made in large vocabulary continuous speech recognition of large read-speech corpora such as the ARPA Wall Street Journal-based CSR corpus (WSJ) for American English and the BREF corpus for French, these tasks remain relatively artificial. In this paper we report on our development work in moving from laboratory read speech data to real-world speech data in order to build a system for the new ARPA broadcast news transcription task. The LIMSI Nov96 speech recognizer makes use of continuous density HMMs with Gaussian mixtures for acoustic modeling and n-gram statistics estimated on newspaper texts. The acoustic models are trained on the WSJO/WSJ1, and adapted using MAP estimation with task-specific training data. The overall word error on the Nov96 partitioned evaluation test was 27.1%.
Jean-Luc Gauvain, Gilles Adda, Lori Lamel, Martine Adda-Decker
ICASSP3
1997 Speaker recognition with the Switchboard corpus
abstract
We present our development work carried out in preparation for the March'96 speaker recognition test on the Switchboard corpus organized by NIST. The speaker verification system evaluated was a Gaussian mixture model (GMM). We provide experimental results on the development test and evaluation test data, and some experiments carried out since the evaluation comparing the GMM with a phone-based approach. Better performance is obtained by training on data from multiple sessions, and with different handsets. High error rates are obtained even using a phone-based approach both with and without the use of orthographic transcriptions of the training data. We also describe a human perceptual test carried out on a subset of the development data, which demonstrates the difficulty human listeners had with this task.
Lori Lamel, Jean-Luc Gauvain
ICASSP1
1997 Text normalization and speech recognition in French
abstract
In this paper we present a quantitative investigation into the impact of text normalization on lexica and language models for speech recognition in French. The text normalization process defines what is considered to be a word by the recognition system. Depending on this definition we can measure different lexical coverages and language model perplexities, both of which are closely related to the speech recognition accuracies obtained on read newspaper texts. Different text normalizations of up to 185M words of newspaper texts are presented along with corresponding lexical coverage and perplexity measures. Some normalizations were found to be necessary to achieve good lexical coverage, while others were more or less equivalent in this regard. The choice of normalization to create language models for use in the recognition experiments with read newspaper texts was based on these findings. Our best system configuration obtained a 11.2% word error rate in the AUPELF `French-speaking' spee...
Gilles Adda, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel
EUROSPEECH4
1997 Language identification with language-independent acoustic models
abstract
In this paper we explore the use of languageindependent acoustic models for language identification (LID). The phone sequence output by a single language-independent phone recognizer is rescored with language-dependent phonotactic models approximated by phone bigrams. The language-independent phoneme inventory was obtained by Agglomerative Hierarchical Clustering, using a measure of similarity between phones. This system is compared with a parallel language-dependent phone architecture, which uses optimally the acoustic log likelihood and the phonotactic score for language identification. Experiments were carried out on the 4-language telephone speech corpus IDEAL, containing calls in British English, Spanish, French and German. Results show that the language-independent approach performs as well as the language-dependent one: 9% versus 10% of error rate on 10 second chunks, for the 4-language task. 1. INTRODUCTION This paper presents some of our recent research on automatic language ...
Cristobal Corredor-Ardoy, Jean-Luc Gauvain, Martine Adda-Decker, Lori Lamel
EUROSPEECH4
1997 Transcription of broadcast news
abstract
In this paper we report on our recent work in transcribing broadcast news shows. Radio and television broadcasts contain signal segments of various linguistic and acoustic natures. The shows contain both prepared and spontaneous speech. The signal may be studio quality or have been transmitted over a telephone or other noisy channel (ie., corrupted by additive noise and nonlinear distorsions), or may contain speech over music. Transcription of this type of data poses challenges in dealing with the continuous stream of data under varying conditions. Our approach to this problem is to segment the data into a set of categories, which are then processed with category specific acoustic models. We describe our 65k speech recognizer and experiments using different sets of acoustic models for transcription of broadcast news data. The use of prior knowledge of the segment boundaries and types is shown to not crucially affect the performance. 1. INTRODUCTION The goal of this research is to au...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
EUROSPEECH2
1997 Multilingual large vocabulary speech recognition: the European SQALE project
Steve J. Young, Martine Adda-Decker, Xavier L. Aubert, Christian Dugast, Jean-Luc Gauvain, Dan J. Kershaw, Lori Lamel, David A. van Leeuwen, David Pye, Anthony J. Robinson, Herman J. M. Steeneken, Philip C. Woodland
Comput. Speech Lang.7
1997 RailTel: Railway Telephone Services
Roberto Billi, Lori Lamel
Speech Commun.2
1997 The LIMSI RailTel System: Field trial of a telephone service for rail travel information
Lori Lamel, Samir Bennacef, Sophie Rosset, Laurence Devillers, S. Foukia, Jean-Jacques Gangolf, Jean-Luc Gauvain
Speech Commun.1
1996 Developments in large vocabulary, continuous speech recognition of German
abstract
We describe our large vocabulary continuous speech recognition system for the German language, the development of which was partly carried out within the context of the European LRE project 62-058 SQALE. The recognition system is the LIMSI recognizer originally developed for French and American English, which has been adapted to German. Specificities of German, as relevant to the recognition system, are presented. These specificities have been accounted for during the recognizer's adaptation process. We present experimental results on a first test set ger-dev95 to measure progress in system development. Results are given with the final system using different acoustic model sets on two test sets ger-dev95 and ger-eval95. This system achieved a word error rate of 17.3% (official word error rate of 16.1% after SQALE adjudication process) on the ger-eval95 test set.
Martine Adda-Decker, Gilles Adda, Lori Lamel, Jean-Luc Gauvain
ICASSP3
1996 Developments in continuous speech dictation using the 1995 ARPA NAB news task
abstract
We report on the LIMSI recognizer evaluated in the ARPA 1995 North American Business (NAB) news benchmark test. In contrast to previous evaluations, the new Hub 3 test aims at improving basic SI, CSR performance on unlimited-vocabulary read speech recorded under more varied acoustical conditions (background environmental noise and unknown microphones). The LIMSI recognizer is an HMM-based system with a Gaussian mixture. Decoding is carried out in multiple forward acoustic passes, where more refined acoustic and language models are used in successive passes and information is transmitted via word graphs. In order to deal with the varied acoustic conditions, channel compensation is performed iteratively, refining the noise estimates before the first three decoding passes. The final decoding pass is carried out with speaker-adapted models obtained via unsupervised adaptation using the MLLR method. On the Sennheiser microphone (average SNR 29 dB) a word error of 9.1% was obtained, which can be compared to 17.5% on the secondary microphone data (average SNR 15 dB) using the same recognition system.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Driss Matrouf
ICASSP2
1996 Dialog in the RAILTEL telephone-based system
Samir Bennacef, Laurence Devillers, Sophie Rosset, Lori Lamel
ICSLP4
1996 Speech recognition for an information kiosk
Jean-Luc Gauvain, Jean-Jacques Gangolf, Lori Lamel
ICSLP3
1996 On designing pronunciation lexicons for large vocabulary, continuous speech recognition
Lori Lamel, Gilles Adda
ICSLP1
1996 Spoken language processing in a multilingual context
abstract
In this paper we overview the spoken language processing activities at LIMSI, which are carried out in a multilingual framework.These activities include speech-to-text conversion, spoken language systems for information retrieval, speaker and language recognition, and speech response.The Spoken Language Processing Group has also been actively involved in corpora development and evaluation.The group has regularly participated in evaluations organized by ARPA, in the LE-SQALE project, and in the AUPELF-UREF program for provision of linguistic resources and evaluation tests for French.
Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain, Gilles Adda
ICSLP1
1996 Data collection for the MASK kiosk: WOz vs prototype system
Andrew Life, Ian Salter, Jean-Noël Temem, Franck Bernard, Sophie Rosset, Samir Bennacef, Lori Lamel
ICSLP7
1996 BABEL: an eastern european multi-language database
Peter Roach, Simon Arnfield, William J. Barry, J. Baltova, Marian Boldea, Adrian Fourcin, Wiktor Gonet, Ryszard Gubrynowicz, E. Hallum, Lori Lamel, Krzysztof Marasek, Alain Marchal, Einar Meister, Klára Vicsi
ICSLP10
1996 Comments on "Towards increasing speech recognition error rates" by H. Bourlard, H. Hermansky, and N. Morgan
Joseph Mariani, Jean-Luc Gauvain, Lori Lamel
Speech Commun.3
1995 Developments in continuous speech dictation using the ARPA WSJ task
abstract
We report on our recent development work in large vocabulary, American English continuous speech dictation. We have experimented with (1) alternative analyses for the acoustic front end, (2) the use of an enlarged vocabulary so as to reduce the number of errors due to out-of-vocabulary words, (3) extensions to the lexical representation, (4) the use of additional acoustic training data, and (5) modification of the acoustic models for telephone speech. The recognizer was evaluated on Hubs 1 and 2 of the fall 1994 ARPA NAB CSR Hub and Spoke Benchmark test. Experimental results for development and evaluation test data are given, as well as an analysis of the errors on the development data.
Jean-Luc Gauvain, Lori Lamel, Martine Adda-Decker
ICASSP2
1995 EUROM - a spoken language resource for the EU - the SAM projects
Dominic S. F. Chan, Adrian Fourcin, Dafydd Gibbon, Björn Granström, Mark A. Huckvale, George K. Kokkinakis, Knut Kvale, Lori Lamel, Børge Lindberg, Asunción Moreno, Jiannis Mouropoulos, Francesco Senia, Isabel Trancoso, Corin 't Veld, Jerome Zeiliger
EUROSPEECH8
1995 Experiments with speaker verification over the telephone
abstract
In this paper we present a study on speaker verification showing achievable performance levels for both high quality speech and telephone speech and for two operational modes, i.e. textdependent and text-independent speaker verification. A statistical modeling approach is taken, where for text independent verification the talker is viewed as a source of phones, modeled by a fully connected Markov chain, where the lexical and syntactic structures of the language are approximated by local phonotactic constraints. A first series of experiments were carried out on high quality speech from the BREF corpus to validate this approach and resulted in an a posteriori equal error rate of 0.3% in textdependent as well as in text-independent mode. A second series of experiments were carried out on a telephone corpus recorded specifically for speaker verification algorithm development. On this data, the lowest equal error rate is 2.9% for the text-dependent mode when 2 trials are allowed per attempt...
Jean-Luc Gauvain, Lori Lamel, B. Prouts
EUROSPEECH2
1995 Issues in Large Vocabulary, Multilingual Speech Recognition
abstract
In this paper we report on our activities in multilingual, speakerindependent, large vocabulary continuous speech recognition. The multilingual aspect of this work is of particular importance in Europe, where each country has its own national language. Our existing recognizer for American English and French, has been ported to British English and German. It has been assessed in the context of the LRESQALE project whose objective was to experiment with installing in Europe a multilingual evaluation paradigm for the assessment of large vocabulary, continuous speech recognition systems. The recognizer makes use of phone-based continuous density HMM for acoustic modeling and n-gram statistics estimated on newspaper texts for language modeling. The system has been evaluated on a dictation task with read, newspaper-based corpora, the ARPA Wall Street Journal corpus of American English, the WSJCAM0 corpus of British English, the BREF-Le Monde corpus of French and the PHONDAT-Frankfurter Runds...
Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain
EUROSPEECH1
1995 Development of spoken language corpora for travel information
abstract
In this paper we report on our ongoing work in developing spoken language corpora in the context of information access in two travel domain tasks, L'ATIS and MASK. The collection of spoken language corpora remains an important research area and represents a significant portion of work in the development of spoken language systems. The use of additional acoustic and language model training data has been shown to almost systematically improve performance in continuous speech recognition. Similarly, progress in spokenlanguage understanding is closely linked to the availability of spoken language corpora. We record subjects on a regular basis using development versions of the spoken language systems for both tasks, obtaining over 1000 queries/month from 20 subjects. To help assess our progress in system development, each subject since March'95 completes a questionnaire addressing the user-friendliness, reliability, ease-of-use of the MASK data collection system. INTRODUCTION The collecti...
Lori Lamel, Sophie Rosset, Samir Bennacef, Hélène Bonneau-Maynard, Laurence Devillers, Jean-Luc Gauvain
EUROSPEECH1
1995 A phone-based approach to non-linguistic speech feature identification
Lori Lamel, Jean-Luc Gauvain
Comput. Speech Lang.1
1994 The LIMSI continuous speech dictation system: evaluation on the ARPA Wall Street Journal task
abstract
We report progress made at LIMSI in speaker-independent large vocabulary speech dictation using the ARPA Wall Street Journal-based CSR corpus. The recognizer makes use of continuous density HMM with Gaussian mixture for acoustic modeling and n-gram statistics estimated on the newspaper texts for language modeling. The recognizer uses a time-synchronous graph-search strategy which is shown to still be viable with vocabularies of up to 20 K words when used with bigram back-off language models. A second forward pass, which makes use of a word graph generated with the bigram, incorporates a trigram language model. Acoustic modeling uses cepstrum-based features, context-dependent phone models (intra and interword), phone duration models, and sex-dependent models. The recognizer has been evaluated in the Nov92 and Nov93 ARPA tests for vocabularies of up to 20,000 words.>
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
ICASSP (1)2
1994 Language identification using phone-based acoustic likelihoods
abstract
Applies the technique of phone-based acoustic likelihoods to the problem of language identification. The basic idea is to process the unknown speech signal by language-specific phone model sets in parallel, and to hypothesize the language associated with the model set having the highest likelihood. Using laboratory quality speech the language can be identified as French or English with better than 99% accuracy with only as little as 2 seconds of speech. On spontaneous telephone speech from the OGI corpus, the language can be identified as French or English with 82% accuracy with 10 seconds of speech. The 10 language identification rate using the OGI corpus is 59.7% with 10 seconds of signal.>
Lori Lamel, Jean-Luc Gauvain
ICASSP (1)1
1994 A spoken language system for information retrieval
Samir Bennacef, Hélène Bonneau-Maynard, Jean-Luc Gauvain, Lori Lamel, Wolfgang Minker
ICSLP4
1994 Continuous speech dictation in French
abstract
A major research activity at LIMSI is multilingual, speakerindependent, large vocabulary speech dictation. In this paper we report on efforts in large vocabulary, speaker-independent continuous speech recognition of French using the BREF corpus. Recognition experiments were carried out with vocabularies containing up to 20k words. The recognizer makes use of continuous density HMM with Gaussian mixture for acoustic modeling and n-gram statistics estimated on 38 million words of newspaper text from Le Monde for language modeling. The recognizer uses a time-synchronous graph-search strategy. When a bigram language model is used, recognition is carried out in a single forward pass. A second forward pass, which makes use of a word graph generated with the bigram language model, incorporates a trigram language model. Acoustic modeling uses cepstrum-based features, contextdependent phone models and phone duration models. An average phone accuracy of 86% was achieved. A word accuracy of 84% h...
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
ICSLP2
1994 The translanguage English database (TED)
Lori Lamel, Florian Schiel, Adrian Fourcin, Joseph Mariani, Hans G. Tillmann
ICSLP1
1994 Speech-To-Text Conversion in French
abstract
Speech-to-text conversion of French necessitates that both the acoustic level recognition and language modeling be tailored to the French language. Work in this area was initiated at LIMSI over 10 years ago. In this paper a summary of the ongoing research in this direction is presented. Included are studies on distributional properties of French text materials; problems specific to speech-to-text conversion particular of French; studies in phoneme-to-grapheme conversion for continuous, error-free phonemic strings; past work on isolated-word speech-to-text conversion; and more recent work on continuous-speech, speech-to-text conversion. Also demonstrated is the use of phone recognition for both language and speaker identification. The continuous speech-to-text conversion for French is based on a speaker-independent, vocabulary-independent recognizer. In this paper phone recognition and word recognition results are reported evaluating this recognizer on read speech taken from the BREF corpus. The recognizer was trained on over 4 hours of speech from 57 speakers, and tested on sentences from an independent set of 19 speakers. A phone accuracy of 78.7% was obtained using a set of 35 phones. The word accuracy was 88% for a 1139 word lexicon and 86% for a 2716 word lexicon, with a word pair grammar with respective perplexities of 100 and 160. Using a bigram grammar, word accuracies of 85.5% and 81.7% were obtained with 5 K and 20 K word vocabularies, with respective perplexities of 122 and 205.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Joseph Mariani
Int. J. Pattern Recognit. Artif. Intell.2
1994 Speaker-independent continuous speech dictation
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
Speech Communication2
1993 Cross-lingual experiments with phone recognition
Lori Lamel, Jean-Luc Gauvain
ICASSP (2)1
1993 A French version of the MIT-ATIS system: portability issues
Hélène Bonneau-Maynard, Jean-Luc Gauvain, David Goodine, Lori Lamel, Joseph Polifroni, Stephanie Seneff
EUROSPEECH4
1993 Speaker-independent continuous speech dictation
abstract
Abstract In this paper we report on progress made at LIMSI in speaker-independent large vocabulary speech dictation using newspaper-based speech corpora in English and French. The recognizer makes use of continuous density HMMs with Gaussian mixtures for acoustic modeling and n -gram statistics estimated on newspaper texts for language modeling. Acoustic modeling uses cepstrum-based features, context-dependent phone models (intra and interword), phone duration models, and sex-dependent models. For English the ARPA Wall Street Journal -based CSR corpus is used and for French the BREF corpus containing recordings of texts from the French newspaper Le Monde is used. Experiments were carried out with both these corpora at the phone level and at the word level with vocabularies containing up to 20,000 words. Word recognition experiments are also described for the ARPA RM task which has been widely used to evaluate and compare systems.
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker
EUROSPEECH2
1993 Identifying non-linguistic speech features
abstract
Over the last decade technological advances have been made which enable us to envision real-world applications of speech technologies. It is possible to foresee applications, for example, information centers in public places such as train stations and airports, where the spoken query is to be recognized without even prior knowledge of the languagebeing spoken. Other applications may require accurate identification of the speaker for security reasons, including control of access to confidential information or for telephone-based transactions.
Lori Lamel, Jean-Luc Gauvain
EUROSPEECH1
1993 High performance speaker-independent phone recognition using CDHMM
abstract
In this paper we report high phone accuracies on three corpora: WSJ0, BREF and TIMIT. The main characteristics of the phone recognizer are: high dimensional feature vector (48), context- and genderdependent phone models with duration distribution, continuous density HMM with Gaussian mixtures, and n-gram probabilities for the phonotatic constraints. These models are trained on speech data that have either phonetic or orthographic transcriptions using maximum likelihood and maximum a posteriori estimation techniques. On the WSJ0 corpus with a 46 phone set we obtain phone accuraciesof 72.4% and 74.4% using 500 and 1600 CD phone units, respectively. Accuracy on BREF with 35 phones is as high as 78.7% with only 428 CD phone units. On TIMIT using the 61 phone symbols and only 500 CD phone units, we obtain a phoneaccuracyof 67.2% which correspond to 73.4% when the recognizer output is mapped to the commonly used 39 phone set. Making reference to our work on large vocabularyCSR, we show that ...
Lori Lamel, Jean-Luc Gauvain
EUROSPEECH1
1993 A knowledge-based system for stop consonant identification based on speech spectrogram reading
Lori Lamel
Comput. Speech Lang.1
1992 Experiments on speaker-independent phone recognition using BREF
abstract
A series of experiments for speaker-independent, continuous speech phone recognition have been carried out using the recently recorded BREF corpus. The authors' experiments were the first to use this database, and are meant to provide a baseline performance evaluation for vocabulary independent phone recognition. The system was trained using hand-verified data from 43 speakers. Using 35 context-dependent phone models, a baseline phone accuracy of 60% (no phone grammar) has been obtained on an independent test set of 7635 phone segments from 19 speakers. Including phone bigram probabilities as phonotactic constraints results in a performance of 63.3%. A phone accuracy of 68.6% (73.3% correct) was obtained with 428 context dependent models.>
Lori Lamel, Jean-Luc Gauvain
ICASSP1
1990 Design considerations and text selection for BREF, a large French read-speech corpus
abstract
BREF, a large read-speech corpus in French has been designed with several aims: to provide enough speech data to develop dictation machines, to provide data for evaluation of continuous speech recognition systems (both speaker-dependent and speaker-independent), and to provide a corpus of continuous speech to study phonological variations. This paper presents some of the design considerations of BREF, focusing on the text analysis and the selection of text materials. The texts to be read were selected from 4.6 million words of the French newspaper, Le Monde. In total, 11,000 texts were selected, with an emphasis on maximizing the number of distinct triphones. Separate text materials were selected for training and test corpora. The goal is to obtain about 10,000 words (approximately 60-70 min.) of speech from each of 100 speakers, from different French dialects. INTRODUCTION One of the main obstacles to progress in continuous speech recognition has been the lack of sufficient speech m...
Jean-Luc Gauvain, Lori Lamel, Maxine Eskénazi
ICSLP2
1986 An expert spectrogram reader: A knowledge-based approach to speech recognition
abstract
Human experts can determine the phonetic identity of unknown utterances from a visual examination of the spectrogram with performance better than available computer systems. The spectrogram-reading process involves the use of multiple sources of knowledge, including articulatory movements, acoustic phonetics, phonotactics, and linguistics. In addition, the experts' performance can be attributed to their ability to deal with partial and/or conflicting information, as well as multiple cues. This paper investigates the feasibility of constructing a knowledge-based system that mimics the process of spectrogram reading by humans. In a task of identifying stop consonants extracted from continuous speech, the system achieved performance that is comparable to that of the experts.
Victor Zue, Lori Lamel
ICASSP2
1984 Properties of consonant sequences within words and across word boundaries
abstract
This paper is concerned with the problem of locating word boundaries from phonemic strings. As a step towards a better understanding of the acoustic phonetic properties of consonant sequences within and across word boundaries, we studied the distributional properties of these sequences. The database consists of several text files from different discourses, ranging in size from 200 to 38,000 words. In each case we tabulated the number of distinct word-initial, word-medial, word-final, and word-boundary consonant sequences and their frequency of occurrence. Word-medial consonant sequences represent a small subset of those found across word boundaries. Of the non-medial consonant sequences, approximately 80% have a unique boundary location. When the consonant sequence cannot uniquely specify the location of a word boundary, different placement of word boundaries often results in allophones with substantially different acoustic characteristics.
Lori Lamel, Victor Zue
ICASSP1
1982 Performance improvement in a dynamic-programming-based isolated word recognition system for the alpha-digit task
abstract
In isolated word recognition, the alpha-digit vocabulary has been recognized as one of the most difficult due to the acoustic similarities of the lexical entries. The purpose of this study is two-fold: a) To see how a dynamic programming approach can be augmented with phonetic information to improve recognition accuracy of the alpha-digits; and b) To minimize computational requirements for the recognition task. Performance improvement is accomplished by dividing the vocabulary into subsets based on the syllabic patterns of the words and by emphasizing the consonant-vowel transitional regions of the words. This algorithm has been tested on 10 speakers, 5 male and 5 female. The division of the vocabulary results in a substantial savings in computation with essentially no decrease in recognition accuracy. In addition to computational savings, emphasizing the transitional portions of the word in some cases results in accuracy improvement. Discussion of the results and suggestions for further improvement are presented.
Lori Lamel, Victor Zue
ICASSP1