VLDB 2026 Research / reviewers in the wild / expert
Martine Adda-Decker
dblp:43/2130
· DBLP profile ↗
96ranked-venue papers
14as first author
10since 2021 · last 2026
0000-0003-2154-7438ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 11 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Added Value of Metadata and Annotations: Evidence from Two Large-Scale, Naturalistic Corpus Studies
Anisia Popescu, Johanna Cronenberg, Ioana Vasilescu, Ioana Chitoran, Lori Lamel, Martine Adda-Decker |
LREC | 6 |
| 2025 | Corpus-Based Insights into Mandarin Neutral Tone: Effects of Tonal Context and Structural Patterns in Spontaneous SpeechabstractAbstract book: https://www.isca-archive.org/interspeech_2025/booklet.pdf /// Full conference proceedings: https://www.isca-archive.org/interspeech_2025/ Nicolas Audibert, Yaru Wu, Martine Adda-Decker |
INTERSPEECH | 4 |
| 2025 | Apical vs. Regular Vowel Duration: A Corpus-based Analysis of Contextual Influences in Standard MandarinabstractInternational audience Bowei Shao, Martine Adda-Decker |
INTERSPEECH | 3 |
| 2024 | Cross-linguistic transfer of phonological assimilation in early and late bilinguals
Sharon Peperkamp, Sonya Kaiser, Lori Lamel, Martine Adda-Decker |
CogSci | 4 |
| 2024 | An introduction to pluricentric languages in speech science and technologyabstractPluricentric languages are languages that are spoken in at least two countries where they have an official function and thus develop national varieties with specific linguistic and pragmatic features. Presently 43 languages have been identified as belonging to this category, for instance, English, Spanish, German, Bengali, Hindi and Urdu. This article forms an introduction to the special issue “Pluricentric Languages in Speech Science and Technology” by giving an overview of current challenges with respect to the development of speech and language resources, annotation and analysis tools, as well as speech technology services for pluricentric languages. The article discusses potential solutions that come from cross-fertilization: on the one hand, how phonetic and linguistic knowledge may contribute to advancements in speech technology, and on the other, how speech technology may facilitate phonetic and linguistic studies on pluricentric languages. In our discussion, we include the research methods and findings of the eight research articles of this special issue and point towards promising paths for future research in the field. Barbara Schuppler, Martine Adda-Decker, Catia Cucchiarini, Rudolf Muhr |
Speech Commun. | 2 |
| 2023 | Mandarin lexical tone duration: Impact of speech style, word length, syllable position and prosodic position
Yaru Wu, Martine Adda-Decker, Lori Lamel |
Speech Commun. | 2 |
| 2022 | When Phonetics Meets Morphology: Intervocalic Voicing Within and Across Words in Romance LanguagesabstractInternational audience Mathilde Hutin, Martine Adda-Decker, Lori Lamel, Ioana Vasilescu |
INTERSPEECH | 2 |
| 2022 | Extracting Linguistic Knowledge from Speech: A Study of Stop Realization in 5 Romance LanguagesabstractThis paper builds upon recent work in leveraging the corpora and tools originally used to develop speech technologies for corpus-based linguistic studies. We address the non-canonical realization of consonants in connected speech and we focus on voicing alternation phenomena of stops in 5 standard varieties of Romance languages (French, Italian, Spanish, Portuguese, Romanian). For these languages, both large scale corpora and speech recognition systems were available for the study. We use forced alignment with pronunciation variants and machine learning techniques to examine to what extent such frequent phenomena characterize languages and what are the most triggering factors. The results confirm that voicing alternations occur in all Romance languages. Automatic classification underlines that surrounding contexts and segment duration are recurring contributing factors for modeling voicing alternation. The results of this study also demonstrate the new role that machine learning techniques such as classification algorithms can play in helping to extract linguistic knowledge from speech and to suggest interesting research directions. Yaru Wu, Mathilde Hutin, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
LREC | 5 |
| 2022 | Using a Knowledge Base to Automatically Annotate Speech Corpora and to Identify Sociolinguistic VariationabstractSpeech characteristics vary from speaker to speaker. While some variation phenomena are due to the overall communication setting, others are due to diastratic factors such as gender, provenance, age, and social background. The analysis of these factors, although relevant for both linguistic and speech technology communities, is hampered by the need to annotate existing corpora or to recruit, categorise, and record volunteers as a function of targeted profiles. This paper presents a methodology that uses a knowledge base to provide speaker-specific information. This can facilitate the enrichment of existing corpora with new annotations extracted from the knowledge base. The method also helps the large scale analysis by automatically extracting instances of speech variation to correlate with diastratic features. We apply our method to an over 120-hour corpus of broadcast speech in French and investigate variation patterns linked to reduction phenomena and/or specific to connected speech such as disfluencies. We find significant differences in speech rate, the use of filler words, and the rate of non-canonical realisations of frequent segments as a function of different professional categories and age groups. Yaru Wu, Fabian M. Suchanek, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
LREC | 5 |
| 2021 | Synchronic Fortition in Five Romance Languages? A Large Corpus-Based Study of Word-Initial DevoicingabstractInternational audience Mathilde Hutin, Yaru Wu, Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
Interspeech | 6 |
| 2020 | Ongoing Phonologization of Word-Final Voicing Alternations in Two Romance Languages: Romanian and FrenchabstractInternational audience Mathilde Hutin, Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker |
INTERSPEECH | 5 |
| 2020 | Mandarin Lexical Tones: A Corpus-Based Study of Word Length, Syllable Position and Prosodic Position on DurationabstractInternational audience Yaru Wu, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 2 |
| 2019 | " Gra[f] e!" Word-Final Devoicing of Obstruents in Standard French: An Acoustic Study Based on Large CorporaabstractInternational audience Adèle Jatteau, Ioana Vasilescu, Lori Lamel, Martine Adda-Decker, Nicolas Audibert |
INTERSPEECH | 4 |
| 2018 | Studying Vowel Variation in French-Algerian Arabic Code-switched SpeechabstractInternational audience Jane Wottawa, Djegdjiga Amazouz, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 3 |
| 2018 | The French-Algerian Code-Switching Triggered audio corpus (FACST)
Djegdjiga Amazouz, Martine Adda-Decker, Lori Lamel |
LREC | 2 |
| 2018 | A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noël Kouarata, Lori Lamel, Hélène Bonneau-Maynard, Markus Müller 0001, Annie Rialland, Sebastian Stüker, François Yvon, Marcely Zanon Boito |
LREC | 3 |
| 2018 | Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville)
Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt |
LREC | 2 |
| 2017 | Addressing Code-Switching in French/Algerian Arabic SpeechabstractInternational audience Djegdjiga Amazouz, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 2 |
| 2017 | Developing an Embosi (Bantu C25) Speech Variant Dictionary to Model Vowel Elision and Morpheme DeletionabstractInternational audience Jamison Cooper-Leavitt, Lori Lamel, Annie Rialland, Martine Adda-Decker, Gilles Adda |
INTERSPEECH | 4 |
| 2017 | Investigating the Effect of ASR Tuning on Named Entity Recognition
Mohamed Ameur Ben Jannet, Olivier Galibert, Martine Adda-Decker, Sophie Rosset |
INTERSPEECH | 3 |
| 2017 | Schwa Realization in French: Using Automatic Speech Processing to Study Phonological and Socio-Linguistic Factors in Large CorporaabstractInternational audience Yaru Wu, Martine Adda-Decker, Cécile Fougeron, Lori Lamel |
INTERSPEECH | 2 |
| 2016 | Lig-Aikuma: A Mobile App to Collect Parallel Speech for Under-Resourced Language Studies
Elodie Gauthier, David Blachon, Laurent Besacier, Guy-Noël Kouarata, Martine Adda-Decker, Annie Rialland, Gilles Adda, Grégoire Bachman |
INTERSPEECH | 5 |
| 2016 | Preliminary Experiments on Unsupervised Word Discovery in MboshiabstractInternational audience Pierre Godard, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Laurent Besacier, Hélène Bonneau-Maynard, Guy-Noël Kouarata, Kevin Löser, Annie Rialland, François Yvon |
INTERSPEECH | 3 |
| 2016 | Putting German [ʃ] and [ç] in Two Different Boxes: Native German vs L2 German of French LearnersabstractInternational audience Jane Wottawa, Martine Adda-Decker, Frédéric Isel |
INTERSPEECH | 2 |
| 2016 | French Learners Audio Corpus of German Speech (FLACGS)
Jane Wottawa, Martine Adda-Decker |
LREC | 2 |
| 2015 | A crosslinguistic study of prosodic focusabstractWe examined the production and perception of (contrastive) prosodic focus, using a paradigm based on digit strings, in which the same material and discourse contexts can be used in different languages. We found a striking difference between languages like English and Mandarin Chinese, where prosodic focus is clearly marked in production and accurately recognized in perception, and languages like Korean, where prosodic focus is neither clearly marked in production nor accurately recognized in perception. We also present comparable production data for Suzhou Wu, Japanese, and French. Yong-cheol Lee, Sisi Chen, Martine Adda-Decker, Angélique Amelot, Satoshi Nambu, Mark Y. Liberman |
ICASSP | 4 |
| 2015 | Comparing journalistic and spontaneous speech: prosodic and spectral analysisabstractInternational audience Cédric Gendrot, Martine Adda-Decker, Yaru Wu |
INTERSPEECH | 2 |
| 2015 | How to evaluate ASR output for named entity recognition?abstractInternational audience Mohamed Ameur Ben Jannet, Olivier Galibert, Martine Adda-Decker, Sophie Rosset |
INTERSPEECH | 3 |
| 2015 | Analysing rhythm in ritual discourse in yucatec maya using automatic speech alignmentabstractInternational audience Valentina Vapnarsky, Claude Barras, Cédric Becquey, David Doukhan, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 5 |
| 2014 | An educational platform to capture, visualize and analyze rare singing
Patrick Chawah, Samer Al Kork, Thibaut Fux, Martine Adda-Decker, Angélique Amelot, Nicolas Audibert, Bruce Denby, Gérard Dreyfus, Aurore Jaumard-Hakoun, Claire Pillot-Loiseau, Pierre Roussel-Ragot, Maureen Stone 0001, Kele Xu, Lise Crevier-Buchman |
INTERSPEECH | 4 |
| 2014 | A CRF-based approach to automatic disfluency detection in a French call-centre corpusabstractInternational audience Camille Dutrey, Chloé Clavel, Sophie Rosset, Ioana Vasilescu, Martine Adda-Decker |
INTERSPEECH | 5 |
| 2014 | Pronunciation variation in read and conversational austrian GermanabstractInternational audience Barbara Schuppler, Martine Adda-Decker, Juan Andres Morales-Cordovilla |
INTERSPEECH | 2 |
| 2014 | 3d tongue motion visualization based on ultrasound image sequences
Kele Xu, Yin Yang 0002, Aurore Jaumard-Hakoun, Martine Adda-Decker, Angélique Amelot, Samer Al Kork, Lise Crevier-Buchman, Patrick Chawah, Gérard Dreyfus, Thibaut Fux, Claire Pillot-Loiseau, Pierre Roussel-Ragot, Maureen Stone 0001, Bruce Denby |
INTERSPEECH | 4 |
| 2014 | ETER : a new metric for the evaluation of hierarchical named entity recognition
Mohamed Ameur Ben Jannet, Martine Adda-Decker, Olivier Galibert, Juliette Kahn, Sophie Rosset |
LREC | 2 |
| 2014 | Automatic language identity tagging on word and sentence-level in multilingual text sources: a case-study on Luxembourgish
Thomas Lavergne, Gilles Adda, Martine Adda-Decker, Lori Lamel |
LREC | 3 |
| 2014 | Human annotation of ASR error regions: Is "gravity" a sharable concept for human annotators?
Daniel Luzzati, Cyril Grouin, Ioana Vasilescu, Martine Adda-Decker, Eric Bilinski, Nathalie Camelin, Juliette Kahn, Carole Lailler, Lori Lamel, Sophie Rosset |
LREC | 4 |
| 2013 | Recent evolution of non-standard consonantal variants in French broadcast newsabstractThis paper investigates sociophonetic questions about global tendencies in contemporaneous European spoken French. The authors argue that automatic alignment allowing targeted variants can provide evidence for current hypotheses about possible ongoing sound changes or about destandardization even in formal contexts as broadcast news. This study focused on the evolution over a decade, in radio or TV news, of three 'non- standard' consonantal variants: consonant cluster reduction, affrication/palatalization of dental stops and voiceless fricative epithesis. Measures obtained by this method showed that the first variant remains almost absent in journalists' speech, exactly as affrication of /d/. In contrast, affrication of /t/ is increasing and the fricative epithesis, partially unpredictable, becomes longer. Our findings support the use of automatic alignment as an aid to validate sociolinguistic hypotheses and to develop pattern-driven studies, gathering more variables. Maria Candea, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 2 |
| 2013 | How are word-final schwas different in the north and south of france?abstractThe aim of this paper is twofold: (i) give a large-scale description in realized word-final schwas of French lexical words for different regions (North vs. South) and different speaking styles (read vs. spontaneous speech); (ii) highlight differences in prosodic features and test these differences via automatic classification techniques. The proposed study relies on a subset of 12.5 hours of the French PFC corpus. Manually transcribed speech was segmented and labeled using automatic speech alignment and a pronunciation dictionary including optional word-final schwas for all words ending in a consonant. F0 and intensity values were extracted and averaged over segments. Our study revealed that, for both speaking styles, word-final schwas of southern French tended to keep relatively high F0 values and longer durations in comparison with northern French where F0 tends to drop on a word-final schwa. On average, spontaneous speech featured smaller F0 drops between final full vowel and subsequent word-final schwa vowel as well as longer durations. The automatic North/South classification of word-final schwas achieved better results for spontaneous speech. As for distinguishing between speaking styles, southern French obtained slightly better scores than the northern varieties. Rena Nemoto, Martine Adda-Decker |
INTERSPEECH | 2 |
| 2013 | A tool to elicit and collect multicultural and multimodal laughter
Mariette Soury, Clément Gossart, Martine Adda-Decker, Laurence Devillers |
INTERSPEECH | 3 |
| 2012 | Designing French Tale Corpora for Entertaining Text To Speech Synthesis
David Doukhan, Sophie Rosset, Albert Rilliard, Christophe d'Alessandro, Martine Adda-Decker |
LREC | 5 |
| 2012 | Cross-lingual studies of ASR errors: paradigms for perceptual evaluations
Ioana Vasilescu, Martine Adda-Decker, Lori Lamel |
LREC | 2 |
| 2011 | Prosodic Analysis of a Corpus of TalesabstractInternational audience David Doukhan, Albert Rilliard, Sophie Rosset, Martine Adda-Decker, Christophe d'Alessandro |
INTERSPEECH | 4 |
| 2011 | Cross-Lingual Study of ASR Errors: On the Role of the Context in Human Perception of Near-HomophonesabstractInternational audience Ioana Vasilescu, Dahbia Yahia, Natalie D. Snoeren, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 4 |
| 2011 | Characterisation and identification of non-native French accents
Bianca Vieru-Dimulescu, Philippe Boula de Mareüil, Martine Adda-Decker |
Speech Commun. | 3 |
| 2010 | Comparing mono- & multilingual acoustic seed models for a low e-resourced language: a case-study of luxembourgishabstractLuxembourgish is embedded in a multilingual context on the divide between Romance and Germanic cultures and has often been viewed as one of Europe’s under-resourced languages. We focus on the acoustic modeling of Luxembourgish. By taking advantage of monolingual acoustic seeds selected from German, French or English model sets via IPA symbol correspondances, we investigated whether Luxembourgish spoken words were globally better represented by one of these languages. Although speech in Luxembourgish is frequently interspersed with French words, forced alignments on these data showed a clear preference for Germanic acoustic models with only a limited usage of French. German models provided the best match with 54% of the data, 35% for English and only 11% for French models. A set of multilingual acoustic models, estimated the pooled German, French, and English audio data, captured 27% to 48% of the data depending on conditions. Index Terms: multilingual alignment, acoustic seed models, under-resourced languages, Luxembourgish, English, French, German. Martine Adda-Decker, Lori Lamel, Natalie D. Snoeren |
INTERSPEECH | 1 |
| 2010 | A Question-answer Distance Measure to Investigate QA System Progress
Guillaume Bernard 0002, Sophie Rosset, Martine Adda-Decker, Olivier Galibert |
LREC | 3 |
| 2010 | Word Boundaries in French: Evidence from Large Speech Corpora
Rena Nemoto, Martine Adda-Decker, Jacques Durand |
LREC | 2 |
| 2010 | Comparison of Spectral Properties of Read, Prepared and Casual Speech in French
Jean-Luc Rouas, Mayumi Beppu, Martine Adda-Decker |
LREC | 3 |
| 2010 | The Study of Writing Variants in an Under-resourced Language: Some Evidence from Mobile N-Deletion in Luxembourgish
Natalie D. Snoeren, Martine Adda-Decker, Gilles Adda |
LREC | 2 |
| 2010 | On the Role of Discourse Markers in Interactive Spoken Question Answering Systems
Ioana Vasilescu, Sophie Rosset, Martine Adda-Decker |
LREC | 3 |
| 2010 | The Nijmegen Corpus of Casual French
Francisco Torreira, Martine Adda-Decker, Mirjam Ernestus |
Speech Commun. | 2 |
| 2009 | A perceptual investigation of speech transcription errors involving frequent near-homophones in French and american EnglishabstractThis article compares the errors made by automatic speech recognizers to those made by humans for near-homophones in American English and French. This exploratory study focuses on the impact of limited word context and the potential resulting ambiguities for automatic speech recognition (ASR) systems and human listeners. Perceptual experiments using 7-gram chunks centered on incorrect or correct words output by an ASR system, show that humans make significantly more transcription errors on the first type of stimuli, thus highlighting the local ambiguity. The long-term aim of this study is to improve the modeling of such ambiguous items in order to reduce ASR errors. Ioana Vasilescu, Martine Adda-Decker, Lori Lamel, Pierre A. Hallé |
INTERSPEECH | 2 |
| 2009 | Linguistically-motivated automatic classification of regional French varietiesabstractThe goal of this study is to automatically differentiate French varieties (standard French and French varieties spoken in the South of France, Alsace, Belgium and Switzerland) by applying a linguistically-motivated approach. We took advantage of automatic phoneme alignment to measure vowel formants, consonant (de)voicing, pronunciation variants as well as prosodic cues. These features were then used to identify French varieties by applying classification techniques. On large corpora of hundreds of speakers, over 80% correct identification scores were obtained. The confusions between varieties and the features used (by decision trees) are linguistically grounded. Index Terms: language variety identification, regional French accents, classification Cécile Woehrling, Philippe Boula de Mareüil, Martine Adda-Decker |
INTERSPEECH | 3 |
| 2008 | A corpus-based prosodic study of Alsatian, Belgian and Swiss FrenchabstractThe object of this paper is a prosodic study of the French language as it is spoken in Alsace, Belgium and Switzerland, also compared with standard French through large corpora (over 100 hours) of scripted and spontaneous speech. The data were segmented into phones by automatic alignment; pitch values were extracted and averaged over segments. Two features are addressed: initial stress (through pitch and duration correlates) and penultimate lengthening. Different patterns enable us to distinguish the three varieties under investigation. Swiss speakers exhibit pitch rise and polysyllabic word onset lengthening in clitic–nonclitic sequences, while Alsatians tend to lengthen the initial vowel of nonclitic words. Belgians show prepausal penultimate lengthening whereas the Swiss tend to lengthen the last two prepausal vowels. Cécile Woehrling, Philippe Boula de Mareüil, Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 3 |
| 2008 | Annotation and analysis of overlapping speech in political interviews
Martine Adda-Decker, Claude Barras, Gilles Adda, Patrick Paroubek, Philippe Boula de Mareüil, Benoit Habert |
LREC | 1 |
| 2008 | Developments of "Lëtzebuergesch" Resources for Automatic Speech Processing and Linguistic Studies
Martine Adda-Decker, Thomas Pellegrini, Eric Bilinski, Gilles Adda |
LREC | 1 |
| 2008 | Speech Errors on Frequently Observed Homophones in French: Perceptual Evaluation vs Automatic Classification
Rena Nemoto, Ioana Vasilescu, Martine Adda-Decker |
LREC | 3 |
| 2006 | Language, gender, speaking style and language proficiency as factors influencing the autonomous vocalic filler production in spontaneous speechabstractAbstract This paper deals with the factors characterizing the production of autonomous vocalic filled pauses in large spontaneous speech corpora, namely language, gender, speaking style and language proficiency. Two types of corpora are analyzed: a corpus of broadcast news in French and American English and a corpus of short talks in a conference in English spoken by native and non-native speakers. Several acoustic and prosodic parameters are evaluated and correlated with each factor, namely timbre, pitch, duration and density. Results presented here show that the timbre is correlated with language and language proficiency, whereas the duration is linked both to gender and speaking style, the latter conditioning also the hesitation density in speech. Index terms: speech disfluencies, autonomous filled pauses, L1/L2, emotional state. 1. Introduction This paper focuses on autonomous vocalic filled pauses in spontaneous speech corpora. Among the phenomena described as “disfluencies”, filled pauses represent one of the most frequently encountered across languages. Autonomous vocalic hesitations as a type of filled pause are widely represented and consist in the insertion “at any moment” in the speech flow of a lengthened vocalic segment, alone or in combination with other segments (such as a nasal coda in English). Its aim is “to announce the initiation of what is expected to be a […] delay in speaking” [1]. Autonomous vocalic hesitations occur without lexical support and are thus to be distinguished from vocal lengthening of segments belonging to lexical items (generally function words). Filled pauses have however other possible realizations, as for instance lengthened nasal consonants (“mm” in Mandarin Chinese) or demonstratives (“ano”, “eto” in Japanese) [2,3]. For the present study we consider vocalic hesitations in French (“euh”) and English (“uh”, “um” in American English; ”er” in British English). Previously autonomous vocalic hesitations have been studied in intra- and inter-language perspectives with no particular consideration of the role of the context on their acoustic and prosodic characteristics. In our former studies, we have compared autonomous vocalic hesitations in 8 languages: American English, Middle Oriental Arabic, Mandarin Chinese, French, Italian, South-American Spanish, and European Portuguese. We have focused on the support vowel of the hesitations in each considered language. The support vowel has been defined as the main vocalic segment of a hesitation, i.e. the longest and most stable realization of each item. This vowel occurs in isolation (as unique realization of the hesitation), in a diphthong or followed by a nasal consonant as in English. Among the parameters characterizing the support vowel, duration, pitch and timbre have received a particular attention. Analysis revealed that the timbre is the most language-dependent parameter characterizing vocalic hesitations. Pitch and duration help both at differentiating the hesitation vowel from vowels with similar timbre within a given language. Pitch and duration seem to show universal patterns, i.e. the main vowel of a hesitation is significantly longer than other similar intra-lexical vowels and exhibits a flat and stable F0 contour [4]. Consequently, the hypothesis has been made that timbre is a language-dependent parameter, whereas pitch and duration could be considered as language-independent features. In this study, we consider 4 factors which may play a role in the production of vocalic hesitation in spontaneous speech corpora: language; gender; spoken style and language proficiency (mother tongue vs. second language). Ioana Vasilescu, Martine Adda-Decker |
INTERSPEECH | 2 |
| 2005 | Do speech recognizers prefer female speakers?
Martine Adda-Decker, Lori Lamel |
INTERSPEECH | 1 |
| 2005 | Where are we in transcribing French broadcast news?abstractInternational audience Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Véronique Gendner, Lori Lamel, Holger Schwenk |
INTERSPEECH | 3 |
| 2005 | Impact of duration on F1/F2 formant values of oral vowels: an automatic analysis of large broadcast news corpora in French and GermanabstractInternational audience Cédric Gendrot, Martine Adda-Decker |
INTERSPEECH | 2 |
| 2005 | Perceptual salience of language-specific acoustic differences in autonomous fillers across eight languagesabstractInternational audience Ioana Vasilescu, Maria Candea, Martine Adda-Decker |
INTERSPEECH | 3 |
| 2005 | Different size multilingual phone inventories and context-dependent acoustic models for language identification
Martine Adda-Decker, Fabien Antoine |
INTERSPEECH | 2 |
| 2005 | Investigating syllabic structures and their variation in spontaneous French
Martine Adda-Decker, Philippe Boula de Mareüil, Gilles Adda, Lori Lamel |
Speech Commun. | 1 |
| 2004 | Speech transcription in multiple languagesabstractThe paper summarizes recent work underway at LIMSI on speech-to-text transcription in multiple languages. The research has been oriented towards the processing of broadcast audio and conversational speech for information access. Broadcast news transcription systems have been developed for seven languages, and it is planned to address several other languages in the near term. Research on conversational speech has mainly focused on the English language, with some initial work on French, Arabic and Spanish. Automatic processing must take into account the characteristics of the audio data, such as needing to deal with the continuous data stream, specificities of the language and the use of an imperfect word transcription for accessing the information content. Our experience thus far indicates that at today's word error rates, the techniques used in one language can be successfully ported to other languages, and most of the language specificities concern lexical and pronunciation modeling. Lori Lamel, Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Leonardo Canseco-Rodriguez, Langzhou Chen, Olivier Galibert, Abdelkhalek Messaoudi, Holger Schwenk |
ICASSP (3) | 4 |
| 2004 | Automatic Audio and Manual Transcripts Alignment, Time-code Transfer and Selection of Exact Transcripts
Claude Barras, Gilles Adda, Martine Adda-Decker, Benoit Habert, Philippe Boula de Mareüil, Patrick Paroubek |
LREC | 3 |
| 2003 | A corpus-based decompounding algorithm for German lexical modeling in LVCSRabstractIn this paper a corpus-based decompounding algorithm is described and applied for German LVCSR. The decompounding algorithm contributes to address two major problems for LVCSR: lexical coverage and letter-to-sound conversion. The idea of the algorithm is simple: given a word start of length # only few different characters can continue an admissible word in the language. But concerning compounds, if word start # reaches a constituent word boundary, the set of successor characters can theoretically include any character. The algorithm has been applied to a 300M word corpus with 2.6M distinct words. 800k decomposition rules have been extracted automatically. OOV (out of vocabulary) word reductions of 25% to 50% relative have been achieved using word lists from 65k to 600k words. Pronunciation dictionaries have been developed for the LIMSI 300k German recognition system. As no language specific knowledge is required beyond the text corpus, the algorithm can apply more generally to any compounding language. Martine Adda-Decker |
INTERSPEECH | 1 |
| 2003 | The 300k LIMSI German broadcast news transcription system
Kevin McTait, Martine Adda-Decker |
INTERSPEECH | 2 |
| 2002 | Studying pronunciation variants in French by using alignment techniques
Philippe Boula de Mareüil, Martine Adda-Decker |
INTERSPEECH | 2 |
| 2001 | Processing Broadcast Audio for Information AccessabstractThis paper addresses recent progress in speaker-independent, large vocabulary, continuous speech recognition, which has opened up a wide range of near and mid-term applications. One rapidly expanding application area is the processing of broadcast audio for information access. At LIMSI, broadcast news transcription systems have been developed for English, French, German, Mandarin and Portuguese, and systems for other languages are under development. Audio indexation must take into account the specificities of audio data, such as needing to deal with the continuous data stream and an imperfect word transcription. Some near-term applications areas are audio data mining, selective dissemination of information and media monitoring. Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker, Claude Barras, Langzhou Chen, Yannick de Kercadio |
ACL | 4 |
| 2001 | Using information retrieval methods for language model adaptationabstractIn this paper we report experiments on language model adaptation using information retrieval methods, drawing upon recent developments in information extraction and topic tracking. One of the problems is extracting reliable topic information with high confidence from the audio signal in the presence of recognition errors. The work in the information retrieval domain on information extraction and topic tracking suggested a new way to solve this problem. In this work, we make use of information retrieval methods to extract topic information in the word recognizer hypotheses, which are then used to automatically select adaptation data from a very large general text corpus. Two adaptive language models, a mixture based model and a MAP based model, have been investigated using the adaptation data. Experiments carried out with the LIMSI Mandarin broadcast news transcription system gives a relative character error rate reduction of 4.3% with this adaptation method. Langzhou Chen, Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker |
INTERSPEECH | 5 |
| 2001 | Towards multilingual interoperability in automatic speech recognition
Martine Adda-Decker |
Speech Commun. | 1 |
| 2000 | Investigating text normalization and pronunciation variants for German broadcast transcriptionabstractIn this paper we describe our ongoing work concerning lexical modeling in the LIMSI broadcast transcription system for German. Lexical decomposition is investigated with a twofold goal: lexical coverage optimization and improved letter-to-sound conversion. A set of about 450 decompounding rules, developed using statistics from a 300M word corpus, reduces the OOV rate from 4.5% to 4.0% on a 30k development text set. Adding partial inflection stripping, the OOV rate drops to 2.9%. For letterto -sound conversion, decompounding reduces cross-lexeme ambiguities and thus contributes to more consistent pronunciation dictionaries. Another point of interest concerns reduced pronunciation modeling. Word error rates, measured on 1.3 hours of ARTE TV broadcast, vary between 18 and 24% depending on the show and the system configuration. Our experiments indicate that using reduced pronunciations slightly decreases word error rates. 1. INTRODUCTION The German language, more than other major western ... Martine Adda-Decker, Gilles Adda, Lori Lamel |
INTERSPEECH | 1 |
| 1999 | Large vocabulary speech recognition in FrenchabstractWe present some design considerations concerning our large vocabulary continuous speech recognition system in French. The impact of the epoch of the text training material on lexical coverage, language model perplexity and recognition performance on newspaper texts is demonstrated. The effectiveness of larger vocabulary sizes and larger text training corpora for language modeling is investigated. French is a highly inflected language producing large lexical variety and a high homophone rate. About 30% of recognition errors are shown to be due to substitutions between inflected forms of a given root form. When word error rates are analysed as a function of word frequency, a significant increase in the error rate can be measured for frequency ranks above 5000. Martine Adda-Decker, Gilles Adda, Jean-Luc Gauvain, Lori Lamel |
ICASSP | 1 |
| 1999 | Comparing different model configurations for language identification using a phonotactic approachabstractIn this paper different model configurations for language identification using a phonotactic approach are explored. Identification experiments were carried out on the 11-language telephone speech corpus OGI-TS, containing calls in French, English, German, Spanish, Japanese, Korean, Mandarin, Tamil, Farsi, Hindi, and Vietnamese. Phone sequences output by one or multiple phone recognizers are rescored with language-dependent phonotactic models approximated by phone bigrams. The parameters of different sets of acoustic phone models were estimated using the 4-language IDEAL corpus. Sets of language-specific phonotactic models were trained using the training portion of the OGITS CORPUS. Error rates are significantly reduced by combining language-dependent and language-independent acoustic decoders, especially for short segments. A 9.9% LID error rate was obtained on the 11-language task using phonotactic models trained on spontaneous speech data. These results show that the phonotactic approach is relative insensitive to an acoustic mismatch between training and test conditions. Driss Matrouf, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel |
EUROSPEECH | 2 |
| 1999 | Pronunciation variants across system configuration, language and speaking style
Martine Adda-Decker, Lori Lamel |
Speech Commun. | 1 |
| 1998 | Multilingual phone recognition of spontaneous telephone speechabstractIn this paper we report on experiments with phone recognition of spontaneous telephone speech. Phone recognizers were trained and assessed on IDEAL, a multilingual corpus containing telephone speech in French, British English, German and Castillan Spanish. We investigated the influence of the training material composition (size and linguistic content) on the recognition performance using context-independent (CI) hidden Markov models (HMMs) and phonotactic bigram models. We found that when testing on spontaneous speech data, using only spontaneous speech training data gave the highest phone accuracies for the four languages, even though this data comprises only 14% of the available training data. The use of context-dependent (CD) HMMs reduced the phone error across the 4 languages, with the average error reduced to 51.9% from the 57.4% obtained with CI models. We suggest a straightforward way of detecting non speech phenomena. The basic idea is to remove sequences of consonants between two silence labels from the recognized phone strings prior to scoring. This simple technique reduces the relative average phone error rate by 5.4%. The lowest phone error with CD models and filtering was obtained for Spanish (39.1%) with 4 language average being 49.1%. Cristobal Corredor-Ardoy, Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain |
ICASSP | 3 |
| 1998 | Language identification incorporating lexical informationabstractIn this paper we explore the use of lexical information for language identification (LID). Our reference LID system uses language-dependent acoustic phone models and phone-based bigram language models. For each language, lexical information is introduced by augmenting the phone vocabulary with the N most frequent words in the training data. Combined phone and word bigram models are used to provide linguistic constraints during acoustic decoding. Experiments were carried out on a 4-language telephone speech corpus. Using lexical information achieves a relative error reduction of about 20% on spontaneous and read speech compared to the reference phone-based system. Identification rates of 92%, 96% and 99% are achieved for spontaneous, read and task-specific speech segments respectively, with prior speech detection. Driss Matrouf, Martine Adda-Decker, Lori Lamel, Jean-Luc Gauvain |
ICSLP | 2 |
| 1998 | On the use of speech and text corpora for speech recognition in French
Martine Adda-Decker, Gilles Adda, Lori Lamel, Jean-Luc Gauvain |
LREC | 1 |
| 1998 | Towards tokenization evaluation
Benoit Habert, Gilles Adda, Martine Adda-Decker, Philippe Boula de Mareüil, Silvana Ferrari, Olivier Ferret, Gabriel Illouz, P. Paraubeck |
LREC | 3 |
| 1998 | A multilingual corpus for language identification
Lori Lamel, Gilles Adda, Martine Adda-Decker, Cristobal Corredor-Ardoy, Jean-Jacques Gangolf, Jean-Luc Gauvain |
LREC | 3 |
| 1997 | Transcribing broadcast news showsabstractWhile significant improvements have been made in large vocabulary continuous speech recognition of large read-speech corpora such as the ARPA Wall Street Journal-based CSR corpus (WSJ) for American English and the BREF corpus for French, these tasks remain relatively artificial. In this paper we report on our development work in moving from laboratory read speech data to real-world speech data in order to build a system for the new ARPA broadcast news transcription task. The LIMSI Nov96 speech recognizer makes use of continuous density HMMs with Gaussian mixtures for acoustic modeling and n-gram statistics estimated on newspaper texts. The acoustic models are trained on the WSJO/WSJ1, and adapted using MAP estimation with task-specific training data. The overall word error on the Nov96 partitioned evaluation test was 27.1%. Jean-Luc Gauvain, Gilles Adda, Lori Lamel, Martine Adda-Decker |
ICASSP | 4 |
| 1997 | Text normalization and speech recognition in FrenchabstractIn this paper we present a quantitative investigation into the impact of text normalization on lexica and language models for speech recognition in French. The text normalization process defines what is considered to be a word by the recognition system. Depending on this definition we can measure different lexical coverages and language model perplexities, both of which are closely related to the speech recognition accuracies obtained on read newspaper texts. Different text normalizations of up to 185M words of newspaper texts are presented along with corresponding lexical coverage and perplexity measures. Some normalizations were found to be necessary to achieve good lexical coverage, while others were more or less equivalent in this regard. The choice of normalization to create language models for use in the recognition experiments with read newspaper texts was based on these findings. Our best system configuration obtained a 11.2% word error rate in the AUPELF `French-speaking' spee... Gilles Adda, Martine Adda-Decker, Jean-Luc Gauvain, Lori Lamel |
EUROSPEECH | 2 |
| 1997 | Language identification with language-independent acoustic modelsabstractIn this paper we explore the use of languageindependent acoustic models for language identification (LID). The phone sequence output by a single language-independent phone recognizer is rescored with language-dependent phonotactic models approximated by phone bigrams. The language-independent phoneme inventory was obtained by Agglomerative Hierarchical Clustering, using a measure of similarity between phones. This system is compared with a parallel language-dependent phone architecture, which uses optimally the acoustic log likelihood and the phonotactic score for language identification. Experiments were carried out on the 4-language telephone speech corpus IDEAL, containing calls in British English, Spanish, French and German. Results show that the language-independent approach performs as well as the language-dependent one: 9% versus 10% of error rate on 10 second chunks, for the 4-language task. 1. INTRODUCTION This paper presents some of our recent research on automatic language ... Cristobal Corredor-Ardoy, Jean-Luc Gauvain, Martine Adda-Decker, Lori Lamel |
EUROSPEECH | 3 |
| 1997 | Transcription of broadcast newsabstractIn this paper we report on our recent work in transcribing broadcast news shows. Radio and television broadcasts contain signal segments of various linguistic and acoustic natures. The shows contain both prepared and spontaneous speech. The signal may be studio quality or have been transmitted over a telephone or other noisy channel (ie., corrupted by additive noise and nonlinear distorsions), or may contain speech over music. Transcription of this type of data poses challenges in dealing with the continuous stream of data under varying conditions. Our approach to this problem is to segment the data into a set of categories, which are then processed with category specific acoustic models. We describe our 65k speech recognizer and experiments using different sets of acoustic models for transcription of broadcast news data. The use of prior knowledge of the segment boundaries and types is shown to not crucially affect the performance. 1. INTRODUCTION The goal of this research is to au... Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker |
EUROSPEECH | 4 |
| 1997 | Multilingual large vocabulary speech recognition: the European SQALE project
Steve J. Young, Martine Adda-Decker, Xavier L. Aubert, Christian Dugast, Jean-Luc Gauvain, Dan J. Kershaw, Lori Lamel, David A. van Leeuwen, David Pye, Anthony J. Robinson, Herman J. M. Steeneken, Philip C. Woodland |
Comput. Speech Lang. | 2 |
| 1996 | Developments in large vocabulary, continuous speech recognition of GermanabstractWe describe our large vocabulary continuous speech recognition system for the German language, the development of which was partly carried out within the context of the European LRE project 62-058 SQALE. The recognition system is the LIMSI recognizer originally developed for French and American English, which has been adapted to German. Specificities of German, as relevant to the recognition system, are presented. These specificities have been accounted for during the recognizer's adaptation process. We present experimental results on a first test set ger-dev95 to measure progress in system development. Results are given with the final system using different acoustic model sets on two test sets ger-dev95 and ger-eval95. This system achieved a word error rate of 17.3% (official word error rate of 16.1% after SQALE adjudication process) on the ger-eval95 test set. Martine Adda-Decker, Gilles Adda, Lori Lamel, Jean-Luc Gauvain |
ICASSP | 1 |
| 1996 | Spoken language processing in a multilingual contextabstractIn this paper we overview the spoken language processing activities at LIMSI, which are carried out in a multilingual framework.These activities include speech-to-text conversion, spoken language systems for information retrieval, speaker and language recognition, and speech response.The Spoken Language Processing Group has also been actively involved in corpora development and evaluation.The group has regularly participated in evaluations organized by ARPA, in the LE-SQALE project, and in the AUPELF-UREF program for provision of linguistic resources and evaluation tests for French. Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain, Gilles Adda |
ICSLP | 2 |
| 1995 | Developments in continuous speech dictation using the ARPA WSJ taskabstractWe report on our recent development work in large vocabulary, American English continuous speech dictation. We have experimented with (1) alternative analyses for the acoustic front end, (2) the use of an enlarged vocabulary so as to reduce the number of errors due to out-of-vocabulary words, (3) extensions to the lexical representation, (4) the use of additional acoustic training data, and (5) modification of the acoustic models for telephone speech. The recognizer was evaluated on Hubs 1 and 2 of the fall 1994 ARPA NAB CSR Hub and Spoke Benchmark test. Experimental results for development and evaluation test data are given, as well as an analysis of the errors on the development data. Jean-Luc Gauvain, Lori Lamel, Martine Adda-Decker |
ICASSP | 3 |
| 1995 | Issues in Large Vocabulary, Multilingual Speech RecognitionabstractIn this paper we report on our activities in multilingual, speakerindependent, large vocabulary continuous speech recognition. The multilingual aspect of this work is of particular importance in Europe, where each country has its own national language. Our existing recognizer for American English and French, has been ported to British English and German. It has been assessed in the context of the LRESQALE project whose objective was to experiment with installing in Europe a multilingual evaluation paradigm for the assessment of large vocabulary, continuous speech recognition systems. The recognizer makes use of phone-based continuous density HMM for acoustic modeling and n-gram statistics estimated on newspaper texts for language modeling. The system has been evaluated on a dictation task with read, newspaper-based corpora, the ARPA Wall Street Journal corpus of American English, the WSJCAM0 corpus of British English, the BREF-Le Monde corpus of French and the PHONDAT-Frankfurter Runds... Lori Lamel, Martine Adda-Decker, Jean-Luc Gauvain |
EUROSPEECH | 2 |
| 1994 | The LIMSI continuous speech dictation system: evaluation on the ARPA Wall Street Journal taskabstractWe report progress made at LIMSI in speaker-independent large vocabulary speech dictation using the ARPA Wall Street Journal-based CSR corpus. The recognizer makes use of continuous density HMM with Gaussian mixture for acoustic modeling and n-gram statistics estimated on the newspaper texts for language modeling. The recognizer uses a time-synchronous graph-search strategy which is shown to still be viable with vocabularies of up to 20 K words when used with bigram back-off language models. A second forward pass, which makes use of a word graph generated with the bigram, incorporates a trigram language model. Acoustic modeling uses cepstrum-based features, context-dependent phone models (intra and interword), phone duration models, and sex-dependent models. The recognizer has been evaluated in the Nov92 and Nov93 ARPA tests for vocabularies of up to 20,000 words.> Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker |
ICASSP (1) | 4 |
| 1994 | Continuous speech dictation in FrenchabstractA major research activity at LIMSI is multilingual, speakerindependent, large vocabulary speech dictation. In this paper we report on efforts in large vocabulary, speaker-independent continuous speech recognition of French using the BREF corpus. Recognition experiments were carried out with vocabularies containing up to 20k words. The recognizer makes use of continuous density HMM with Gaussian mixture for acoustic modeling and n-gram statistics estimated on 38 million words of newspaper text from Le Monde for language modeling. The recognizer uses a time-synchronous graph-search strategy. When a bigram language model is used, recognition is carried out in a single forward pass. A second forward pass, which makes use of a word graph generated with the bigram language model, incorporates a trigram language model. Acoustic modeling uses cepstrum-based features, contextdependent phone models and phone duration models. An average phone accuracy of 86% was achieved. A word accuracy of 84% h... Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker |
ICSLP | 4 |
| 1994 | Speaker-independent continuous speech dictation
Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker |
Speech Communication | 4 |
| 1993 | Speaker-independent continuous speech dictationabstractAbstract In this paper we report on progress made at LIMSI in speaker-independent large vocabulary speech dictation using newspaper-based speech corpora in English and French. The recognizer makes use of continuous density HMMs with Gaussian mixtures for acoustic modeling and n -gram statistics estimated on newspaper texts for language modeling. Acoustic modeling uses cepstrum-based features, context-dependent phone models (intra and interword), phone duration models, and sex-dependent models. For English the ARPA Wall Street Journal -based CSR corpus is used and for French the BREF corpus containing recordings of texts from the French newspaper Le Monde is used. Experiments were carried out with both these corpora at the phone level and at the word level with vocabularies containing up to 20,000 words. Word recognition experiments are also described for the ARPA RM task which has been widely used to evaluate and compare systems. Jean-Luc Gauvain, Lori Lamel, Gilles Adda, Martine Adda-Decker |
EUROSPEECH | 4 |
| 1992 | Experiments on stress-dependent phone modelling for continuous speech recognitionabstractStress is an important feature for speech recognition. The acoustic realization of a stressed phone trends to be rather different from its unstressed counterpart, and stress may be a distinctive feature for lexical classification, as in the case for English. In most cases stress is important for understanding the meaning of an utterance, and is thus related to syntax and semantics. The authors focus on the acoustic modeling of stressed phones, and, particularly, on stressed vowels. The recognition system is based on discrete HMM phone models, and has been developed within the European ESPRIT-POLYGLOT 2041 project. The influence of the stress feature on acoustic modeling is being assessed on two different continuous speech databases in two different languages (French and American-English), as the effect of stress on the acoustic evidence may be language-dependent. First results concerning phonetic and lexical evaluations are given.> Martine Adda-Decker, Gilles Adda |
ICASSP | 1 |
| 1989 | Continuous speech recognition using phone-based anchor point detection and diphone-based dp-matching
Martine Adda-Decker |
EUROSPEECH | 1 |