Laurent Besacier

dblp:67/3881 · DBLP profile ↗
← Back
146ranked-venue papers
14as first author
21since 2021 · last 2026
0000-0001-7411-9125ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 110 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 86 · 14 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Surfacing Governing Principles for Chatbots: A Workbench and Comparative Study
abstract
Trust in Large Language Model chatbots depends not only on what these systems do but also on how their behavior is governed and communicated. We present Trust Mediator, a workbench that supports service owners in authoring and assessing principle sets for LLM-driven chatbots through persona-based exploration and structured scaffolds. To examine this workflow, we use three analytic lenses—specificity, coverage, and coherence—to characterize the principles produced. In an exploratory between-subjects study, we compared manual and assisted principle authoring. Participants in both conditions viewed principles as useful for governing and assessing chatbot behavior. Assisted authoring was generally perceived as more supportive and tended to broaden coverage. Manual authoring required more effort but yielded principles that were significantly more specific. These findings highlight complementary strengths of assisted and manual pathways, illustrating the value of treating principle sets as design objects within governance workflows. Beyond their analytic role in this study, the lenses also suggest opportunities for supporting the construction and inspection of principle sets.
Antonietta Grasso, Jisun Park 0005, Jutta Willamowski, Laurent Besacier, Jos Rozen
CHI4
2025 Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
abstract
While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final data quality. To reduce the costs of these essential controls, this paper investigates the use of Speech Foundation Models (SFMs) to automate the validation process, examining for the first time the cost/quality trade-off in data acquisition. Experiments conducted on French, German, and Korean data demonstrate that SFM-based validation has the potential to reduce reliance on human validation, resulting in an estimated cost saving of over 40.0% without degrading final data quality. These findings open new opportunities for more efficient, cost-effective, and scalable speech data acquisition.
Beomseok Lee, Marco Gaido, Ioan Calapodescu, Laurent Besacier, Matteo Negri
COLING4
2025 ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
abstract
Research on Large Language Models (LLMs) has recently witnessed an increasing interest in extending the models’ context size to better capture dependencies within long documents. While benchmarks have been proposed to assess long-range abilities, existing efforts primarily considered generic tasks that are not necessarily aligned with real-world applications. In contrast, we propose a new benchmark for long-context LLMs focused on a practical meeting assistant scenario in which the long contexts consist of transcripts obtained by automatic speech recognition, presenting unique challenges for LLMs due to the inherent noisiness and oral nature of such data. Our benchmark, ELITR-Bench, augments the existing ELITR corpus by adding 271 manually crafted questions with their ground-truth answers, as well as noisy versions of meeting transcripts altered to target different Word Error Rate levels. Our experiments with 12 long-context LLMs on ELITR-Bench confirm the progress made across successive generations of both proprietary and open models, and point out their discrepancies in terms of robustness to transcript noise. We also provide a thorough analysis of our GPT-4-based evaluation, including insights from a crowdsourcing study. Our findings indicate that while GPT-4’s scores align with human judges, its ability to distinguish beyond three score levels may be limited.
Thibaut Thonet, Laurent Besacier, Jos Rozen
COLING2
2024 mHuBERT-147: A Compact Multilingual HuBERT Model
Marcely Zanon Boito, Vivek Iyer, Nikolaos Lagos, Laurent Besacier, Ioan Calapodescu
INTERSPEECH4
2024 Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri, Laurent Besacier
INTERSPEECH5
2024 LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Comput. Speech Lang.22
2022 Learning From Failure: Data Capture in an Australian Aboriginal Community
abstract
Most low resource language technology development is premised on the need to collect data for training statistical models.When we follow the typical process of recording and transcribing text for small Indigenous languages, we hit up against the so-called "transcription bottleneck."Therefore it is worth exploring new ways of engaging with speakers which generate data while avoiding the transcription bottleneck.We have deployed a prototype app for speakers to use for confirming system guesses in an approach to transcription based on word spotting.However, in the process of testing the app we encountered many new problems for engagement with speakers.This paper presents a close-up study of the process of deploying data capture technology on the ground in an Australian Aboriginal community.We reflect on our interactions with participants and draw lessons that apply to anyone seeking to develop methods for language data collection in an Indigenous community.
Éric Le Ferrand, Steven Bird, Laurent Besacier
ACL (1)3
2022 Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models
abstract
Multi-encoder models are a broad family of context-aware neural machine translation systems that aim to improve translation quality by encoding document-level contextual information alongside the current sentence.The context encoding is undertaken by contextual parameters, trained on document-level data.In this work, we discuss the difficulty of training these parameters effectively, due to the sparsity of the words in need of context (i.e., the training signal), and their relevant context.We propose to pre-train the contextual parameters over split sentence pairs, which makes an efficient use of the available data for two reasons.Firstly, it increases the contextual training signal by breaking intra-sentential syntactic relations, and thus pushing the model to search the context for disambiguating clues more frequently.Secondly, it eases the retrieval of relevant context, since context segments become shorter.We propose four different splitting methods, and evaluate our approach with BLEU and contrastive test sets.Results show that it consistently improves learning of contextual parameters, both in low and high resource settings.
Lorenzo Lupo, Marco Dinarelli, Laurent Besacier
ACL (1)3
2022 Weakly Supervised Word Segmentation for Computational Language Documentation
abstract
Word and morpheme segmentation are fundamental steps of language documentation as they allow to discover lexical units in a language for which the lexicon is unknown.However, in most language documentation scenarios, linguists do not start from a blank page: they may already have a pre-existing dictionary or have initiated manual segmentation of a small part of their data.This paper studies how such a weak supervision can be taken advantage of in Bayesian non-parametric models of segmentation.Our experiments on two very low resource languages (Mboshi and Japhug), whose documentation is still in progress, show that weak supervision can be beneficial to the segmentation quality.In addition, we investigate an incremental learning scenario where manual segmentations are provided in a sequential manner.This work opens the way for interactive annotation tools for documentary linguists.
Shu Okabe, Laurent Besacier, François Yvon
ACL (1)2
2022 Fashioning Local Designs from Generic Speech Technologies in an Australian Aboriginal Community
abstract
An increasing number of papers have been addressing issues related to low-resource languages and the transcription bottleneck paradigm. After several years spent in Northern Australia, where some of the strongest Aboriginal languages are spoken, we could observe a gap between the motivations depicted in research contributions in this space and the Northern Australian context. In this paper, we address this gap in research by exploring the potential of speech recognition in an Aboriginal community. We describe our work from training a spoken term detection system to its implementation in an activity with Aboriginal participants. We report here on one side how speech recognition technologies can find their place in an Aboriginal context and, on the other, methodological paths that allowed us to reach better comprehension and engagement from Aboriginal participants.
Éric Le Ferrand, Steven Bird, Laurent Besacier
COLING3
2022 SMaLL-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource Languages
abstract
Alireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson, Laurent Besacier. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Alireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson 0001, Laurent Besacier
EMNLP6
2022 A Study of Gender Impact in Self-supervised Models for Speech-to-Text Systems
abstract
Self-supervised models for speech processing emerged recently as popular foundation blocks in speech processing pipelines.These models are pre-trained on unlabeled audio data and then used in speech processing downstream tasks such as automatic speech recognition (ASR) or speech translation (ST).Since these models are now used in research and industrial systems alike, it becomes necessary to understand the impact caused by some features such as gender distribution within pre-training data.Using French as our investigation language, we train and compare gender-specific wav2vec 2.0 models against models containing different degrees of gender balance in their pretraining data.The comparison is performed by applying these models to two speech-to-text downstream tasks: ASR and ST.Results show the type of downstream integration matters.We observe lower overall performance using gender-specific pretraining before fine-tuning an end-to-end ASR system.However, when self-supervised models are used as feature extractors, the overall ASR and ST results follow more complex patterns in which the balanced pre-trained model does not necessarily lead to the best results.Lastly, our crude 'fairness' metric, the relative performance difference measured between female and male test sets, does not display a strong variation from balanced to gender-specific pre-trained wav2vec 2.0 models.
Marcely Zanon Boito, Laurent Besacier, Natalia A. Tomashenko, Yannick Estève
INTERSPEECH2
2022 ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks
abstract
We aim at improving spoken language modeling (LM) using very large amount of automatically transcribed speech.We leverage the INA (French National Audiovisual Institute 1 ) collection and obtain 19GB of text after applying ASR on 350,000 hours of diverse TV shows.From this, spoken language models are trained either by fine-tuning an existing LM (FlauBERT 2 ) or through training a LM from scratch.New models (FlauBERT-Oral) are shared with the community and evaluated for 3 downstream tasks: spoken language understanding, classification of TV shows and speech syntactic parsing.Results show that FlauBERT-Oral can be beneficial compared to its initial FlauBERT version demonstrating that, despite its inherent noisy nature, ASR-generated text can be used to build spoken language models.
Valentin Pelloin, Franck Dary, Nicolas Hervé, Benoît Favre, Nathalie Camelin, Antoine Laurent, Laurent Besacier
INTERSPEECH7
2022 BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model
abstract
International audience
Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber
INTERSPEECH2
2021 Controlling Prosody in End-to-End TTS: A Case Study on Contrastive Focus Generation
abstract
We are also grateful to our
Siddique Latif, Inyoung Kim, Ioan Calapodescu, Laurent Besacier
CoNLL4
2021 Multilingual Unsupervised Neural Machine Translation with Denoising Adapters
abstract
We consider the problem of multilingual unsupervised machine translation, translating to and from languages that only have monolingual data by using auxiliary parallel language pairs.For this problem the standard procedure so far to leverage the monolingual data is back-translation, which is computationally costly and hard to tune.In this paper we propose instead to use denoising adapters, adapter layers with a denoising objective, on top of pre-trained mBART-50.In addition to the modularity and flexibility of such an approach we show that the resulting translations are on-par with back-translating as measured by BLEU, and furthermore it allows adding unseen languages incrementally.
Ahmet Üstün, Alexandre Berard, Laurent Besacier, Matthias Gallé
EMNLP (1)3
2021 An Empirical Study of End-To-End Simultaneous Speech Translation Decoding Strategies
abstract
This paper proposes a decoding strategy for end-to-end simultaneous speech translation. We leverage end-to-end models trained in offline mode and conduct an empirical study for two language pairs (English-to-German and English-to-Portuguese). We also investigate different output token granularities including characters and Byte Pair Encoding (BPE) units. The results show that the proposed decoding approach allows to control BLEU/Average Lagging trade-off along different latency regimes. Our best decoding settings achieve comparable results with a strong cascade model evaluated on the simultaneous translation track of IWSLT 2020 shared task.
Yannick Estève, Laurent Besacier
ICASSP3
2021 LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech
abstract
Self-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve performance on downstream tasks such as automatic speech recognition (ASR). While these works suggest it is possible to reduce dependence on labeled data for building efficient speech systems, their evaluation was mostly made on ASR and using multiple and heterogeneous experimental settings (most of them for English). This questions the objective comparison of SSL approaches and the evaluation of their impact on building speech systems. In this paper, we propose LeBenchmark: a reproducible framework for assessing SSL from speech. It not only includes ASR (high and low resource) tasks but also spoken language understanding, speech translation and emotion recognition. We also focus on speech technologies in a language different than English: French. SSL models of different sizes are trained from carefully sourced and documented datasets. Experiments show that SSL is beneficial for most but not all tasks which confirms the need for exhaustive and reliable benchmarks to evaluate its real impact. LeBenchmark is shared with the scientific community for reproducible research in SSL from speech.
Solène Evain, Hang Le 0001, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia A. Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Estève, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Interspeech18
2021 Impact of Encoding and Segmentation Strategies on End-to-End Simultaneous Speech Translation
abstract
Boosted by the simultaneous translation shared task at IWSLT 2020, promising end-to-end online speech translation approaches were recently proposed.They consist in incrementally encoding a speech input (in a source language) and decoding the corresponding text (in a target language) with the best possible trade-off between latency and translation quality.This paper investigates two key aspects of end-to-end simultaneous speech translation: (a) how to encode efficiently the continuous speech flow, and (b) how to segment the speech flow in order to alternate optimally between reading (R: encoding input) and writing (W: decoding output) operations.We extend our previously proposed end-to-end online decoding strategy and show that while replacing BLSTM by ULSTM encoding degrades performance in offline mode, it actually improves both efficiency and performance in online mode.We also measure the impact of different methods to segment the speech signal (using fixed interval boundaries, oracle word boundaries or randomly set boundaries) and show that our best end-to-end online decoding strategy is surprisingly the one that alternates R/W operations on fixed size blocks on our English-German speech translation setup.
Yannick Estève, Laurent Besacier
Interspeech3
2021 Alternate Endings: Improving Prosody for Incremental Neural TTS with Predicted Future Text Input
abstract
The prosody of a spoken word is determined by its surrounding context. In incremental text-to-speech synthesis, where the synthesizer produces an output before it has access to the complete input, the full context is often unknown which can result in a loss of naturalness in the synthesized speech. In this paper, we investigate whether the use of predicted future text can attenuate this loss. We compare several test conditions of next future word: (a) unknown (zero-word), (b) language model predicted, (c) randomly predicted and (d) ground-truth. We measure the prosodic features (pitch, energy and duration) and find that predicted text provides significant improvements over a zero-word lookahead, but only slight gains over random-word lookahead. We confirm these results with a perceptive test.
Brooke Stephenson, Thomas Hueber, Laurent Girin, Laurent Besacier
Interspeech4
2021 Joint source-target encoding with pervasive attention
Maha Elbayad, Laurent Besacier, Jakob Verbeek
Mach. Transl.2
2020 Online Versus Offline NMT Quality: An In-depth Analysis on English-German and German-English
abstract
Maha Elbayad, Michael Ustaszewski, Emmanuelle Esperança-Rodier, Francis Brunet-Manquat, Jakob Verbeek, Laurent Besacier. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Maha Elbayad, Michael Ustaszewski, Emmanuelle Esperança-Rodier, Francis Brunet-Manquat, Jakob Verbeek, Laurent Besacier
COLING6
2020 Enabling Interactive Transcription in an Indigenous Community
abstract
We propose a novel transcription workflow which combines spoken term detection and humanin-the-loop, together with a pilot experiment.This work is grounded in an almost zero-resource scenario where only a few terms have so far been identified, involving two endangered languages.We show that in the early stages of transcription, when the available data is insufficient to train a robust ASR system, it is possible to take advantage of the transcription of a small number of isolated words in order to bootstrap the transcription of a speech collection.
Éric Le Ferrand, Steven Bird, Laurent Besacier
COLING3
2020 Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation
abstract
We introduce dual-decoder Transformer, a new model architecture that jointly performs automatic speech recognition (ASR) and multilingual speech translation (ST).Our models are based on the original Transformer architecture (Vaswani et al., 2017) but consist of two decoders, each responsible for one task (ASR or ST).Our major contribution lies in how these decoders interact with each other: one decoder can attend to different information sources from the other via a dual-attention mechanism.We propose two variants of these architectures corresponding to two different levels of dependencies between the decoders, called the parallel and cross dual-decoder Transformers, respectively.Extensive experiments on the MuST-C dataset show that our models outperform the previously-reported highest translation performance in the multilingual settings, and outperform as well bilingual one-to-one results.Furthermore, our parallel models demonstrate no trade-off between ASR and ST compared to the vanilla multi-task architecture.Our code and pre-trained models are available at https://
Hang Le 0001, Juan Pino 0001, Changhan Wang, Jiatao Gu, Didier Schwab, Laurent Besacier
COLING6
2020 Catplayinginthesnow: Impact of Prior Segmentation on a Model of Visually Grounded Speech
abstract
The language acquisition literature shows that children do not build their lexicon by segmenting the spoken input into phonemes and then building up words from them, but rather adopt a top-down approach and start by segmenting word-like units and then break them down into smaller units.This suggests that the ideal way of learning a language is by starting from full semantic units.In this paper, we investigate if this is also the case for a neural model of Visually Grounded Speech trained on a speech-image retrieval task.We evaluated how well such a network is able to learn a reliable speech-to-image mapping when provided with phone, syllable, or word boundary information.We present a simple way to introduce such information into an RNN-based model and investigate which type of boundary is the most efficient.We also explore at which level of the network's architecture such information should be introduced so as to maximise its performances.Finally, we show that using multiple boundary types at once in a hierarchical structure, by which low-level segments are used to recompose high-level segments, is beneficial and yields better results than using lowlevel or high-level segments in isolation.
William Havard, Laurent Besacier, Jean-Pierre Chevrot
CoNLL2
2020 Monolingual Adapters for Zero-Shot Neural Machine Translation
abstract
We propose a novel adapter layer formalism for adapting multilingual models.They are more parameter-efficient than existing adapter layers while obtaining as good or better performance.The layers are specific to one language (as opposed to bilingual adapters) allowing to compose them and generalize to unseen language-pairs.In this zero-shot setting, they obtain a median improvement of +2.77BLEU points over a strong 20-language multilingual Transformer baseline trained on TED talks.
Jerin Philip, Alexandre Berard, Matthias Gallé, Laurent Besacier
EMNLP (1)4
2020 A Data Efficient End-to-End Spoken Language Understanding Architecture
abstract
End-to-end architectures have been recently proposed for spoken language understanding (SLU) and semantic parsing. Based on a large amount of data, those models learn jointly acoustic and linguistic-sequential features. Such architectures give very good results in the context of domain, intent and slot detection, their application in a more complex semantic chunking and tagging task is less easy. For that, in many cases, models are combined with an external a language model to enhance their performance.In this paper we introduce a data efficient system which is trained end-to-end, with no additional, pre-trained external module. One key feature of our approach is an incremental training procedure where acoustic, language and semantic models are trained sequentially one after the other. The proposed model has a reasonable size and achieves competitive results with respect to state-of-the-art while using a small training dataset. In particular, we reach 24.02% Concept Error Rate (CER) on MEDIA/test while training on MEDIA/train without any additional data.
Marco Dinarelli, Nikita Kapoor, Bassam Jabaian, Laurent Besacier
ICASSP4
2020 The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units
abstract
International audience
Ewan Dunbar, Julien Karadayi, Mathieu Bernard, Xuan-Nga Cao, Robin Algayres, Lucas Ondel Yang, Laurent Besacier, Sakriani Sakti, Emmanuel Dupoux
INTERSPEECH7
2020 Efficient Wait-k Models for Simultaneous Machine Translation
abstract
Simultaneous machine translation consists in starting output generation before the entire input sequence is available. Wait-k decoders offer a simple but efficient approach for this problem. They first read k source tokens, after which they alternate between producing a target token and reading another source token. We investigate the behavior of wait-k decoding in low resource settings for spoken corpora using IWSLT datasets. We improve training of these models using unidirectional encoders, and training across multiple values of k. Experiments with Transformer and 2D-convolutional architectures show that our wait-k models generalize well across a wide range of latency levels. We also show that the 2D-convolution architecture is competitive with Transformers for simultaneous translation of spoken language.
Maha Elbayad, Laurent Besacier, Jakob Verbeek
INTERSPEECH2
2020 Investigating Self-Supervised Pre-Training for End-to-End Speech Translation
abstract
International audience
Fethi Bougares, Natalia A. Tomashenko, Yannick Estève, Laurent Besacier
INTERSPEECH5
2020 Modeling ASR Ambiguity for Neural Dialogue State Tracking
abstract
Spoken dialogue systems typically use a list of top-N ASR hypotheses for inferring the semantic meaning and tracking the state of the dialogue. However ASR graphs, such as confusion networks (confnets), provide a compact representation of a richer hypothesis space than a top-N ASR list. In this paper, we study the benefits of using confusion networks with a state-of-the-art neural dialogue state tracker (DST). We encode the 2-dimensional confnet into a 1-dimensional sequence of embeddings using an attentional confusion network encoder which can be used with any DST system. Our confnet encoder is plugged into the state-of-the-art 'Global-locally Self-Attentive Dialogue State Tacker' (GLAD) model for DST and obtains significant improvements in both accuracy and inference time compared to using top-N ASR hypotheses.
Vaishali Pal, Fabien Guillot, Manish Shrivastava 0001, Jean-Michel Renders, Laurent Besacier
INTERSPEECH5
2020 What the Future Brings: Investigating the Impact of Lookahead for Incremental Neural TTS
abstract
International audience
Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber
INTERSPEECH2
2020 MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible
abstract
The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for potentially 700 languages. However, the fact that the source content (the Bible) is the same for all the languages is not exploited to date. Therefore, this article proposes to add multilingual links between speech segments in different languages, and shares a large and clean dataset of 8,130 parallel spoken utterances across 8 languages (56 language pairs). We name this corpus MaSS (Multilingual corpus of Sentence-aligned Spoken utterances). The covered languages (Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish) allow researches on speech-to-speech alignment as well as on translation for typologically different language pairs. The quality of the final corpus is attested by human evaluation performed on a corpus subset (100 utterances, 8 language pairs). Lastly, we showcase the usefulness of the final product on a bilingual speech retrieval task.
Marcely Zanon Boito, William Havard, Mahault Garnerin, Éric Le Ferrand, Laurent Besacier
LREC5
2020 Gender Representation in Open Source Speech Resources
abstract
With the rise of artificial intelligence (AI) and the growing use of deep-learning architectures, the question of ethics, transparency and fairness of AI systems has become a central concern within the research community. We address transparency and fairness in spoken language systems by proposing a study about gender representation in speech resources available through the Open Speech and Language Resource platform. We show that finding gender information in open source corpora is not straightforward and that gender balance depends on other corpus characteristics (elicited/non elicited speech, low/high resource language, speech task targeted). The paper ends with recommendations about metadata and gender information for researchers in order to assure better transparency of the speech systems built using such corpora.
Mahault Garnerin, Solange Rossato, Laurent Besacier
LREC3
2020 FlauBERT: Unsupervised Language Model Pre-training for French
abstract
Language models have become a key step to achieve state-of-the art results in many different Natural Language Processing (NLP) tasks. Leveraging the huge amount of unlabeled texts nowadays available, they provide an efficient way to pre-train continuous word representations that can be fine-tuned for a downstream task, along with their contextualization at the sentence level. This has been widely demonstrated for English using contextualized representations (Dai and Le, 2015; Peters et al., 2018; Howard and Ruder, 2018; Radford et al., 2018; Devlin et al., 2019; Yang et al., 2019b). In this paper, we introduce and share FlauBERT, a model learned on a very large and heterogeneous French corpus. Models of different sizes are trained using the new CNRS (French National Centre for Scientific Research) Jean Zay supercomputer. We apply our French language models to diverse NLP tasks (text classification, paraphrasing, natural language inference, parsing, word sense disambiguation) and show that most of the time they outperform other pre-training approaches. Different versions of FlauBERT as well as a unified evaluation protocol for the downstream tasks, called FLUE (French Language Understanding Evaluation), are shared to the research community for further reproducible experiments in French NLP.
Hang Le 0001, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, Didier Schwab
LREC9
2020 Investigating alignment interpretability for low-resource NMT
Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier
Mach. Transl.3
2020 Speech Technology for Unwritten Languages
abstract
Speech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible.
Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001
IEEE ACM Trans. Audio Speech Lang. Process.12
2019 Word Recognition, Competition, and Activation in a Model of Visually Grounded Speech
abstract
In this paper, we study how word-like units are represented and activated in a recurrent neural model of visually grounded speech.The model used in our experiments is trained to project an image and its spoken description in a common representation space.We show that a recurrent model trained on spoken sentences implicitly segments its input into word-like units and reliably maps them to their correct visual referents.We introduce a methodology originating from linguistics to analyse the representation learned by neural networks -the gating paradigm -and show that the correct representation of a word is only activated if the network has access to first phoneme of the target word, suggesting that the network does not rely on a global acoustic pattern.Furthermore, we find out that not all speech frames (MFCC vectors in our case) play an equal role in the final encoded representation of a given word, but that some frames have a crucial effect on it.Finally, we suggest that word representation could be activated through a process of lexical competition.
William Havard, Jean-Pierre Chevrot, Laurent Besacier
CoNLL3
2019 Models of Visually Grounded Speech Signal Pay Attention to Nouns: A Bilingual Experiment on English and Japanese
abstract
We investigate the behaviour of attention in neural models of visually grounded speech trained on two languages: English and Japanese. Experimental results show that attention focuses on nouns and this behaviour holds true for two very typologically different languages. We also draw parallels between artificial neural attention and human attention and show that neural attention focuses on word endings as it has been theorised for human attention. Finally, we investigate how two visually grounded monolingual models can be used to perform cross-lingual speech-to-speech retrieval. For both languages, the enriched bilingual (speech-image) corpora with part-of-speech tags and forced alignments are distributed to the community for reproducible research.
William Havard, Jean-Pierre Chevrot, Laurent Besacier
ICASSP3
2019 Empirical Evaluation of Sequence-to-Sequence Models for Word Discovery in Low-Resource Settings
abstract
International audience
Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier
INTERSPEECH3
2019 The Zero Resource Speech Challenge 2019: TTS Without T
abstract
We present the Zero Resource Speech Challenge 2019, which proposes to build a speech synthesizer without any text or phonetic labels: hence, TTS without T (text-to-speech without text). We provide raw audio for a target voice in an unknown language (the Voice dataset), but no alignment, text or labels. Participants must discover subword units in an unsupervised way (using the Unit Discovery dataset) and align them to the voice recordings in a way that works best for the purpose of synthesizing novel utterances from novel speakers, similar to the target speaker's voice. We describe the metrics used for evaluation, a baseline system consisting of unsupervised subword unit discovery plus a standard TTS system, and a topline TTS using gold phoneme transcriptions. We present an overview of the 19 submitted systems from 10 teams and discuss the main results.
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel Yang, Alan W. Black, Laurent Besacier, Sakriani Sakti, Emmanuel Dupoux
INTERSPEECH11
2019 A neural approach for inducing multilingual resources and natural language processing tools for low-resource languages
abstract
Abstract This work focuses on the rapid development of linguistic annotation tools for low-resource languages (languages that have no labeled training data). We experiment with several cross-lingual annotation projection methods using recurrent neural networks (RNN) models. The distinctive feature of our approach is that our multilingual word representation requires only a parallel corpus between source and target languages. More precisely, our approach has the following characteristics: (a) it does not use word alignment information, (b) it does not assume any knowledge about target languages (one requirement is that the two languages (source and target) are not too syntactically divergent), which makes it applicable to a wide range of low-resource languages, (c) it provides authentic multilingual taggers (one tagger forNlanguages). We investigate both uni and bidirectional RNN models and propose a method to include external information (for instance, low-level information from part-of-speech tags) in the RNN to train higher level taggers (for instance, Super Sense taggers). We demonstrate the validity and genericity of our model by using parallel corpora (obtained by manual or automatic translation). Our experiments are conducted to induce cross-lingual part-of-speech and Super Sense taggers. We also use our approach in a weakly supervised context, and it shows an excellent potential for very low-resource settings (less than 1k training utterances).
Othman Zennaki, Nasredine Semmar, Laurent Besacier
Nat. Lang. Eng.3
2018 Token-level and sequence-level loss smoothing for RNN language models
abstract
Despite the effectiveness of recurrent neural network language models, their maximum likelihood estimation suffers from two limitations.It treats all sentences that do not match the ground truth as equally poor, ignoring the structure of the output space.Second, it suffers from "exposure bias": during training tokens are predicted given ground-truth sequences, while at test time prediction is conditioned on generated output sequences.To overcome these limitations we build upon the recent reward augmented maximum likelihood approach i.e. sequence-level smoothing that encourages the model to predict sentences close to the ground truth according to a given performance metric.We extend this approach to token-level loss smoothing, and propose improvements to the sequence-level smoothing approach.Our experiments on two different tasks, image captioning and machine translation, show that token-level and sequence-level loss smoothing are complementary, and significantly improve results.
Maha Elbayad, Laurent Besacier, Jakob Verbeek
ACL (1)2
2018 Unsupervised Learning of Word Segmentation: Does Tone Matter?
Pierre Godard, Kevin Löser, Alexandre Allauzen, Laurent Besacier, François Yvon
CICLing (1)4
2018 Pervasive Attention: 2D Convolutional Neural Networks for Sequence-to-Sequence Prediction
abstract
Current state-of-the-art machine translation systems are based on encoder-decoder architectures, that first encode the input sequence, and then generate an output sequence based on the input encoding.Both are interfaced with an attention mechanism that recombines a fixed encoding of the source tokens based on the decoder state.We propose an alternative approach which instead relies on a single 2D convolutional neural network across both sequences.Each layer of our network recodes source tokens on the basis of the output sequence produced so far.Attention-like properties are therefore pervasive throughout the network.Our model yields excellent results, outperforming state-of-the-art encoderdecoder systems, while being conceptually simpler and having fewer parameters.
Maha Elbayad, Laurent Besacier, Jakob Verbeek
CoNLL2
2018 End-to-End Automatic Speech Translation of Audiobooks
abstract
We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmented for this task. Previous works investigated the extreme case where source language transcription is not available during learning nor decoding, but we also study a midway case where source language transcription is available at training time only. In this case, a single model is trained to decode source speech into target text in a single pass. Experimental results show that it is possible to train compact and efficient end-to-end speech translation models in this setup. We also distribute the corpus and hope that our speech translation baseline on this corpus will be challenged in the future.
Alexandre Berard, Laurent Besacier, Ali Can Kocabiyikoglu, Olivier Pietquin
ICASSP2
2018 ASR Performance Prediction on Unseen Broadcast Programs Using Convolutional Neural Networks
abstract
In this paper, we address a relatively new task: prediction of ASR performance on unseen broadcast programs. We first propose an heterogenous French corpus dedicated to this task. Two prediction approaches are compared: a state-of-the-art performance prediction based on regression (engineered features) and a new strategy based on convolutional neural networks (learnt features). We particularly focus on the combination of both textual (ASR transcription) and signal inputs. While the joint use of textual and signal features did not work for the regression baseline, the combination of inputs for CNNs leads to the best WER prediction performance. We also show that our CNN prediction remarkably predicts the WER distribution on a collection of speech recordings.
Zied Elloumi, Laurent Besacier, Olivier Galibert, Juliette Kahn, Benjamin Lecouteux
ICASSP2
2018 Bayesian Models for Unit Discovery on a Very Low Resource Language
abstract
Developing speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the I-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus.
Lucas Ondel Yang, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Lukás Burget, François Yvon, Sanjeev Khudanpur
ICASSP3
2018 Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop
abstract
We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech.
Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux
ICASSP2
2018 Automatic Recognition of Affective Laughter in Spontaneous Dyadic Interactions from Audiovisual Signals
abstract
Laughter is a highly spontaneous behavior that frequently occurs during social interactions. It serves as an expressive-communicative social signal which conveys a large spectrum of affect display. Even though many studies have been performed on the automatic recognition of laughter -- or emotion -- from audiovisual signals, very little is known about the automatic recognition of emotion conveyed by laughter. In this contribution, we provide insights on emotional laughter by extensive evaluations carried out on a corpus of dyadic spontaneous interactions, annotated with dimensional labels of emotion (arousal and valence). We evaluate, by automatic recognition experiments and correlation based analysis, how different categories of laughter, such as unvoiced laughter, voiced laughter, speech laughter, and speech (non-laughter) can be differentiated from audiovisual features, and to which extent they might convey different emotions. Results show that voiced laughter performed best in the automatic recognition of arousal and valence for both audio and visual features. The context of production is further analysed and results show that, acted and spontaneous expressions of laughter produced by a same person can be differentiated from audiovisual signals, and multilingual induced expressions can be differentiated from those produced during interactions.
Reshmashree B. Kantharaju, Fabien Ringeval, Laurent Besacier
ICMI3
2018 Unsupervised Word Segmentation from Speech with Attention
abstract
International audience
Pierre Godard, Marcely Zanon Boito, Lucas Ondel Yang, Alexandre Berard, François Yvon, Aline Villavicencio, Laurent Besacier
INTERSPEECH7
2018 A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noël Kouarata, Lori Lamel, Hélène Bonneau-Maynard, Markus Müller 0001, Annie Rialland, Sebastian Stüker, François Yvon, Marcely Zanon Boito
LREC5
2018 Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
Ali Can Kocabiyikoglu, Laurent Besacier, Olivier Kraif
LREC2
2018 Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville)
Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt
LREC5
2018 Automatic quality estimation for speech translation using joint ASR and MT features
Ngoc-Tien Le, Benjamin Lecouteux, Laurent Besacier
Mach. Transl.3
2017 Unwritten languages demand attention too! Word discovery with encoder-decoder models
abstract
Word discovery is the task of extracting words from un-segmented text. In this paper we examine to what extent neural networks can be applied to this task in a realistic unwritten language scenario, where only small corpora and limited annotations are available. We investigate two scenarios: one with no supervision and another with limited supervision with access to the most frequent words. Obtained results show that it is possible to retrieve at least 27% of the gold standard vocabulary by training an encoder-decoder neural machine translation system with only 5,157 sentences. This result is close to those obtained with a task-specific Bayesian nonparametric model. Moreover, our approach has the advantage of generating translation alignments, which could be used to create a bilingual lexicon. As a future perspective, this approach is also well suited to work directly from speech.
Marcely Zanon Boito, Alexandre Berard, Aline Villavicencio, Laurent Besacier
ASRU4
2017 The zero resource speech challenge 2017
abstract
We describe a new challenge aimed at discovering subword and word units from raw speech. This challenge is the followup to the Zero Resource Speech Challenge 2015. It aims at constructing systems that generalize across languages and adapt to new speakers. The design features and evaluation metrics of the challenge are presented and the results of seventeen models are discussed.
Ewan Dunbar, Xuan-Nga Cao, Juan Benjumea, Julien Karadayi, Mathieu Bernard, Laurent Besacier, Xavier Anguera Miró, Emmanuel Dupoux
ASRU6
2017 Machine Assisted Analysis of Vowel Length Contrasts in Wolof
abstract
Growing digital archives and improving algorithms for automatic analysis of text and speech create new research opportunities for fundamental research in phonetics. Such empirical approaches allow statistical evaluation of a much larger set of hypothesis about phonetic variation and its conditioning factors (among them geographical / dialectal variants). This paper illustrates this vision and proposes to challenge automatic methods for the analysis of a not easily observable phenomenon: vowel length contrast. We focus on Wolof, an under-resourced language from Sub-Saharan Africa. In particular, we propose multiple features to make a fine evaluation of the degree of length contrast under different factors such as: read vs semi spontaneous speech ; standard vs dialectal Wolof. Our measures made fully automatically on more than 20k vowel tokens show that our proposed features can highlight different degrees of contrast for each vowel considered. We notably show that contrast is weaker in semi-spontaneous speech and in a non standard semi-spontaneous dialect.
Elodie Gauthier, Laurent Besacier, Sylvie Voisin
INTERSPEECH2
2017 Disentangling ASR and MT Errors in Speech Translation
Ngoc-Tien Le, Benjamin Lecouteux, Laurent Besacier
MTSummit (1)3
2017 Find the errors, get the better: Enhancing machine translation via word confidence estimation
abstract
Abstract This paper presents two novel ideas of improving the Machine Translation (MT) quality by applying the word-level quality prediction for the second pass of decoding. In this manner, the word scores estimated by word confidence estimation systems help to reconsider the MT hypotheses for selecting a better candidate rather than accepting the current sub-optimal one. In the first attempt, the selection scope is limited to the MTN-best list, in which our proposed re-ranking features are combined with those of the decoder for re-scoring. Then, the search space is enlarged over the entire search graph, storing many more hypotheses generated during the first pass of decoding. Over all paths containing words of theN-best list, we propose an algorithm to strengthen or weaken them depending on the estimated word quality. In both methods, the highest score candidate after the search becomes the official translation. The results obtained show that both approaches advance the MT quality over the one-pass baseline, and the search graph re-decoding achieves more gains (in BLEU score) thanN-best List Re-ranking method.
Ngoc-Quang Luong, Laurent Besacier, Benjamin Lecouteux
Nat. Lang. Eng.2
2016 Word2Vec vs DBnary: Augmenting METEOR using Vector Representations or Lexical Resources?
abstract
This paper presents an approach combining lexico-semantic resources and distributed representations of words applied to the evaluation in machine translation (MT). This study is made through the enrichment of a well-known MT evaluation metric: METEOR. METEOR enables an approximate match (synonymy or morphological similarity) between an automatic and a reference translation. Our experiments are made in the framework of the Metrics task of WMT 2014. We show that distributed representations are a good alternative to lexico-semanticresources for MT evaluation and they can even bring interesting additional information. The augmented versions of METEOR, using vector representations, are made available on our Github page.
Christophe Servan, Alexandre Berard, Zied Elloumi, Hervé Blanchon, Laurent Besacier
COLING5
2016 Inducing Multilingual Text Analysis Tools Using Bidirectional Recurrent Neural Networks
abstract
This work focuses on the development of linguistic analysis tools for resource-poor languages. We use a parallel corpus to produce a multilingual word representation based only on sentence level alignment. This representation is combined with the annotated source side (resource-rich language) of the parallel corpus to train text analysis tools for resource-poor languages. Our approach is based on Recurrent Neural Networks (RNN) and has the following advantages: (a) it does not use word alignment information, (b) it does not assume any knowledge about foreign languages, which makes it applicable to a wide range of resource-poor languages, (c) it provides truly multilingual taggers. In a previous study, we proposed a method based on Simple RNN to automatically induce a Part-Of-Speech (POS) tagger. In this paper, we propose an improvement of our neural model. We investigate the Bidirectional RNN and the inclusion of external information (for instance low level information from Part-Of-Speech tags) in the RNN to train a more complex tagger (for instance, a multilingual super sense tagger). We demonstrate the validity and genericity of our method by using parallel corpora (obtained by manual or automatic translation). Our experiments are conducted to induce cross-lingual POS and super sense taggers.
Othman Zennaki, Nasredine Semmar, Laurent Besacier
COLING3
2016 First Automatic Fongbe Continuous Speech Recognition System: Development of Acoustic Models and Language Models
abstract
This paper reports our efforts toward an ASR system for a new under-resourced language (Fongbe).The aim of this work is to build acoustic models and language models for continuous speech decoding in Fongbe.The problem encountered with Fongbe (an African language spoken especially in Benin, Togo, and Nigeria) is that it does not have any language resources for an ASR system.As part of this work, we have first collected Fongbe text and speech corpora that are described in the following sections.Acoustic modeling has been worked out at a graphemic level and language modeling has provided two language models for performance comparison purposes.We also performed a vowel simplification by removing tones diacritics in order to investigate their impact on the language models.
Fréjus A. A. Laleye, Laurent Besacier, Eugène C. Ezin, Cina Motamed
FedCSIS2
2016 OCR-aided person annotation and label propagation for speaker modeling in TV shows
abstract
In this paper, we present an approach for minimizing human effort in manual speaker annotation. Label propagation is used at each iteration of an active learning cycle. More precisely, a selection strategy for choosing the most suitable speech track to be labeled is proposed. Four different selection strategies are evaluated and all the tracks in a corresponding cluster are gathered using agglomerative clustering in order to propagate human annotations. To further reduce the manual labor required, an optical character recognition system is used to bootstrap annotations. At each step of the cycle, annotations are used to build speaker models. The quality of the generated speaker models is evaluated at each step using an i-vector based speaker identification system. The presented approach shows promising results on the REPERE corpus with a minimum amount of human effort for annotation.
Mateusz Budnik, Laurent Besacier, Ali Khodabakhsh 0001, Cenk Demiroglu
ICASSP2
2016 Lig-Aikuma: A Mobile App to Collect Parallel Speech for Under-Resourced Language Studies
Elodie Gauthier, David Blachon, Laurent Besacier, Guy-Noël Kouarata, Martine Adda-Decker, Annie Rialland, Gilles Adda, Grégoire Bachman
INTERSPEECH3
2016 Speed Perturbation and Vowel Duration Modeling for ASR in Hausa and Wolof Languages
abstract
Automatic Speech Recognition (ASR) for (under-resourced) Sub-Saharan African languages faces several challenges: small amount of transcribed speech, written language normalization issues, few text resources available for language modeling, as well as specific features (tones, morphology, etc.) that need to be taken into account seriously to optimize ASR performance.This paper tries to address some of the above challenges through the development of ASR systems for two Sub-Saharan African languages: Hausa and Wolof.First, we investigate data augmentation technique (through speed perturbation) to overcome the lack of resources.Secondly, the main contribution is our attempt to model vowel length contrast existing in both languages.For reproducible experiments, the ASR systems developed for Hausa and Wolof are made available to the research community on github.To our knowledge, the Wolof ASR system presented in this paper is the first large vocabulary continuous speech recognition system ever developed for this language.
Elodie Gauthier, Laurent Besacier, Sylvie Voisin
INTERSPEECH2
2016 Preliminary Experiments on Unsupervised Word Discovery in Mboshi
abstract
International audience
Pierre Godard, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Laurent Besacier, Hélène Bonneau-Maynard, Guy-Noël Kouarata, Kevin Löser, Annie Rialland, François Yvon
INTERSPEECH5
2016 Better Evaluation of ASR in Speech Translation Context Using Word Embeddings
abstract
International audience
Ngoc-Tien Le, Christophe Servan, Benjamin Lecouteux, Laurent Besacier
INTERSPEECH4
2016 MultiVec: a Multilingual and Multilevel Representation Learning Toolkit for NLP
Alexandre Berard, Christophe Servan, Olivier Pietquin, Laurent Besacier
LREC4
2016 A Multilingual, Multi-style and Multi-granularity Dataset for Cross-language Textual Similarity Detection
Jérémy Ferrero, Frédéric Agnès, Laurent Besacier, Didier Schwab
LREC3
2016 Collecting Resources in Sub-Saharan African Languages for Automatic Speech Recognition: a Case Study of Wolof
Elodie Gauthier, Laurent Besacier, Sylvie Voisin, Michael Melese Woldeyohannis, Uriel Pascal Elingui
LREC2
2016 The CAMOMILE Collaborative Annotation Platform for Multi-modal, Multi-lingual and Multi-media Documents
Johann Poignant, Mateusz Budnik, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Gilles Adda, Laurent Besacier, Hazim Kemal Ekenel, Gil Francopoulo, Javier Hernando, Joseph Mariani, Ramon Morros, Georges Quénot, Sophie Rosset, Thomas Tamisier
LREC8
2016 A unified framework for translation and understanding allowing discriminative joint decoding for multilingual speech semantic interpretation
Bassam Jabaian, Fabrice Lefèvre, Laurent Besacier
Comput. Speech Lang.3
2016 Naming multi-modal clusters to identify persons in TV broadcast
Johann Poignant, Guillaume Fortier, Laurent Besacier, Georges Quénot
Multim. Tools Appl.3
2015 Spoken language translation graphs re-decoding using automatic quality assessment
abstract
This paper investigates how automatic quality assessment of spoken language translation (SLT), also named confidence estimation (CE), can help re-decoding SLT output graphs and improve the overall speech translation performance. Our graph redecoding method can be seen as a second-pass of translation. For this, a robust word confidence estimator for SLT is required. We propose several estimators based on our estimation of transcription (ASR) quality, translation (MT) quality, or both (combined ASR+MT). Using these word confidence measures to re-decode the spoken language translation graph leads to a significant BLEU improvement (more than 2 points) compared to our SLT baseline, for a French-English SLT task. These results could be applied to interactive speech translation or computer-assisted translation of speeches and lectures.
Laurent Besacier, Benjamin Lecouteux, Ngoc-Quang Luong, Ngoc-Tien Le
ASRU1
2015 Speech technologies for african languages: example of a multilingual calculator for education
Laurent Besacier, Elodie Gauthier, Mathieu Mangeot, Philippe Bretier, Paul C. Bagshaw, Olivier Rosec, Thierry Moudenc, François Pellegrino, Sylvie Voisin, Egidio Marsico, Pascal Nocera
INTERSPEECH1
2015 Collaborative annotation for person identification in TV shows
Mateusz Budnik, Laurent Besacier, Johann Poignant, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Thomas Tamisier
INTERSPEECH2
2015 Using resources from a closely-related language to develop ASR for a very under-resourced language: a case study for iban
abstract
This paper presents our strategies for developing an automatic speech recognition system for Iban, an under-resourced language. We faced several challenges such as no pronunciation dictionary and lack of training material for building acoustic models. To overcome these problems, we proposed approaches which exploit resources from a closely-related language (Malay). We developed a semi-supervised method for building the pronunciation dictionary and applied cross-lingual strategies for improving acoustic models trained with very limited training data. Both approaches displayed very encouraging results, which show that data from a closely-related language, if available, can be exploited to build ASR for a new language. In the final part of the paper, we present a zero-shot ASR using Malay resources that can be used as an alternative method for transcribing Iban speech.
Sarah Flora Samson Juan, Laurent Besacier, Benjamin Lecouteux, Mohamed Dyab
INTERSPEECH2
2015 METEOR for multiple target languages using DBnary
Zied Elloumi, Hervé Blanchon, Gilles Sérasset, Laurent Besacier
MTSummit4
2015 Unsupervised and Lightly Supervised Part-of-Speech Tagging Using Recurrent Neural Networks
Othman Zennaki, Nasredine Semmar, Laurent Besacier
PACLIC3
2015 Towards accurate predictors of word quality for Machine Translation: Lessons learned on French-English and English-Spanish systems
Ngoc-Quang Luong, Laurent Besacier, Benjamin Lecouteux
Data Knowl. Eng.2
2015 Unsupervised Speaker Identification in TV Broadcast Based on Written Names
abstract
Identifying speakers in TV broadcast in an unsupervised way (i.e., without biometric models) is a solution for avoiding costly annotations. Existing methods usually use pronounced names, as a source of names, for identifying speech clusters provided by a diarization step but this source is too imprecise for having sufficient confidence. To overcome this issue, another source of names can be used: the names written in a title block in the image track. We first compared these two sources of names on their abilities to provide the name of the speakers in TV broadcast. This study shows that it is more interesting to use written names for their high precision for identifying the current speaker. We also propose two approaches for finding speaker identity based only on names written in the image track. With the “late naming” approach, we propose different propagations of written names onto clusters. Our second proposition, “Early naming,” modifies the speaker diarization module (agglomerative clustering) by adding constraints preventing two clusters with different associated written names to be merged together. These methods were tested on the REPERE corpus phase 1, containing 3 hours of annotated videos. Our best “late naming” system reaches an F-measure of 73.1%. “early naming” improves over this result both in terms of identification error rate and of stability of the clustering stopping criterion. By comparison, a mono-modal, supervised speaker identification system with 535 speaker models trained on matching development data and additional TV and radio data only provided a 57.2% F-measure.
Johann Poignant, Laurent Besacier, Georges Quénot
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 An efficient two-pass decoder for SMT using word confidence estimation
Ngoc-Quang Luong, Laurent Besacier, Benjamin Lecouteux
EAMT2
2014 Introduction to the special issue on processing under-resourced languages
Laurent Besacier, Etienne Barnard, Alexey Karpov 0001, Tanja Schultz
Speech Commun.1
2014 Automatic speech recognition for under-resourced languages: A survey
Laurent Besacier, Etienne Barnard, Alexey Karpov 0001, Tanja Schultz
Speech Commun.1
2014 SMT-based ASR domain adaptation methods for under-resourced languages: Application to Romanian
Horia Cucu, Andi Buzo, Laurent Besacier, Corneliu Burileanu
Speech Commun.3
2014 Using different acoustic, lexical and language modeling units for ASR of an under-resourced language - Amharic
Martha Yifiru Tachbelie, Solomon Teferra Abate, Laurent Besacier
Speech Commun.3
2013 Unsupervised naming of speakers in broadcast TV: using written names, pronounced names or both?
abstract
International audience
Johann Poignant, Laurent Besacier, Viet Bac Le, Sophie Rosset, Georges Quénot
INTERSPEECH2
2013 Comparison and Combination of Lightly Supervised Approaches for Language Portability of a Spoken Language Understanding System
abstract
Portability of a spoken dialogue system (SDS) to a new domain or a new language is a hot topic as it may imply gains in time and cost for building new SDSs. In particular in this paper we investigate several fast and efficient approaches for language portability of the spoken language understanding (SLU) module of a dialogue system. We show that the use of statistical machine translation (SMT) can reduce the time and the cost of porting a system from a source to a target language. For conceptual decoding, a state-of-the-art module based on conditional random fields (CRF) is used and a new approach based on phrase-based statistical machine translation (PB-SMT) is also evaluated. The experimental results show the efficiency of the proposed methods for a fast and low cost SLU language portability. In addition, we propose two methods to increase SLU robustness to translation errors. Overall, it is shown that the combination of all these approaches can further reduce the concept error rate. While most of the experiments in this paper deal with portability from French to Italian (given the availability of the Media French corpus and its subset manually translated into Italian), a validation of our methodology is eventually proposed in Arabic.
Bassam Jabaian, Laurent Besacier, Fabrice Lefèvre
IEEE Trans. Speech Audio Process.2
2012 From Text Detection in Videos to Person Identification
abstract
We present in this article a video OCR system that detects and recognizes overlaid texts in video as well as its application to person identification in video documents. We proceed in several steps. First, text detection and temporal tracking are performed. After adaptation of images to a standard OCR system, a final post-processing combines multiple transcriptions of the same text box. The semi-supervised adaptation of this system to a particular video type (video broadcast from a French TV) is proposed and evaluated. The system is efficient as it runs 3 times faster than real time (including the OCR step) on a desktop Linux box. Both text detection and recognition are evaluated individually and through a person recognition task where it is shown that the combination of OCR and audio (speaker) information can greatly improve the performances of a state of the art audio based person identification system.
Johann Poignant, Laurent Besacier, Georges Quénot, Franck Thollard
ICME2
2012 Portability of Semantic Annotations for Fast Development of Dialogue Corpora
Bassam Jabaian, Fabrice Lefèvre, Laurent Besacier
INTERSPEECH3
2012 Unsupervised Speaker Identification using Overlaid Texts in TV Broadcast
abstract
Poster Session: Speaker Recognition III
Johann Poignant, Hervé Bredin, Viet Bac Le, Laurent Besacier, Claude Barras, Georges Quénot
INTERSPEECH4
2012 Leveraging study of robustness and portability of spoken language understanding systems across languages and domains: the PORTMEDIA corpora
Fabrice Lefèvre, Djamel Mostefa, Laurent Besacier, Yannick Estève, Matthieu Quignard, Nathalie Camelin, Benoît Favre, Bassam Jabaian, Lina Maria Rojas-Barahona
LREC3
2012 Collection of a Large Database of French-English SMT Output Corrections
Marion Potet, Emmanuelle Esperança-Rodier, Laurent Besacier, Hervé Blanchon
LREC3
2011 Investigating the role of machine translated text in ASR domain adaptation: Unsupervised and semi-supervised methods
abstract
This study investigates the use of machine translated text for ASR domain adaptation. The proposed methodology is applicable when domain-specific data is available in language X only, whereas the goal is to develop a domain-specific system in language Y. Two semi-supervised methods are introduced and compared with a fully unsupervised approach, which represents the baseline. While both unsupervised and semi-supervised approaches allow to quickly develop an accurate domain-specific ASR system, the semi-supervised approaches overpass the unsupervised one by 10% to 29% relative, depending on the amount of human post-processed data available. An in-depth analysis, to explain how the machine translated text improves the performance of the domain-specific ASR, is also given at the end of this paper.
Horia Cucu, Laurent Besacier, Corneliu Burileanu, Andi Buzo
ASRU2
2011 Oracle-based Training for Phrase-based Statistical Machine Translation
Marion Potet, Emmanuelle Esperança-Rodier, Hervé Blanchon, Laurent Besacier
EAMT4
2011 Combination of stochastic understanding and machine translation systems for language portability of dialogue systems
abstract
In this paper, several approaches for language portability of dialogue systems are investigated with a focus on the spoken language understanding (SLU) component. We show that the use of statistical machine translation (SMT) can greatly reduce the time and cost of porting an existing system from a source to a target language. Using automatically translated training data we study phrase-based machine translation as an alternative to conditional random fields for conceptual decoding to compensate for the loss of a precise concept-word alignment. Also two ways to increase SLU robustness to translation errors (smeared training data and translation post editing) are shown to improve performance when test data are translated then decoded in the source language. Overall the combination of all these approaches allows to reduce even further the concept error rate. Experiments were carried out on the French MEDIA dialogue corpus with a subset manually translated into Italian.
Bassam Jabaian, Laurent Besacier, Fabrice Lefèvre
ICASSP2
2011 Quality Assessment of Crowdsourcing Transcriptions for African Languages
abstract
International audience
Hadrien Gelas, Solomon Teferra Abate, Laurent Besacier, François Pellegrino
INTERSPEECH3
2011 Speech Modulation Features for Robust Nonnative Speech Accent Detection
Sam Sethserey, Laurent Besacier, Eric Castelli, Haizhou Li 0001, Chng Eng Siong
INTERSPEECH3
2010 A fully unsupervised approach for mining parallel data from comparable corpora
Thi-Ngoc-Diep Do, Laurent Besacier, Eric Castelli
EAMT2
2010 Investigating multiple approaches for SLU portability to a new language
abstract
International audience
Bassam Jabaian, Laurent Besacier, Fabrice Lefèvre
INTERSPEECH2
2010 Unsupervised acoustic model adaptation for multi-origin non native ASR
Sam Sethserey, Eric Castelli, Laurent Besacier
INTERSPEECH3
2010 Automatic Identification of Arabic Dialects
Mohamed Belgacem, Georges Antoniadis, Laurent Besacier
LREC3
2010 Content-based search in multilingual audiovisual documents using the International Phonetic Alphabet
Georges Quénot, Tien Ping Tan, Viet Bac Le, Stéphane Ayache, Laurent Besacier, Philippe Mulhem
Multim. Tools Appl.5
2009 Multiple text segmentation for statistical language modeling
abstract
International audience
Sopheap Seng, Laurent Besacier, Brigitte Bigi, Eric Castelli
INTERSPEECH2
2009 Human translations guided language discovery for ASR systems
abstract
International audience
Sebastian Stüker, Laurent Besacier, Alex Waibel
INTERSPEECH2
2009 Automatic Speech Recognition for Under-Resourced Languages: Application to Vietnamese Language
abstract
This paper presents our work in automatic speech recognition (ASR) in the context of under-resourced languages with application to Vietnamese. Different techniques for bootstrapping acoustic models are presented. First, we present the use of acoustic-phonetic unit distances and the potential of crosslingual acoustic modeling for under-resourced languages. Experimental results on Vietnamese showed that with only a few hours of target language speech data, crosslingual context independent modeling worked better than crosslingual context dependent modeling. However, it was outperformed by the latter one, when more speech data were available. We concluded, therefore, that in both cases, crosslingual systems are better than monolingual baseline systems. The proposal of grapheme-based acoustic modeling, which avoids building a phonetic dictionary, is also investigated in our work. Finally, since the use of sub-word units (morphemes, syllables, characters, etc.) can reduce the high out-of-vocabulary rate and improve the lack of text resources in statistical language modeling for under-resourced languages, we propose several methods to decompose, normalize and combine word and sub-word lattices generated from different ASR systems. The proposed lattice combination scheme results in a relative syllable error rate reduction of 6.6% over the sentence MAP baseline method for a Vietnamese ASR task.
Viet Bac Le, Laurent Besacier
IEEE Trans. Speech Audio Process.2
2008 Word/sub-word lattices decomposition and combination for speech recognition
abstract
This paper presents the benefit of using multiple lexical units in the post-processing stage of an ASR system. Since the use of sub-word units can reduce the high out-of-vocabulary rate and improve the lack of text resources in statistical language modeling, we propose several methods to decompose, normalize and combine word and sub-word lattices generated from different ASR systems. By using a sub-word information table, every word in a lattice can be decomposed into sub-word units. These decomposed lattices can be combined into a common lattice in order to generate a confusion network. This lattices combination scheme results in an absolute syllable error rate reduction of about 1.4% over the sentence MAP baseline method for a Vietnamese ASR task. By comparing with the N-best lists combination and voting method, the proposed method works better.
Viet Bac Le, Sopheap Seng, Laurent Besacier, Brigitte Bigi
ICASSP3
2008 Feature adaptation of hearing-impaired lip shapes: the vowel case in the cued speech context
abstract
4
Noureddine Aboutabit, Denis Beautemps, Olivier Mathieu, Laurent Besacier
INTERSPEECH4
2008 Improving pronunciation modeling for non-native speech recognition
abstract
In this paper, three different approaches to pronunciation modeling are investigated. Two existing pronunciation modeling approaches, namely the pronunciation dictionary and n-best rescoring approach are modified to work with little amount of non-native speech. We also propose a speaker clustering approach, which capable of grouping the speakers based on their pronunciation habits. Given some speech, the approach can also be used for pronunciation adaptation. This approach is called latent pronunciation analysis. The results show that conventional pronunciation dictionary perform slightly better than n-best list rescoring, while the latent pronunciation analysis has shown to be beneficial for speaker clustering, and it can produce nearly the same improvement as the pronunciation dictionary approach, without the need to know the origin of the speaker. Index Terms: non-native ASR, decision trees, n-best list rescoring, latent phonemic analysis 1.
Tien Ping Tan, Laurent Besacier
INTERSPEECH2
2008 First Broadcast News Transcription System for Khmer Language
Sopheap Seng, Sam Sethserey, Laurent Besacier, Brigitte Bigi, Eric Castelli
LREC3
2007 Acoustic Model Interpolation for Non-Native Speech Recognition
abstract
This paper proposes three interpolation techniques which use the target language and the speaker's native language to improve non-native speech recognition system. These interpolation techniques are manual interpolation, weighted least square and eigenvoices. Each of them can be used under different situation and constraints. In contrast to weighted least square and eigenvoices methods, manual interpolation can be achieved offline without any adaptation data. These methods can also be combined with MLLR to improve the recognition rate. Experiments presented in this paper show that the best non native adaptation method, combined with MLLR can give 10% WER absolute reduction on a French automatic speech recognition system for both Chinese and Vietnamese native speakers.
Tien Ping Tan, Laurent Besacier
ICASSP (4)2
2007 On Efficient Coupling of ASR and SMT for Speech Translation
abstract
This paper presents an efficient tightly integrated approach for improved speech translation performance. The proposed approach combines the automatic speech recognition (ASR) and statistical machine translation (SMT) components in a bi-directional fashion. First, our SMT decoder takes the speech recognition lattice to perform an integrated search for the optimal translation by combining various ASR scores and translation models. Our approach is implemented within the recently proposed Folsom SMT framework that employs a multilayer search algorithm to conduct efficient operations on multiple graphs, which not only achieves memory efficiency and fast speed that is critical for real time speech translation applications, but also provides significant accuracy improvements. Secondly, we also report our experiments where the ASR is customized by reinforcing the language model to favor downstream translation component. We evaluated our approach on a large vocabulary speech translation task, and we obtain more than 2 point BLEU improvement over standard cascaded 1-best speech translation.
Bowen Zhou 0006, Laurent Besacier
ICASSP (4)2
2007 A HMM recognition of consonant-vowel syllables from lip contours: the cued speech case
abstract
International audience
Noureddine Aboutabit, Denis Beautemps, Jeanne Clarke, Laurent Besacier
INTERSPEECH4
2007 Automatic question detection: prosodic-lexical features and crosslingual experiments
abstract
In this paper, we present our work on automatic question detection from the speech signal. We are interested in developing automatic detection system and investigate the portability of such system to a new language. The first goal of this paper is to propose and evaluate a combined approach for automatic question detection where prosodic features are augmented by the use of lexical features. It is shown that both early and late integration of theses features in a decision treebased classifier improves the question detection performance compared to a baseline system using prosodic features only. The second goal of this paper is to conduct a crosslingual (French / Vietnamese) evaluation concerning the use of prosodic features. It is shown that our first system developed for French which uses an initial prosodic feature set can be improved using a new feature set that takes into account some specific prosodic characteristics of the Vietnamese tonal language. Both Vietnamese and French question detection systems obtain Fratio performance around 80% on pre-segmented meeting and dialog utterances.
Minh-Quang Vu, Laurent Besacier, Eric Castelli
INTERSPEECH2
2007 Modeling context and language variation for non-native speech recognition
abstract
Non-native speakers often face difficulty in pronouncing like the native speakers. This paper proposes to model pronunciation variation in non-native speaker’s speech using only acoustics models, without the need for the corpus. Variation in term of context and language will be modeled. The combination of both modeling resulted in the reduction of absolute WER as much as 16 % and 6 % for native Vietnamese and Chinese speakers of French. Index Terms: non-native ASR, context modeling,
Tien Ping Tan, Laurent Besacier
INTERSPEECH2
2006 Hand and Lip Desynchronization Analysis in French Cued Speech: Automatic Temporal Segmentation of Hand Flow
abstract
In the context of cued speech gesture phonetic translation, the automatic recognition of lip and hand movements is a key factor. The hand and the lip parameters are not synchronized, thus the fusion of the two channels (hand and lips) needs the knowledge of the desynchronized delay. This contribution focuses on the presentation of an automatic algorithm for temporal segmentation of the hand cue information based on Gaussian modeling of the hand position and minimum of velocity. The segmentation delivers the beginning of the hand transition and the instant of attained position. The hand segmentation is used to calculate the delay between hand and lip targets, in relation with the corresponding acoustic realization in the case of French CV syllables extracted from a corpus of phrases uttered and coded by a cued speech speaker. This study confirms in a more complex context the importance of the instant of attained hand position as pointed out by Attina and colleagues, in terms of control and for the fusion process
Noureddine Aboutabit, Denis Beautemps, Laurent Besacier
ICASSP (1)3
2006 ASR and Translation for Under-Resourced Languages
abstract
There are more than 6000 languages in the world but only a small number possess the resources required for implementation of human language technologies (HLT). Thus, HLT are mostly concerned by languages for which large resources are available or which have suddenly become of interest because of the economic or political scene. On the contrary, languages from developing countries or minorities have been less worked on in the past years. One way of improving this "language divide" is do more research on portability of HLT for multilingual applications. In this paper, we concentrate on speech-to-speech translation. We present here our methodology for fast development of ASR systems for under-resourced languages or, as they are called now, pi-languages (poorly equipped). We present the resources collected for Vietnamese, and the experimental results of our first Vietnamese ASR system. The current validation of our methodology for Khmer is described next. We also discuss some issues related to machine translation and present first contributions of our laboratory in this context of "pi-languages"
Laurent Besacier, Viet Bac Le, Christian Boitet, Vincent Berment
ICASSP (5)1
2006 Acoustic-Phonetic Unit Similarities For Context Dependent Acoustic Model Portability
abstract
This paper addresses particularly the use of acoustic-phonetic unit similarities for portability of context dependent acoustic models to new languages. Since the IPA-based method is limited to a source/target phoneme mapping table construction, an estimation method of the similarity between two phonemes is proposed in this paper. Based on these phoneme similarities, some estimation methods for polyphone similarity and clustered polyphonic model similarity are investigated. For a new language, first a polyphonic decision tree is built with a small amount of speech data. Then, clustered models in the target language are duplicated from the nearest clustered models in the source language and adapted with limited data to the target language. Results obtained from the experiments demonstrate the feasibility of these methods.
Viet Bac Le, Laurent Besacier, Tanja Schultz
ICASSP (1)2
2006 Characterization of cued speech vowels from the inner lip contour
abstract
4 pages
Noureddine Aboutabit, Denis Beautemps, Laurent Besacier
INTERSPEECH3
2006 On the use of morphological analysis for dialectal Arabic speech recognition
abstract
Arabic has a large number of affixes that can modify a stem to form words. In automatic speech recognition (ASR) this leads to a high out-of-vocabulary (OOV) rate for typical lexicon size, and hence a potential increase in WER. This is even more pronounced for dialects of Arabic where additional affixes are often introduced and the available data is typically sparse. To address this problem we introduce a simple word decomposition algorithm which only requires a text corpus and a predefined list of affixes. Using this al-gorithm to create the lexicon for Iraqi Arabic ASR results in about 10 % relative improvement in word error rate (WER). Also using the union of the segmented and unsegmented vocabularies and in-terpolating the corresponding language models results in further WER reduction. The net WER improvement is about 13%. 1.
Mohamed Afify, Ruhi Sarikaya, Hong-Kwang Jeff Kuo, Laurent Besacier
INTERSPEECH4
2006 Comparison of acoustic modeling techniques for Vietnamese and Khmer ASR
abstract
This paper presents a comparison of some different acoustic modeling strategies for under-resourced languages. When only limited speech data are available for under-resourced languages, we propose some crosslingual acoustic modeling techniques. We apply and compare these techniques in Vietnamese ASR. Since there is no pronunciation dictionary for some under-resourced languages, we investigate grapheme-based acoustic modeling. Some initialization techniques for context independent modeling and some question generation techniques for context dependent modeling are applied and compared for Khmer ASR. Index Terms: ASR, acoustic modeling, Vietnamese, Khmer.
Viet Bac Le, Laurent Besacier
INTERSPEECH2
2006 A French Non-Native Corpus for Automatic Speech Recognition
Tien Ping Tan, Laurent Besacier
LREC2
2006 Towards speech Translation of Non Written Languages
abstract
A large amount of languages in the world do not have an acknowledged written form. However, for a task like speech to speech translation, the written form of a language may be considered as secondary and it might be possible, under certain conditions, to bypass it. This paper is our first attempt to show that such an approach is possible. We propose a phone-based speech translation approach where translation models are learned on a parallel corpus made of foreign phone sequences and their corresponding English translation. Our experiments show that using our so-called phone-based approach leads almost to the same performance as the baseline approach, while being theoretically applicable to any non written language.
Laurent Besacier, Bowen Zhou 0006
SLT1
2006 Step-by-step and integrated approaches in broadcast news speaker diarization
Sylvain Meignier, Daniel Moraru, Corinne Fredouille, Jean-François Bonastre, Laurent Besacier
Comput. Speech Lang.5
2006 Information Extraction From Sound for Medical Telemonitoring
abstract
Today, the growth of the aging population in Europe needs an increasing number of health care professionals and facilities for aged persons. Medical telemonitoring at home (and, more generally, telemedicine) improves the patient's comfort and reduces hospitalization costs. Using sound surveillance as an alternative solution to video telemonitoring, this paper deals with the detection and classification of alarming sounds in a noisy environment. The proposed sound analysis system can detect distress or everyday sounds everywhere in the monitored apartment, and is connected to classical medical telemonitoring sensors through a data fusion process. The sound analysis system is divided in two stages: sound detection and classification. The first analysis stage (sound detection) must extract significant sounds from a continuous signal flow. A new detection algorithm based on discrete wavelet transform is proposed in this paper, which leads to accurate results when applied to nonstationary signals (such as impulsive sounds). The algorithm presented in this paper was evaluated in a noisy environment and is favorably compared to the state of the art algorithms in the field. The second stage of the system is sound classification, which uses a statistical approach to identify unknown sounds. A statistical study was done to find out the most discriminant acoustical parameters in the input of the classification module. New wavelet based parameters, better adapted to noise, are proposed in this paper. The telemonitoring system validation is presented through various real and simulated test sets. The global sound based system leads to a 3% missed alarm rate and could be fused with other medical sensors to improve performance.
Dan Istrate, Eric Castelli, Michel Vacher, Laurent Besacier, Jean-François Serignat
IEEE Trans. Inf. Technol. Biomed.4
2005 First Steps in Fast Acoustic Modeling for a New Target Language: Application to Vietnamese
abstract
This paper presents our first steps in fast acoustic modeling for a new target language. Both knowledge-based and data-driven methods were used to obtain phone mapping tables between a source language (French) and a target language (Vietnamese). While acoustic models borrowed directly from the source language did not perform very well, we have shown that using a small amount of adaptation data in the target language (one or two hours) lead to very acceptable automatic speech recognition (ASR) performance. Our best continuous Vietnamese recognition system, adapted with only two hours of Vietnamese data, obtains a word accuracy of 63.9% on one hour of Vietnamese speech dialog for instance.
Viet Bac Le, Laurent Besacier
ICASSP (1)2
2005 Audio, video and audio-visual signatures for short video clip detection: experiments on Trecvid2003
abstract
In this paper, we present the association of audio and video signatures for short video clip detection. First, we present an audio signature based on the spectral flatness measure. Then we describe a spatio-temporal video signature, based on the evolution of gray level centroids over time. The major contribution of this work is the association of these two signatures in a so-called audiovisual signature by late integration of similarity measures obtained on both modalities. Our experiments conducted on a large video database (28 Gb/34 h extracted from TRECVID2003) show that our audio-visual signature is more robust than the audio-only or video-only signatures, and also permits better detection of video clips of shorter duration (about 2 seconds).
Benjamin Senechal, Denis Pellerin, Laurent Besacier, Isabelle Simand, Stéphane Bres
ICME3
2005 A speaker independent "liveness" test for audio-visual biometrics
abstract
In biometrics, it is crucial to detect impostors and thwart replay attacks. However, few researches have focused yet on the “liveness” verifi cation. This test ensures that biometric cues being acquired are actual measurements from a live person who is present at the time of capture. Here, we propose a speaker independent “liveness” verifi cation method for audiovideo identifi cation systems. It uses the correlation that exists between the lip movements and the speech produced. Two data analysis methods are considered to model this statistical link. Finally, according to tests carried out on the XM2VTS database, the best liveness verifi cation EER achieved is 14.5% .
Nicolas Eveno, Laurent Besacier
INTERSPEECH2
2004 Benefits of prior acoustic segmentation for automatic speaker segmentation
abstract
The paper investigates the interest of segmentation in acoustic macro classes (like gender or bandwidth) as front-end processing for the segmentation/diarization task. The impact of this prior acoustic segmentation is evaluated in terms of speaker diarization performance in the particular context of NIST RT'03 evaluation (done on the HUB4 broadcast news corpora). It is rarely discussed in the literature, but our work shows that the application of prior acoustic segmentation, in a similar way to the automatic speech recognition task, may be very useful to the speaker segmentation task. Experiments were conducted using two different kinds of speaker segmentation systems developed individually by the LIA and CLIPS laboratories in the framework of the ELISA consortium. For both systems, improvement was observed when combined with prior acoustic segmentation. However, a larger impact, in terms of performance, is observed on the LIA system based on an ascending/HMM approach compared to the CLIPS system based on speaker turn detection.
Sylvain Meignier, Daniel Moraru, Corinne Fredouille, Laurent Besacier, Jean-François Bonastre
ICASSP (1)4
2004 The ELISA consortium approaches in broadcast news speaker segmentation during the NIST 2003 rich transcription evaluation
abstract
The paper presents the ELISA consortium activities in automatic speaker segmentation, also known as speaker diarization, during the NIST rich transcription (RT), 2003, evaluation. The experiments were conducted on real broadcast news data (HUB4). Two different approaches from the CLIPS and LIA laboratories are presented and different possibilities of combining them are investigated, in the framework of the ELISA consortium. The system submitted as an ELISA primary system obtained the second lowest segmentation error rate compared to the other RT03-participant primary systems. Another ELISA system submitted as a secondary system outperformed the best primary system and obtained the lowest speaker segmentation error rate.
Daniel Moraru, Sylvain Meignier, Corinne Fredouille, Laurent Besacier, Jean-François Bonastre
ICASSP (1)4
2004 Spoken and Written Language Resources for Vietnamese
Viet Bac Le, Do Dat Tran, Eric Castelli, Laurent Besacier, Jean-François Serignat
LREC4
2003 The ELISA consortium approaches in speaker segmentation during the NIST 2002 speaker recognition evaluation
abstract
This paper presents the ELISA consortium activities in automatic speaker segmentation during last NIST 2002 evaluation: two different approaches from CLIPS and LIA laboratories are presented and the possibility of combining them either by applying them consecutively, or by fusing the decisions made by each of them, is investigated. Various types of data were available for NIST 2002. The ELISA systems obtained the lower error rates for two corpora: the CLIPS system obtained the best performance on the Meeting data, the LIA system obtained the best performance on the Switchboard data. The combining strategies proposed in this paper allowed us to improve the performance of the best single system on both data types (up to 30 % of error rate reduction).
Daniel Moraru, Sylvain Meignier, Laurent Besacier, Jean-François Bonastre, Ivan Magrin-Chagnolleau
ICASSP (2)3
2003 Using the web for fast language model construction in minority languages
abstract
International audience
Viet Bac Le, Brigitte Bigi, Laurent Besacier, Eric Castelli
INTERSPEECH3
2003 The NESPOLE! voIP multilingual corpora in tourism and medical domains
abstract
In this paper we present the multilingual VoIP (Voice over Internet Protocol networks) corpora collected for the second showcase of the Nespole! project in the tourism and medical domains. The corpora comprise over 20 hours of human-to-human monolingual dialogues in English, French, German and Italian: 66 dialogues in the tourism domain and 49 in the medical domain. We describe in detail the data collection (technical set-up, scenarios for each domain, recording procedure and data transcription), as well as statistically illustrated corpora and a preliminary data analysis. 1.
Nadia Mana, Susanne Burger, Roldano Cattoni, Laurent Besacier, Victoria MacLaren, John W. McDonough, Florian Metze
INTERSPEECH4
2002 Speech-to-speech translation system evaluation: results for French for the NESPOLE! project first showcase
abstract
In this paper we give the results of a set of evaluations conducted in the context of a speech to speech translation project (NESPOLE!). The chosen situation involves a client (French, German, American) talking to an Italian travel agent (both using their own language) to organize a stay in Italy. Fives series of evaluation were conducted on the same data set. The first series concerned the Automatic Speech Recognition alone. Two other series were about monolingual (back-) translation from ASR outputs on the data set and form transcriptions of the data set. The last ones were about bilingual translation from both the ASR outputs and the transcriptions. The goal of the evaluation was to check the performances of the system at the end of the second year of the project. The fives sets of results concerning the French modules are given and commented.
Solange Rossato, Hervé Blanchon, Laurent Besacier
INTERSPEECH3
2001 Speech translation for French in the NESPOLE! European project
abstract
International audience
Laurent Besacier, Hervé Blanchon, Yannick Fouquet, Jean-Philippe Guilbaud, Stéphane Helme, Sylviane Mazenot, Daniel Moraru, Dominique Vaufreydaz
INTERSPEECH1
2001 The nespole! voIP dialogue database
abstract
This paper presents the status of the NESPOLE! data collection as of end of February, 2001. A multilingual VoIP (Voice over Internet Protocol networks) database consisting of 200 dialogues in 4 languages (English, German, Italian and French) was recorded and transcribed. Dialogue speakers were connected via a H323 video-conferencing terminal. We describe the task, the technical architecture, the recording procedure and the transcription process of the NESPOLE! data collection. We provide some statistics concerning the data and, finally, we address problems that arose during the collection and annotation process. 1.
Susanne Burger, Laurent Besacier, Paolo Coletti, Florian Metze, Céline Morel
INTERSPEECH2
2001 The effect of speech and audio compression on speech recognition performance
abstract
This paper proposes an in-depth look at the influence of different speech and audio codecs on the performance of our continuous speech recognition engine. GSM full rate, G711, G723.1 and MPEG coders are investigated. It is shown that MPEG transcoding degrades the speech recognition performance for low bitrates whereas performance remains acceptable for specialized speech coders like GSM or G711. A new strategy is proposed to cope with degradation due to low bitrate coding. The acoustic models of the speech recognition system are trained with transcoded speech (one acoustic model for each speech/audio codec). First results show that this strategy allows one to recover acceptable performance.
Laurent Besacier, Carole Bergamini, Dominique Vaufreydaz, Eric Castelli
MMSP1
2000 GSM speech coding and speaker recognition
abstract
This paper investigates the influence of GSM speech coding on text independent speaker recognition performance. The three existing GSM speech coder standards were considered. The whole TIMIT database was passed through these coders, obtaining three transcoded databases. In a first experiment, it was found that the use of GSM coding degrades significantly the identification and verification performance (performance in correspondence with the perceptual speech quality of each coder). In a second experiment, the features for the speaker recognition system were calculated directly from the information available in the encoded bit stream. It was found that a low LPC order in GSM coding is responsible for most performance degradations. By extracting the features directly from the encoded bit-stream, we also managed to obtain a speaker recognition system equivalent in performance to the original one which decodes and reanalyzes speech before performing recognition.
Laurent Besacier, Sara Grassi, Alain Dufaux, Michael Ansorge, Fausto Pellandini
ICASSP1
2000 A New Methodology for Speech Corpora Definition from Internet Documents
Dominique Vaufreydaz, Carole Bergamini, Jean-François Serignat, Laurent Besacier, Mohammad Akbar 0002
LREC4
2000 Subband architecture for automatic speaker recognition
Laurent Besacier, Jean-François Bonastre
Signal Process.1
2000 Localization and selection of speaker-specific information with statistical modeling
Laurent Besacier, Jean-François Bonastre, Corinne Fredouille
Speech Commun.1
1999 Experimental evaluation of text-independent speaker verification on laboratory and field test databases in the M2VTS project
abstract
This paper describes the Philips Large Vocabulary Continuous Mandarin speech recognition system for the 1999 Taiwan benchmark. The basic system architecture is based on the Philips LVCSR technology developed for Western languages. However, several modifications are made in order to better suitted processing Chinese spoken languages. In the paper, we present some experimental results on the two tasks we participated in this benchmark: digit and continuous syllable. For the development set, we were able to obtain a digit/syllable error rate of 2.9%/23.9%. At the final evaluation, our system achieves the lowest error rate of 3.1%/24.3% among all participating sites.
Laurent Besacier, Jürgen Lüttin, Gilbert Maître, Eric Meurville
EUROSPEECH1
1998 Frame pruning for speaker recognition
abstract
In this paper, we propose a frame selection procedure for text-independent speaker identification. Instead of averaging the frame likelihoods along the whole test utterance, some of these are rejected (pruning) and the final score is computed with a limited number of frames. This pruning stage requires a prior frame level likelihood normalization in order to make comparison between frames meaningful. This normalization procedure alone leads to a significant performance enhancement. As far as pruning is concerned, the optimal number of frames pruned is learned on a tuning data set for normal and telephone speech. Validation of the pruning procedure on 567 speakers leads to a 27% identification rate improvement on TIMIT, and to 17% on NTIMIT.
Laurent Besacier, Jean-François Bonastre
ICASSP1
1998 Time and frequency pruning for speaker identification
abstract
This work is an attempt to refine decisions in speaker identification. A test utterance is divided into multiple time-frequency blocks on which a normalized likelihood score is calculated. Instead of averaging the block-likelihoods along the whole test utterance, some of them are rejected (pruning) and the final score is computed with a limited number of time-frequency blocks. The results obtained in the special case of time pruning lead the authors to experiment a joint time and frequency pruning approach. The optimal percentage of blocks pruned is learned on a tuning data set with the minimum identification error criterion. Validation of the time-frequency pruning process on 567 speakers leads to a significant error rate reduction for short training and test duration.
Laurent Besacier, Jean-François Bonastre
ICPR1