VLDB 2026 Research / reviewers in the wild / expert
Yannick Estève
dblp:66/7159
· DBLP profile ↗
107ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-3656-8883ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 6 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WhiteHouse: Translation of the Casablanca Corpus for Multi-dialectal Arabic Speech Translation
Fethi Bougares, Salima Mdhaffar, Yannick Estève |
LREC | 3 |
| 2026 | SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
Haroun Elleuch, Salima Mdhaffar, Yannick Estève, Fethi Bougares |
LREC | 3 |
| 2026 | Using Multimodal and Language-Agnostic Sentence Embeddings for Abstractive SummarizationabstractInternational audience Chaimae Chellaf, Salima Mdhaffar, Yannick Estève, Stéphane Huet |
LREC | 3 |
| 2026 | Pantagruel: Unified Self-Supervised Encoders for French Text and SpeechabstractInternational audience Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane, Mohammed Ghennai, Qianwen Guan, Kirill Milintsevich, Salima Mdhaffar, Aidan Mannion, Nils Defauw, Shuyue Gu, Alexandre Audibert, Marco Dinarelli, Yannick Estève, Lorraine Goeuriot, Steffen Lalande, Nicolas Hervé, Maximin Coavoux, François Portet, Étienne Ollion, Marie Candito, Maxime Peyrard, Solange Rossato, Benjamin Lecouteux, Aurélie Nardy, Gilles Sérasset, Vincent Segonne, Solène Evain, Diandra Fabre, Didier Schwab |
LREC | 14 |
| 2026 | Improving End-to-End Speech Translation for the Low Resource Language Fongbe to FrenchabstractThis study addresses the challenges of end-to-end (E2E) Speech-to-Text Translation (STT) for the low-resource Fongbe-to-French language pair using a transfer learning approach. We first establish robust baselines by integrating state-of-the-art pretrained speech encoders (HuBERT-147, AfriHuBERT, XLS-R, Whisper) with powerful text decoders (mBART, NLLB). This initial phase identified the AfriHuBERT–NLLB and XLS-R–NLLB combinations as the most competitive E2E configurations. To further enhance performance, we propose and evaluate three hybrid feature fusion strategies, the Bidirectional Co-Attention (BCOAT), the Feature-wise Linear Modulation (FiLM), and the Feature Sum (SUM). These methods are designed to strategically fuse intermediate representations extracted from two distinct and powerful encoders, AfriHuBERT and XLS-R, within the E2E architecture. The fusion process aims to enrich the linguistic and tonal information critical for accurate translation of the tonal Fongbe language. Experimental results demonstrate significant performance gains over the baselines. The BLEU score improved from 26.32 to a peak of 27.78 for the AfriHuBERT–NLLB configuration (using FiLM), and from 26.27 to a maximum of 28.05 for the XLS-R–NLLB configuration (using SUM). These findings confirm that translation quality for tonal languages like Fongbe can be substantially improved by extracting and combining high-quality, complementary features through advanced encoder fusion. Our hybrid feature fusion methods present a substantial advance in speech translation quality within resource-scarce linguistic environments. D. Fortune Kponou, Fréjus A. A. Laleye, Eugène C. Ezin, Yannick Estève |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2025 | SENSE models: an open source solution for multilingual and multimodal semantic-based tasksabstractThis paper introduces SENSE (Shared Embedding for Nlingual Speech and tExt), an open-source solution inspired by the SAMUXLSR framework and conceptually similar to Meta AI’s SONAR models. These approaches rely on a teacher-student framework to align a self-supervised speech encoder with the language-agnostic continuous representations of a text encoder at the utterance level. We describe how the original SAMU-XLSR method has been updated by selecting a stronger teacher text model and a better initial speech encoder. The source code for training and using SENSE models has been integrated into the SpeechBrain toolkit, and the first SENSE model we trained has been publicly released. We report experimental results on multilingual and multimodal semantic tasks, where our SENSE model achieves highly competitive performance. Finally, this study offers new insights into how semantics are captured in such semantically aligned speech encoders. Salima Mdhaffar, Haroun Elleuch, Chaimae Chellaf, Yannick Estève |
ASRU | 5 |
| 2025 | ADI-20: Arabic Dialect Identification dataset and modelsabstractPublished in Interspeech 2025 Haroun Elleuch, Salima Mdhaffar, Yannick Estève, Fethi Bougares |
INTERSPEECH | 3 |
| 2025 | Beyond Similarity Scoring: Detecting Entailment and Contradiction in Multilingual and Multimodal ContextsabstractInternational audience Othman Istaiteh, Salima Mdhaffar, Yannick Estève |
INTERSPEECH | 3 |
| 2025 | Extending the Fongbe to French Speech Translation Corpus: resources, models and benchmark
D. Fortune Kponou, Salima Mdhaffar, Fréjus A. A. Laleye, Eugène C. Ezin, Yannick Estève |
INTERSPEECH | 5 |
| 2025 | Towards Early Prediction of Self-Supervised Speech Model Performance
Ryan Whetten, Lucas Maison, Titouan Parcollet, Marco Dinarelli, Yannick Estève |
INTERSPEECH | 5 |
| 2024 | TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingabstractIn recent years, there has been a significant increase in interest in developing Spoken Language Understanding (SLU) systems. SLU involves extracting a list of semantic information from the speech signal. A major issue for SLU systems is the lack of sufficient amount of bi-modal (audio and textual semantic annotation) training data. Existing SLU resources are mainly available in high-resource languages such as English, Mandarin and French. However, one of the current challenges concerning low-resourced languages is data collection and annotation. In this work, we present a new freely available corpus, named TARIC-SLU, composed of railway transport conversations in Tunisian dialect that is continuously annotated in dialogue acts and slots. We describe the semantic model of the dataset, the data and experiments conducted to build ASR-based and SLU-based baseline models. To facilitate its use, a complete recipe, including data preparation, training and evaluation scripts, has been built and will be integrated to SpeechBrain, a popular open-source conversational AI toolkit based on PyTorch. Salima Mdhaffar, Fethi Bougares, Renato De Mori, Mohamed Salah Zaïem, Mirco Ravanelli, Yannick Estève |
LREC/COLING | 6 |
| 2024 | Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice AssistantsabstractRecent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled demographic tags. This paper introduces the Sonos Voice Control Bias Assessment Dataset, an open dataset composed of voice assistant requests for North American English in the music domain (1,038 speakers, 166 hours, 170k audio samples, with 9,040 unique labelled transcripts) with a controlled demographic diversity (gender, age, dialectal region and ethnicity). We also release a statistical demographic bias assessment methodology, at the univariate and multivariate levels, tailored to this specific use case and leveraging spoken language understanding metrics rather than transcription accuracy, which we believe is a better proxy for user experience. To demonstrate the capabilities of this dataset and statistical method to detect demographic bias, we consider a pair of state-of-the-art Automatic Speech Recognition and Spoken Language Understanding models. Results show statistically significant differences in performance across age, dialectal region and ethnicity. Multivariate tests are crucial to shed light on mixed effects between dialectal region, gender and age. Chloé Sekkat, Fanny Leroy, Salima Mdhaffar, Blake Perry Smith, Yannick Estève, Joseph Dureau, Alice Coucke |
LREC/COLING | 5 |
| 2024 | A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language UnderstandingabstractSelf-Supervised Learning is vastly used to efficiently represent speech for Spoken Language Understanding, gradually replacing conventional approaches. Meanwhile, textual SSL models are proposed to encode language-agnostic semantics. SAMU-XLSR framework employed this semantic information to enrich multilingual speech representations. A recent study investigated SAMU-XLSR in-domain semantic enrichment by specializing it on downstream transcriptions, leading to state-of-the-art results on a challenging SLU task. This study's interest lies in the loss of multilingual performances and lack of specific-semantics training induced by such specialization in close languages without any SLU implication. We also consider SAMU-XLSR's loss of initial cross-lingual abilities due to a separate SLU fine-tuning. Therefore, this paper proposes a dual task learning approach to improve SAMU-XLSR semantic enrichment while considering distant languages for multilingual and language portability experiments. Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Yannick Estève |
INTERSPEECH | 4 |
| 2024 | An Analysis of Linear Complexity Attention Substitutes With Best-RQabstractSelf-Supervised Learning (SSL) has proven to be effective in various domains, including speech processing. However, SSL is computationally and memory expensive. This is in part due the quadratic complexity of multi-head self-attention (MHSA). Alternatives for MHSA have been proposed and used in the speech domain, but have yet to be investigated properly in an SSL setting. In this work, we study the effects of replacing MHSA with recent state-of-the-art alternatives that have linear complexity, namely, HyperMixing, Fastformer, SummaryMixing, and Mamba. We evaluate these methods by looking at the speed, the amount of VRAM consumed, and the performance on the SSL MP3S benchmark. Results show that these linear alternatives maintain competitive performance compared to MHSA while, on average, decreasing VRAM consumption by around 20% to 60% and increasing speed from 7% to 65% for input sequences ranging from 20 to 80 seconds. Ryan Whetten, Titouan Parcollet, Adel Moumen, Marco Dinarelli, Yannick Estève |
SLT | 5 |
| 2024 | LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Comput. Speech Lang. | 14 |
| 2024 | Open-Source Conversational AI with SpeechBrain 1.0abstractSpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete recipes of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks. Mirco Ravanelli, Titouan Parcollet, Adel Moumen, Sylvain de Langen, Cem Subakan, Peter Plantinga, Yingzhi Wang 0002, Pooneh Mousavi, Luca Della Libera, Artem Ploujnikov, Francesco Paissan, Davide Borra, Mohamed Salah Zaïem, Zeyu Zhao 0004, Shucong Zhang, Georgios Karakasidis, Sung-Lin Yeh, Pierre Champion, Aku Rouhe, Rudolf Braun, Florian Mai, Juan Zuluaga-Gomez, Seyed Mahed Mousavi, Andreas Nautsch, Xuechen Liu 0001, Sangeet Sagar, Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière, Mickael Rouvier, Renato De Mori, Yannick Estève |
J. Mach. Learn. Res. | 33 |
| 2023 | Enhancing Expressivity Transfer in Textless Speech-to-Speech TranslationabstractTextless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and transferring expressivity accurately across different languages. Expressivity plays a vital role in conveying emotions, nuances, and cultural subtleties, thereby enhancing communication across diverse languages. To address this issue this study presents a novel method that operates at the discrete speech unit level and leverages multilingual emotion embeddings to capture language-agnostic information. Specifically, we demonstrate how these embeddings can be used to effectively predict the pitch and duration of speech units in the target language. Through objective and subjective experiments conducted on a French-to-English translation task, our findings highlight the superior expressivity transfer achieved by our approach compared to current state-of-the-art systems. Jarod Duret, Benjamin O'Brien, Yannick Estève, Titouan Parcollet |
ASRU | 3 |
| 2023 | Improving Accented Speech Recognition with Multi-Domain TrainingabstractThanks to the rise of self-supervised learning, automatic speech recognition (ASR) systems now achieve near human performance on a wide variety of datasets. However, they still lack generalization capability and are not robust to domain shifts like accent variations. In this work, we use speech audio representing four different French accents to create fine-tuning datasets that improve the robustness of pre-trained ASR models. By incorporating various accents in the training set, we obtain both in-domain and out-of-domain improvements. Our numerical experiments show that we can reduce error rates by up to 25% (relative) on African and Belgian accents compared to single-domain training while keeping a good performance on standard French. Lucas Maison, Yannick Estève |
ICASSP | 2 |
| 2023 | Federated Learning for ASR Based on wav2vec 2.0abstractThis paper presents a study on the use of federated learning to train an ASR model based on a wav2vec 2.0 model pre-trained by self supervision. Carried out on the well-known TED-LIUM 3 dataset, our experiments show that such a model can obtain, with no use of a language model, a word error rate of 10.92% on the official TEDLIUM 3 test set, without sharing any data from the different users. We also analyse the ASR performance for speakers depending to their participation to the federated learning. Since federated learning was first introduced for privacy purposes, we also measure its ability to protect speaker identity. To do that, we exploit an approach to analyze information contained in exchanged models based on a neural network footprint on an indicator dataset. This analysis is made layer-wise and shows which layers in an exchanged wav2vec 2.0based model bring the speaker identity information. Salima Mdhaffar, Natalia A. Tomashenko, Jean-François Bonastre, Yannick Estève |
ICASSP | 5 |
| 2023 | Semantic Enrichment Towards Efficient Speech RepresentationsabstractOver the past few years, self-supervised learned speech representations have emerged as fruitful replacements for conventional surface representations when solving Spoken Language Understanding (SLU) tasks.Simultaneously, multilingual models trained on massive textual data were introduced to encode language agnostic semantics.Recently, the SAMU-XLSR approach introduced a way to make profit from such textual models to enrich multilingual speech representations with language agnostic semantics.By aiming for better semantic extraction on a challenging Spoken Language Understanding task and in consideration with computation costs, this study investigates a specific in-domain semantic enrichment of the SAMU-XLSR model by specializing it on a small amount of transcribed data from the downstream task.In addition, we show the benefits of the use of same-domain French and Italian benchmarks for low-resource language portability and explore cross-domain capacities of the enriched SAMU-XLSR. Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Yannick Estève |
INTERSPEECH | 5 |
| 2023 | Some Voices are Too Common: Building Fair Speech Recognition Systems Using the CommonVoice Dataset
Lucas Maison, Yannick Estève |
INTERSPEECH | 2 |
| 2023 | Speech and multilingual natural language framework for speaker change detection and diarization
Or Haim Anidjar, Yannick Estève, Chen Hajaj, Amit Dvir, Itshak Lapidot |
Expert Syst. Appl. | 2 |
| 2022 | Retrieving Speaker Information from Personalized Acoustic Models for Speech RecognitionabstractThe widespread of powerful personal devices capable of collecting voice of their users has opened the opportunity to build speaker adapted speech recognition system (ASR) or to participate to collaborative learning of ASR. In both cases, personalized acoustic models (AM), i.e. fine-tuned AM with specific speaker data, can be built. A question that naturally arises is whether the dissemination of personalized acoustic models can leak personal information. In this paper, we show that it is possible to retrieve the gender of the speaker, but also his identity, by just exploiting the weight matrix changes of a neural acoustic model locally adapted to this speaker. Incidentally we observe phenomena that may be useful towards explainability of deep neural networks in the context of speech processing. Gender can be identified almost surely using only the first layers and speaker verification performs well when using middle-up layers. Our experimental study on the TED-LIUM 3 dataset with HMM/TDNN models shows a purity of 95% for gender detection, and an Equal Error Rate of 9.07% for a speaker verification task by only exploiting the weights from personalized models that could be exchanged instead of user data. Salima Mdhaffar, Jean-François Bonastre, Marc Tommasi, Natalia A. Tomashenko, Yannick Estève |
ICASSP | 5 |
| 2022 | Privacy Attacks for Automatic Speech Recognition Acoustic Models in A Federated Learning FrameworkabstractThis paper investigates methods to effectively retrieve speaker information from the personalized speaker adapted neural network acoustic models (AMs) in automatic speech recognition (ASR). This problem is especially important in the context of federated learning of ASR acoustic models where a global model is learnt on the server based on the updates received from multiple clients. We propose an approach to analyze information in neural network AMs based on a neural network footprint on the so-called Indicator dataset. Using this method, we develop two attack models that aim to infer speaker identity from the updated personalized models without access to the actual users’ speech data. Experiments on the TED-LIUM 3 corpus demonstrate that the proposed approaches are very effective and can provide equal error rate (EER) of 1–2%. Natalia A. Tomashenko, Salima Mdhaffar, Marc Tommasi, Yannick Estève, Jean-François Bonastre |
ICASSP | 4 |
| 2022 | A Study of Gender Impact in Self-supervised Models for Speech-to-Text SystemsabstractSelf-supervised models for speech processing emerged recently as popular foundation blocks in speech processing pipelines.These models are pre-trained on unlabeled audio data and then used in speech processing downstream tasks such as automatic speech recognition (ASR) or speech translation (ST).Since these models are now used in research and industrial systems alike, it becomes necessary to understand the impact caused by some features such as gender distribution within pre-training data.Using French as our investigation language, we train and compare gender-specific wav2vec 2.0 models against models containing different degrees of gender balance in their pretraining data.The comparison is performed by applying these models to two speech-to-text downstream tasks: ASR and ST.Results show the type of downstream integration matters.We observe lower overall performance using gender-specific pretraining before fine-tuning an end-to-end ASR system.However, when self-supervised models are used as feature extractors, the overall ASR and ST results follow more complex patterns in which the balanced pre-trained model does not necessarily lead to the best results.Lastly, our crude 'fairness' metric, the relative performance difference measured between female and male test sets, does not display a strong variation from balanced to gender-specific pre-trained wav2vec 2.0 models. Marcely Zanon Boito, Laurent Besacier, Natalia A. Tomashenko, Yannick Estève |
INTERSPEECH | 4 |
| 2022 | End-to-end model for named entity recognition from speech without paired training dataabstractRecent works showed that end-to-end neural approaches tend to become very popular for spoken language understanding (SLU).Through the term end-to-end, one considers the use of a single model optimized to extract semantic information directly from the speech signal.A major issue for such models is the lack of paired audio and textual data with semantic annotation.In this paper, we propose an approach to build an end-to-end neural model to extract semantic information in a scenario in which zero paired audio data is available.Our approach is based on the use of an external model trained to generate a sequence of vectorial representations from text.These representations mimic the hidden representations that could be generated inside an end-to-end automatic speech recognition (ASR) model by processing a speech signal.A SLU neural module is then trained using these representations as input and the annotated text as output.Last, the SLU module replaces the top layers of the ASR model to achieve the construction of the end-to-end model.Our experiments on named entity recognition, carried out on the QUAERO corpus, show that this approach is very promising, getting better results than a comparable cascade approach or than the use of synthetic voices. Salima Mdhaffar, Jarod Duret, Titouan Parcollet, Yannick Estève |
INTERSPEECH | 4 |
| 2022 | Speech Resources in the Tamasheq LanguageabstractIn this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of radio recordings from daily broadcast news in Niger (Studio Kalangou) and Mali (Studio Tamani). We share (i) a massive amount of unlabeled audio data (671 hours) in five languages: French from Niger, Fulfulde, Hausa, Tamasheq and Zarma, and (ii) a smaller 17 hours parallel corpus of audio recordings in Tamasheq, with utterance-level translations in the French language. All this data is shared under the Creative Commons BY-NC-ND 3.0 license. We hope these resources will inspire the speech community to develop and benchmark models using the Tamasheq language. Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche-Braham, Loïc Barrault, Mickael Rouvier, Yannick Estève |
LREC | 7 |
| 2022 | The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation toolsabstractWith the emergence of neural end-to-end approaches for spoken language understanding (SLU), a growing number of studies have been presented during these last three years on this topic. The major part of these works addresses the spoken language understanding domain through a simple task like speech intent detection. In this context, new benchmark datasets have also been produced and shared with the community related to this task. In this paper, we focus on the French MEDIA SLU dataset, distributed since 2005 and used as a benchmark dataset for a large number of research works. This dataset has been shown as being the most challenging one among those accessible to the research community. Distributed by ELRA, this corpus is free for academic research since 2019. Unfortunately, the MEDIA dataset is not really used beyond the French research community. To facilitate its use, a complete recipe, including data preparation, training and evaluation scripts, has been built and integrated to SpeechBrain, an already popular open-source and all-in-one conversational AI toolkit based on PyTorch. This recipe is presented in this paper. In addition, based on the feedback of some researchers who have worked on this dataset for several years, some corrections have been brought to the initial manual annotation: the new version of the data will also be integrated into the ELRA catalogue, as the original one. More, a significant amount of data collected during the construction of the MEDIA corpus in the 2000s was never used until now: we present the first results reached on this subset — also included in the MEDIA SpeechBrain recipe — , that will be used for now as the MEDIA test2. Last, we discuss evaluation issues. Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar, Nathalie Camelin, Sahar Ghannay, Bassam Jabaian, Yannick Estève |
LREC | 8 |
| 2022 | Impact Analysis of the Use of Speech and Language Models Pretrained by Self-Supersivion for Spoken Language UnderstandingabstractPretrained models through self-supervised learning have been recently introduced for both acoustic and language modeling. Applied to spoken language understanding tasks, these models have shown their great potential by improving the state-of-the-art performances on challenging benchmark datasets. In this paper, we present an error analysis reached by the use of such models on the French MEDIA benchmark dataset, known as being one of the most challenging benchmarks for the slot filling task among all the benchmarks accessible to the entire research community. One year ago, the state-of-art system reached a Concept Error Rate (CER) of 13.6% through the use of a end-to-end neural architecture. Some months later, a cascade approach based on the sequential use of a fine-tuned wav2vec2.0 model and a fine-tuned BERT model reaches a CER of 11.2%. This significant improvement raises questions about the type of errors that remain difficult to treat, but also about those that have been corrected using these models pre-trained through self-supervision learning on a large amount of data. This study brings some answers in order to better understand the limits of such models and open new perspectives to continue improving the performance. Salima Mdhaffar, Valentin Pelloin, Antoine Caubrière, Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Nathalie Camelin, Yannick Estève |
LREC | 8 |
| 2022 | On the Use of Semantically-Aligned Speech Representations for Spoken Language UnderstandingabstractIn this paper we examine the use of semantically-aligned speech representations for end-to-end spoken language understanding (SLU). We employ the recently-introduced SAMU-XLSR model, which is designed to generate a single embedding that captures the semantics at the utterance level, semantically aligned across different languages. This model combines the acoustic frame-level speech representation learning model (XLS-R) with the Language Agnostic BERT Sentence Embedding (LaBSE) model. We show that the use of the SAMU-XLSR model instead of the initial XLS-R model improves significantly the performance in the framework of end-to-end SLU. Finally, we present the benefits of using this model towards language portability in SLU. Gaëlle Laperrière, Valentin Pelloin, Mickael Rouvier, Themos Stafylakis, Yannick Estève |
SLT | 5 |
| 2021 | An Empirical Study of End-To-End Simultaneous Speech Translation Decoding StrategiesabstractThis paper proposes a decoding strategy for end-to-end simultaneous speech translation. We leverage end-to-end models trained in offline mode and conduct an empirical study for two language pairs (English-to-German and English-to-Portuguese). We also investigate different output token granularities including characters and Byte Pair Encoding (BPE) units. The results show that the proposed decoding approach allows to control BLEU/Average Lagging trade-off along different latency regimes. Our best decoding settings achieve comparable results with a strong cascade model evaluated on the simultaneous translation track of IWSLT 2020 shared task. Yannick Estève, Laurent Besacier |
ICASSP | 2 |
| 2021 | End2End Acoustic to Semantic TransductionabstractIn this paper, we propose a novel end-to-end sequence-to-sequence spoken language understanding model using an attention mechanism. It reliably selects contextual acoustic features in order to hypothesize semantic contents. An initial architecture capable of extracting all pronounced words and concepts from acoustic spans is designed and tested. With a shallow fusion language model, this system reaches a 13.6 concept error rate (CER) and an 18.5 concept value error rate (CVER) on the French MEDIA corpus, achieving an absolute 2.8 points reduction compared to the state-of-the-art. Then, an original model is proposed for hypothesizing concepts and their values. This transduction reaches a 15.4 CER and a 21.6 CVER without any new type of context. Valentin Pelloin, Nathalie Camelin, Antoine Laurent, Renato De Mori, Antoine Caubrière, Yannick Estève, Sylvain Meignier |
ICASSP | 6 |
| 2021 | LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from SpeechabstractSelf-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve performance on downstream tasks such as automatic speech recognition (ASR). While these works suggest it is possible to reduce dependence on labeled data for building efficient speech systems, their evaluation was mostly made on ASR and using multiple and heterogeneous experimental settings (most of them for English). This questions the objective comparison of SSL approaches and the evaluation of their impact on building speech systems. In this paper, we propose LeBenchmark: a reproducible framework for assessing SSL from speech. It not only includes ASR (high and low resource) tasks but also spoken language understanding, speech translation and emotion recognition. We also focus on speech technologies in a language different than English: French. SSL models of different sizes are trained from carefully sourced and documented datasets. Experiments show that SSL is beneficial for most but not all tasks which confirms the need for exhaustive and reliable benchmarks to evaluate its real impact. LeBenchmark is shared with the scientific community for reproducible research in SSL from speech. Solène Evain, Hang Le 0001, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia A. Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Estève, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Interspeech | 12 |
| 2021 | Impact of Encoding and Segmentation Strategies on End-to-End Simultaneous Speech TranslationabstractBoosted by the simultaneous translation shared task at IWSLT 2020, promising end-to-end online speech translation approaches were recently proposed.They consist in incrementally encoding a speech input (in a source language) and decoding the corresponding text (in a target language) with the best possible trade-off between latency and translation quality.This paper investigates two key aspects of end-to-end simultaneous speech translation: (a) how to encode efficiently the continuous speech flow, and (b) how to segment the speech flow in order to alternate optimally between reading (R: encoding input) and writing (W: decoding output) operations.We extend our previously proposed end-to-end online decoding strategy and show that while replacing BLSTM by ULSTM encoding degrades performance in offline mode, it actually improves both efficiency and performance in online mode.We also measure the impact of different methods to segment the speech signal (using fixed interval boundaries, oracle word boundaries or randomly set boundaries) and show that our best end-to-end online decoding strategy is surprisingly the one that alternates R/W operations on fixed size blocks on our English-German speech translation setup. Yannick Estève, Laurent Besacier |
Interspeech | 2 |
| 2021 | On the Use of Self-Supervised Pre-Trained Acoustic and Linguistic Features for Continuous Speech Emotion RecognitionabstractPre-training for feature extraction is an increasingly studied approach to get better continuous representations of audio and text content. In the present work, we use wav2vec and camemBERT as self-supervised learned models to represent our data in order to perform continuous emotion recognition from speech (SER) on AlloSat, a large French emotional database describing the satisfaction dimension, and on the state of the art corpus SEWA focusing on valence, arousal and liking dimensions. To the authors' knowledge, this paper presents the first study showing that the joint use of wav2vec and BERT-like pre-trained features is very relevant to deal with continuous SER task, usually characterized by a small amount of labeled training data. Evaluated by the well-known concordance correlation coefficient (CCC), our experiments show that we can reach a CCC value of 0.825 instead of 0.592 when using MFCC in conjunction with word2vec word embedding on the AlloSat dataset. Manon Macary, Marie Tahon, Yannick Estève, Anthony Rousseau |
SLT | 3 |
| 2020 | Error Analysis Applied to End-to-End Spoken Language UnderstandingabstractThis paper presents a qualitative study of errors produced by an end-to-end spoken language understanding (SLU) system (speech signal to concepts) that reaches state of the art performance. Different studies are proposed to better understand the weaknesses of such systems: comparison to a classical pipeline SLU system, a study on the cause of concept deletions (the most frequent error), observation of a problem in the capability of the end-to-end SLU system to segment correctly concepts, analysis of the system behavior to process unseen concept/value pairs, analysis of the benefit of the curriculum-based transfer learning approach. Last, we proposed a way to compute embeddings of sub-sequences that seem to contain relevant information for future work. Antoine Caubrière, Sahar Ghannay, Natalia A. Tomashenko, Renato De Mori, Antoine Laurent, Emmanuel Morin, Yannick Estève |
ICASSP | 7 |
| 2020 | Dialogue History Integration into End-to-End Signal-to-Concept Spoken Language Understanding SystemsabstractThis work investigates the embeddings for representing dialog history in spoken language understanding (SLU) systems. We focus on the scenario when the semantic information is extracted directly from the speech signal by means of a single end-to-end neural network model. We proposed to integrate dialogue history into an end-to-end signal-to-concept SLU system. The dialog history is represented in the form of dialog history embedding vectors (so-called h-vectors) and is provided as an additional information to end-to-end SLU models in order to improve the system performance. Three following types of h-vectors are proposed and experimentally evaluated in this paper: (1) supervised-all embeddings predicting bag-of-concepts expected in the answer of the user from the last dialog system response; (2) supervised-freq embeddings focusing on predicting only a selected set of semantic concept (corresponding to the most frequent errors in our experiments); and (3) unsupervised embeddings. Experiments on the MEDIA corpus for the semantic slot filling task demonstrate that the proposed h-vectors improve the model performance. Natalia A. Tomashenko, Christian Raymond, Antoine Caubrière, Renato De Mori, Yannick Estève |
ICASSP | 5 |
| 2020 | Confidence Measure for Speech-to-Concept End-to-End Spoken Language UnderstandingabstractArticle soumis et accepté à la conférence Interspeech - Octobre 2020 - Shangai. Antoine Caubrière, Yannick Estève, Antoine Laurent, Emmanuel Morin |
INTERSPEECH | 2 |
| 2020 | Investigating Self-Supervised Pre-Training for End-to-End Speech TranslationabstractInternational audience Fethi Bougares, Natalia A. Tomashenko, Yannick Estève, Laurent Besacier |
INTERSPEECH | 4 |
| 2020 | Toward Qualitative Evaluation of Embeddings for Arabic Sentiment AnalysisabstractIn this paper, we propose several protocols to evaluate specific embeddings for Arabic sentiment analysis (SA) task. In fact, Arabic language is characterized by its agglutination and morphological richness contributing to great sparsity that could affect embedding quality. This work presents a study that compares embeddings based on words and lemmas in SA frame. We propose first to study the evolution of embedding models trained with different types of corpora (polar and non polar) and explore the variation between embeddings by observing the sentiment stability of neighbors in embedding spaces. Then, we evaluate embeddings with a neural architecture based on convolutional neural network (CNN). We make available our pre-trained embeddings to Arabic NLP research community with free to use. We provide also for free resources used to evaluate our embeddings. Experiments are done on the Large Arabic-Book Reviews (LABR) corpus in binary (positive/negative) classification frame. Our best result reaches 91.9%, that is higher than the best previous published one (91.5%). Amira Barhoumi, Nathalie Camelin, Chafik Aloulou, Yannick Estève, Lamia Hadrich Belguith |
LREC | 4 |
| 2020 | Where are we in Named Entity Recognition from Speech?abstractNamed entity recognition (NER) from speech is usually made through a pipeline process that consists in (i) processing audio using an automatic speech recognition system (ASR) and (ii) applying a NER to the ASR outputs. The latest data available for named entity extraction from speech in French were produced during the ETAPE evaluation campaign in 2012. Since the publication of ETAPE’s campaign results, major improvements were done on NER and ASR systems, especially with the development of neural approaches for both of these components. In addition, recent studies have shown the capability of End-to-End (E2E) approach for NER / SLU tasks. In this paper, we propose a study of the improvements made in speech recognition and named entity recognition for pipeline approaches. For this type of systems, we propose an original 3-pass approach. We also explore the capability of an E2E system to do structured NER. Finally, we compare the performances of ETAPE’s systems (state-of-the-art systems in 2012) with the performances obtained using current technologies. The results show the interest of the E2E approach, which however remains below an updated pipeline approach. Antoine Caubrière, Sophie Rosset, Yannick Estève, Antoine Laurent, Emmanuel Morin |
LREC | 3 |
| 2020 | AlloSat: A New Call Center French Corpus for Satisfaction and Frustration AnalysisabstractWe present a new corpus, named AlloSat, composed of real-life call center conversations in French that is continuously annotated in frustration and satisfaction. This corpus has been set up to develop new systems able to model the continuous aspect of semantic and paralinguistic information at the conversation level. The present work focuses on the paralinguistic level, more precisely on the expression of emotions. In the call center industry, the conversation usually aims at solving the caller’s request. As far as we know, most emotional databases contain static annotations in discrete categories or in dimensions such as activation or valence. We hypothesize that these dimensions are not task-related enough. Moreover, static annotations do not enable to explore the temporal evolution of emotional states. To solve this issue, we propose a corpus with a rich annotation scheme enabling a real-time investigation of the axis frustration / satisfaction. AlloSat regroups 303 conversations with a total of approximately 37 hours of audio, all recorded in real-life environments collected by Allo-Media (an intelligent call tracking company). First regression experiments, with audio features, show that the evolution of frustration / satisfaction axis can be retrieved automatically at the conversation level. Manon Macary, Marie Tahon, Yannick Estève, Anthony Rousseau |
LREC | 3 |
| 2020 | A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case StudyabstractThis corpus is part of the PASTEL (Performing Automated Speech Transcription for Enhancing Learning) project aiming to explore the potential of synchronous speech transcription and application in specific teaching situations. It includes 10 hours of different lectures, manually transcribed and segmented. The main interest of this corpus lies in its multimodal aspect: in addition to speech, the courses were filmed and the written presentation supports (slides) are made available. The dataset may then serve researches in multiple fields, from speech and language to image and video processing. The dataset will be freely available to the research community. In this paper, we first describe in details the annotation protocol, including a detailed analysis of the manually labeled data. Then, we propose some possible use cases of the corpus with baseline results. The use cases concern scientific fields from both speech and text processing, with language model adaptation, thematic segmentation and transcription to slide alignment. Salima Mdhaffar, Yannick Estève, Antoine Laurent, Nicolas Hernandez, Richard Dufour, Delphine Charlet, Géraldine Damnati, Solen Quiniou, Nathalie Camelin |
LREC | 2 |
| 2020 | Align then Summarize: Automatic Alignment Methods for Summarization Corpus CreationabstractSummarizing texts is not a straightforward task. Before even considering text summarization, one should determine what kind of summary is expected. How much should the information be compressed? Is it relevant to reformulate or should the summary stick to the original phrasing? State-of-the-art on automatic text summarization mostly revolves around news articles. We suggest that considering a wider variety of tasks would lead to an improvement in the field, in terms of generalization and robustness. We explore meeting summarization: generating reports from automatic transcriptions. Our work consists in segmenting and aligning transcriptions with respect to reports, to get a suitable dataset for neural summarization. Using a bootstrapping approach, we provide pre-alignments that are corrected by human annotators, making a validation set against which we evaluate automatic models. This consistently reduces annotators’ efforts by providing iteratively better pre-alignment and maximizes the corpus size by using annotations from our automatic alignment models. Evaluation is conducted on publicmeetings, a novel corpus of aligned public meetings. We report automatic alignment and summarization performances on this corpus and show that automatic alignment is relevant for data annotation since it leads to large improvement of almost +4 on all ROUGE scores on the summarization task. Paul Tardy, David Janiszek, Yannick Estève |
LREC | 3 |
| 2020 | A study of continuous space word and sentence representations applied to ASR error detection
Sahar Ghannay, Yannick Estève, Nathalie Camelin |
Speech Commun. | 2 |
| 2019 | Curriculum-Based Transfer Learning for an Effective End-to-End Spoken Language Understanding and Domain PortabilityabstractWe present an end-to-end approach to extract semantic concepts directly from the speech audio signal. To overcome the lack of data available for this spoken language understanding approach, we investigate the use of a transfer learning strategy based on the principles of curriculum learning. This approach allows us to exploit out-of-domain data that can help to prepare a fully neural architecture. Experiments are carried out on the French MEDIA and PORTMEDIA corpora and show that this end-to-end SLU approach reaches the best results ever published on this task. We compare our approach to a classical pipeline approach that uses ASR, POS tagging, lemmatizer, chunker... and other NLP tools that aim to enrich ASR outputs that feed an SLU text to concepts system. Last, we explore the promising capacity of our end-to-end SLU approach to address the problem of domain portability. Antoine Caubrière, Natalia A. Tomashenko, Antoine Laurent, Emmanuel Morin, Nathalie Camelin, Yannick Estève |
INTERSPEECH | 6 |
| 2019 | Qualitative Evaluation of ASR Adaptation in a Lecture Context: Application to the PASTEL CorpusabstractInternational audience Salima Mdhaffar, Yannick Estève, Nicolas Hernandez, Antoine Laurent, Richard Dufour, Solen Quiniou |
INTERSPEECH | 2 |
| 2019 | Investigating Adaptation and Transfer Learning for End-to-End Spoken Language Understanding from SpeechabstractInternational audience Natalia A. Tomashenko, Antoine Caubrière, Yannick Estève |
INTERSPEECH | 3 |
| 2018 | Multifaceted Engagement in Social Interaction with a Machine: The JOKER ProjectabstractThis paper addresses the problem of evaluating engagement of the human participant by combining verbal and nonverbal behaviour along with contextual information. This study will be carried out through four different corpora. Four different systems designed to explore essential and complementary aspects of the JOKER system in terms of paralinguistic/linguistic inputs were used for the data collection. An annotation scheme dedicated to the labeling of verbal and non-verbal behavior have been designed. From our experiment, engagement in HRI should be multifaceted. Laurence Devillers, Sophie Rosset, Guillaume Dubuisson Duplessis, Lucile Bechade, Yücel Yemez, Bekir Berker Türker, Tevfik Metin Sezgin, Engin Erzin, Kevin El Haddad, Stéphane Dupont, Paul Deléglise, Yannick Estève, Carole Lailler, Emer Gilmartin, Nick Campbell 0001 |
FG | 12 |
| 2018 | Task Specific Sentence Embeddings for ASR Error DetectionabstractInternational audience Sahar Ghannay, Yannick Estève, Nathalie Camelin |
INTERSPEECH | 2 |
| 2018 | Speaker Adaptive Training and Mixup Regularization for Neural Network Acoustic Models in Automatic Speech RecognitionabstractInternational audience Natalia A. Tomashenko, Yuri Y. Khokhlov, Yannick Estève |
INTERSPEECH | 3 |
| 2018 | Acoustic-dependent Phonemic Transcription for Text-to-speech SynthesisabstractInternational audience Kévin Vythelingum, Yannick Estève, Olivier Rosec |
INTERSPEECH | 2 |
| 2018 | FrNewsLink : a corpus linking TV Broadcast News Segments and Press Articles
Nathalie Camelin, Géraldine Damnati, Abdessalam Bouchekif, Anaïs Landeau, Delphine Charlet, Yannick Estève |
LREC | 6 |
| 2018 | Simulating ASR errors for training SLU systems
Edwin Simonnet, Sahar Ghannay, Nathalie Camelin, Yannick Estève |
LREC | 4 |
| 2018 | Evaluation of Feature-Space Speaker Adaptation for End-to-End Acoustic Models
Natalia A. Tomashenko, Yannick Estève |
LREC | 2 |
| 2018 | End-To-End Named Entity And Semantic Concept Extraction From SpeechabstractNamed entity recognition (NER) is among SLU tasks that usually extract semantic information from textual documents. Until now, NER from speech is made through a pipeline process that consists in processing first an automatic speech recognition (ASR) on the audio and then processing a NER on the ASR outputs. Such approach has some disadvantages (error propagation, metric to tune ASR systems sub-optimal in regards to the final task, reduced space search at the ASR output level,...) and it is known that more integrated approaches outperform sequential ones, when they can be applied. In this paper, we explore an end-to-end approach that directly extracts named entities from speech, though a unique neural architecture. On a such way, a joint optimization is possible for both ASR and NER. Experiments are carried on French data easily accessible, composed of data distributed in several evaluation campaigns. The results are promising since this end-to-end approach provides similar results (F-measure=0.66 on test data) than a classical pipeline approach to detect named entity categories (F-measure=0.64). Last, we also explore this approach applied to semantic concept extraction, through a slot filling task known as a spoken language understanding problem, and also observe an improvement in comparison to a pipeline approach. Sahar Ghannay, Antoine Caubrière, Yannick Estève, Nathalie Camelin, Edwin Simonnet, Antoine Laurent, Emmanuel Morin |
SLT | 3 |
| 2017 | Error detection of grapheme-to-phoneme conversion in text-to-speech synthesis using speech signal and lexical contextabstractIn unit selection text-to-speech synthesis, voice creation involved a phonemic transcription of read speech. This is produced by an automatic grapheme-to-phoneme conversion of the text read, followed by a manual correction. Although grapheme-to-phoneme conversion makes few errors, the manual correction is time consuming as every generated phoneme should be checked. We propose a method to automatically detect grapheme-to-phoneme conversion errors by comparing contrastives phonemisation hypothesis. A lattice-based forced alignment system is implemented, allowing for signal-dependent phonemisation. We implement also a sequence-to-sequence neural network model to obtain a context-dependent grapheme-to-phoneme conversion. On a French dataset, we show that we can detect to 86.3% of the errors made by a commercial grapheme-to-phoneme system. Moreover, the amount of data annotated as erroneous is kept under 10% of the total evaluation data. The time spent for phoneme manual checking can thus been drastically reduced without decreasing significantly the phonemic transcription quality. Kévin Vythelingum, Yannick Estève, Olivier Rosec |
ASRU | 2 |
| 2017 | Evaluating Automatic Topic Segmentation as a Segment Retrieval TaskabstractInternational audience Abdessalam Bouchekif, Delphine Charlet, Géraldine Damnati, Nathalie Camelin, Yannick Estève |
INTERSPEECH | 5 |
| 2017 | ASR Error Management for Improving Spoken Language UnderstandingabstractThis paper addresses the problem of automatic speech recognition (ASR) error detection and their use for improving spoken language understanding (SLU) systems. In this study, the SLU task consists in automatically extracting, from ASR transcriptions , semantic concepts and concept/values pairs in a e.g touristic information system. An approach is proposed for enriching the set of semantic labels with error specific labels and by using a recently proposed neural approach based on word embeddings to compute well calibrated ASR confidence measures. Experimental results are reported showing that it is possible to decrease significantly the Concept/Value Error Rate with a state of the art system, outperforming previously published results performance on the same experimental data. It also shown that combining an SLU approach based on conditional random fields with a neural encoder/decoder attention based architecture , it is possible to effectively identifying confidence islands and uncertain semantic output segments useful for deciding appropriate error handling actions by the dialogue manager strategy . Edwin Simonnet, Sahar Ghannay, Nathalie Camelin, Yannick Estève, Renato De Mori |
INTERSPEECH | 4 |
| 2016 | Title assignment for automatic topic segments in TV broadcast newsabstractThis paper addresses the task of assigning a title to topic segments automatically extracted from TV Broadcast News video recordings. We propose to associate a topic segment with the title of a newspaper article collected on the web at the same date. The task implies pairing newspaper articles and topic segments by maximising a given similarity measure. This approach raises several issues, such as the selection of candidate newspaper articles, the vectorial representation of both the segment and the articles, the choice of a suitable similarity measure, and the robustness to automatic segmentation errors. Experiments were conducted on various French TV Broadcast News shows recorded during one week, in conjunction with text articles collected through the Google News homepage at the same period. We introduce a full evaluation framework allowing the measurement of the quality of topic segment retrieval, topic title assignment and also joint retrieval and titling. The approach yields good titling performance and reveals to be robust to automatic segmentation. Abdessalam Bouchekif, Géraldine Damnati, Delphine Charlet, Nathalie Camelin, Yannick Estève |
ICASSP | 5 |
| 2016 | Acoustic Word Embeddings for ASR Error DetectionabstractInternational audience Sahar Ghannay, Yannick Estève, Nathalie Camelin, Paul Deléglise |
INTERSPEECH | 2 |
| 2016 | Conditional Random Fields for the Tunisian Dialect Grapheme-to-Phoneme Conversion
Abir Masmoudi 0001, Mariem Ellouze, Fethi Bougares, Yannick Estève, Lamia Hadrich Belguith |
INTERSPEECH | 4 |
| 2016 | On the Use of Gaussian Mixture Model Framework to Improve Speaker Adaptation of Deep Neural Network Acoustic ModelsabstractInternational audience Natalia A. Tomashenko, Yuri Y. Khokhlov, Yannick Estève |
INTERSPEECH | 3 |
| 2016 | Word Embedding Evaluation and Combination
Sahar Ghannay, Benoît Favre, Yannick Estève, Nathalie Camelin |
LREC | 3 |
| 2016 | Enhancing The RATP-DECODA Corpus With Linguistic Annotations For Performing A Large Range Of NLP Tasks
Carole Lailler, Anaïs Landeau, Frédéric Béchet, Yannick Estève, Paul Deléglise |
LREC | 4 |
| 2016 | LIUM ASR systems for the 2016 Multi-Genre Broadcast Arabic challengeabstractThis paper describes the automatic speech recognition (ASR) systems developed by LIUM in the framework of the 2016 Multi-Genre Broadcast (MGB-2) Challenge in the Arabic language. LIUM participated in the first of the two proposed tasks, namely the speech-to-text transcription of Aljazeera recordings. We present the approaches and details found in our systems, as well as our results in the evaluation campaign: the primary LIUM ASR system attained the second position. The main aspects come from the use of GMM-derived features for training a DNN, combined with the use of time-delay neural networks for acoustic models, the use of two different approaches in order to automatically phonetize Arabic words, and finally, the training data selection strategy for acoustic and language models. Natalia A. Tomashenko, Kévin Vythelingum, Anthony Rousseau, Yannick Estève |
SLT | 4 |
| 2015 | Multimodal data collection of human-robot humorous interactions in the Joker projectabstractThanks to a remarkably great ability to show amusement and engagement, laughter is one of the most important social markers in human interactions. Laughing together can actually help to set up a positive atmosphere and favors the creation of new relationships. This paper presents a data collection of social interaction dialogs involving humor between a human participant and a robot. In this work, interaction scenarios have been designed in order to study social markers such as laughter. They have been implemented within two automatic systems developed in the Joker project: a social dialog system using paralinguistic cues and a task-based dialog system using linguistic content. One of the major contributions of this work is to provide a context to study human laughter produced during a human-robot interaction. The collected data will be used to build a generic intelligent user interface which provides a multimodal dialog system with social communication skills including humor and other informal socially oriented behaviors. This system will emphasize the fusion of verbal and non-verbal channels for emotional and social behavior perception, interaction and generation capabilities. Laurence Devillers, Sophie Rosset, Guillaume Dubuisson Duplessis, Mohamed El Amine Sehili, Lucile Bechade, Agnès Delaborde, Clément Gossart, Vincent Letard, Fan Yang 0017, Yücel Yemez, Bekir Berker Türker, Tevfik Metin Sezgin, Kevin El Haddad, Stéphane Dupont, Daniel Luzzati, Yannick Estève, Emer Gilmartin, Nick Campbell 0001 |
ACII | 16 |
| 2015 | CRIM and LIUM approaches for multi-genre broadcast media transcriptionabstractThe Multi-Genre Broadcast Challenge at ASRU 2015 is a controlled evaluation of speech recognition, speaker diarization, and lightly supervised alignment using BBC TV recordings. CRIM and LIUM teams participated in the speech recognition part of the challenge with a joint submission. This paper presents the CRIM and LIUM's contributions. Each team made different choices to develop its ASR system. By the way, it was expected to compare and to evaluate different approaches to diarization and acoustic modeling, and to get complementary ASR systems for effective merging. CRIM's main contributions are the use of a training scenario similar to multi-lingual training to estimate the deep neural net (DNN) acoustic models with most of the data, the use of a pruned trigram model for search, in addition to the use of a genre-dependent quadgram language model for rescoring the lattice from the search. For LIUM, the focus was on fast decoding with high accuracy. The final word error rates (WER) after merging show that it is possible to get reasonable WER with automatically aligned files. The final global WER of 25.1% corresponds to a WER reduction of about 20% absolute in comparison to the ASR baseline system provided by the organizers. Vishwa Gupta, Paul Deléglise, Gilles Boulianne, Yannick Estève, Sylvain Meignier, Anthony Rousseau |
ASRU | 4 |
| 2015 | Arabic Transliteration of Romanized Tunisian Dialect Text: A Preliminary Investigation
Abir Masmoudi 0001, Nizar Habash, Mariem Ellouze, Yannick Estève, Lamia Hadrich Belguith |
CICLing (1) | 4 |
| 2015 | Diachronic semantic cohesion for topic segmentation of TV broadcast newsabstractInternational audience Abdessalam Bouchekif, Géraldine Damnati, Yannick Estève, Delphine Charlet, Nathalie Camelin |
INTERSPEECH | 3 |
| 2015 | Nao is doing humour in the CHIST-ERA joker project
Guillaume Dubuisson Duplessis, Lucile Bechade, Mohamed El Amine Sehili, Agnès Delaborde, Vincent Letard, Anne-Laure Ligozat, Paul Deléglise, Yannick Estève, Sophie Rosset, Laurence Devillers |
INTERSPEECH | 8 |
| 2014 | Is incremental cross-show speaker diarization efficient for processing large volumes of data?abstractInternational audience Grégor Dupuy, Sylvain Meignier, Yannick Estève |
INTERSPEECH | 3 |
| 2014 | A Corpus and Phonetic Dictionary for Tunisian Arabic Speech Recognition
Abir Masmoudi 0001, Mariem Ellouze, Yannick Estève, Lamia Hadrich Belguith, Nizar Habash |
LREC | 3 |
| 2014 | Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks
Anthony Rousseau, Paul Deléglise, Yannick Estève |
LREC | 3 |
| 2014 | Characterizing and detecting spontaneous speech: Application to speaker role recognition
Richard Dufour, Yannick Estève, Paul Deléglise |
Speech Commun. | 2 |
| 2013 | Blip10000: a social video dataset containing SPUG content for tagging and retrievalabstractThe increasing amount of digital multimedia content available is inspiring potential new types of user interaction with video data. Users want to easily find the content by searching and browsing. For this reason, techniques are needed that allow automatic categorisation, searching the content and linking to related information. In this work, we present a dataset that contains comprehensive semi-professional user-generated (SPUG) content, including audiovisual content, user-contributed metadata, automatic speech recognition transcripts, automatic shot boundary files, and social information for multiple 'social levels'. We describe the principal characteristics of this dataset and present results that have been achieved on different tasks. Sebastian Schmiedeke, Isabelle Ferrané, Maria Eskevich, Christoph Kofler, Martha A. Larson, Yannick Estève, Lori Lamel, Gareth J. F. Jones, Thomas Sikora |
MMSys | 7 |
| 2013 | Dynamic Combination of Automatic Speech Recognition Systems by Driven DecodingabstractCombining automatic speech recognition (ASR) systems generally relies on the posterior merging of the outputs or on acoustic cross-adaptation. In this paper, we propose an integrated approach where outputs of secondary systems are integrated in the search algorithm of a primary one. In this driven decoding algorithm (DDA), the secondary systems are viewed as observation sources that should be evaluated and combined to others by a primary search algorithm. DDA is evaluated on a subset of the ESTER I corpus consisting of 4 hours of French radio broadcast news. Results demonstrate DDA significantly outperforms vote-based approaches: we obtain an improvement of 14.5% relative word error rate over the best single-systems, as opposed to the the 6.7% with a ROVER combination. An in-depth analysis of the DDA shows its ability to improve robustness (gains are greater in adverse conditions) and a relatively low dependency on the search algorithm. The application of DDA to both and beam-search-based decoder yields similar performances. Benjamin Lecouteux, Georges Linarès, Yannick Estève, Guillaume Gravier |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Low latency combination of parallelized single-pass LVCSR systemsabstractInternational audience Fethi Bougares, Mickael Rouvier, Yannick Estève, Georges Linarès |
INTERSPEECH | 3 |
| 2012 | I-vectors and ILP clustering adapted to cross-show speaker diarizationabstractInternational audience Grégor Dupuy, Mickael Rouvier, Sylvain Meignier, Yannick Estève |
INTERSPEECH | 4 |
| 2012 | Leveraging study of robustness and portability of spoken language understanding systems across languages and domains: the PORTMEDIA corpora
Fabrice Lefèvre, Djamel Mostefa, Laurent Besacier, Yannick Estève, Matthieu Quignard, Nathalie Camelin, Benoît Favre, Bassam Jabaian, Lina Maria Rojas-Barahona |
LREC | 4 |
| 2012 | TED-LIUM: an Automatic Speech Recognition dedicated corpus
Anthony Rousseau, Paul Deléglise, Yannick Estève |
LREC | 3 |
| 2011 | Bag of n-gram driven decoding for LVCSR system harnessingabstractThis paper focuses on automatic speech recognition systems combination based on driven decoding paradigms. The driven decoding algorithm (DDA) involves the use of a 1-best hypothesis provided by an auxiliary system as another knowledge source in the search algorithm of a primary system. In previous studies, it was shown that DDA outperforms ROVER when the primary system is guided by a more accurate system. In this paper we propose a new method to manage auxiliary transcriptions which are presented as a bag-of-n-grams (BONG) without temporal matching. These modifications allow to make easier the combination of several hypotheses given by different auxiliary systems. Using BONG combination with hypotheses provided by two auxiliary systems, each of which obtained more than 23% of WER on the same data, our experiments show that a CMU Sphinx based ASR system can reduce its WER from 19.85% to 18.66% which is better than the results reached with DDA or classical ROVER combination. Fethi Bougares, Yannick Estève, Paul Deléglise, Georges Linarès |
ASRU | 2 |
| 2011 | Investigation of Spontaneous Speech Characterization Applied to Speaker Role RecognitionabstractExtracting information from large data is a challenging task. In this paper, we investigate the link between speech spontaneity levels and speaker roles, and the relevance to use an automatic spontaneous speech characterization as a speaker role identification feature. Applying this automatic spontaneous speech characterization system to a broadcast news corpus containing ten manually labeled speaker roles allowed us to highlight this relationship. So, we propose to directly apply the spontaneous speech characterization approach in order to automatically recognize speaker roles. Experimental results show that characteristics used to detect speech spontaneity could be very useful to recognize speaker roles, as we reached an overall classification precision of 74.4%. Richard Dufour, Yannick Estève, Paul Deléglise |
INTERSPEECH | 2 |
| 2010 | Unsupervised model adaptation on targeted speech segments for LVCSR system combinationabstractIn context of Large-Vocabulary Continuous Speech Recognition, systems can reach a high level of performance when dealing with prepared speech, while their performance drops on spontaneous speech. This decrease is due to the fact that these two kinds of speech are marked by strong acoustic and linguistic differences. Previous research works had been done to detect and repair some peculiarities of spontaneous speech, as disfluencies, and to create specific models to improve recognition accuracy: a large amount of data is needed to see improvements and is expensive to collect. In this paper, we present a solution to create specialized acoustic and language models, by automatically extracting a data subset from the initial training corpus containing spontaneous speech, and adapting initial acoustic and linguistic models on it. As we assume these models can be complementary, we propose to combine general and adapted ASR system outputs. Experimental results show statistically significant gain, for a negligible cost (no additional training data and no human intervention). Richard Dufour, Fethi Bougares, Yannick Estève, Paul Deléglise |
INTERSPEECH | 3 |
| 2010 | A language-identification inspired method for spontaneous speech detectionabstractInternational audience Mickael Rouvier, Richard Dufour, Georges Linarès, Yannick Estève |
INTERSPEECH | 4 |
| 2010 | Identification of Speakers by Name Using Belief Functions
Simon Petit-Renaud, Vincent Jousse, Sylvain Meignier, Yannick Estève |
IPMU (1) | 4 |
| 2010 | The EPAC Corpus: Manual and Automatic Annotations of Conversational Speech in French Broadcast News
Yannick Estève, Thierry Bazillon, Jean-Yves Antoine, Frédéric Béchet, Jérôme Farinas |
LREC | 1 |
| 2009 | Local and global models for spontaneous speech segment detection and characterizationabstractProcessing spontaneous speech is one of the many challenges that automatic speech recognition (ASR) systems have to deal with. The main evidences characterizing spontaneous speech are disfluencies (filled pause, repetition, repair and false start) and many studies have focused on the detection and the correction of these disfluencies. In this study we define spontaneous speech as unprepared speech, in opposition to prepared speech where utterances contain well-formed sentences close to those that can be found in written documents. Disfluencies are of course very good indicators of unprepared speech, however they are not the only ones: ungrammaticality and language register are also important as well as prosodic patterns. This paper proposes a set of acoustic and linguistic features that can be used for characterizing and detecting spontaneous speech segments from large audio databases. More, we introduce a strategy that takes advantage of a global classification procfalseess using a probabilistic model which significantly improves the spontaneous speech detection. Richard Dufour, Yannick Estève, Paul Deléglise, Frédéric Béchet |
ASRU | 2 |
| 2009 | Automatic named identification of speakers using diarization and ASR systemsabstractIn this paper, we consider the extraction of speaker identity from audio records of broadcast news without a priori acoustic information about speakers. Using an automatic speech recognition system and an automatic speaker diarization system, we present improvements for a method which allows to extract speaker identities from automatic transcripts and to assign them to speech segments. Experiments are carried out on French broadcast news records from the ESTER 1 evaluation campaign. Experimental results using outputs of automatic speech recognition and automatic diarization are presented. Vincent Jousse, Simon Petit-Renaud, Sylvain Meignier, Yannick Estève, Christine Jacquin |
ICASSP | 4 |
| 2009 | Iterative filtering of phonetic transcriptions of proper nounsabstractThis paper focuses on an approach to enhancing automatic phonetic transcription of proper nouns by using an iterative filter to retain only the most relevant part of a large set of phonetic variants, obtained by combining rule-based generation with extraction from actual audio signals. Using this technique, we were able to reduce the error rate affecting proper nouns during automatic speech transcription of the ESTER corpus of French broadcast news. The role of the filtering was to ensure that the new phonetic variants of proper nouns would not induce new errors in the transcription of the rest of the words. Antoine Laurent, Téva Merlin, Sylvain Meignier, Yannick Estève, Paul Deléglise |
ICASSP | 4 |
| 2009 | Improvements to the LIUM French ASR system based on CMU sphinx: what helps to significantly reduce the word error rate?abstractInternational audience Paul Deléglise, Yannick Estève, Sylvain Meignier, Téva Merlin |
INTERSPEECH | 2 |
| 2008 | Generalized driven decoding for speech recognition system combinationabstractDriven decoding algorithm (DDA) is initially an integrated approach for the combination of 2 speech recognition (ASR) systems. It consists in guiding the search algorithm of a primary ASR system by the one-best hypothesis of an auxiliary system. In this paper, we generalize DDA to confusion-network driven decoding and we propose new combination schemes for multiple system combination. Since previous experiments involved 2 ASR systems on broadcast news data, the proposed extended DDA is evaluated using 3 ASR systems from different labs. Results show that generalized- DDA outperforms significantly ROVER method: we obtain a 15.7% relative word error rate improvement with respect to the best single system, as opposed to 8.5% with the ROVER combination. Benjamin Lecouteux, Georges Linarès, Yannick Estève, Guillaume Gravier |
ICASSP | 3 |
| 2008 | Data selection and smoothing in an open-source system for the 2008 NIST machine translation evaluationabstractInternational audience Holger Schwenk, Yannick Estève |
INTERSPEECH | 2 |
| 2008 | Manual vs Assisted Transcription of Prepared and Spontaneous Speech
Thierry Bazillon, Yannick Estève, Daniel Luzzati |
LREC | 2 |
| 2008 | Combined Systems for Automatic Phonetic Transcription of Proper Nouns
Antoine Laurent, Téva Merlin, Sylvain Meignier, Yannick Estève, Paul Deléglise |
LREC | 4 |
| 2008 | Correcting asr outputs: Specific solutions to specific errors in FrenchabstractAutomatic speech recognition (ASR) systems are used in a large number of applications, in spite of the inevitable recognition errors. In this study we propose a pragmatic approach to automatically repair ASR outputs by taking into account linguistic and acoustic information, using formal rules or stochastic methods. The proposed strategy consists in developing a specific correction solution for each specific kind of errors. In this paper, we apply this strategy on two case studies specific to French language. We show that it is possible, on automatic transcriptions of French broadcast news, to decrease the error rate of a specific error by 11.4% in one of two the case studies, and 86.4% in the other one. These results are encouraging and show the interest of developing more specific solutions to cover a wider set of errors in a future work. Richard Dufour, Yannick Estève |
SLT | 2 |
| 2007 | System Combination by Driven DecodingabstractThe combination of automatic speech recognition (ASR) systems generally relies on a posteriori merge of system outputs or on a cross-adaptation. In this paper, we propose an integrated approach where the search of a primary system is driven by the outputs of a secondary one. This method allows to drive the primary system search by using the one-best hypotheses and the word posteriors gathered from the secondary system. Experiments are carried out within the experimental framework of the ESTER evaluation campaign (S. Galliano et al. 2005). Results show that the driven decoding algorithm significantly outperforms the two single ASR systems (-8% of relative WER, -1.7% absolute). Finally, we investigate the interactions between driven decoding and cross-adaptations. The best cross-adaptation strategy in combination with the driven decoding process brings to a final absolute gain of about 1.9% WER. Benjamin Lecouteux, Georges Linarès, Yannick Estève, Julie Mauclair |
ICASSP (4) | 3 |
| 2007 | Extracting true speaker identities from transcriptionsabstractInternational audience Yannick Estève, Sylvain Meignier, Paul Deléglise, Julie Mauclair |
INTERSPEECH | 1 |
| 2006 | Automatic Detection of Well Recognized Words in Automatic Speech Transcriptions
Julie Mauclair, Yannick Estève, Simon Petit-Renaud, Paul Deléglise |
LREC | 2 |
| 2005 | The LIUM speech transcription system: a CMU Sphinx III-based system for French broadcast newsabstractInternational audience Paul Deléglise, Yannick Estève, Sylvain Meignier, Téva Merlin |
INTERSPEECH | 2 |
| 2004 | Automatic learning of interpretation strategies for spoken dialogue systemsabstractThe paper proposes a new application of automatically trained decision trees to derive the interpretation of a spoken sentence. A new strategy for building structured cohorts of candidates is also described. By evaluating predicates related to the acoustic confidence of the words expressing a concept, the linguistic and semantic consistency of candidates in the cohort and the rank of a candidate within a cohort, the decision tree automatically learns a decision strategy for rescoring or rejecting an n-best list of candidates representing a user's utterance. A relative reduction of 18.6% in the understanding error rate is obtained by our rescoring strategy with no utterance rejection and a relative reduction of 43.1% of the same error rate is achieve with a rejection rate of only 8% of the utterances. Christian Raymond, Frédéric Béchet, Renato De Mori, Géraldine Damnati, Yannick Estève |
ICASSP (1) | 5 |
| 2003 | Conceptual decoding for spoken dialog systemsabstractInternational audience Yannick Estève, Christian Raymond, Frédéric Béchet, Renato De Mori |
INTERSPEECH | 1 |
| 2003 | On the use of linguistic consistency in systems for human-computer dialoguesabstractThis paper introduces new recognition strategies based on reasoning about results obtained with different Language Models (LMs). Strategies are built following the conjecture that the consensus among the results obtained with different models gives rise to different situations in which hypothesized sentences have different word error rates (WER) and may be further processed with other LMs. New LMs are built by data augmentation using ideas from latent semantic analysis and trigram analogy. Situations are defined by expressing the consensus among the recognition results produced with different LMs and by the amount of unobserved trigrams in the hypothesized sentence. The diagnostic power of the use of observed trigrams or their corresponding class trigrams is compared with that of situations based on values of sentence posterior probabilities. In order to avoid or correct errors due to syntactic inconsistence of the recognized sentence, automata, obtained by explanation-based learning, are introduced and used in certain conditions. Semantic Classification Trees are introduced to provide sentence patterns expressing constraints of long distance syntactic coherence. Results on a dialogue corpus provided by France Telecom R&D have shown that starting with a WER of 21.87% on a test set of 1422 sentences, it is possible to subdivide the sentences into three sets characterized by automatically recognized situations. The first one has a coverage of 68% with a WER of 7.44%. The second one has various types of sentences with a WER around 20%. The third one contains 13% of the sentences that should be rejected with a WER around 49%. The second set characterizes sentences that should be processed with particular care by the dialogue interpreter with the possibility of asking a confirmation from the user. Yannick Estève, Christian Raymond, Renato De Mori, David Janiszek |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | On the use of structures in language models for dialogueabstractInternational audience Renato De Mori, Yannick Estève, Christian Raymond |
INTERSPEECH | 2 |
| 2001 | Stochastic finite state automata language model triggered by dialogue statesabstractWithin the framework of Natural Spoken Dialogue systems, this paper describes a method for dynamically adapting a Language Model (LM) to the dialogue states detected. This LM combines a standard n-gram model with Stochastic Finite State Automata (SFSAs). During the training process, the sentence corpus used to train the LM is split into several hierarchical clusters in a 2-step process which involves both explicit knowledge and statistical criteria. All the clusters are stored in a binary tree where the whole corpus is attached to the root node. Each level of the tree corresponds to a higher specialization of the sub-corpora attached to the nodes and each node corresponds to a different dialogue state. From the same sentence corpus, SFSAs are extracted in order to model longer contexts than the ones used in the standard n-gram model. A set of SFSAs is attached to each node of the tree as well as a sub-LM which combines a bigram trained on the sub-corpus of the node and the SFSAs selected. A first decoding process calculates a word-graph as well as a first sentence hypothesis. This first hypothesis will be used to find the optimal node in the LM tree. Then, a rescoring process of the word graph using the LM attached to the node selected is performed. By adapting the LM to the dialogue state detected, we show a statistically significant gain in WER on a dialogue corpus collected by France Telecom R&D . Yannick Estève, Frédéric Béchet, Alexis Nasr, Renato De Mori |
INTERSPEECH | 1 |
| 2000 | Dynamic selection of language models in a dialogue systemabstractThis paper describes a method for building statistical Language Models (LMs) dedicated to specific dialogue situations. The architecture of the speech recognition system proposed uses several LMs. The first stage of this system, consists of producing a word-lattice from a given sentence uttered by a speaker. A general LM calculates a sentence-hypothesis. Then, in a second stage, the system chooses a specialized LM according to the word-lattice and the previous hypothesis. Another decoding process is performed using this specialized LM in order to produce a new sentence- hypothesis. Finally, a decision-module processes these two hypotheses in order to assign three confidence levels to the sentence-hypothesis produced. These confidence levels can be used by the dialogue manager in order to improve the dialogue, by asking a confirmation to the speaker when a sentence is labeled ambiguous. This research is supported by France Telecom's R&D under the contract 971B427. Yannick Estève, Frédéric Béchet, Renato De Mori |
INTERSPEECH | 1 |
| 1999 | A language model combining n-grams and stochastic finite state automataabstractMaximum a posteriori adaptation method combines the prior knowledge with adaptation data from a new speaker, which has a nice asymptotical property, but has a slow adaptation rate for not modifying unseen models. In a strictly Bayesian approach, prior parameters are assumed known, based on common or subjective knowledge. But a practical solution is to adopt an empirical Bayesian approach, where the prior parameters are estimated directly from training speech data itself. So there is a problem of mismatches between training and testing conditions. In this paper we propose a prior parameter transformation (PPT) adaptation approach that transforms the prior parameters to be more representative of the new speaker. It can influence unseen models by tying prior parameter transformations across different models according to amount of adaptation data available. Based on the improved prior information better model parameters can be obtained even with small amount of adaptation data. Alexis Nasr, Yannick Estève, Frédéric Béchet, Thierry Spriet, Renato De Mori |
EUROSPEECH | 2 |