Frédéric Béchet

dblp:54/4944 · DBLP profile ↗
← Back
100ranked-venue papers
24as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 16 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 63 · 19 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CareMedEval Dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
abstract
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliability remains limited, particularly for critical reasoning in specialized domains. We introduce CareMedEval, an original dataset designed to evaluate LLMs on biomedical critical appraisal and reasoning tasks. Derived from authentic exams taken by French medical students, the dataset contains 534 questions based on 37 scientific articles. Unlike existing benchmarks, CareMedEval explicitly evaluates critical reading and reasoning grounded in scientific papers. Benchmarking state-of-the-art generalist and biomedical-specialized LLMs under various context conditions reveals the difficulty of the task: open and commercial models fail to exceed an Exact Match Rate of 0.5 even though generating intermediate reasoning tokens considerably improves the results. Yet, models remain challenged especially on questions about study limitations and statistical analysis. CareMedEval provides a challenging benchmark for grounded reasoning, exposing current LLM limitations and paving the way for future development of automated support for critical appraisal.
Doria Bonzi, Alexandre Guiggi, Frédéric Béchet, Carlos Ramisch, Benoît Favre
LREC3
2025 Statistical Deficiency for Task Inclusion Estimation
abstract
Loïc Fosse, Frederic Bechet, Benoit Favre, Géraldine Damnati, Gwénolé Lecorvé, Maxime Darrin, Philippe Formont, Pablo Piantanida. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Loïc Fosse, Frédéric Béchet, Benoît Favre, Géraldine Damnati, Gwénolé Lecorvé, Maxime Darrin, Philippe Formont, Pablo Piantanida
ACL (1)2
2025 Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
abstract
This study investigates the behavior of model-integrated routers in Mixture of Experts (MoE) models, focusing on how tokens are routed based on their linguistic features, specifically Part-of-Speech (POS) tags. The goal is to explore across different MoE architectures whether experts specialize in processing tokens with similar linguistic traits. By analyzing token trajectories across experts and layers, we aim to uncover how MoE models handle linguistic information. Findings from six popular MoE models reveal expert specialization for specific POS categories, with routing paths showing high predictive accuracy for POS, highlighting the value of routing paths in characterizing tokens.
Elie Antoine, Frédéric Béchet, Philippe Langlais
COLING2
2025 Factual Knowledge Assessment of Language Models Using Distractors
abstract
Language models encode extensive factual knowledge within their parameters. The accurate assessment of this knowledge is crucial for understanding and improving these models. In the literature, factual knowledge assessment often relies on cloze sentences, which can lead to erroneous conclusions due to the complexity of natural language (out-of-subject continuations, the existence of many correct answers and the several ways of expressing them). In this paper, we introduce a new interpretable knowledge assessment method that mitigates these issues by leveraging distractors—incorrect but plausible alternatives to the correct answer. We propose several strategies for retrieving distractors and determine the most effective one through experimentation. Our method is evaluated against existing approaches, demonstrating solid alignment with human judgment and stronger robustness to verbalization artifacts. The code and data to reproduce our experiments are available on GitHub.
Hichem Ammar Khodja, Abderrahmane Ait gueni ssaid, Frédéric Béchet, Quentin Brabant, Alexis Nasr, Gwénolé Lecorvé
COLING3
2024 WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
abstract
The factuality of large language model (LLMs) tends to decay over time since events posterior to their training are “unknown” to them. One way to keep models up-to-date could be factual update: the task of inserting, replacing, or removing certain simple (atomic) facts within the model. To study this task, we present WikiFactDiff, a dataset that describes the evolution of factual knowledge between two dates as a collection of simple facts divided into three categories: new, obsolete, and static. We describe several update scenarios arising from various combinations of these three types of basic update. The facts are represented by subject-relation-object triples; indeed, WikiFactDiff was constructed by comparing the state of the Wikidata knowledge base at 4 January 2021 and 27 February 2023. Those fact are accompanied by verbalization templates and cloze tests that enable running update algorithms and their evaluation metrics. Contrary to other datasets, such as zsRE and CounterFact, WikiFactDiff constitutes a realistic update setting that involves various update scenarios, including replacements, archival, and new entity insertions. We also present an evaluation of existing update algorithms on WikiFactDiff.
Hichem Ammar Khodja, Frédéric Béchet, Quentin Brabant, Alexis Nasr, Gwénolé Lecorvé
LREC/COLING2
2024 A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
abstract
We introduce an evaluation methodology for reading comprehension tasks based on the intuition that certain examples, by the virtue of their linguistic complexity, consistently yield lower scores regardless of model size or architecture.We capitalize on semantic frame annotation for characterizing this complexity, and study seven complexity factors that may account for model's difficulty.We first deploy this methodology on a carefully annotated French reading comprehension benchmark showing that two of those complexity factors are indeed good predictors of models' failure, while others are less so.We further deploy our methodology on a well studied English benchmark by using Chat-GPT as a proxy for semantic annotation.Our study reveals that fine-grained linguisticallymotivated automatic evaluation of a reading comprehension task is not only possible, but helps understand models' abilities to handle specific linguistic characteristics of input examples.It also shows that current state-of-the-art models fail with some for those characteristics which suggests that adequately handling them requires more than merely increasing model size.
Elie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe Langlais
EMNLP2
2024 Unified Framework for Spoken Language Understanding and Summarization in Task-Based Human Dialog processing
abstract
International audience
Eunice Akani, Frédéric Béchet, Benoît Favre, Romain Gemignani
INTERSPEECH2
2023 Investigating the Effect of Relative Positional Embeddings on AMR-to-Text Generation with Structural Adapters
abstract
Sebastien Montella, Alexis Nasr, Johannes Heinecke, Frederic Bechet, Lina M. Rojas Barahona. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Sébastien Montella, Alexis Nasr, Johannes Heinecke, Frédéric Béchet, Lina Maria Rojas-Barahona
EACL4
2023 Abstract Representation for Multi-Intent Spoken Language Understanding
abstract
Current sequence tagging models based on Deep Neural Network models with pretrained language models achieve almost perfect results on many SLU benchmarks with a flat semantic annotation at the token level such as ATIS or SNIPS. When dealing with more complex human-machine interactions (multi-domain, multi-intent, dialog context), relational semantic structures are needed in order to encode the links between slots and intents within an utterance and through dialog history. We propose in this study a new way to project annotation in an abstract structure with more compositional expressive power and a model to directly generate this abstract structure. We evaluate it on the MultiWoz dataset in a contextual SLU experimental setup. We show that this projection can be used to extend the existing flat annotations towards graph-based structures.
Rim Abrougui, Géraldine Damnati, Johannes Heinecke, Frédéric Béchet
ICASSP4
2023 Reducing named entity hallucination risk to ensure faithful summary generation
abstract
The faithfulness of abstractive text summarization at the named entities level is the focus of this study.We propose to add a new criterion to the summary selection method based on the "risk" of generating entities that do not belong to the source document.This method is based on the assumption that Out-Of-Document entities are more likely to be hallucinations.This assumption was verified by a manual annotation of the entities occurring in a set of generated summaries on the CNN/DM corpus.This study showed that only 29% of the entities outside the source document were inferrable by the annotators, leading to 71% of hallucinations among OOD entities.We test our selection method on the CNN/DM corpus and show that it significantly reduces the hallucination risk on named entities while maintaining competitive results with respect to automatic evaluation metrics like ROUGE.
Eunice Akani, Benoît Favre, Frédéric Béchet, Romain Gemignani
INLG3
2022 Question Generation and Answering for exploring Digital Humanities collections
abstract
This paper introduces the question answering paradigm as a way to explore digitized archive collections for Social Science studies. In particular, we are interested in evaluating largely studied question generation and question answering approaches on a new type of documents, as a step forward beyond traditional benchmark evaluations. Question generation can be used as a way to provide enhanced training material for Machine Reading Question Answering algorithms but also has its own purpose in this paradigm, where relevant questions can be used as a way to create explainable links between documents. To this end, generating large amounts of question is not the only motivation, but we need to include qualitative and semantic control to the generation process. We propose a new approach for question generation, relying on a BART Transformer based generative model, for which input data are enriched by semantic constraints. Question generation and answering are evaluated on several French corpora, and the whole approach is validated on a new corpus of digitized archive collection of a French Social Science journal.
Frédéric Béchet, Elie Antoine, Jérémy Auguste, Géraldine Damnati
LREC1
2021 Predicting Links on Wikipedia with Anchor Text Information
abstract
Wikipedia, the largest open-collaborative online encyclopedia, is a corpus of documents bound together by internal hyperlinks. These links form the building blocks of a large network whose structure contains important information on the concepts covered in this encyclopedia. The presence of a link between two articles, materialised by an anchor text in the source page pointing to the target page, can increase readers' understanding of a topic. However, the process of linking follows specific editorial rules to avoid both under-linking and over-linking. In this paper, we study the transductive and the inductive tasks of link prediction on several subsets of the English Wikipedia and identify some key challenges behind automatic linking based on anchor text information. We propose an appropriate evaluation sampling methodology and compare several algorithms. Moreover, we propose baseline models that provide a good estimation of the overall difficulty of the tasks.
Robin Brochier, Frédéric Béchet
SIGIR2
2021 Multimodal Machine Learning for Natural Language Processing: Disambiguating Prepositional Phrase Attachments with Images
Sebastien Delecraz, Leonor Becerra-Bonache, Benoît Favre, Alexis Nasr, Frédéric Béchet
Neural Process. Lett.5
2020 Cross-lingual and Cross-domain Evaluation of Machine Reading Comprehension with Squad and CALOR-Quest Corpora
abstract
Machine Reading received recently a lot of attention thanks to both the availability of very large corpora such as SQuAD or MS MARCO containing triplets (document, question, answer), and the introduction of Transformer Language Models such as BERT which obtain excellent results, even matching human performance according to the SQuAD leaderboard. One of the key features of Transformer Models is their ability to be jointly trained across multiple languages, using a shared subword vocabulary, leading to the construction of cross-lingual lexical representations. This feature has been used recently to perform zero-shot cross-lingual experiments where a multilingual BERT model fine-tuned on a machine reading comprehension task exclusively for English was directly applied to Chinese and French documents with interesting performance. In this paper we study the cross-language and cross-domain capabilities of BERT on a Machine Reading Comprehension task on two corpora: SQuAD and a new French Machine Reading dataset, called CALOR-QUEST. The semantic annotation available on CALOR-QUEST allows us to give a detailed analysis on the kinds of questions that are properly handled through the cross-language process. We will try to answer this question: which factor between language mismatch and domain mismatch has the strongest influence on the performances of a Machine Reading Comprehension task?
Delphine Charlet, Géraldine Damnati, Frédéric Béchet, Gabriel Marzinotto, Johannes Heinecke
LREC3
2019 Can We Predict Self-reported Customer Satisfaction from Interactions?
abstract
In the context of contact centers, customers' satisfaction after a conversation with an agent is a critical issue which has to be collected in order to detect problems and improve quality of service. Automatically predicting customer satisfaction directly from system logs, without any survey or manual annotation is a challenging task of a great interest for the field of human-human conversation understanding and for improving contact center quality of service. Unlike previous studies that have focused on questions directly related to the content of a conversation, we look at a more general opinion about a service which is called the "Net Promoter Score" (NPS) where customers are considered either as promoters, detractors or neutral. On a very large corpus of chat-conversations with customer satisfaction surveys, we explore several classification scheme in order to achieve this prediction task, only using conversation logs.
Jérémy Auguste, Delphine Charlet, Géraldine Damnati, Frédéric Béchet, Benoît Favre
ICASSP4
2019 Benchmarking Benchmarks: Introducing New Automatic Indicators for Benchmarking Spoken Language Understanding Corpora
abstract
Empirical evaluation is nowadays the main evaluation paradigm in Natural Language Processing for assessing the relevance of a new machine-learning based model. If large corpora are available for tasks such as Automatic Speech Recognition , this is not the case for other tasks such as Spoken Language Understanding (SLU), consisting in translating spoken transcriptions into a formal representation often based on semantic frames. Corpora such as ATIS or SNIPS are widely used to compare systems, however differences in performance among systems are often very small, not statistically significant , and can be produced by biases in the data collection or the annotation scheme, as we presented on the ATIS corpus (Is ATIS too shallow?, IS2018). We propose in this study a new methodology for assessing the relevance of an SLU corpus. We claim that only taking into account systems performance does not provide enough insight about what is covered by current state-of-the-art models and what is left to be done. We apply our methodology on a set of 4 SLU systems and 5 benchmark corpora (ATIS, SNIPS, M2M, MEDIA) and automatically produce several indicators assessing the relevance (or not) of each corpus for benchmarking SLU models.
Frédéric Béchet, Christian Raymond
INTERSPEECH1
2019 Adapting a FrameNet Semantic Parser for Spoken Language Understanding Using Adversarial Learning
abstract
International audience
Gabriel Marzinotto, Géraldine Damnati, Frédéric Béchet
INTERSPEECH3
2018 Is ATIS Too Shallow to Go Deeper for Benchmarking Spoken Language Understanding Models?
abstract
The ATIS (Air Travel Information Service) corpus will be soon celebrating its 30th birthday. Designed originally to benchmark spoken language systems, it still represents the most well-known corpus for benchmarking Spoken Language Understanding (SLU) systems. In 2010, in a paper titled What is left to be understood in ATIS? [1], Tur et al. discussed the relevance of this corpus after more than 10 years of research on statistical models for performing SLU tasks. Nowadays, in the Deep Neural Network (DNN) era, ATIS is still used as the main benchmark corpus for evaluating all kinds of DNN models, leading to further improvements, although rather limited, in SLU accuracy compared to previous state-of-the-art models. We propose in this paper to investigate these results obtained on ATIS from a qualitative point of view rather than just a quantitative point of view and answer the two following questions: what kind of qualitative improvement brought DNN models to SLU on the ATIS corpus? Is there anything left, from a qualitative point of view, in the remaining 5% of errors made by current state-of-the-art models?
Frédéric Béchet, Christian Raymond
INTERSPEECH1
2018 Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text
Géraldine Damnati, Jérémy Auguste, Alexis Nasr, Delphine Charlet, Johannes Heinecke, Frédéric Béchet
LREC6
2018 Adding Syntactic Annotations to Flickr30k Entities Corpus for Multimodal Ambiguous Prepositional-Phrase Attachment Resolution
Sebastien Delecraz, Alexis Nasr, Frédéric Béchet, Benoît Favre
LREC3
2018 Semantic Frame Parsing for Information Extraction : the CALOR corpus
Gabriel Marzinotto, Jérémy Auguste, Frédéric Béchet, Géraldine Damnati, Alexis Nasr
LREC3
2016 Joint Syntactic and Semantic Analysis with a Multitask Deep Learning Framework for Spoken Language Understanding
abstract
International audience
Jérémie Tafforeau, Frédéric Béchet, Thierry Artières, Benoît Favre
INTERSPEECH2
2016 Beyond Utterance Extraction: Summary Recombination for Speech Summarization
abstract
International audience
Jérémy Trione, Benoît Favre, Frédéric Béchet
INTERSPEECH3
2016 Summarizing Behaviours: An Experiment on the Annotation of Call-Centre Conversations
Morena Danieli, A. R. Balamurali, Evgeny A. Stepanov, Benoît Favre, Frédéric Béchet, Giuseppe Riccardi
LREC5
2016 Enhancing The RATP-DECODA Corpus With Linguistic Annotations For Performing A Large Range Of NLP Tasks
Carole Lailler, Anaïs Landeau, Frédéric Béchet, Yannick Estève, Paul Deléglise
LREC3
2016 Syntactic parsing of chat language in contact center conversation corpus
abstract
Chat language is often referred to as Computer-mediated communication (CMC).Most of the previous studies on chat language has been dedicated to collecting "chat room" data as it is the kind of data which is the most accessible on the WEB.This kind of data falls under the informal register whereas we are interested in this paper in understanding the mechanisms of a more formal kind of CMC: dialog chat in contact centers.The particularities of this type of dialogs and the type of language used by customers and agents is the focus of this paper towards understanding this new kind of CMC data.The challenges for processing chat data comes from the fact that Natural Language Processing tools such as syntactic parsers and part of speech taggers are typically trained on mismatched conditions, we describe in this study the impact of such a mismatch for a syntactic parsing task.
Alexis Nasr, Géraldine Damnati, Aleksandra Guerraz, Frédéric Béchet
SIGDIAL Conference4
2015 Multimodal embedding fusion for robust speaker role recognition in video broadcast
abstract
Person role recognition in video broadcasts consists in classifying people into roles such as anchor, journalist, guest, etc. Existing approaches mostly consider one modality, either audio (speaker role recognition) or image (shot role recognition), firstly because of the non-synchrony between both modalities, and secondly because of the lack of a video corpus annotated in both modalities. Deep Neural Networks (DNN) approaches offer the ability to learn simultaneously feature representations (embeddings) and classification functions. This paper presents a multimodal fusion of audio, text and image embeddings spaces for speaker role recognition in asynchronous data. Monomodal embeddings are trained on exogenous data and fine-tuned using a DNN on 70 hours of French Broadcasts corpus for the target task. Experiments on the REPERE corpus show the benefit of the embeddings level fusion compared to the monomodal embeddings systems and to the standard late fusion method.
Mickael Rouvier, Sebastien Delecraz, Benoît Favre, Meriem Bendris, Frédéric Béchet
ASRU5
2015 "speech is silver, but silence is golden": improving speech-to-speech translation performance by slashing users input
abstract
Speech-to-speech translation is a challenging task mixing two of the most ambitious Natural Language Processing challenges: Machine Translation (MT) and Automatic Speech Recognition (ASR). Recent advances in both fields have led to operational systems achieving good performance when used in matching conditions with those of ASR and MT models training. Regardless of the quality of these models, errors are inevitable due to some technical limitations of the systems (e.g. closed vocabulary) and intrinsic ambiguities of spoken languages. However all ASR and MT errors don’t have the same impact on the usability of a given speech-to-speech dialog system: some can be very benign, unconsciously corrected by users, some can damage the understanding between users and eventually lead the dialog to a failure. We present in this paper a strategy focusing on ASR error segments that have a high negative impact on MT performance. We propose a method that consists firstly in automatically detecting these erroneous segments then secondly estimating their impact on MT. We show that removing such segments prior to translation can lead to a significant decrease in translation error rate, even without any correction strategy.
Frédéric Béchet, Benoît Favre, Mickael Rouvier
INTERSPEECH1
2015 Adapting lexical representation and OOV handling from written to spoken language with word embedding
abstract
Word embeddings have become ubiquitous in NLP, especially when using neural networks. One of the assumptions of such representations is that words with similar properties have similar representation, allowing for better generalization from subsequent models. In the standard setting, two kinds of training corpora are used: a very large unlabeled corpus for learning the word embedding representations; and an in-domain training corpus with gold labels for training classifiers on the target NLP task. Because of the amount of data required to learn embeddings, they are trained on large corpus of written text. This can be an issue when dealing with non-canonical language, such as spontaneous speech: embeddings have to be adapted to fit the particularities of spoken transcriptions. However the adaptation corpus available for a given speech application can be limited, resulting in a high number of words from the embedding space not occurring in the adaptation space. We present in this paper a method for adapting an embedding space trained on written text to a spoken corpus of limited size. In particular we deal with words from the embedding space not occurring in the adaptation data. We report experiments done on a Part-OfSpeech task on spontaneous speech transcriptions collected in a call-centre. We show that our word embedding adaptation approach outperforms state-of-the-art Conditional Random Field approach when little in-domain adaptation data is available.
Jérémie Tafforeau, Thierry Artières, Benoît Favre, Frédéric Béchet
INTERSPEECH4
2015 Call Centre Conversation Summarization: A Pilot Task at Multiling 2015
abstract
This paper describes the results of the Call Centre Conversation Summarization task at Multiling'15.The CCCS task consists in generating abstractive synopses from call centre conversations between a caller and an agent.Synopses are summaries of the problem of the caller, and how it is solved by the agent.Generating them is a very challenging task given that deep analysis of the dialogs and text generation are necessary.Three languages were addressed: French, Italian and English translations of conversations from those two languages.The official evaluation metric was ROUGE-2.Two participants submitted a total of four systems which had trouble beating the extractive baselines.The datasets released for the task will allow more research on abstractive dialog summarization.
Benoît Favre, Evgeny A. Stepanov, Jérémy Trione, Frédéric Béchet, Giuseppe Riccardi
SIGDIAL Conference4
2014 Retrieving the syntactic structure of erroneous ASR transcriptions for open-domain Spoken Language Understanding
abstract
Retrieving the syntactic structure of erroneous ASR transcriptions can be of great interest for open-domain Spoken Language Understanding tasks in order to correct or at least reduce the impact of ASR errors on final applications. Most of the previous works on ASR and syntactic parsing have addressed this problem by using syntactic features during ASR to help reducing Word Error Rate (WER). The improvement obtained is often rather small, however the structure and the relations between words obtained through parsing can be of great interest for the SLU processes, even without a significant decrease of WER. That is why we adopt another point of view in this paper: considering that ASR transcriptions contain inevitably some errors, we show in this study that it is possible to improve the syntactic analysis of these erroneous transcriptions by performing a joint error detection / syntactic parsing process. The applicative framework used in this study is a speech-to-speech system developed through the DARPA BOLT project.
Frédéric Béchet, Benoît Favre, Alexis Nasr, Mathieu Morey
ICASSP1
2014 Reranked aligners for interactive transcript correction
abstract
Clarification dialogs can help address ASR errors in speech-to-speech translation systems and other interactive applications. We propose to use variants of Levenshtein alignment for merging an er-rorful utterance with a targeted rephrase of an error segment. ASR errors that might harm the alignment are addressed through phonetic matching, and a word embedding distance is used to account for the use of synonyms outside targeted segments. These features lead to a relative improvement of 30% of word error rate on sentences with ASR errors compared to not performing the clarification. Twice as many utterances are completely corrected compared to using basic word alignment. Furthermore, we generate a set of potential merges and train a neural network on crowd-sourced rephrases in order to select the best merger, leading to 24% more instances completely corrected. The system is deployed in the framework of the BOLT project.
Benoît Favre, Mickael Rouvier, Frédéric Béchet
ICASSP3
2014 Multimodal understanding for person recognition in video broadcasts
abstract
International audience
Frédéric Béchet, Meriem Bendris, Delphine Charlet, Géraldine Damnati, Benoît Favre, Mickael Rouvier, Rémi Auguste, Benjamin Bigot, Richard Dufour, Corinne Fredouille, Georges Linarès, Jean Martinet, Grégory Senay, Pierre Tirilly
INTERSPEECH1
2014 Adapting dependency parsing to spontaneous speech for open domain spoken language understanding
abstract
Parsing human-human conversations consists in automatically enriching text transcription with semantic structure information. We use in this paper a FrameNet-based approach to semantics that, without needing a full semantic parse of a message, goes further than a simple flat translation of a message into basic concepts. FrameNet-based semantic parsing may follow a syntactic parsing step, however spoken conversations in customer service telephone call centers present very specific characteristics such as non-canonical language, noisy messages (disfluencies, repetitions, truncated words or automatic speech transcription errors) and the presence of superfluous information. For syntactic parsing the traditional view based on context-free grammars is not suitable for processing non-canonical text. New approaches to parsing based on dependency structures and discriminative machine learning techniques are more adapted to process spontaneous speech for two main reasons: (a) they need less training data and (b) the annotation with syntactic dependencies of conversation transcripts is simpler than with syntactic constituents. Another advantage is that partial annotation can be performed. This paper presents the adaptation of a syntactic dependency parser to process very spontaneous speech recorded in a callcentre environment. This parser is used in order to produce FrameNet candidates for characterizing conversations between an operator and a caller.
Frédéric Béchet, Alexis Nasr, Benoît Favre
INTERSPEECH1
2014 A Collection of Scholarly Book Reviews from the Platforms of electronic sources in Humanities and Social Sciences OpenEdition.org
Chahinez Benkoussas, Hussam Hamdan, Patrice Bellot, Frédéric Béchet, Elodie Faath
LREC4
2014 Automatically enriching spoken corpora with syntactic information for linguistic studies
Alexis Nasr, Frédéric Béchet, Benoît Favre, Thierry Bazillon, José Deulofeu, André Valli
LREC2
2014 Joint decoding of complementary utterances
abstract
Errors in open-domain ASR can be corrected by asking the speaker to rephrase targeted segments in utterances where they have been detected. The utterance merging problem consists in generating a better transcript from the utterance where errors have been detected and a clarification utterance. We introduce an alignment-decoding algorithm for jointly processing the two utterances and benefit from the complementary information they contain. The algorithm aligns word lattices in the WFST framework with a probabilistic cost model. Results on the BOLT-BC speech-to-speech translation task show an improvement of 2.84 points of accuracy compared to aligning the one best without joint decoding.
Mickael Rouvier, Benoît Favre, Frédéric Béchet
SLT3
2013 "Can you give me another word for hyperbaric?": Improving speech translation using targeted clarification questions
abstract
We present a novel approach for improving communication success between users of speech-to-speech translation systems by automatically detecting errors in the output of automatic speech recognition (ASR) and statistical machine translation (SMT) systems. Our approach initiates system-driven targeted clarification about errorful regions in user input and repairs them given user responses. Our system has been evaluated by unbiased subjects in live mode, and results show improved success of communication between users of the system.
Necip Fazil Ayan, Arindam Mandal, Michael W. Frandsen, Jing Zheng 0001, Peter Blasco, Andreas Kathol, Frédéric Béchet, Benoît Favre, Alex Marin, Tom Kwiatkowski, Mari Ostendorf, Luke Zettlemoyer, Philipp Salletmayr, Julia Hirschberg, Svetlana Stoyanchev
ICASSP7
2013 ASR error segment localization for spoken recovery strategy
abstract
Even though small ASR errors might not impact downstream processes that make use of the transcript, larger error segments like those generated by OOVs can have a considerable impact on applications such as speech-to-speech translation and can eventually lead to communication failure between users of the system. This work focuses on error detection in ASR output targeted towards significant error segments that can be recovered using a dialog system. We propose a CRF system trained to recognize error segments with ASR confidence-based, lexical and syntactic features. The most significant error segment is passed to a dialog system for interactive recovery in which rephrased words are reinserted in the original. 22% of utterances can be fully recovered and an interesting by-product is that rewriting error segments as a single token reduces WER by 17% on an adverse corpus.
Frédéric Béchet, Benoît Favre
ICASSP1
2012 Detecting person presence in TV shows with linguistic and structural features
abstract
Person detection and recognition in videos is a hard problem due to the intrinsic ambiguities of the sound and image channels and their interaction. Whatever method is used to extract person hypotheses from the audio or the image channels, person recognition in videos relies on a multimodal decision process that merges the different hypotheses produced in order to decide, for each frame, who is present in the video at the audio level, at the image level or at the content level (person mention in speech or inserted text boxes). In this framework the focus of this paper is to produce a list of person presence hypotheses from the audio channel of a video document only, to be used in addition to person presence detected at the image level by a multimodal fusion process. In this study we focus on the audio channel only, using two kinds of features: linguistic features corresponding to the way a person is mentioned by a speaker; structural features corresponding to the context of occurrence of a name in a show. We show that both sets of features are complementary and that good results can be achieved on a TV show corpus annotated with person presence labels.
Frédéric Béchet, Benoît Favre, Géraldine Damnati
ICASSP1
2012 Automatic transcription error recovery for Person Name Recognition
abstract
Person Name Recognition from transcriptions of TV shows spoken content is a crucial step towards multimedia document indexing. Recognizing Person Names implies the combination of three main modules: Automatic Speech Recognition, Named-Entity Recognition and Entity Linking to associate the recognized surface form to a normalized Person Name. The three modules are potentially error prone. Hence, beyond each module's intrinsic complexity, the Person Names issue suffers from the highly dynamic evolution of vocabularies and occurrence contexts that are correlated to various dimensions (such as actuality, topic of the show…). This paper focuses on the first module and proposes an approach to recover from transcription errors made on Person Names. An error correction method is applied on the textual ASR output and we show that it is all the more efficient that it is coupled with a specific error region detection system. Experiments on the French REPERE database show that Person Names transcription can be efficiently corrected while preserving the overall transcription quality and thus increasing the performance of the whole Person Name Recognition process.
Richard Dufour, Géraldine Damnati, Delphine Charlet, Frédéric Béchet
INTERSPEECH4
2012 Applying multiview learning algorithms to human-human conversation classification
Sokol Koço, Cécile Capponi, Frédéric Béchet
INTERSPEECH3
2012 Syntactic annotation of spontaneous speech: application to call-center conversation data
Thierry Bazillon, Melanie Deplano, Frédéric Béchet, Alexis Nasr, Benoît Favre
LREC3
2012 DECODA: a call-centre human-human spoken conversation corpus
Frédéric Béchet, Benjamin Maza, Nicolas Bigouroux, Thierry Bazillon, Marc El-Bèze, Renato De Mori, Eric Arbillot
LREC1
2011 Applying Multiclass Bandit algorithms to call-type classification
abstract
We analyze the problem of call-type classification using data that is weakly labelled. The training data is not systematically annotated, but we consider we have a weak or lazy oracle able to answer the question “Is sample x of class q?” by a simple `yes' or `no' answer. This situation of learning might be encountered in many real-world problems where the cost of labelling data is very high. We prove that it is possible to learn linear classifiers in this setting, by estimating adequate expectations inspired by the Multiclass Bandit paradgim. We propose a learning strategy that builds on Kessler's construction to learn multiclass perceptrons. We test our learning procedure against two real-world datasets from spoken langage understanding and provide compelling results.
Liva Ralaivola, Benoît Favre, Pierre Gotab, Frédéric Béchet, Géraldine Damnati
ASRU4
2011 Speaker Role Recognition Using Question Detection and Characterization
Thierry Bazillon, Benjamin Maza, Mickael Rouvier, Frédéric Béchet, Alexis Nasr
INTERSPEECH4
2010 Unsupervised knowledge acquisition for Extracting Named Entities from speech
abstract
This paper presents a Named Entity Recognition (NER) method dedicated to process speech transcriptions. The main principle behind this method is to collect in an unsupervised way lexical knowledge for all entries in the ASR lexicon. This knowledge is gathered with two methods: by automatically extracting NEs on a very large set of textual corpora and by exploiting directly the structure contained in the Wikipedia resource. This lexical knowledge is used to update the statistical models of our NER module based on a mixed approach with generative models (Hidden Markov Models - HMM) and discriminative models (Conditional Random Field - CRF). This approach has been evaluated within the French ESTER 2 evaluation program and obtained the best results at the NER task on ASR transcripts.
Frédéric Béchet, Eric Charton
ICASSP1
2010 On the use of machine translation for spoken language understanding portability
abstract
Across language portability of a spoken language understanding system (SLU) deals with the possibility of reusing with moderate effort in a new language knowledge and data acquired for another language. The approach proposed in this paper is motivated by the availability of the fairly large MEDIA corpus carefully transcribed in French and semantically annotated in terms of constituents. A method is proposed for manually translating a portion of the training set for training an automatic machine translation (MT) system to be used for translating the remaining data. As the source language is annotated in terms of concept tags, a solution is presented for automatically transferring these tags to the translated corpus. Experimental results are presented on the accuracy of the translation expressed with the BLEU score as function of the size of the training corpus. It is shown that the process leads to comparable concept error rates in the two languages making the proposed approach suitable for SLU portability across languages.
Christophe Servan, Nathalie Camelin, Christian Raymond, Frédéric Béchet, Renato De Mori
ICASSP4
2010 Online SLU model adaptation with a partial oracle
Pierre Gotab, Géraldine Damnati, Frédéric Béchet, Lionel Delphin-Poulat
INTERSPEECH3
2010 The EPAC Corpus: Manual and Automatic Annotations of Conversational Speech in French Broadcast News
Yannick Estève, Thierry Bazillon, Jean-Yves Antoine, Frédéric Béchet, Jérôme Farinas
LREC4
2010 Frame based interpretation of conversational speech
abstract
Two approaches to Spoken Language Understanding based on frames describing chunked knowledge are described. They are applied to the MEDIA corpus annotated in terms of concepts expressing chunks of spoken sentences. General rules of knowledge composition and inference appear to be adequate to effectively applying the application ontology for obtaining frame based representations of dialogue turns. The main difficulty appears to be the characterization of the syntactic knowledge expressing semantic links between knowledge chunks. This knowledge can be hand-crafted or automatically learned from examples. It is shown that the latter approach outperforms the former if applied to ASR error prone transcriptions.
Frédéric Béchet, Christian Raymond, Frédéric Duvert, Renato De Mori
SLT1
2010 Detection and Interpretation of Opinion Expressions in Spoken Surveys
abstract
This paper describes a system for automatic opinion analysis from spoken messages collected in the context of a user satisfaction survey. Opinion analysis is performed from the perspective of opinion monitoring. A process is outlined for detecting segments expressing opinions in a speech signal. Methods are proposed for accepting or rejecting segments from messages that are not reliably analyzed due to the limitations of automatic speech recognition processes, for assigning opinion hypotheses to segments and for evaluating hypothesis opinion proportions. Specific language models are introduced for representing opinion concepts. These models are used for hypothesizing opinion carrying segments in a spoken message. Each segment is interpreted by a classifier based on the Adaboost algorithm which associates a pair of topic and polarity labels to each segment. The different processes are trained and evaluated on a telephone corpus collected in a deployed customer care service. The use of conditional random fields (CRFs) is also considered for detecting segments and results are compared for different types of data and approaches. By optimizing the choice of the strategy parameters, it is possible to estimate user opinion proportions with a Kullback-Leibler divergence of 0.047 bits with respect to the true proportions obtained with a manual annotation of the spoken messages. The proportions estimated with such a low divergence are accurate enough for monitoring user satisfaction over time.
Nathalie Camelin, Frédéric Béchet, Géraldine Damnati, Renato De Mori
IEEE Trans. Speech Audio Process.2
2009 Local and global models for spontaneous speech segment detection and characterization
abstract
Processing spontaneous speech is one of the many challenges that automatic speech recognition (ASR) systems have to deal with. The main evidences characterizing spontaneous speech are disfluencies (filled pause, repetition, repair and false start) and many studies have focused on the detection and the correction of these disfluencies. In this study we define spontaneous speech as unprepared speech, in opposition to prepared speech where utterances contain well-formed sentences close to those that can be found in written documents. Disfluencies are of course very good indicators of unprepared speech, however they are not the only ones: ungrammaticality and language register are also important as well as prosodic patterns. This paper proposes a set of acoustic and linguistic features that can be used for characterizing and detecting spontaneous speech segments from large audio databases. More, we introduce a strategy that takes advantage of a global classification procfalseess using a probabilistic model which significantly improves the spontaneous speech detection.
Richard Dufour, Yannick Estève, Paul Deléglise, Frédéric Béchet
ASRU4
2009 Active learning for rule-based and corpus-based Spoken Language Understanding models
abstract
Active learning can be used for the maintenance of a deployed spoken dialog system (SDS) that evolves with time and when large collection of dialog traces can be collected on a daily basis. At the spoken language understanding (SLU) level this maintenance process is crucial as a deployed SDS evolves quickly when services are added, modified or dropped. Knowledge-based approaches, based on manually written grammars or inference rules, are often preferred as system designers can modify directly the SLU models in order to take into account such a modification in the service, even if no or very little related data has been collected. However as new examples are added to the annotated corpus, corpus-based methods can then be applied, replacing or in addition to the initial knowledge-based models. This paper describes an active learning scheme, based on an SLU criterion, which is used for automatically updating the SLU models of a deployed SDS. Two kind of SLU models are going to be compared: rule-based ones, used in the deployed system and consisting of several thousands of hand-crafted rules; corpus-based ones, based on the automatic learning of classifiers on an annotated corpus.
Pierre Gotab, Frédéric Béchet, Géraldine Damnati
ASRU2
2009 Robust dependency parsing for spoken language understanding of spontaneous speech
abstract
We describe in this paper a syntactic parser for spontaneous speech geared towards the identification of verbal subcategorization frames. The parser proceeds in two stages. The first stage is based on generic syntactic resources for French. The second stage is a reranker which is specially trained for a given application. The parser is evaluated on the MEDIA corpus. 1.
Frédéric Béchet, Alexis Nasr
INTERSPEECH1
2009 Error correction of proportions in spoken opinion surveys
abstract
International audience
Nathalie Camelin, Renato De Mori, Frédéric Béchet, Géraldine Damnati
INTERSPEECH3
2008 Semantic composition process in a speech understanding system
abstract
A knowledge representation formalism for SLU is introduced. It is used for incremental and partially automated annotation of the Media corpus in terms of semantic structures. An automatic interpretation process is described for composing semantic structures from basic semantic constituents using patterns involving constituents and words. The process has procedures for obtaining semantic compositions and for generating frame hypotheses by inference. This process is evaluated on a dialogue corpus manually annotated at the word and semantic constituent levels.
Frédéric Duvert, Marie-Jean Meurs, Christophe Servan, Frédéric Béchet, Fabrice Lefèvre, Renato De Mori
ICASSP4
2008 On-demand new word learning using world wide web
abstract
Most of the Web-based methods for lexicon augmenting consist in capturing global semantic features of the targeted domain in order to collect relevant documents from the Web. We suggest that the local context of the out-of-vocabulary (OOV) words contains relevant information on the OOV words. With this information, we propose to use the Web to build locally-augmented lexicons which are used in a final local decoding pass. Our experiments confirm the relevance of the Web for the OOV word retrieval. Different methods are proposed to retrieve the hypothesis words. Finally we present the integration of new words in the transcription process based on part-of-speech models. This technique allows to recover 7.6% of the significant OOV words and the accuracy of the system is improved.
Stanislas Oger, Georges Linarès, Frédéric Béchet, Pascal Nocera
ICASSP3
2008 Automatic customer feedback processing: alarm detection in open question spoken messages
abstract
This paper describes an alarm detection system dedicated to process automatically customer feedbacks in call-centers. Previous studies presented a strategy that consists in the robust detection of subjective opinions about a particular topic in a spoken message. In the present study, we focus on the alarm detection problem in a customer spoken feedback application. We want to characterize each customer's survey with a degree of emergency. All the messages considered as urgent need a quick answer from the call-center service in order to satisfy the customer. The strategy proposed is based on a classification scheme that takes into account all the features that can characterize a survey: answers to the closed questions, topics and opinions detected in the open question spoken message, confidence scores from the Automatic Speech Recognition (ASR) and Spoken Language Understanding (SLU) modules. A field trial realized among real customers has shown that despite the ASR robustness issues, our system efficiently ranks the most urgent messages and brings a finer analysis on the surveys than the one provided by processing the closed questions alone.
Nathalie Camelin, Géraldine Damnati, Frédéric Béchet, Renato De Mori
INTERSPEECH3
2008 Fast call-classification system development without in-domain training data
abstract
This paper presents a new method for the fast development of call-routing systems based on pre-existing corpora and knowledge databases.This method pushes forward the reduction of specific data collection and annotation for developing a new call-classification system.No specific data collection is needed for training both for the Automatic Speech Recognition (ASR) and classification models.The main idea is to re-use existing data to train the models, according to a priori knowledge on the task targeted.The experimental framework used in this study is a call-routing system applied to a civil service information telephone application.All the a priori knowledge used to develop the system is extracted from the civil service information website as well as pre-existing corpora.The evaluation of our strategy has been made on a test corpus containing 216 utterances recorded by 10 different speakers.
Christophe Servan, Frédéric Béchet
INTERSPEECH2
2008 Semantic Frame Annotation on the French MEDIA corpus
Marie-Jean Meurs, Frédéric Duvert, Frédéric Béchet, Fabrice Lefèvre, Renato De Mori
LREC3
2008 Local Methods for On-Demand Out-of-Vocabulary Word Retrieval
Stanislas Oger, Georges Linarès, Frédéric Béchet
LREC3
2008 Speaker turn characterization for spoken dialog system monitoring and adaptation
abstract
This paper describes an utterance classification method based on a multiple decoding scheme. We use the Spoken Language Understanding (SLU) strategy proposed within the European project LUNA. The goal of this classification process is to characterize each speaker's turn, in a dialog context, according to different categories relevant from an SLU point of view: out-of-domain messages, requests not covered by the interpretation module, frequent requests,.... These categories are used for two purposes in an off-line mode: system monitoring for detecting changes in users' behaviour and system adaptation by selecting dialogs likely to contain some phenomenon poorly covered by the models for an active learning scheme. All the models and the evaluations are performed on the France Telecom FT3000 corpus.
Géraldine Damnati, Frédéric Béchet, Renato De Mori
SLT2
2007 Spoken Language Understanding Strategies on the France Telecom 3000 Voice Agency Corpus
abstract
Telephone services are now deployed that allow users to react to telephone prompts in spoken natural language. These systems have limited domain semantics and dialogue strategies which are represented by finite state diagrams. Most of these systems adopt a sequential approach where the automatic speech recognition (ASR) process, the spoken language understanding (SLU) process and the dialogue management (DM) are separate processes. In the framework of the France Telecom 3000 voice service, we propose in this paper to study several strategies in order to integrate more closely these three processes: ASR, SLU, and DM. By means of a finite state machine paradigm encoding the different models used by these three levels we show how the search for the best sequence of dialogue states can be done simultaneously at the word, concept, interpretation and dialogue state levels.
Géraldine Damnati, Frédéric Béchet, Renato De Mori
ICASSP (4)2
2007 Speech mining in noisy audio message corpus
abstract
Within the framework of automatic analysis of spoken telephone surveys we propose a robust Speech Mining strategy that selects, from a large database of spoken messages, the ones likely to be correctly processed by the Automatic Speech Recognition and Classification processes. The problem considered in this paper is the analysis of messages uttered by the users of a telephone service in response to a recorded message that asks if a problem they had was satisfactorily solved. Very often in these cases, subjective information is combined with factual information. The purpose of this type of analysis is the extraction of the distribution of users opinions. Therefore it is very important to check the representativeness of the subset of messages kept by the rejection strategies. Several measures, based on the Kullback-Leibler divergence, are proposed in order to evaluate the correctness of the information extracted as well as its representativeness .
Nathalie Camelin, Frédéric Béchet, Géraldine Damnati, Renato De Mori
INTERSPEECH2
2007 Conditional use of word lattices, confusion networks and 1-best string hypotheses in a sequential interpretation strategy
abstract
Within the context of a deployed spoken dialog service, this study presents a new interpretation strategy based on the sequential use of different ASR output representations: 1-best strings, word lattices and confusion networks. The goal is to reject as early as possible in the decoding process the nonrelevant messages containing non-speech or out-of-domain content. This is done through the 1-pass of the ASR decoding process thanks to specific acoustic and language models. A confusion network (CN) is then calculated for the remaining messages and another rejection process is applied with the confidence measures obtained in the CN. The messages kept at this stage are considered relevant; therefore the search for the best interpretation is applied to a richer search space than just the 1-best word string: either the whole CN or the whole word lattice. An improved, SLU oriented, CN generation algorithm is also proposed that significantly reduces the size of the CN obtained while improving the recognition performance. This strategy is evaluated on a large corpus of real users’ messages obtained from a deployed service. 1
Bogdan Minescu, Géraldine Damnati, Frédéric Béchet, Renato De Mori
INTERSPEECH3
2007 Information retrieval strategies for accessing african audio corpora
abstract
In this paper we present a first approach to access African oral corpora, combining automatic speech recognition and information retrieval. Firstly, we present the principal characteristics of our Somali speech recognizer [8] and the results obtained on real audio archives gathered from Djibouti Radio. Secondly, we present a Hybrid Language Model (HLM) including words and sub-words to improve the robustness against OOV words. We proceed to Information Retrieval experiments with various strategies. We search on the different outputs of the ASR system (words, sub-words and hybrid). We finally present a new strategy combining sub-words and words to enhance the information retrieval results. Index Terms: speech recognition, information retrieval, hybrid language model, Somali language.
Abdillahi Nimaan, Pascal Nocera, Frédéric Béchet, Jean-François Bonastre
INTERSPEECH3
2007 Sequential Decision Strategies for Machine Interpretation of Speech
abstract
Recognition errors made by automatic speech recognition (ASR) systems may not prevent the development of useful dialogue applications if the interpretation strategy has an introspection capability for evaluating the reliability of the results. This paper proposes an interpretation strategy which is particularly effective when applications are developed with a training corpus of moderate size. From the lattice of word hypotheses generated by an ASR system, a short list of conceptual structures is obtained with a set of finite state machines (FSM). Interpretation or a rejection decision is then performed by a tree-based strategy. The nodes of the tree correspond to elaboration-decision units containing a redundant set of classifiers. A decision tree based and two large margin classifiers are trained with a development set to become interpretation knowledge sources. Discriminative training of the classifiers selects linguistic and confidence-based features for contributing to a cooperative assessment of the reliability of an interpretation. Such an assessment leads to the definition of a limited number of reliability states. The probability that a proposed interpretation is correct is provided by its reliability state and transmitted to the dialogue manager. Experimental results are presented for a telephone service application
Christian Raymond, Frédéric Béchet, Nathalie Camelin, Renato De Mori, Géraldine Damnati
IEEE Trans. Speech Audio Process.2
2006 Opinion mining in a telephone survey corpus
abstract
International audience
Nathalie Camelin, Géraldine Damnati, Frédéric Béchet, Renato De Mori
INTERSPEECH3
2006 Conceptual decoding from word lattices: application to the spoken dialogue corpus MEDIA
abstract
Within the framework of the French evaluation program MEDIA on spoken dialogue systems, this paper presents the methods proposed at the LIA for the robust extraction of basic conceptual constituents (or concepts) from an audio message. The conceptual decoding model proposed follows a stochastic paradigm and is directly integrated into the Automatic Speech Recognition (ASR) process. This approach allows us to keep the probabilistic search space on sequences of words produced by the ASR module and to project it to a probabilistic search space of sequences of concepts. This paper presents the first ASR results on the French spoken dialogue corpus MEDIA, available through ELDA. The experiments made on this corpus show that the performance reached by our approach is better than the traditional sequential approach that looks first for the best sequence of words before looking for the best sequence of concepts. Index Terms: Automatic Speech Recognition, Spoken Dialogue, Spoken Language Understanding.
Christophe Servan, Christian Raymond, Frédéric Béchet, Pascal Nocera
INTERSPEECH3
2006 Results of the French Evalda-Media evaluation campaign for literal understanding
Hélène Bonneau-Maynard, Christelle Ayache, Frédéric Béchet, Alexandre Denis 0002, Anne Kuhn, Fabrice Lefèvre, Djamel Mostefa, Matthieu Quignard, Sophie Rosset, Christophe Servan, Jeanne Villaneau
LREC3
2006 Spoken Opinion Extraction for Detecting Variations in User Satisfaction
abstract
In a previous paper, extensions of the 2-level stochastic speech understanding system have been proposed. Firstly the 3-level system is obtained through the introduction of a stochastic concept value normalization module. Then the 2+1-level system is obtained as a degraded 3-level system where the conceptual decoding and value normalization steps are decoupled, thus allowing to greatly reduce the model complexity and improve its trainability. In this paper, a multi-level spoken language understanding system is presented. This stochastic module is for the first time based on dynamic Bayesian networks. Factored language models with a generalized parallel backoff procedure are used as edge implementation to provide efficiently smoothed conditional probability estimates. This framework allows a great flexibility in terms of probability representation facilitating the development of the stochastic levels of the system. The proposed approaches, 3-level and 2+1-level, are evaluated on the French MEDIA task (tourist information and hotel booking). The MEDIA 10k-utterance training corpus is segmentally annotated, allowing a direct training of the various levels of the conceptual models. The best DBN-based system obtains performance comparable to those of the MEDIA'05 evaluation campaign best system (H. Bonneau-Maynard et al., 2005).
Frédéric Béchet, Géraldine Damnati, Nathalie Camelin, Renato De Mori
SLT1
2006 Beyond ASR 1-best: Using word confusion networks in spoken language understanding
Dilek Hakkani-Tür, Frédéric Béchet, Giuseppe Riccardi, Gökhan Tür
Comput. Speech Lang.2
2006 On the use of finite state transducers for semantic interpretation
Christian Raymond, Frédéric Béchet, Renato De Mori, Géraldine Damnati
Speech Commun.2
2005 Semantic Interpretation With Error Correction
abstract
The paper presents a semantic interpretation strategy, for spoken dialogue systems, including an error correction process. Semantic interpretations output by the spoken understanding module may be incorrect, but some semantic components may be correct. A set of situations are introduced, describing semantic confidence based on the agreement of semantic interpretations proposed by different classification methods. The interpretation strategy considers, with the highest priority, the validation of the interpretation arising from the most likely sequence of words. If our confidence score model gives a high probability that this interpretation is not correct, then possible corrections of it are considered using the other sequences in the N-best lists of possible interpretations. This strategy is evaluated on a dialogue corpus provided by France Telecom R&D and collected for a tourism telephone service. Significant reduction in understanding error rate are obtained as well as powerful new confidence measures.
Christian Raymond, Frédéric Béchet, Nathalie Camelin, Renato De Mori, Géraldine Damnati
ICASSP (1)2
2005 Mining broadcast news data: robust information extraction from word lattices
abstract
International audience
Benoît Favre, Frédéric Béchet, Pascal Nocera
INTERSPEECH2
2005 Evaluating the pronunciation of proper names by four French grapheme-to-phoneme converters
abstract
International Speech Communication Association (Isca) - International Astronautical Federation. ISBN : 13 9781604234480.
Philippe Boula de Mareüil, Christophe d'Alessandro, Gérard Bailly, Frédéric Béchet, Marie-Neige Garcia, Michel Morel, Romain Prudon, Jean Véronis
INTERSPEECH4
2004 Tagging with Hidden Markov Models Using Ambiguous Tags
Alexis Nasr, Frédéric Béchet, Alexandra Volanschi
COLING2
2004 Mining Spoken Dialogue Corpora for System Evaluation and Modelin
Frédéric Béchet, Giuseppe Riccardi, Dilek Hakkani-Tür
EMNLP1
2004 Automatic learning of interpretation strategies for spoken dialogue systems
abstract
The paper proposes a new application of automatically trained decision trees to derive the interpretation of a spoken sentence. A new strategy for building structured cohorts of candidates is also described. By evaluating predicates related to the acoustic confidence of the words expressing a concept, the linguistic and semantic consistency of candidates in the cohort and the rank of a candidate within a cohort, the decision tree automatically learns a decision strategy for rescoring or rejecting an n-best list of candidates representing a user's utterance. A relative reduction of 18.6% in the understanding error rate is obtained by our rescoring strategy with no utterance rejection and a relative reduction of 43.1% of the same error rate is achieve with a rejection rate of only 8% of the utterances.
Christian Raymond, Frédéric Béchet, Renato De Mori, Géraldine Damnati, Yannick Estève
ICASSP (1)2
2004 The French MEDIA/EVALDA Project: the Evaluation of the Understanding Capability of Spoken Language Dialogue Systems
Laurence Devillers, Hélène Bonneau-Maynard, Sophie Rosset, Patrick Paroubek, Kevin McTait, Djamel Mostefa, Khalid Choukri, Laurent Charnay, Caroline Bousquet-Vernhettes, Nadine Vigouroux, Frédéric Béchet, Laurent Romary, Jean-Yves Antoine, Jeanne Villaneau, Myriam Vergnes, Jérôme Goulian
LREC11
2004 Data augmentation and language model adaptation using singular value decomposition
Frédéric Béchet, Renato De Mori, David Janiszek
Pattern Recognit. Lett.1
2004 Detecting and extracting named entities from spontaneous speech in a mixed-initiative spoken dialogue context: How May I Help You?sm, tm
Frédéric Béchet, Allen L. Gorin, Jeremy H. Wright, Dilek Hakkani-Tür
Speech Commun.1
2003 Dynamic scheduling of decoding processes for directory assistance
abstract
This paper deals with the difficult task of recognition of a large vocabulary of proper names in a directory assistance application. Research on the European project SMADA has shown that there is a need of an elaborate and effective decision strategy that limits the risk of false automation. This paper proposes a new strategy which integrates, as well as a general decoder, a set of decoders specialized in some specific situations. Specialized recognition processes do not need to be applicable for every input, but they have to be scheduled and performed only under certain conditions. A first implementation of such a model is proposed here, through a rejection strategy of the hypotheses output by a general decoder. This strategy leads to a very significant improvement over the results obtained by a standard rejection method based on acoustic confidence scores only.
Renato De Mori, Frédéric Béchet, Gérard Subsol, Dominique Massonié
ICASSP (1)2
2003 Multi-channel sentence classification for spoken dialogue language modeling
abstract
In traditional language modeling word prediction is based on the local context (e.g. n-gram). In spoken dialog, language statistics are affected by the multidimensional structure of the human-machine interaction. In this paper we investigate the statistical dependencies of users’ responses with respect to the system’s and user’s channel. The system channel components are the prompts’ text, dialogue history, dialogue state. The user channel components are the Automatic Speech Recognition (ASR) transcriptions, the semantic classifier output and the sentence length. We describe an algorithm for language model rescoring using users’ response classification. The user’s response is first mapped into a multidimensional state and the state specific language model is applied for ASR rescoring. We present perplexity and ASR results on the How May I Help You ? sm 100K spoken dialogs.
Frédéric Béchet, Giuseppe Riccardi, Dilek Hakkani-Tür
INTERSPEECH1
2003 Conceptual decoding for spoken dialog systems
abstract
International audience
Yannick Estève, Christian Raymond, Frédéric Béchet, Renato De Mori
INTERSPEECH3
2002 Dynamic generation of proper name pronunciations for directory assistance
abstract
This paper deals with the difficult task of recognition of a large vocabulary of proper names in a directory assistance application. Rather than augmenting the lexicon with alternate pronunciations. which is unsuitable for very large vocabularies of proper names, a class of distortions of the canonical form is used as knowledge source (KS) for a new evaluation of the N best hypotheses generated in a first recognition phase, in which new probability distributions are used. The KS is the result of applying constraints inspired by speech science to distortions obtained by automatic learning. Experiments on a very large French directory document the validity of the approach.
Frédéric Béchet, Renato De Mori, Gérard Subsol
ICASSP1
2002 Named entity extraction from spontaneous speech in how may i help you?
Frédéric Béchet, Allen L. Gorin, Jeremy H. Wright, Dilek Hakkani-Tür
INTERSPEECH1
2001 Data augmentation and language model adaptation
abstract
A method is presented for augmenting word n-gram counts in a matrix which represents a 2-gram language model (LM) This method is based on numerical distances in a reduced space obtained by singular value decomposition. Rescoring word lattices in a spoken dialogue application using an LM containing augmented counts has lead to a word error rate (WER) reduction of 6.5%. By further interpolating augmented counts with the counts extracted from a very large newspaper corpus, but only for selected histories, a total WER reduction of 11.7% was obtained. We show that this approach gives better results than a global count interpolation for all histories of the LM.
David Janiszek, Renato De Mori, Frédéric Béchet
ICASSP3
2001 Stochastic finite state automata language model triggered by dialogue states
abstract
Within the framework of Natural Spoken Dialogue systems, this paper describes a method for dynamically adapting a Language Model (LM) to the dialogue states detected. This LM combines a standard n-gram model with Stochastic Finite State Automata (SFSAs). During the training process, the sentence corpus used to train the LM is split into several hierarchical clusters in a 2-step process which involves both explicit knowledge and statistical criteria. All the clusters are stored in a binary tree where the whole corpus is attached to the root node. Each level of the tree corresponds to a higher specialization of the sub-corpora attached to the nodes and each node corresponds to a different dialogue state. From the same sentence corpus, SFSAs are extracted in order to model longer contexts than the ones used in the standard n-gram model. A set of SFSAs is attached to each node of the tree as well as a sub-LM which combines a bigram trained on the sub-corpus of the node and the SFSAs selected. A first decoding process calculates a word-graph as well as a first sentence hypothesis. This first hypothesis will be used to find the optimal node in the LM tree. Then, a rescoring process of the word graph using the LM attached to the node selected is performed. By adapting the LM to the dialogue state detected, we show a statistically significant gain in WER on a dialogue corpus collected by France Telecom R&D .
Yannick Estève, Frédéric Béchet, Alexis Nasr, Renato De Mori
INTERSPEECH2
2000 Tagging Unknown Proper Names Using Decision Trees
abstract
This paper describes a supervised learning method to automatically select from a set of noun phrases, embedding proper names of different semantic classes, their most distinctive features. The result of the learning process is a decision tree which classifies an unknown proper name on the basis of its context of occurrence. This classifier is used to estimate the probability distribution of an out of vocabulary proper name over a tagset. This probability distribution is itself used to estimate the parameters of a stochastic part of speech tagger.
Frédéric Béchet, Alexis Nasr, Franck Genet
ACL1
2000 Introduction to the IST-HLT project speech-driven multimodal automatic directory assistance (SMADA)
abstract
\n Contains fulltext :\n 75039.pdf (author's version ) (Open Access)\n
Frédéric Béchet, Els den Os, Lou Boves, Jürgen Sienel
INTERSPEECH1
2000 Dynamic selection of language models in a dialogue system
abstract
This paper describes a method for building statistical Language Models (LMs) dedicated to specific dialogue situations. The architecture of the speech recognition system proposed uses several LMs. The first stage of this system, consists of producing a word-lattice from a given sentence uttered by a speaker. A general LM calculates a sentence-hypothesis. Then, in a second stage, the system chooses a specialized LM according to the word-lattice and the previous hypothesis. Another decoding process is performed using this specialized LM in order to produce a new sentence- hypothesis. Finally, a decision-module processes these two hypotheses in order to assign three confidence levels to the sentence-hypothesis produced. These confidence levels can be used by the dialogue manager in order to improve the dialogue, by asking a confirmation to the speaker when a sentence is labeled ambiguous. This research is supported by France Telecom's R&D under the contract 971B427.
Yannick Estève, Frédéric Béchet, Renato De Mori
INTERSPEECH2
2000 Integrating MAP and linear transformation for language model adaptation
David Janiszek, Frédéric Béchet, Renato De Mori
INTERSPEECH2
1999 Large Span statistical language models: application to homophone disambiguation for large vocabulary speech recognition in French
Frédéric Béchet, Alexis Nasr, Thierry Spriet, Renato De Mori
EUROSPEECH1
1999 A language model combining n-grams and stochastic finite state automata
abstract
Maximum a posteriori adaptation method combines the prior knowledge with adaptation data from a new speaker, which has a nice asymptotical property, but has a slow adaptation rate for not modifying unseen models. In a strictly Bayesian approach, prior parameters are assumed known, based on common or subjective knowledge. But a practical solution is to adopt an empirical Bayesian approach, where the prior parameters are estimated directly from training speech data itself. So there is a problem of mismatches between training and testing conditions. In this paper we propose a prior parameter transformation (PPT) adaptation approach that transforms the prior parameters to be more representative of the new speaker. It can influence unseen models by tying prior parameter transformations across different models according to amount of adaptation data available. Based on the improved prior information better model parameters can be obtained even with small amount of adaptation data.
Alexis Nasr, Yannick Estève, Frédéric Béchet, Thierry Spriet, Renato De Mori
EUROSPEECH3
1998 Evaluation of grapheme-to phoneme conversion for text-to-speech synthesis in French
Philippe Boula de Mareüil, François Yvon, Christophe d'Alessandro, V. Auberg, Michel Bagein, Gérard Bailly, Frédéric Béchet, S. Fonkia, Jean-Philippe Goldman, Eric Keller, Douglas D. O'Shaughnessy, Steve Pagel, F. Sannier, Jean Véronis, Brigitte Zellner Keller
LREC7
1998 Objective evaluation of grapheme to phoneme conversion for text-to-speech synthesis in French
François Yvon, Philippe Boula de Mareüil, Christophe d'Alessandro, Véronique Aubergé, Michel Bagein, Gérard Bailly, Frédéric Béchet, S. Foukia, J.-F. Goldman, Eric Keller, Douglas D. O'Shaughnessy, Vincent Pagel, Fred Sannier, Jean Véronis, Brigitte Zellner
Comput. Speech Lang.7
1997 Automatic assignment of part-of-speech to out-of-vocabulary words for text-to-speech processing
Frédéric Béchet, Marc El-Bèze
EUROSPEECH1
1991 Bottom-up acoustic-phonetic decoding for the selection of word cohorts from a large vocabulary
Henri Meloni, Frédéric Béchet, Philippe Gilles
EUROSPEECH2