Benoît Favre

dblp:91/1442 · DBLP profile ↗
← Back
72ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-9777-4613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning
Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils, Benoît Favre, Richard Dufour
LREC4
2026 CareMedEval Dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
abstract
Critical appraisal of scientific literature is an essential skill in the biomedical field. While large language models (LLMs) can offer promising support in this task, their reliability remains limited, particularly for critical reasoning in specialized domains. We introduce CareMedEval, an original dataset designed to evaluate LLMs on biomedical critical appraisal and reasoning tasks. Derived from authentic exams taken by French medical students, the dataset contains 534 questions based on 37 scientific articles. Unlike existing benchmarks, CareMedEval explicitly evaluates critical reading and reasoning grounded in scientific papers. Benchmarking state-of-the-art generalist and biomedical-specialized LLMs under various context conditions reveals the difficulty of the task: open and commercial models fail to exceed an Exact Match Rate of 0.5 even though generating intermediate reasoning tokens considerably improves the results. Yet, models remain challenged especially on questions about study limitations and statistical analysis. CareMedEval provides a challenging benchmark for grounded reasoning, exposing current LLM limitations and paving the way for future development of automated support for critical appraisal.
Doria Bonzi, Alexandre Guiggi, Frédéric Béchet, Carlos Ramisch, Benoît Favre
LREC5
2025 Statistical Deficiency for Task Inclusion Estimation
abstract
Loïc Fosse, Frederic Bechet, Benoit Favre, Géraldine Damnati, Gwénolé Lecorvé, Maxime Darrin, Philippe Formont, Pablo Piantanida. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Loïc Fosse, Frédéric Béchet, Benoît Favre, Géraldine Damnati, Gwénolé Lecorvé, Maxime Darrin, Philippe Formont, Pablo Piantanida
ACL (1)3
2024 Analysing Communicative Intent Coordination in Child-Caregiver Interactions
Abhishek Agrawal, Benoît Favre, Abdellah Fourtassi
CogSci2
2024 Automatic Coding of Contingency in Child-Caregiver Conversations
abstract
One of the most important communicative skills children have to learn is to engage in meaningful conversations with people around them. At the heart of this learning lies the mastery of contingency, i.e., the ability to contribute to an ongoing exchange in a relevant fashion (e.g., by staying on topic). Current research on this question relies on the manual annotation of a small sample of children, which limits our ability to draw general conclusions about development. Here, we propose to mitigate the limitations of manual labor by relying on automatic tools for contingency judgment in children’s early natural interactions with caregivers. Drawing inspiration from the field of dialogue systems evaluation, we built and compared several automatic classifiers. We found that a Transformer-based pre-trained language model – when fine-tuned on a relatively small set of data we annotated manually (around 3,500 turns) – provided the best predictions. We used this model to automatically annotate, new and large-scale data, almost two orders of magnitude larger than our fine-tuning set. It was able to replicate existing results and generate new data-driven hypotheses. The broad impact of the work is to provide resources that can help the language development community study communicative development at scale, leading to more robust theories.
Abhishek Agrawal, Mitja Nikolaus, Benoît Favre, Abdellah Fourtassi
LREC/COLING3
2024 CHICA: A Developmental Corpus of Child-Caregiver's Face-to-face vs. Video Call Conversations in Middle Childhood
abstract
Existing studies of naturally occurring language-in-interaction have largely focused on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood. The current work contributes to filling this gap by introducing CHICA (for Child Interpersonal Communication Analysis), a developmental corpus of child-caregiver conversations at home, involving groups of French-speaking children aged 7, 9, and 11 years old. Each dyad was recorded twice: once in a face-to-face setting and once using computer-mediated video calls. For the face-to-face settings, we capitalized on recent advances in mobile, lightweight eye-tracking and head motion detection technology to optimize the naturalness of the recordings, allowing us to obtain both precise and ecologically valid data. Further, we mitigated the challenges of manual annotation by relying – to the extent possible – on automatic tools in speech processing and computer vision. Finally, to demonstrate the richness of this corpus for the study of child communicative development, we provide preliminary analyses comparing several measures of child-caregiver conversational dynamics across developmental age, modality, and communicative medium. We hope the current corpus will allow new discoveries into the properties and mechanisms of multimodal communicative development across middle childhood.
Dhia-Elhak Goumri, Abhishek Agrawal, Mitja Nikolaus, Hong Duc Thang Vu, Kübra Bodur, Elias Emmar, Cassandre Armand, Chiara Mazzocconi, Shreejata Gupta, Laurent Prévot 0001, Benoît Favre, Leonor Becerra-Bonache, Abdellah Fourtassi
LREC/COLING11
2024 Unified Framework for Spoken Language Understanding and Summarization in Task-Based Human Dialog processing
abstract
International audience
Eunice Akani, Frédéric Béchet, Benoît Favre, Romain Gemignani
INTERSPEECH3
2024 Investigating self-supervised speech models' ability to classify animal vocalizations: The case of gibbon's vocal signatures
abstract
International audience
Jules Cauzinille, Benoît Favre, Ricard Marxer, Dena J. Clink, Abdul Hamid Ahmad, Arnaud Rey
INTERSPEECH2
2024 Text summary evaluation based on interpretable semantic textual similarity
Goutam Majumder, Vikrant Rajput, Partha Pakray, Sivaji Bandyopadhyay, Benoît Favre
Multim. Tools Appl.5
2023 Development of Multimodal Turn Coordination in Conversations: Evidence for Adult-like behavior in Middle Childhood
Abhishek Agrawal, Kübra Bodur, Benoît Favre, Abdellah Fourtassi
CogSci4
2023 Reducing named entity hallucination risk to ensure faithful summary generation
abstract
The faithfulness of abstractive text summarization at the named entities level is the focus of this study.We propose to add a new criterion to the summary selection method based on the "risk" of generating entities that do not belong to the source document.This method is based on the assumption that Out-Of-Document entities are more likely to be hallucinations.This assumption was verified by a manual annotation of the entities occurring in a set of generated summaries on the CNN/DM corpus.This study showed that only 29% of the entities outside the source document were inferrable by the annotators, leading to 71% of hallucinations among OOD entities.We test our selection method on the CNN/DM corpus and show that it significantly reduces the hallucination risk on named entities while maintaining competitive results with respect to automatic evaluation metrics like ROUGE.
Eunice Akani, Benoît Favre, Frédéric Béchet, Romain Gemignani
INLG2
2022 Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective
abstract
In recent years, joint text-image embeddings have significantly improved thanks to the development of transformer-based Vision-Language models. Despite these advances, we still need to better understand the representations produced by those models. In this paper, we compare pre-trained and fine-tuned representations at a vision, language and multimodal level. To that end, we use a set of probing tasks to evaluate the performance of state-of-the-art Vision-Language models and introduce new datasets specifically for multimodal probing. These datasets are carefully designed to address a range of multimodal capabilities while minimizing the potential for models to rely on bias. Although the results confirm the ability of Vision-Language models to understand color at a multimodal level, the models seem to prefer relying on bias in text data for object position and size. On semantically adversarial examples, we find that those models are able to pinpoint fine-grained multimodal differences. Finally, we also notice that fine-tuning a Vision-Language model on multimodal tasks does not necessarily improve its multimodal ability. We make all datasets and code available to replicate experiments.
Emmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît Favre
AAAI4
2022 Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?
abstract
Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks.Yet, the exact capabilities of these black-box models are still poorly understood.While much of previous work has focused on studying their ability to learn meaning at the word-level, their ability to track syntactic dependencies between words has received less attention.We take a first step in closing this gap by creating a new multimodal task targeted at evaluating understanding of predicate-noun dependencies in a controlled setup.We evaluate a range of state-of-the-art models and find that their performance on the task varies considerably, with some models performing relatively well and others at chance level.In an effort to explain this variability, our analyses indicate that the quality (and not only sheer quantity) of pretraining data is essential.Additionally, the best performing models leverage fine-grained multimodal pretraining objectives in addition to the standard image-text matching objectives.This study highlights that targeted and controlled evaluations are a crucial step for a precise and rigorous test of the multimodal knowledge of vision-and-language models.
Mitja Nikolaus, Emmanuelle Salin, Stéphane Ayache, Abdellah Fourtassi, Benoît Favre
EMNLP5
2022 ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks
abstract
We aim at improving spoken language modeling (LM) using very large amount of automatically transcribed speech.We leverage the INA (French National Audiovisual Institute 1 ) collection and obtain 19GB of text after applying ASR on 350,000 hours of diverse TV shows.From this, spoken language models are trained either by fine-tuning an existing LM (FlauBERT 2 ) or through training a LM from scratch.New models (FlauBERT-Oral) are shared with the community and evaluated for 3 downstream tasks: spoken language understanding, classification of TV shows and speech syntactic parsing.Results show that FlauBERT-Oral can be beneficial compared to its initial FlauBERT version demonstrating that, despite its inherent noisy nature, ASR-generated text can be used to build spoken language models.
Valentin Pelloin, Franck Dary, Nicolas Hervé, Benoît Favre, Nathalie Camelin, Antoine Laurent, Laurent Besacier
INTERSPEECH4
2022 "Do you follow me?": A Survey of Recent Approaches in Dialogue State Tracking
abstract
While communicating with a user, a taskoriented dialogue system has to track the user's needs at each turn according to the conversation history.This process called dialogue state tracking (DST) is crucial because it directly informs the downstream dialogue policy.DST has received a lot of interest in recent years with the text-to-text paradigm emerging as the favored approach.In this review paper, we first present the task and its associated datasets.Then, considering a large number of recent publications, we identify highlights and advances of research in 2021-2022.Although neural approaches have enabled significant progress, we argue that some critical aspects of dialogue systems such as generalizability are still underexplored.To motivate future studies, we propose several research avenues.
Léo Jacqmin, Lina Maria Rojas-Barahona, Benoît Favre
SIGDIAL3
2021 Multimodal Machine Learning for Natural Language Processing: Disambiguating Prepositional Phrase Attachments with Images
Sebastien Delecraz, Leonor Becerra-Bonache, Benoît Favre, Alexis Nasr, Frédéric Béchet
Neural Process. Lett.3
2020 Designing an IIR Research Apparatus with Users with Severe Intellectual Disability
abstract
Traditional methods of engagement with pre-defined queries, verbal instruction and interviewing do not provide necessary means to address information-seeking behavior and visual browsing for participants with severe autism and intellectual disability. In this paper, we identify challenges and characteristics of providing effective methods to explore visual browsing and video recommender systems with one non-verbal participant with autism and intellectual disability. We contribute a case study and a reflection on a) how iterative design approaches that builds on special interests and strengths of one individual with disability can support experimental IIR research in becoming more inclusive, b) some of the ethical consideration that arise in the tensions between participation in the research and other interests and c) how flexible experimental and apparatus design can further allow participant's terms to prevail.
Filip Bircanin, Laurianne Sitbon, Benoît Favre, Margot Brereton
CHIIR3
2020 Engaging the Abilities of Participants with Intellectual Disabilityin IIR Research
abstract
At CHIIR 2019, Berget and MacFarlane [4] pointed out the need for ethical methodologies when involving participants with dyslexia. In this paper, we further propose that a stance of ability based design and participatory design approaches can further involve, engage and support people with intellectual disability in interactive information retrieval (IIR) research. Through a case study with an accessible prototype designed to access instructional videos, we demonstrate how an approach building on participant's interests and providing them support as part of the study design leads to ecologically valid observations. The accessible prototype makes use of images as prompts and query support, and includes social aspects. Our observations confirm that users with intellectual disability favour a visual approach to information access and interaction. The contributions of this work are primarily 1) a 2 step approach with supported participatory design approaches involving early prototypes 2) a case study of this approach to investigate information access interfaces with people with intellectual disability and 3) a reflection on the case study and applicability of the method in IIR evaluation.
Laurianne Sitbon, Benoît Favre, Margot Brereton, Stewart Koplick, Lauren Fell
CHIIR2
2020 A Framework for Information Accessibility in Large Video Repositories
abstract
Online videos are a medium of choice for young adults to access or receive information, and recent work has highlighted that it is a particularly effective medium for adults with intellectual disability, by its visual nature. Reflecting on a case study presenting fieldwork observations of how adults with intellectual disability engage with videos on the Youtube platform, we propose a framework to define and evaluate the accessibility of such large video repositories, from an informational perspective. The proposed framework nuances the concept of information accessibility from that of the accessibility of information access interfaces themselves (generally catered for under web accessibility guidelines), or that of the documents (generally covered in general accessibility guidelines). It also includes a notion of search (or browsing) accessibility, which reflects the ability to reach the document containing the information. In the context of large information repositories, this concept goes beyond how the documents are organized into how automated processes (browsing or searching) can support users. In addition to the framework we also detail specifics of document accessibility for videos. The framework suggests a multi-dimensional approach to information accessibility evaluation which includes both cognitive and sensory aspects. This framework can serve as a basis for practitioners when designing video information repositories accessible to people with intellectual disability, and extends on the information presentation guidelines such as suggested by the WCAG.
Laurianne Sitbon, Benoît Favre, Jinglan Zhang, Andy Bayor, Stewart Koplick, Filip Bircanin, Margot Brereton
CHIIR2
2020 Neural Representations of Dialogical History for Improving Upcoming Turn Acoustic Parameters Prediction
abstract
International audience
Simone Fuscone, Benoît Favre, Laurent Prévot 0001
INTERSPEECH2
2020 Filtering conversations through dialogue acts labels for improving corpus-based convergence studies
abstract
During an interaction the tendency of speakers to change their speech production to make it more similar to their interlocutor's speech is called convergence.Convergence had been studied due to its relevance for cognitive models of communication as well as for dialogue system adaptation to the user.Convergence effects have been established on controlled data sets while tracking its dynamics on generic corpora has provided positive but more contrasted outcomes.We propose to enrich large conversational corpora with dialogue acts information and to use these acts as filters to create subsets of homogeneous conversational activity.Those subsets allow a more precise comparison between speakers' speech variables.We compare convergence on acoustic variables (Energy, Pitch and Speech Rate) measured on raw data sets, with human and automatically data sets labelled with dialog acts type.We found that such filtering helps in observing convergence suggesting that future studies should consider such high level dialogue activity types and the related NLP techniques as important tools for analyzing conversational interpersonal dynamics.
Simone Fuscone, Benoît Favre, Laurent Prévot 0001
SIGdial2
2019 Can We Predict Self-reported Customer Satisfaction from Interactions?
abstract
In the context of contact centers, customers' satisfaction after a conversation with an agent is a critical issue which has to be collected in order to detect problems and improve quality of service. Automatically predicting customer satisfaction directly from system logs, without any survey or manual annotation is a challenging task of a great interest for the field of human-human conversation understanding and for improving contact center quality of service. Unlike previous studies that have focused on questions directly related to the content of a conversation, we look at a more general opinion about a service which is called the "Net Promoter Score" (NPS) where customers are considered either as promoters, detractors or neutral. On a very large corpus of chat-conversations with customer satisfaction surveys, we explore several classification scheme in order to achieve this prediction task, only using conversation logs.
Jérémy Auguste, Delphine Charlet, Géraldine Damnati, Frédéric Béchet, Benoît Favre
ICASSP5
2018 Adding Syntactic Annotations to Flickr30k Entities Corpus for Multimodal Ambiguous Prepositional-Phrase Attachment Resolution
Sebastien Delecraz, Alexis Nasr, Frédéric Béchet, Benoît Favre
LREC4
2016 Investigation of speaker embeddings for cross-show speaker diarization
abstract
This paper proposes to investigate speaker embeddings, a representation extracted from hidden layers of deep neural networks trained on a speaker identification task, on cross-show diarization. The new representation brings an improvement over i-vectors, and we show that while shallow hidden layers give best results on the single-show condition, deeper layers yield better performance on cross-show diarization. This confirms that deep representations model higher level features which help generalizing to different acoustic conditions. Experiments, conducted on the French corpus of REPERE, show that the deep speaker embeddings technique decreases DER by 0.82 points.
Mickael Rouvier, Benoît Favre
ICASSP2
2016 Joint Syntactic and Semantic Analysis with a Multitask Deep Learning Framework for Spoken Language Understanding
abstract
International audience
Jérémie Tafforeau, Frédéric Béchet, Thierry Artières, Benoît Favre
INTERSPEECH4
2016 Beyond Utterance Extraction: Summary Recombination for Speech Summarization
abstract
International audience
Jérémy Trione, Benoît Favre, Frédéric Béchet
INTERSPEECH2
2016 Summarizing Behaviours: An Experiment on the Annotation of Call-Centre Conversations
Morena Danieli, A. R. Balamurali, Evgeny A. Stepanov, Benoît Favre, Frédéric Béchet, Giuseppe Riccardi
LREC4
2016 A Document Repository for Social Media and Speech Conversations
Adam Funk, Robert J. Gaizauskas, Benoît Favre
LREC3
2016 Word Embedding Evaluation and Combination
Sahar Ghannay, Benoît Favre, Yannick Estève, Nathalie Camelin
LREC2
2015 Multimodal embedding fusion for robust speaker role recognition in video broadcast
abstract
Person role recognition in video broadcasts consists in classifying people into roles such as anchor, journalist, guest, etc. Existing approaches mostly consider one modality, either audio (speaker role recognition) or image (shot role recognition), firstly because of the non-synchrony between both modalities, and secondly because of the lack of a video corpus annotated in both modalities. Deep Neural Networks (DNN) approaches offer the ability to learn simultaneously feature representations (embeddings) and classification functions. This paper presents a multimodal fusion of audio, text and image embeddings spaces for speaker role recognition in asynchronous data. Monomodal embeddings are trained on exogenous data and fine-tuned using a DNN on 70 hours of French Broadcasts corpus for the target task. Experiments on the REPERE corpus show the benefit of the embeddings level fusion compared to the monomodal embeddings systems and to the standard late fusion method.
Mickael Rouvier, Sebastien Delecraz, Benoît Favre, Meriem Bendris, Frédéric Béchet
ASRU3
2015 Concept-based Summarization using Integer Linear Programming: From Concept Pruning to Multiple Optimal Solutions
abstract
In concept-based summarization, sentence selection is modelled as a budgeted maximum coverage problem.As this problem is NP-hard, pruning low-weight concepts is required for the solver to find optimal solutions efficiently.This work shows that reducing the number of concepts in the model leads to lower ROUGE scores, and more importantly to the presence of multiple optimal solutions.We address these issues by extending the model to provide a single optimal solution, and eliminate the need for concept pruning using an approximation algorithm that achieves comparable performance to exact inference.
Florian Boudin, Hugo Mougard, Benoît Favre
EMNLP3
2015 "speech is silver, but silence is golden": improving speech-to-speech translation performance by slashing users input
abstract
Speech-to-speech translation is a challenging task mixing two of the most ambitious Natural Language Processing challenges: Machine Translation (MT) and Automatic Speech Recognition (ASR). Recent advances in both fields have led to operational systems achieving good performance when used in matching conditions with those of ASR and MT models training. Regardless of the quality of these models, errors are inevitable due to some technical limitations of the systems (e.g. closed vocabulary) and intrinsic ambiguities of spoken languages. However all ASR and MT errors don’t have the same impact on the usability of a given speech-to-speech dialog system: some can be very benign, unconsciously corrected by users, some can damage the understanding between users and eventually lead the dialog to a failure. We present in this paper a strategy focusing on ASR error segments that have a high negative impact on MT performance. We propose a method that consists firstly in automatically detecting these erroneous segments then secondly estimating their impact on MT. We show that removing such segments prior to translation can lead to a significant decrease in translation error rate, even without any correction strategy.
Frédéric Béchet, Benoît Favre, Mickael Rouvier
INTERSPEECH2
2015 Adapting lexical representation and OOV handling from written to spoken language with word embedding
abstract
Word embeddings have become ubiquitous in NLP, especially when using neural networks. One of the assumptions of such representations is that words with similar properties have similar representation, allowing for better generalization from subsequent models. In the standard setting, two kinds of training corpora are used: a very large unlabeled corpus for learning the word embedding representations; and an in-domain training corpus with gold labels for training classifiers on the target NLP task. Because of the amount of data required to learn embeddings, they are trained on large corpus of written text. This can be an issue when dealing with non-canonical language, such as spontaneous speech: embeddings have to be adapted to fit the particularities of spoken transcriptions. However the adaptation corpus available for a given speech application can be limited, resulting in a high number of words from the embedding space not occurring in the adaptation space. We present in this paper a method for adapting an embedding space trained on written text to a spoken corpus of limited size. In particular we deal with words from the embedding space not occurring in the adaptation data. We report experiments done on a Part-OfSpeech task on spontaneous speech transcriptions collected in a call-centre. We show that our word embedding adaptation approach outperforms state-of-the-art Conditional Random Field approach when little in-domain adaptation data is available.
Jérémie Tafforeau, Thierry Artières, Benoît Favre, Frédéric Béchet
INTERSPEECH3
2015 Call Centre Conversation Summarization: A Pilot Task at Multiling 2015
abstract
This paper describes the results of the Call Centre Conversation Summarization task at Multiling'15.The CCCS task consists in generating abstractive synopses from call centre conversations between a caller and an agent.Synopses are summaries of the problem of the caller, and how it is solved by the agent.Generating them is a very challenging task given that deep analysis of the dialogs and text generation are necessary.Three languages were addressed: French, Italian and English translations of conversations from those two languages.The official evaluation metric was ROUGE-2.Two participants submitted a total of four systems which had trouble beating the extractive baselines.The datasets released for the task will allow more research on abstractive dialog summarization.
Benoît Favre, Evgeny A. Stepanov, Jérémy Trione, Frédéric Béchet, Giuseppe Riccardi
SIGDIAL Conference1
2015 MultiLing 2015: Multilingual Summarization of Single and Multi-Documents, On-line Fora, and Call-center Conversations
abstract
George Giannakopoulos, Jeff Kubina, John Conroy, Josef Steinberger, Benoit Favre, Mijail Kabadjov, Udo Kruschwitz, Massimo Poesio. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015.
George Giannakopoulos, Jeff Kubina, John M. Conroy, Josef Steinberger, Benoît Favre, Mijail A. Kabadjov, Udo Kruschwitz, Massimo Poesio
SIGDIAL Conference5
2014 Retrieving the syntactic structure of erroneous ASR transcriptions for open-domain Spoken Language Understanding
abstract
Retrieving the syntactic structure of erroneous ASR transcriptions can be of great interest for open-domain Spoken Language Understanding tasks in order to correct or at least reduce the impact of ASR errors on final applications. Most of the previous works on ASR and syntactic parsing have addressed this problem by using syntactic features during ASR to help reducing Word Error Rate (WER). The improvement obtained is often rather small, however the structure and the relations between words obtained through parsing can be of great interest for the SLU processes, even without a significant decrease of WER. That is why we adopt another point of view in this paper: considering that ASR transcriptions contain inevitably some errors, we show in this study that it is possible to improve the syntactic analysis of these erroneous transcriptions by performing a joint error detection / syntactic parsing process. The applicative framework used in this study is a speech-to-speech system developed through the DARPA BOLT project.
Frédéric Béchet, Benoît Favre, Alexis Nasr, Mathieu Morey
ICASSP2
2014 Multiple-view constrained clustering for unsupervised face identification in TV-broadcast
abstract
Our goal is to automatically identify faces in TV broadcast without a pre-defined dictionary of identities. Most methods are based on identity detection (from OCR and ASR) and require a propagation strategy based on visual clustering. In TV content, people appear with many variations making the clustering difficult. In this case, speaker clustering can be a reliable link for face clustering. We propose in this paper to build automatically an incomplete speaker-face mapping based on local evidence of OCR and Lip activity links. Then, we propose schemes of speaker constraints propagation to the face constrained-clustering problem. Experiments performed on the REPERE corpus show an improvement of face identification by propagating names to face clusters (+3.7% F-measure compared to the baseline).
Meriem Bendris, Benoît Favre, Delphine Charlet, Géraldine Damnati, Rémi Auguste
ICASSP2
2014 Reranked aligners for interactive transcript correction
abstract
Clarification dialogs can help address ASR errors in speech-to-speech translation systems and other interactive applications. We propose to use variants of Levenshtein alignment for merging an er-rorful utterance with a targeted rephrase of an error segment. ASR errors that might harm the alignment are addressed through phonetic matching, and a word embedding distance is used to account for the use of synonyms outside targeted segments. These features lead to a relative improvement of 30% of word error rate on sentences with ASR errors compared to not performing the clarification. Twice as many utterances are completely corrected compared to using basic word alignment. Furthermore, we generate a set of potential merges and train a neural network on crowd-sourced rephrases in order to select the best merger, leading to 24% more instances completely corrected. The system is deployed in the framework of the BOLT project.
Benoît Favre, Mickael Rouvier, Frédéric Béchet
ICASSP1
2014 Multimodal understanding for person recognition in video broadcasts
abstract
International audience
Frédéric Béchet, Meriem Bendris, Delphine Charlet, Géraldine Damnati, Benoît Favre, Mickael Rouvier, Rémi Auguste, Benjamin Bigot, Richard Dufour, Corinne Fredouille, Georges Linarès, Jean Martinet, Grégory Senay, Pierre Tirilly
INTERSPEECH5
2014 Adapting dependency parsing to spontaneous speech for open domain spoken language understanding
abstract
Parsing human-human conversations consists in automatically enriching text transcription with semantic structure information. We use in this paper a FrameNet-based approach to semantics that, without needing a full semantic parse of a message, goes further than a simple flat translation of a message into basic concepts. FrameNet-based semantic parsing may follow a syntactic parsing step, however spoken conversations in customer service telephone call centers present very specific characteristics such as non-canonical language, noisy messages (disfluencies, repetitions, truncated words or automatic speech transcription errors) and the presence of superfluous information. For syntactic parsing the traditional view based on context-free grammars is not suitable for processing non-canonical text. New approaches to parsing based on dependency structures and discriminative machine learning techniques are more adapted to process spontaneous speech for two main reasons: (a) they need less training data and (b) the annotation with syntactic dependencies of conversation transcripts is simpler than with syntactic constituents. Another advantage is that partial annotation can be performed. This paper presents the adaptation of a syntactic dependency parser to process very spontaneous speech recorded in a callcentre environment. This parser is used in order to produce FrameNet candidates for characterizing conversations between an operator and a caller.
Frédéric Béchet, Alexis Nasr, Benoît Favre
INTERSPEECH3
2014 Speaker adaptation of DNN-based ASR with i-vectors: does it actually adapt models to speakers?
abstract
Deep neural networks (DNN) are currently very successful for acoustic modeling in ASR systems. One of the main challenges with DNNs is unsupervised speaker adaptation from an initial speaker clustering, because DNNs have a very large number of parameters. Recently, a method has been proposed to adapt DNNs to speakers by combining speaker-specific information (in the form of i-vectors computed at the speaker-cluster level) with fMLLR-transformed acoustic features. In this paper we try to gain insight on what kind of adaptation is performed on DNNs when stacking i-vectors with acoustic features and what information exactly is carried by i-vectors. We observe on REPERE corpus that DNNs trained on i-vector features concatenated with fMLLR-transformed acoustic features lead to a gain of 0.7 points. The experiments shows that using ivector stacking in DNN acoustic models is not only performing speaker adaptation, but also adaptation to acoustic conditions.
Mickael Rouvier, Benoît Favre
INTERSPEECH2
2014 A Repository of State of the Art and Competitive Baseline Summaries for Generic News Summarization
Kai Hong, John M. Conroy, Benoît Favre, Alex Kulesza, Ani Nenkova
LREC3
2014 Automatically enriching spoken corpora with syntactic information for linguistic studies
Alexis Nasr, Frédéric Béchet, Benoît Favre, Thierry Bazillon, José Deulofeu, André Valli
LREC3
2014 Joint decoding of complementary utterances
abstract
Errors in open-domain ASR can be corrected by asking the speaker to rephrase targeted segments in utterances where they have been detected. The utterance merging problem consists in generating a better transcript from the utterance where errors have been detected and a clarification utterance. We introduce an alignment-decoding algorithm for jointly processing the two utterances and benefit from the complementary information they contain. The algorithm aligns word lattices in the WFST framework with a probabilistic cost model. Results on the BOLT-BC speech-to-speech translation task show an improvement of 2.84 points of accuracy compared to aligning the one best without joint decoding.
Mickael Rouvier, Benoît Favre, Frédéric Béchet
SLT2
2013 "Can you give me another word for hyperbaric?": Improving speech translation using targeted clarification questions
abstract
We present a novel approach for improving communication success between users of speech-to-speech translation systems by automatically detecting errors in the output of automatic speech recognition (ASR) and statistical machine translation (SMT) systems. Our approach initiates system-driven targeted clarification about errorful regions in user input and repairs them given user responses. Our system has been evaluated by unbiased subjects in live mode, and results show improved success of communication between users of the system.
Necip Fazil Ayan, Arindam Mandal, Michael W. Frandsen, Jing Zheng 0001, Peter Blasco, Andreas Kathol, Frédéric Béchet, Benoît Favre, Alex Marin, Tom Kwiatkowski, Mari Ostendorf, Luke Zettlemoyer, Philipp Salletmayr, Julia Hirschberg, Svetlana Stoyanchev
ICASSP8
2013 ASR error segment localization for spoken recovery strategy
abstract
Even though small ASR errors might not impact downstream processes that make use of the transcript, larger error segments like those generated by OOVs can have a considerable impact on applications such as speech-to-speech translation and can eventually lead to communication failure between users of the system. This work focuses on error detection in ASR output targeted towards significant error segments that can be recovered using a dialog system. We propose a CRF system trained to recognize error segments with ASR confidence-based, lexical and syntactic features. The most significant error segment is passed to a dialog system for interactive recovery in which rephrased words are reinserted in the original. 22% of utterances can be fully recovered and an interesting by-product is that rewriting error segments as a single token reduces WER by 17% on an adverse corpus.
Frédéric Béchet, Benoît Favre
ICASSP2
2013 Automatic human utility evaluation of ASR systems: does WER really predict performance?
abstract
International audience
Benoît Favre, Kyla Cheung, Siavash Kazemian, Adam Lee, Yang Liu 0004, Cosmin Munteanu, Ani Nenkova, Dennis Ochei, Gerald Penn, Stephen Tratz, Clare R. Voss, Frauke Zeller
INTERSPEECH1
2012 Detecting person presence in TV shows with linguistic and structural features
abstract
Person detection and recognition in videos is a hard problem due to the intrinsic ambiguities of the sound and image channels and their interaction. Whatever method is used to extract person hypotheses from the audio or the image channels, person recognition in videos relies on a multimodal decision process that merges the different hypotheses produced in order to decide, for each frame, who is present in the video at the audio level, at the image level or at the content level (person mention in speech or inserted text boxes). In this framework the focus of this paper is to produce a list of person presence hypotheses from the audio channel of a video document only, to be used in addition to person presence detected at the image level by a multimodal fusion process. In this study we focus on the audio channel only, using two kinds of features: linguistic features corresponding to the way a person is mentioned by a speaker; structural features corresponding to the context of occurrence of a name in a show. We show that both sets of features are complementary and that good results can be achieved on a TV show corpus annotated with person presence labels.
Frédéric Béchet, Benoît Favre, Géraldine Damnati
ICASSP2
2012 Syntactic annotation of spontaneous speech: application to call-center conversation data
Thierry Bazillon, Melanie Deplano, Frédéric Béchet, Alexis Nasr, Benoît Favre
LREC5
2012 Leveraging study of robustness and portability of spoken language understanding systems across languages and domains: the PORTMEDIA corpora
Fabrice Lefèvre, Djamel Mostefa, Laurent Besacier, Yannick Estève, Matthieu Quignard, Nathalie Camelin, Benoît Favre, Bassam Jabaian, Lina Maria Rojas-Barahona
LREC7
2011 Applying Multiclass Bandit algorithms to call-type classification
abstract
We analyze the problem of call-type classification using data that is weakly labelled. The training data is not systematically annotated, but we consider we have a weak or lazy oracle able to answer the question “Is sample x of class q?” by a simple `yes' or `no' answer. This situation of learning might be encountered in many real-world problems where the cost of labelling data is very high. We prove that it is possible to learn linear classifiers in this setting, by estimating adequate expectations inspired by the Multiclass Bandit paradgim. We propose a learning strategy that builds on Kessler's construction to learn multiclass perceptrons. We test our learning procedure against two real-world datasets from spoken langage understanding and provide compelling results.
Liva Ralaivola, Benoît Favre, Pierre Gotab, Frédéric Béchet, Géraldine Damnati
ASRU2
2010 Evaluation of semantic role labeling and dependency parsing of automatic speech recognition output
abstract
Semantic role labeling (SRL) is an important module of spoken language understanding systems. This work extends the standard evaluation metrics for joint dependency parsing and SRL of text in order to be able to handle speech recognition output with word errors and sentence segmentation errors. We propose metrics based on word alignments and bags of relations, and compare their results on the output of several SRL systems on broadcast news and conversations of the OntoNotes corpus. We evaluate and analyze the relation between the performance of the subtasks that lead to SRL, including ASR, part-of-speech tagging or sentence segmentation. The tools are made available to the community.
Benoît Favre, Bernd Bohnet, Dilek Hakkani-Tür
ICASSP1
2010 The UMUS System for Named Entity Generation at GREC 2010
Benoît Favre, Bernd Bohnet
INLG1
2010 Semi-supervised part-of-speech tagging in speech applications
abstract
When no training or adaptation data is available, semisupervised training is a good alternative for processing new domains. We perform Bayesian training of a part-of-speech (POS) tagger from unannotated text and a dictionary of possible tags for each word. We complement that method with supervised prediction of possible tags for out-of-vocabulary words and study the impact of both semi-supervision and starting dictionary size on three representative downstream tasks (named entity tagging, semantic role labeling, ASR output postprocessing) that use POS tags as features. The outcome is no impact or a small decrease in performance compared to using a fully supervised tagger, with even potential gains in case of domain mismatch for the supervised tagger. Tasks that trust the tags completely (like ASR post-processing) are more affected by a reduction of the starting dictionary, but still yield positive outcome.
Richard Dufour, Benoît Favre
INTERSPEECH2
2010 Long story short - Global unsupervised models for keyphrase based meeting summarization
Korbinian Riedhammer, Benoît Favre, Dilek Hakkani-Tür
Speech Commun.2
2010 The CALO Meeting Assistant System
abstract
The CALO Meeting Assistant (MA) provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper presents the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, topic identification and segmentation, question-answer pair identification, action item recognition, decision extraction, and summarization.
Gökhan Tür, Andreas Stolcke, L. Lynn Voss, Stanley Peters, Dilek Hakkani-Tür, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri
IEEE Trans. Speech Audio Process.7
2009 Any questions? Automatic question detection in meetings
abstract
In this paper, we describe our efforts toward the automatic detection of English questions in meetings. We analyze the utility of various features for this task, originating from three distinct classes: lexico-syntactic, turn-related, and pitch-related. Of particular interest is the use of parse tree information in classification, an approach as yet unexplored. Results from experiments on the ICSI MRDA corpus demonstrate that lexico-syntactic features are most useful for this task, with turn-and pitch-related features providing complementary information in combination. In addition, experiments using reference parse trees on the broadcast conversation portion of the OntoNotes release 2.9 data set illustrate the potential of parse trees to outperform word lexical features.
Kofi Boakye, Benoît Favre, Dilek Hakkani-Tür
ASRU2
2009 Integrating prosodic features in extractive meeting summarization
abstract
Speech contains additional information than text that can be valuable for automatic speech summarization. In this paper, we evaluate how to effectively use acoustic/prosodic features for extractive meeting summarization, and how to integrate prosodic features with lexical and structural information for further improvement. To properly represent prosodic features, we propose different normalization methods based on speaker, topic, or local context information. Our experimental results show that using only the prosodic features we achieve better performance than using the non-prosodic information on both the human transcripts and recognition output. In addition, a decision-level combination of the prosodic and non-prosodic features yields further gain, outperforming the individual models.
Shasha Xie, Dilek Hakkani-Tür, Benoît Favre, Yang Liu 0004
ASRU3
2009 Syntactically-informed models for comma prediction
abstract
Providing punctuation in speech transcripts not only improves readability, but it also helps downstream text processing such as information extraction or machine translation. In this paper, we improve by 7% the accuracy of comma prediction in English broadcast news by introducing syntactic features inspired by the role of commas as described in linguistics studies. We conduct an analysis of the impact of those features on other subsets of features (prosody, words...) when combined through CRFs. The syntactic cues can help characterizing large syntactic patterns such as appositions and lists which are not necessarily marked by prosody.
Benoît Favre, Dilek Hakkani-Tür, Elizabeth Shriberg
ICASSP1
2009 A global optimization framework for meeting summarization
abstract
We introduce a model for extractive meeting summarization based on the hypothesis that utterances convey bits of information, or concepts. Using keyphrases as concepts weighted by frequency, and an integer linear program to determine the best set of utterances, that is, covering as many concepts as possible while satisfying a length constraint, we achieve ROUGE scores at least as good as a ROUGE-based oracle derived from human summaries. This brings us to a critical discussion of ROUGE and the future of extractive meeting summarization.
Daniel Gillick, Korbinian Riedhammer, Benoît Favre, Dilek Hakkani-Tür
ICASSP3
2009 Phrase and word level strategies for detecting appositions in speech
abstract
Appositions are grammatical constructs in which two noun phrases are placed side-by-side, one modifying the other. Detecting them in speech can help extract semantic information useful, for instance, for co-reference resolution and question answering. We compare and combine three approaches: wordlevel and phrase-level classifiers, and a syntactic parser trained to generate appositions. On reference parses, the phrase-level classifier outperforms the other approaches while on automatic parses and ASR output, the combination of the appositiongenerating parser and the word-level classifier works best. An analysis of the system errors reveals that parsing accuracy and world knowledge are very important for this task.
Benoît Favre, Dilek Hakkani-Tür
INTERSPEECH1
2009 Clusterrank: a graph based method for meeting summarization
abstract
This paper presents an unsupervised, graph based approach for extractive summarization of meetings. Graph based methods such as TextRank have been used for sentence extraction from news articles. These methods model text as a graph with sentences as nodes and edges based on word overlap. A sentence node is then ranked according to its similarity with other nodes. The spontaneous speech in meetings leads to incomplete, informed sentences with high redundancy and calls for additional measures to extract relevant sentences. We propose an extension of the TextRank algorithm that clusters the meeting utterances and uses these clusters to construct the graph. We evaluate this method on the AM I meeting corpus and show a significant improvement over TextRank and other baseline methods.
Nikhil Garg 0005, Benoît Favre, Korbinian Riedhammer, Dilek Hakkani-Tür
INTERSPEECH2
2009 Combined low level and high level features for out-of-vocabulary word detection
abstract
This paper addresses the issue of Out-Of-Vocabulary (OOV) word detection in Large Vocabulary Continuous Speech Recognition (LVCSR) systems. We propose a method inspired by confidence measures, that consists in analyzing the recognition system outputs in order to automatically detect errors due to OOV words. This method combines various features based on acoustic, linguistic, decoding graph and semantics. We evaluate separately each feature and we estimate their complementarity. Experiments are conducted on a large French broadcast news corpus from the ESTER evaluation campaign. Results show good performance in real conditions: the method obtains an OOV word detection rate of 43%-90% with 2.5%-17.5% of false detection. Index Terms: OOV word detection, confidence measures, speech recognition
Benjamin Lecouteux, Georges Linarès, Benoît Favre
INTERSPEECH3
2009 Leveraging sentence weights in a concept-based optimization framework for extractive meeting summarization
abstract
International audience
Shasha Xie, Benoît Favre, Dilek Hakkani-Tür, Yang Liu 0004
INTERSPEECH2
2009 Generative and Discriminative Methods Using Morphological Information for Sentence Segmentation of Turkish
abstract
This paper presents novel methods for generative, discriminative, and hybrid sequence classification for segmentation of Turkish word sequences into sentences. In the literature, this task is generally solved using statistical models that take advantage of lexical information among others. However, Turkish has a productive morphology that generates a very large vocabulary, making the task much harder. In this paper, we introduce a new set of morphological features, extracted from words and their morphological analyses. We also extend the established method of hidden event language modeling (HELM) to factored hidden event language modeling (fHELM) to handle morphological information. In order to capture non-lexical information, we extract a set of prosodic features, which are mainly motivated from our previous work for other languages. We then employ discriminative classification techniques, boosting and conditional random fields (CRFs), combined with fHELM, for the task of Turkish sentence segmentation.
Ümit Güz, Benoît Favre, Dilek Hakkani-Tür, Gökhan Tür
IEEE Trans. Speech Audio Process.2
2008 Punctuating speech for information extraction
abstract
This paper studies the effect of automatic sentence boundary detection and comma prediction on entity and relation extraction in speech. We show that punctuating the machine generated transcript according to maximum F-measure of period and comma annotation results in suboptimal information extraction. Precisely, period and comma decision thresholds can be chosen in order to improve the entity value score and the relation value score by 4% relative. Error analysis shows that preventing noun-phrase splitting by generating longer sentences and fewer commas can be harmful for IE performance. Indeed, it seems that missed punctuation allows syntactic parsers to merge noun-phrases and prevent the extraction of correct information.
Benoît Favre, Ralph Grishman, Dustin Hillard, Heng Ji 0001, Dilek Hakkani-Tür, Mari Ostendorf
ICASSP1
2008 Packing the meeting summarization knapsack
abstract
Despite considerable work in automatic meeting summarization over the last few years, comparing results remains difficult due to varied task conditions and evaluations. To address this issue, we present a method for determining the best possible extractive summary given an evaluation metric like ROUGE. Our oracle system is based on a knapsack-packing framework, and though NP-Hard, can be solved nearly optimally by a genetic algorithm. To frame new research results in a meaningful context, we suggest presenting our oracle results alongside two simple baselines. We show oracle and baseline results for a variety of evaluation scenarios that have recently appeared in this field.
Korbinian Riedhammer, Daniel Gillick, Benoît Favre, Dilek Hakkani-Tür
INTERSPEECH3
2008 Efficient sentence segmentation using syntactic features
abstract
To enable downstream language processing,automatic speech recognition output must be segmented into its individual sentences. Previous sentence segmentation systems have typically been very local,using low-level prosodic and lexical features to independently decide whether or not to segment at each word boundary position. In this work,we leverage global syntactic information from a syntactic parser, which is better able to capture long distance dependencies. While some previous work has included syntactic features, ours is the first to do so in a tractable, lattice-based way, which is crucial for scaling up to long-sentence contexts. Specifically, an initial hypothesis lattice is constructed using local features. Candidate sentences are then assigned syntactic language model scores. These global syntactic scores are combined with local low-level scores in a log-linear model. The resulting system significantly outperforms the most popular long-span model for sentence segmentation (the hidden event language model) on both reference text and automatic speech recognizer output from news broadcasts.
Benoît Favre, Dilek Hakkani-Tür, Slav Petrov, Daniel Klein 0001
SLT1
2008 A keyphrase based approach to interactive meeting summarization
abstract
Rooted in multi-document summarization, maximum marginal relevance (MMR) is a widely used algorithm for meeting summarization (MS). A major problem in extractive MS using MMR is finding a proper query: the centroid based query which is commonly used in the absence of a manually specified query, can not significantly outperform a simple baseline system. We introduce a simple yet robust algorithm to automatically extract keyphrases (KP) from a meeting which can then be used as a query in the MMR algorithm. We show that the KP based system significantly outperforms both baseline and centroid based systems. As human refined KPs show even better summarization performance, we outline how to integrate the KP approach into a graphical user interface allowing interactive summarization to match the user's needs in terms of summary length and topic focus.
Korbinian Riedhammer, Benoît Favre, Dilek Hakkani-Tür
SLT2
2008 The CALO meeting speech recognition and understanding system
abstract
The CALO Meeting Assistant provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper summarizes the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, question-answer pair identification, action item recognition, decision extraction, and summarization.
Gökhan Tür, Andreas Stolcke, L. Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Dilek Hakkani-Tür, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Stanley Peters, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri
SLT5
2007 An interactive timeline for speech database browsing
abstract
Speech databases lack efficient interfaces to explore information along time. We introduce an interactive timeline that helps the user in browsing an audio stream on a large time scale and recontextualize targeted information. Time can be explored at different granularities using synchronized scales. We try to take advantage of automatic transcription to generate a conceptual structure of the database. The timeline is annotated with two elements to reflect the information distribution relevant to a user need. Information density is computed using an information retrieval model and displayed as a continuous shade on the timeline whereas anchorage points are expected to provide a stronger structure and to guide the user through his exploration. These points are generated using an extractive summarization algorithm. We present a prototype implementing the interactive timeline to browse broadcast news recordings.
Benoît Favre, Jean-François Bonastre, Patrice Bellot
INTERSPEECH1
2005 Mining broadcast news data: robust information extraction from word lattices
abstract
International audience
Benoît Favre, Frédéric Béchet, Pascal Nocera
INTERSPEECH1