EDBT 2026 Demo / reviewers in the wild / expert
Sahar Ghannay
dblp:171/1109
· DBLP profile ↗
31ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-7531-2522ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Multimodal LLMs Generate Pedagogical Questions?abstractInternational audience Thomas Gerald, Sahar Ghannay, Julie Lascar, Paul Lerner, Anne Vilnat |
LREC | 2 |
| 2026 | Format Matters: A Critical Evaluation of Output Formats for Prompting LLMs in SLU and NERabstractInternational audience Pierre Lepagnol, Sahar Ghannay, Thomas Gerald, Christophe Servan, Sophie Rosset |
LREC | 2 |
| 2025 | Evaluating the Confidentiality of Synthetic Clinical Texts Generated by Language Models
Foucauld Estignard, Sahar Ghannay, Julien Girard-Satabin, Nicolas Hiebel, Aurélie Névéol |
AIME (1) | 2 |
| 2025 | Entity-Aware Cross-Modal Pretraining for Knowledge-Based Visual Question Answering
Omar Adjali, Olivier Ferret, Sahar Ghannay, Hervé Le Borgne |
ECIR (3) | 3 |
| 2025 | XAI for Gender Representation in Media AnalysisabstractIn many countries, studies have highlighted the under-representation of women in the media. But beyond quantitative imbalance is the question of the qualitative asymmetry of men and women portrayals. How to help the evaluation of content and salient features specific to male and female discourse? We propose in this study to leverage the knowledge acquired by a classification model trained for gender detection on automatic transcripts, to highlight patterns distinctive of male or female speech. Results show the relevance of coupling explainable AI with classifier confidence to compute consistent attributions. François Buet, Camille Guinaudeau, Cyril Grouin, Sahar Ghannay, Shin'ichi Satoh 0001 |
ICASSP | 4 |
| 2025 | Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot LearningabstractInternational audience Pierre Lepagnol, Sahar Ghannay, Thomas Gerald, Christophe Servan, Sophie Rosset |
INTERSPEECH | 2 |
| 2025 | RNA-TorsionBERT: leveraging language models for RNA 3D torsion angles predictionabstractMOTIVATION: Predicting the 3D structure of RNA is an ongoing challenge that has yet to be completely addressed despite continuous advancements. RNA 3D structures rely on distances between residues and base interactions but also backbone torsional angles. Knowing the torsional angles for each residue could help reconstruct its global folding, which is what we tackle in this work. This paper presents a novel approach for directly predicting RNA torsional angles from raw sequence data. Our method draws inspiration from the successful application of language models in various domains and adapts them to RNA. RESULTS: We have developed a language-based model, RNA-TorsionBERT, incorporating better sequential interactions for predicting RNA torsional and pseudo-torsional angles from the sequence only. Through extensive benchmarking, we demonstrate that our method improves the prediction of torsional angles compared to state-of-the-art methods. In addition, by using our predictive model, we have inferred a torsion angle-dependent scoring function, called TB-MCQ, that replaces the true reference angles by our model prediction. We show that it accurately evaluates the quality of near-native predicted structures, in terms of RNA backbone torsion angle values. Our work demonstrates promising results, suggesting the potential utility of language models in advancing RNA 3D structure prediction. AVAILABILITY AND IMPLEMENTATION: Source code is freely available on the EvryRNA platform: https://evryrna.ibisc.univ-evry.fr/evryrna/RNA-TorsionBERT. Clement Bernard, Guillaume Postic, Sahar Ghannay, Fariza Tahi |
Bioinform. | 3 |
| 2024 | New Semantic Task for the French Spoken Language Understanding MEDIA BenchmarkabstractIntent classification and slot-filling are essential tasks of Spoken Language Understanding (SLU). In most SLU systems, those tasks are realized by independent modules, but for about fifteen years, models achieving both of them jointly and exploiting their mutual enhancement have been proposed. A multilingual module using a joint model was envisioned to create a touristic dialogue system for a European project, HumanE-AI-Net. A combination of multiple datasets, including the MEDIA dataset, was suggested for training this joint model. The MEDIA SLU dataset is a French dataset distributed since 2005 by ELRA, mainly used by the French research community and free for academic research since 2020. Unfortunately, it is annotated only in slots but not intents. An enhanced version of MEDIA annotated with intents has been built to extend its use to more tasks and use cases. This paper presents the semi-automatic methodology used to obtain this enhanced version. In addition, we present the first results of SLU experiments on this enhanced dataset using joint models for intent classification and slot-filling. Nadège Alavoine, Gaëlle Laperrière, Christophe Servan, Sahar Ghannay, Sophie Rosset |
LREC/COLING | 4 |
| 2024 | Small Language Models Are Good Too: An Empirical Study of Zero-Shot ClassificationabstractThis study is part of the debate on the efficiency of large versus small language models for text classification by prompting. We assess the performance of small language models in zero-shot text classification, challenging the prevailing dominance of large models. Across 15 datasets, our investigation benchmarks language models from 77M to 40B parameters using different architectures and scoring functions. Our findings reveal that small models can effectively classify texts, getting on par with or surpassing their larger counterparts. We developed and shared a comprehensive open-source repository that encapsulates our methodologies. This research underscores the notion that bigger isn’t always better, suggesting that resource-efficient small models may offer viable solutions for specific data classification challenges. Pierre Lepagnol, Thomas Gerald, Sahar Ghannay, Christophe Servan, Sophie Rosset |
LREC/COLING | 3 |
| 2024 | mALBERT: Is a Compact Multilingual BERT Model Still Worth It?abstractWithin the current trend of Pretained Language Models (PLM), emerge more and more criticisms about the ethical and ecological impact of such models. In this article, considering these critical remarks, we propose to focus on smaller models, such as compact models like ALBERT, which are more ecologically virtuous than these PLM. However, PLMs enable huge breakthroughs in Natural Language Processing tasks, such as Spoken and Natural Language Understanding, classification, Question–Answering tasks. PLMs also have the advantage of being multilingual, and, as far as we know, a multilingual version of compact ALBERT models does not exist. Considering these facts, we propose the free release of the first version of a multilingual compact ALBERT model, pre-trained using Wikipedia data, which complies with the ethical aspect of such a language model. We also evaluate the model against classical multilingual PLMs in classical NLP tasks. Finally, this paper proposes a rare study on the subword tokenization impact on language performances. Christophe Servan, Sahar Ghannay, Sophie Rosset |
LREC/COLING | 2 |
| 2024 | Multi-Level Information Retrieval Augmented Generation for Knowledge-based Visual Question AnsweringabstractThe Knowledge-Aware Visual Question Answering about Entity task aims to disambiguate entities using textual and visual information, as well as knowledge.It usually relies on two independent steps, information retrieval then reading comprehension, that do not benefit each other.Retrieval Augmented Generation (RAG) offers a solution by using generated answers as feedback for retrieval training.RAG usually relies solely on pseudo-relevant passages retrieved from external knowledge bases which can lead to ineffective answer generation.In this work, we propose a multi-level information RAG approach that enhances answer generation through entity retrieval and query expansion.We formulate a joint-training RAG loss such that answer generation is conditioned on both entity and passage retrievals.We show through experiments new state-of-the-art performance on the VIQuAE KB-VQA benchmark and demonstrate that our approach can help retrieve more actual relevant knowledge to generate accurate answers. Omar Adjali, Olivier Ferret, Sahar Ghannay, Hervé Le Borgne |
EMNLP | 3 |
| 2024 | A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language UnderstandingabstractSelf-Supervised Learning is vastly used to efficiently represent speech for Spoken Language Understanding, gradually replacing conventional approaches. Meanwhile, textual SSL models are proposed to encode language-agnostic semantics. SAMU-XLSR framework employed this semantic information to enrich multilingual speech representations. A recent study investigated SAMU-XLSR in-domain semantic enrichment by specializing it on downstream transcriptions, leading to state-of-the-art results on a challenging SLU task. This study's interest lies in the loss of multilingual performances and lack of specific-semantics training induced by such specialization in close languages without any SLU implication. We also consider SAMU-XLSR's loss of initial cross-lingual abilities due to a separate SLU fine-tuning. Therefore, this paper proposes a dual task learning approach to improve SAMU-XLSR semantic enrichment while considering distant languages for multilingual and language portability experiments. Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Yannick Estève |
INTERSPEECH | 2 |
| 2024 | RNAdvisor: a comprehensive benchmarking tool for the measure and prediction of RNA structural model qualityabstractRNA is a complex macromolecule that plays central roles in the cell. While it is well known that its structure is directly related to its functions, understanding and predicting RNA structures is challenging. Assessing the real or predictive quality of a structure is also at stake with the complex 3D possible conformations of RNAs. Metrics have been developed to measure model quality while scoring functions aim at assigning quality to guide the discrimination of structures without a known and solved reference. Throughout the years, many metrics and scoring functions have been developed, and no unique assessment is used nowadays. Each developed assessment method has its specificity and might be complementary to understanding structure quality. Therefore, to evaluate RNA 3D structure predictions, it would be important to calculate different metrics and/or scoring functions. For this purpose, we developed RNAdvisor, a comprehensive automated software that integrates and enhances the accessibility of existing metrics and scoring functions. In this paper, we present our RNAdvisor tool, as well as state-of-the-art existing metrics, scoring functions and a set of benchmarks we conducted for evaluating them. Source code is freely available on the EvryRNA platform: https://evryrna.ibisc.univ-evry.fr. Clement Bernard, Guillaume Postic, Sahar Ghannay, Fariza Tahi |
Briefings Bioinform. | 3 |
| 2023 | Semantic Enrichment Towards Efficient Speech RepresentationsabstractOver the past few years, self-supervised learned speech representations have emerged as fruitful replacements for conventional surface representations when solving Spoken Language Understanding (SLU) tasks.Simultaneously, multilingual models trained on massive textual data were introduced to encode language agnostic semantics.Recently, the SAMU-XLSR approach introduced a way to make profit from such textual models to enrich multilingual speech representations with language agnostic semantics.By aiming for better semantic extraction on a challenging Spoken Language Understanding task and in consideration with computation costs, this study investigates a specific in-domain semantic enrichment of the SAMU-XLSR model by specializing it on a small amount of transcribed data from the downstream task.In addition, we show the benefits of the use of same-domain French and Italian benchmarks for low-resource language portability and explore cross-domain capacities of the enriched SAMU-XLSR. Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Yannick Estève |
INTERSPEECH | 3 |
| 2023 | Explicit Knowledge Integration for Knowledge-Aware Visual Question Answering about Named EntitiesabstractRecent years have shown unprecedented growth of interest in Vision-Language related tasks, with the need to address the inherent challenges of integrating linguistic and visual information to solve real-world applications. Such a typical task is Visual Question Answering (VQA), which aims to answer questions about visual content. The limitations of the VQA task in terms of question redundancy and poor linguistic variability encouraged researchers to propose Knowledge-aware Visual Question Answering tasks as a natural extension of VQA. In this paper, we tackle the KVQAE (Knowledge-based Visual Question Answering about named Entities) task, which proposes to answer questions about named entities defined in a knowledge base and grounded in visual content. In particular, besides the textual and visual information, we propose to leverage the structural information extracted from syntactic dependency trees and external knowledge graphs to help answer questions about a large spectrum of entities of various types. Thus, by combining contextual and graph-based representations using Graph Convolutional Networks (GCNs), we are able to learn meaningful embeddings for Information Retrieval tasks. Experiments on the ViQuAE public dataset show how our approach improves the state-of-the-art baselines while demonstrating the interest of injecting external knowledge to enhance multimodal information retrieval. Omar Adjali, Paul Grimal, Olivier Ferret, Sahar Ghannay, Hervé Le Borgne |
ICMR | 4 |
| 2022 | Benchmarking Transformers-based models on French Spoken Language Understanding tasksabstractInternational audience Oralie Cattan, Sahar Ghannay, Christophe Servan, Sophie Rosset |
INTERSPEECH | 2 |
| 2022 | The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation toolsabstractWith the emergence of neural end-to-end approaches for spoken language understanding (SLU), a growing number of studies have been presented during these last three years on this topic. The major part of these works addresses the spoken language understanding domain through a simple task like speech intent detection. In this context, new benchmark datasets have also been produced and shared with the community related to this task. In this paper, we focus on the French MEDIA SLU dataset, distributed since 2005 and used as a benchmark dataset for a large number of research works. This dataset has been shown as being the most challenging one among those accessible to the research community. Distributed by ELRA, this corpus is free for academic research since 2019. Unfortunately, the MEDIA dataset is not really used beyond the French research community. To facilitate its use, a complete recipe, including data preparation, training and evaluation scripts, has been built and integrated to SpeechBrain, an already popular open-source and all-in-one conversational AI toolkit based on PyTorch. This recipe is presented in this paper. In addition, based on the feedback of some researchers who have worked on this dataset for several years, some corrections have been brought to the initial manual annotation: the new version of the data will also be integrated into the ELRA catalogue, as the original one. More, a significant amount of data collected during the construction of the MEDIA corpus in the 2000s was never used until now: we present the first results reached on this subset — also included in the MEDIA SpeechBrain recipe — , that will be used for now as the MEDIA test2. Last, we discuss evaluation issues. Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar, Nathalie Camelin, Sahar Ghannay, Bassam Jabaian, Yannick Estève |
LREC | 6 |
| 2022 | Impact Analysis of the Use of Speech and Language Models Pretrained by Self-Supersivion for Spoken Language UnderstandingabstractPretrained models through self-supervised learning have been recently introduced for both acoustic and language modeling. Applied to spoken language understanding tasks, these models have shown their great potential by improving the state-of-the-art performances on challenging benchmark datasets. In this paper, we present an error analysis reached by the use of such models on the French MEDIA benchmark dataset, known as being one of the most challenging benchmarks for the slot filling task among all the benchmarks accessible to the entire research community. One year ago, the state-of-art system reached a Concept Error Rate (CER) of 13.6% through the use of a end-to-end neural architecture. Some months later, a cascade approach based on the sequential use of a fine-tuned wav2vec2.0 model and a fine-tuned BERT model reaches a CER of 11.2%. This significant improvement raises questions about the type of errors that remain difficult to treat, but also about those that have been corrected using these models pre-trained through self-supervision learning on a large amount of data. This study brings some answers in order to better understand the limits of such models and open new perspectives to continue improving the performance. Salima Mdhaffar, Valentin Pelloin, Antoine Caubrière, Gaëlle Laperrière, Sahar Ghannay, Bassam Jabaian, Nathalie Camelin, Yannick Estève |
LREC | 5 |
| 2022 | Continual Self-Supervised Domain Adaptation for End-to-End Speaker DiarizationabstractIn conventional domain adaptation for speaker diarization, a large collection of annotated conversations from the target domain is required. In this work, we propose a novel continual training scheme for domain adaptation of an end-to-end speaker diarization system, which processes one conversation at a time and benefits from full self-supervision thanks to pseudo-labels. The qualities of our method allow for autonomous adaptation (e.g. of a voice assistant to a new house-hold), while also avoiding permanent storage of possibly sensitive user conversations. We experiment extensively on the 11 domains of the DIHARD III corpus and show the effectiveness of our approach with respect to a pre-trained base-line, achieving a relative 17% performance improvement. We also find that data augmentation and a well-defined target domain are key factors to avoid divergence and to benefit from transfer. Juan Manuel Coria, Hervé Bredin, Sahar Ghannay, Sophie Rosset |
SLT | 3 |
| 2021 | Overlap-Aware Low-Latency Online Speaker Diarization Based on End-to-End Local SegmentationabstractWe propose to address online speaker diarization as a combination of incremental clustering and local diarization applied to a rolling buffer updated every 500ms. Every single step of the proposed pipeline is designed to take full advantage of the strong ability of a recently proposed end-to-end overlap-aware segmentation to detect and separate overlapping speakers. In particular, we propose a modified version of the statistics pooling layer (initially introduced in the x-vector architecture) to give less weight to frames where the segmentation model predicts simultaneous speakers. Furthermore, we derive cannot-link constraints from the initial segmentation step to prevent two local speakers from being wrongfully merged during the incremental clustering step. Finally, we show how the latency of the proposed approach can be adjusted between 500ms and 5s to match the requirements of a particular use case, and we provide a systematic analysis of the influence of latency on the overall performance (on AMI, DIHARD and VoxConverse). Juan Manuel Coria, Hervé Bredin, Sahar Ghannay, Sophie Rosset |
ASRU | 3 |
| 2020 | Neural Networks approaches focused on French Spoken Language Understanding: application to the MEDIA Evaluation TaskabstractIn this paper, we present a study on a French Spoken Language Understanding (SLU) task: the MEDIA task.Many works and studies have been proposed for many tasks, but most of them are focused on English language and tasks.The exploration of a richer language like French within the framework of a SLU task implies to recent approaches to handle this difficulty.Since the MEDIA task seems to be one of the most difficult, according to several previous studies, we propose to explore Neural Networks approaches focusing of three aspects: firstly, the Neural Network inputs and more specifically the word embeddings; secondly, we compared French version of BERT against the best setup through different ways; Finally, the comparison against State-of-the-Art approaches.Results show that the word embeddings trained on a small corpus need to be updated during SLU model training.Furthermore, the French BERT fine-tuned approaches outperform the classical Neural Network Architectures and achieves state of the art results.However, the contextual embeddings extracted from one of the French BERT approaches achieve comparable results in comparison to word embedding, when integrated into the proposed neural architecture. Sahar Ghannay, Christophe Servan, Sophie Rosset |
COLING | 1 |
| 2020 | Error Analysis Applied to End-to-End Spoken Language UnderstandingabstractThis paper presents a qualitative study of errors produced by an end-to-end spoken language understanding (SLU) system (speech signal to concepts) that reaches state of the art performance. Different studies are proposed to better understand the weaknesses of such systems: comparison to a classical pipeline SLU system, a study on the cause of concept deletions (the most frequent error), observation of a problem in the capability of the end-to-end SLU system to segment correctly concepts, analysis of the system behavior to process unseen concept/value pairs, analysis of the benefit of the curriculum-based transfer learning approach. Last, we proposed a way to compute embeddings of sub-sequences that seem to contain relevant information for future work. Antoine Caubrière, Sahar Ghannay, Natalia A. Tomashenko, Renato De Mori, Antoine Laurent, Emmanuel Morin, Yannick Estève |
ICASSP | 2 |
| 2020 | What is best for spoken language understanding: small but task-dependant embeddings or huge but out-of-domain embeddings?abstractWord embeddings are shown to be a great asset for several Natural Language and Speech Processing tasks. While they are already evaluated on various NLP tasks, their evaluation on spoken or natural language understanding (SLU) is less studied. The goal of this study is two-fold: firstly, it focuses on semantic evaluation of common word embeddings approaches for SLU task; secondly, it investigates the use of two different data sets to train the embeddings: small and task-dependent corpus or huge and out-of-domain corpus. Experiments are carried out on 5 benchmark corpora (ATIS, SNIPS, SNIPS70, M2M, MEDIA), on which a relevance ranking was proposed in the literature. Interestingly, the performance of the embeddings is independent of the difficulty of the corpora. Moreover, the embeddings trained on huge and out-of-domain corpus yields to better results than the ones trained on small and task-dependent corpus. Sahar Ghannay, Antoine Neuraz, Sophie Rosset |
ICASSP | 1 |
| 2020 | A Cooking Knowledge Graph and Benchmark for Question Answering Evaluation in Lifelong Learning Scenarios
Mathilde Veron, Anselmo Peñas, Guillermo Echegoyen, Somnath Banerjee 0001, Sahar Ghannay, Sophie Rosset |
NLDB | 5 |
| 2020 | A study of continuous space word and sentence representations applied to ASR error detection
Sahar Ghannay, Yannick Estève, Nathalie Camelin |
Speech Commun. | 1 |
| 2018 | Task Specific Sentence Embeddings for ASR Error DetectionabstractInternational audience Sahar Ghannay, Yannick Estève, Nathalie Camelin |
INTERSPEECH | 1 |
| 2018 | Simulating ASR errors for training SLU systems
Edwin Simonnet, Sahar Ghannay, Nathalie Camelin, Yannick Estève |
LREC | 2 |
| 2018 | End-To-End Named Entity And Semantic Concept Extraction From SpeechabstractNamed entity recognition (NER) is among SLU tasks that usually extract semantic information from textual documents. Until now, NER from speech is made through a pipeline process that consists in processing first an automatic speech recognition (ASR) on the audio and then processing a NER on the ASR outputs. Such approach has some disadvantages (error propagation, metric to tune ASR systems sub-optimal in regards to the final task, reduced space search at the ASR output level,...) and it is known that more integrated approaches outperform sequential ones, when they can be applied. In this paper, we explore an end-to-end approach that directly extracts named entities from speech, though a unique neural architecture. On a such way, a joint optimization is possible for both ASR and NER. Experiments are carried on French data easily accessible, composed of data distributed in several evaluation campaigns. The results are promising since this end-to-end approach provides similar results (F-measure=0.66 on test data) than a classical pipeline approach to detect named entity categories (F-measure=0.64). Last, we also explore this approach applied to semantic concept extraction, through a slot filling task known as a spoken language understanding problem, and also observe an improvement in comparison to a pipeline approach. Sahar Ghannay, Antoine Caubrière, Yannick Estève, Nathalie Camelin, Edwin Simonnet, Antoine Laurent, Emmanuel Morin |
SLT | 1 |
| 2017 | ASR Error Management for Improving Spoken Language UnderstandingabstractThis paper addresses the problem of automatic speech recognition (ASR) error detection and their use for improving spoken language understanding (SLU) systems. In this study, the SLU task consists in automatically extracting, from ASR transcriptions , semantic concepts and concept/values pairs in a e.g touristic information system. An approach is proposed for enriching the set of semantic labels with error specific labels and by using a recently proposed neural approach based on word embeddings to compute well calibrated ASR confidence measures. Experimental results are reported showing that it is possible to decrease significantly the Concept/Value Error Rate with a state of the art system, outperforming previously published results performance on the same experimental data. It also shown that combining an SLU approach based on conditional random fields with a neural encoder/decoder attention based architecture , it is possible to effectively identifying confidence islands and uncertain semantic output segments useful for deciding appropriate error handling actions by the dialogue manager strategy . Edwin Simonnet, Sahar Ghannay, Nathalie Camelin, Yannick Estève, Renato De Mori |
INTERSPEECH | 2 |
| 2016 | Acoustic Word Embeddings for ASR Error DetectionabstractInternational audience Sahar Ghannay, Yannick Estève, Nathalie Camelin, Paul Deléglise |
INTERSPEECH | 1 |
| 2016 | Word Embedding Evaluation and Combination
Sahar Ghannay, Benoît Favre, Yannick Estève, Nathalie Camelin |
LREC | 1 |