VLDB 2026 Research / reviewers in the wild / expert
Camille Guinaudeau
dblp:44/7135
· DBLP profile ↗
18ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-7249-8715ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | XAI for Gender Representation in Media AnalysisabstractIn many countries, studies have highlighted the under-representation of women in the media. But beyond quantitative imbalance is the question of the qualitative asymmetry of men and women portrayals. How to help the evaluation of content and salient features specific to male and female discourse? We propose in this study to leverage the knowledge acquired by a classification model trained for gender detection on automatic transcripts, to highlight patterns distinctive of male or female speech. Results show the relevance of coupling explainable AI with classifier confidence to compute consistent attributions. François Buet, Camille Guinaudeau, Cyril Grouin, Sahar Ghannay, Shin'ichi Satoh 0001 |
ICASSP | 2 |
| 2025 | AD2AT: Audio Description to Alternative Text, a Dataset of Alternative Text from Movies
Élise Lincker, Camille Guinaudeau, Shin'ichi Satoh 0001 |
MMM (1) | 2 |
| 2025 | Evaluating VQA Models' Consistency in the Scientific Domain
Khanh-An C. Quan, Camille Guinaudeau, Shin'ichi Satoh 0001 |
MMM (4) | 2 |
| 2025 | Towards Inclusive Education: Multimodal Classification of Textbook Images for Accessibility
Saumya Yadav, Élise Lincker, Caroline Huron, Stéphanie Martin, Camille Guinaudeau, Shin'ichi Satoh 0001, Jainendra Shukla |
MMM (4) | 5 |
| 2024 | Cross-Modal Retrieval for Knowledge-Based Visual Question Answering
Paul Lerner, Olivier Ferret, Camille Guinaudeau |
ECIR (1) | 3 |
| 2024 | Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free MetricsabstractAutomatic metrics are used as proxies to evaluate abstractive summarization systems when human annotations are too expensive.To be useful, these metrics should be fine-grained, show a high correlation with human annotations, and ideally be independent of reference quality; however, most standard evaluation metrics for summarization are reference-based, and existing reference-free metrics correlate poorly with relevance, especially on summaries of longer documents.In this paper, we introduce a reference-free metric that correlates well with human evaluated relevance, while being very cheap to compute.We show that this metric can also be used alongside reference-based metrics to improve their robustness in low quality reference settings. Théo Gigant, Camille Guinaudeau, Marc Décombas, Frédéric Dufaux |
EMNLP | 2 |
| 2023 | TIB: A Dataset for Abstractive Summarization of Long Multimodal Videoconference RecordsabstractLarge language models and multimodal language-vision models give impressive results on current available summarization benchmarks, but are not designed to handle long multimodal documents. Most summarization datasets are composed of either mono-modal documents or short multimodal documents. In order to develop models designed for understanding and summarizing real-world videoconference records that are typically around 1 hour long, we propose a dataset of 9,103 videoconference records extracted from the German National Library of Science and Technology (TIB) archive, along with their abstract. Additionally, we process the content using automatic tools in order to provide the transcripts and key frames. Finally, we present experiments for abstractive summarization, to serve as baseline for future research work in multimodal approaches. Théo Gigant, Frédéric Dufaux, Camille Guinaudeau, Marc Décombas |
CBMI | 3 |
| 2023 | Noisy and Unbalanced Multimodal Document Classification: Textbook Exercises as a Use CaseabstractIn order to foster inclusive education, automatic systems that can adapt textbooks to make them accessible to children with Developmental Coordination Disorder (DCD) are necessary. In this context, we propose a task to classify exercises according to their DCD adaptation type. We introduce a challenging exercise dataset extracted from French textbooks, with two major difficulties: limited and unbalanced, noisy data. To set a baseline on the dataset, we use state-of-the-art models combined through early and late fusion techniques to take advantage of text and vision/layout modalities. Our approach achieves an overall accuracy of 0.802. However, the experiments show the difficulty of the task, especially for minority classes, where the accuracy drops to 0.583. Élise Lincker, Camille Guinaudeau, Olivier Pons, Jérôme Dupire, Céline Hudelot, Vincent Mousseau, Isabelle Barbet, Caroline Huron |
CBMI | 2 |
| 2023 | Multimodal Inverse Cloze Task for Knowledge-Based Visual Question Answering
Paul Lerner, Olivier Ferret, Camille Guinaudeau |
ECIR (1) | 3 |
| 2022 | Bazinga! A Dataset for Multi-Party Dialogues StructuringabstractWe introduce a dataset built around a large collection of TV (and movie) series. Those are filled with challenging multi-party dialogues. Moreover, TV series come with a very active fan base that allows the collection of metadata and accelerates annotation. With 16 TV and movie series, Bazinga! amounts to 400+ hours of speech and 8M+ tokens, including 500K+ tokens annotated with the speaker, addressee, and entity linking information. Along with the dataset, we also provide a baseline for speaker diarization, punctuation restoration, and person entity recognition. The results demonstrate the difficulty of the tasks and of transfer learning from models trained on mono-speaker audio or written text, which is more widely available. This work is a step towards better multi-party dialogue structuring and understanding. Bazinga! is available at hf.co/bazinga. Because (a large) part of Bazinga! is only partially annotated, we also expect this dataset to foster research towards self- or weakly-supervised learning methods. Paul Lerner, Juliette Bergoënd, Camille Guinaudeau, Hervé Bredin, Benjamin Maurice, Sharleyne Lefevre, Martin Bouteiller, Aman Berhe, Léo Galmant, Ruiqing Yin, Claude Barras |
LREC | 3 |
| 2022 | ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named EntitiesabstractWhether to retrieve, answer, translate, or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in answering questions about named entities grounded in a visual context using a Knowledge Base (KB). To benchmark this task, called KVQAE (Knowledge-based Visual Question Answering about named Entities), we provide ViQuAE, a dataset of 3.7K questions paired with images. This is the first KVQAE dataset to cover a wide range of entity types (e.g. persons, landmarks, and products). The dataset is annotated using a semi-automatic method. We also propose a KB composed of 1.5M Wikipedia articles paired with images. To set a baseline on the benchmark, we address KVQAE as a two-stage problem: Information Retrieval and Reading Comprehension, with both zero- and few-shot learning methods. The experiments empirically demonstrate the difficulty of the task, especially when questions are not about persons. This work paves the way for better multimodal entity representations and question answering. The dataset, KB, code, and semi-automatic annotation pipeline are freely available at https://github.com/PaulLerner/ViQuAE. Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, José G. Moreno 0001, Jesús Lovón-Melgarejo |
SIGIR | 3 |
| 2014 | TVD: A Reproducible and Multiply Aligned TV Series Dataset
Anindya Roy, Camille Guinaudeau, Hervé Bredin, Claude Barras |
LREC | 2 |
| 2013 | Graph-based Local Coherence Modeling
Camille Guinaudeau, Michael Strube 0001 |
ACL (1) | 1 |
| 2013 | Multimedia information seeking through search and hyperlinkingabstractSearching for relevant webpages and following hyperlinks to related content is a widely accepted and effective approach to information seeking on the textual web. Existing work on multimedia information retrieval has focused on search for individual relevant items or on content linking without specific attention to search results. We describe our research exploring integrated multimodal search and hyperlinking for multimedia data. Our investigation is based on the MediaEval 2012 Search and Hyperlinking task. This includes a known-item search task using the Blip10000 internet video collection, where automatically created hyperlinks link each relevant item to related items within the collection. The search test queries and link assessment for this task was generated using the Amazon Mechanical Turk crowdsourcing platform. Our investigation examines a range of alternative methods which seek to address the challenges of search and hyperlinking using multimodal approaches. The results of our experiments are used to propose a research agenda for developing effective techniques for search and hyperlinking of multimedia content. Maria Eskevich, Gareth J. F. Jones, Robin Aly, Roeland Ordelman, Danish Nadeem, Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot, Tom De Nies, Pedro Debevere, Rik Van de Walle, Petra Galuscáková, Pavel Pecina, Martha A. Larson |
ICMR | 7 |
| 2012 | Enhancing lexical cohesion measure with confidence measures, semantic relations and language model interpolation for multimedia spoken content topic segmentation
Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot |
Comput. Speech Lang. | 1 |
| 2011 | Accounting for Prosodic Information to Improve ASR-Based Topic Tracking for TV Broadcast NewsabstractThe increasing quantity of video material available on line requires improved methods to help users navigate such data, among which are topic tracking techniques. The goal of this paper is to show that prosodic information can improve an ASR based topic tracking system for French TV Broadcast News. To this end, two kinds of prosodic information--extracted with and without a learning phase--are integrated in the system. This integration shows significant improvements in the F1-measure, by 13 and 8 points for the two techniques compared with the baseline system. Camille Guinaudeau, Julia Hirschberg |
INTERSPEECH | 1 |
| 2010 | Improving ASR-based topic segmentation of TV programs with confidence measures and semantic relationsabstractInternational audience Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot |
INTERSPEECH | 1 |
| 2009 | Can Automatic Speech Transcripts Be Used for Large Scale TV Stream Description and Structuring?abstractThe increasing quantity of TV material requires methods to help users navigate such data streams. Automatically associating a short textual description to each program in a stream, is a first stage to navigating or structuring tasks. Speech contained in TV broadcasts---accessible by means of automatic speech recognition systems in the absence of closed caption---is a highly valuable semantic clue that might be used to link existing textual description such as program guides, with video segments corresponding to program. However, high word error rates are to be expected on some programs, likely to jeopardize the usefulness of transcripts. The goal of this article is to determine to what extent automatic transcripts of TV streams, for various types of programs, can be used for structuring or navigating tasks. To this end, word-based and phonetic-based automatic association between video segments and program descriptions is used as a case study. We show that descriptions from a program guide can be associated with video segments with an accuracy of up to 65% and provide a valuable description to validate existing program labels. Such associations constitute a first stage for structuring task as they enable video segment textual characterization. Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot |
ISM | 1 |