Florian Boudin

dblp:54/2762 · DBLP profile ↗
← Back
26ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0001-5849-2261ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 9 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
abstract
With the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are a core component of scientific work, often presented in varying formats such as tables or charts. Understanding how robust current multimodal large language models (multimodal LLMs) are at verifying scientific claims across different evidence formats remains an important and underexplored challenge. In this paper, we design and conduct a series of experiments to assess the ability of multimodal LLMs to verify scientific claims using both tables and charts as evidence. To enable this evaluation, we adapt two existing datasets of scientific papers by incorporating annotations and structures necessary for a multimodal claim verification task. Using this adapted dataset, we evaluate 12 multimodal LLMs and find that current models perform better with table-based evidence while struggling with chart-based evidence. We further conduct human evaluations and observe that humans maintain strong performance across both formats, unlike the models. Our analysis also reveals that smaller multimodal LLMs (under 8B) show weak correlation in performance between table-based and chart-based tasks, indicating limited cross-modal generalization. These findings highlight a critical gap in current models' multimodal reasoning capabilities. We suggest that future multimodal LLMs should place greater emphasis on improving chart understanding to better support scientific claim verification.
Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa
AAAI4
2026 SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, André Greiner-Petter, Akiko Aizawa
LREC5
2026 Evaluating the Homogeneity of Keyphrase Prediction Models
Maël Houbre, Florian Boudin, Béatrice Daille
LREC2
2025 Identifying Reliable Evaluation Metrics for Scientific Text Revision
abstract
International audience
Léane Jourdan, Nicolas Hernandez, Florian Boudin, Richard Dufour
ACL (1)3
2025 ACL-rlg: A Dataset for Reading List Generation
abstract
Familiarizing oneself with a new scientific field and its existing literature can be daunting due to the large amount of available articles. Curated lists of academic references, or reading lists, compiled by experts, offer a structured way to gain a comprehensive overview of a domain or a specific scientific challenge. In this work, we introduce ACL-rlg, the largest open expert-annotated reading list dataset. We also provide multiple baselines for evaluating reading list generation and formally define it as a retrieval task. Our qualitative study highlights that traditional scholarly search engines and indexing methods perform poorly on this task, and GPT-4o, despite showing better results, exhibits signs of potential data contamination.
Julien Aubert-Béduchaud, Florian Boudin, Béatrice Daille, Richard Dufour
COLING2
2025 Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
Kon Woo Kim, Rezarta Islamaj Dogan, Jin-Dong Kim, Florian Boudin, Akiko Aizawa
NLDB (2)4
2024 CASIMIR: A Corpus of Scientific Articles Enhanced with Multiple Author-Integrated Revisions
abstract
Writing a scientific article is a challenging task as it is a highly codified and specific genre, consequently proficiency in written communication is essential for effectively conveying research findings and ideas. In this article, we propose an original textual resource on the revision step of the writing process of scientific articles. This new dataset, called CASIMIR, contains the multiple revised versions of 15,646 scientific articles from OpenReview, along with their peer reviews. Pairs of consecutive versions of an article are aligned at sentence-level while keeping paragraph location information as metadata for supporting future revision studies at the discourse level. Each pair of revised sentences is enriched with automatically extracted edits and associated revision intention. To assess the initial quality on the dataset, we conducted a qualitative study of several state-of-the-art text revision approaches and compared various evaluation metrics. Our experiments led us to question the relevance of the current evaluation methods for the text revision task.
Léane Jourdan, Florian Boudin, Nicolas Hernandez, Richard Dufour
LREC/COLING2
2022 From Fundamentals to Recent Advances: A Tutorial on Keyphrasification
Debanjan Mahata, Florian Boudin
ECIR (2)3
2022 Cross-lingual and Cross-domain Transfer Learning for Automatic Term Extraction from Low Resource Data
abstract
Automatic Term Extraction (ATE) is a key component for domain knowledge understanding and an important basis for further natural language processing applications. Even with persistent improvements, ATE still exhibits weak results exacerbated by small training data inherent to specialized domain corpora. Recently, transformers-based deep neural models, such as BERT, have proven to be efficient in many downstream NLP tasks. However, no systematic evaluation of ATE has been conducted so far. In this paper, we run an extensive study on fine-tuning pre-trained BERT models for ATE. We propose strategies that empirically show BERT’s effectiveness using cross-lingual and cross-domain transfer learning to extract single and multi-word terms. Experiments have been conducted on four specialized domains in three languages. The obtained results suggest that BERT can capture cross-domain and cross-lingual terminologically-marked contexts shared by terms, opening a new design-pattern for ATE.
Amir Hazem, Mérième Bouhandi, Florian Boudin, Béatrice Daille
LREC3
2022 Extraction and evaluation of formulaic expressions used in scholarly papers
abstract
Formulaic expressions, such as ‘in this paper we propose’, are helpful for authors of scholarly papers because they convey communicative functions; in the above, it is ‘showing the aim of this paper’. Thus, resources of formulaic expressions, such as a dictionary, that could be looked up easily would be useful. However, forms of formulaic expressions can often vary to a great extent. For example, ‘in this paper we propose’, ‘in this study we propose’ and ‘in this paper we propose a new method to’ are all regarded as formulaic expressions. Such a diversity of spans and forms causes problems in both extraction and evaluation of formulaic expressions. In this paper, we propose a new approach that is robust to variation of spans and forms of formulaic expressions. Our approach regards a sentence as consisting of a formulaic part and non-formulaic part. Then, instead of trying to extract formulaic expressions from a whole corpus, by extracting them from each sentence, different forms can be dealt with at once. Based on this formulation, to avoid the diversity problem, we propose evaluating extraction methods by how much they convey specific communicative functions rather than by comparing extracted expressions to an existing lexicon. We also propose a new extraction method that utilises named entities and dependency structures to remove the non-formulaic part from a sentence. Experimental results show that the proposed extraction method achieved the best performance compared to other existing methods.
Kenichi Iwatsuki, Florian Boudin, Akiko Aizawa
Expert Syst. Appl.2
2021 Redefining Absent Keyphrases and their Effect on Retrieval Effectiveness
abstract
Neural keyphrase generation models have recently attracted much interest due to their ability to output absent keyphrases, that is, keyphrases that do not appear in the source text.In this paper, we discuss the usefulness of absent keyphrases from an Information Retrieval (IR) perspective, and show that the commonly drawn distinction between present and absent keyphrases is not made explicit enough.We introduce a finer-grained categorization scheme that sheds more light on the impact of absent keyphrases on scientific document retrieval.Under this scheme, we find that only a fraction (around 20%) of the words that make up keyphrases actually serves as document expansion, but that this small fraction of words is behind much of the gains observed in retrieval effectiveness.We also discuss how the proposed scheme can offer a new angle to evaluate the output of neural keyphrase generation models.
Florian Boudin, Ygor Gallina
NAACL-HLT1
2020 Keyphrase Generation for Scientific Document Retrieval
abstract
Sequence-to-sequence models have lead to significant progress in keyphrase generation, but it remains unknown whether they are reliable enough to be beneficial for document retrieval.This study provides empirical evidence that such models can significantly improve retrieval performance, and introduces a new extrinsic evaluation framework that allows for a better understanding of the limitations of keyphrase generation models.Using this framework, we point out and discuss the difficulties encountered with supplementing documents with -not present in textkeyphrases, and generalizing models across domains.Our code is available at https:// github.com/boudinfl/ir-using-kg.
Florian Boudin, Ygor Gallina, Akiko Aizawa
ACL1
2020 An Evaluation Dataset for Identifying Communicative Functions of Sentences in English Scholarly Papers
abstract
Formulaic expressions, such as ‘in this paper we propose’, are used by authors of scholarly papers to perform communicative functions; the communicative function of the present example is ‘stating the aim of the paper’. Collecting such expressions and pairing them with their communicative functions would be highly valuable for various tasks, particularly for writing assistance. However, such collection and paring in a principled and automated manner would require high-quality annotated data, which are not available. In this study, we address this shortcoming by creating a manually annotated dataset for detecting communicative functions in sentences. Starting from a seed list of labelled formulaic expressions, we retrieved new sentences from scholarly papers in the ACL Anthology and asked multiple human evaluators to label communicative functions. To show the usefulness of our dataset, we conducted a series of experiments that determined to what extent sentence representations acquired by recent models, such as word2vec and BERT, can be employed to detect communicative functions in sentences.
Kenichi Iwatsuki, Florian Boudin, Akiko Aizawa
LREC2
2019 KPTimes: A Large-Scale Dataset for Keyphrase Generation on News Documents
abstract
Keyphrase generation is the task of predicting a set of lexical units that conveys the main content of a source text.Existing datasets for keyphrase generation are only readily available for the scholarly domain and include nonexpert annotations.In this paper we present KPTimes, a large-scale dataset of news texts paired with editor-curated keyphrases.Exploring the dataset, we show how editors tag documents, and how their annotations differ from those found in existing datasets.We also train and evaluate state-of-the-art neural keyphrase generation models on KPTimes to gain insights on how well they perform on the news domain.
Ygor Gallina, Florian Boudin, Béatrice Daille
INLG2
2016 Keyphrase Annotation with Graph Co-Ranking
abstract
Keyphrase annotation is the task of identifying textual units that represent the main content of a document. Keyphrase annotation is either carried out by extracting the most important phrases from a document, keyphrase extraction, or by assigning entries from a controlled domain-specific vocabulary, keyphrase assignment. Assignment methods are generally more reliable. They provide better-formed keyphrases, as well as keyphrases that do not occur in the document. But they are often silent on the contrary of extraction methods that do not depend on manually built resources. This paper proposes a new method to perform both keyphrase extraction and keyphrase assignment in an integrated and mutual reinforcing manner. Experiments have been carried out on datasets covering different domains of humanities and social sciences. They show statistically significant improvements compared to both keyphrase extraction and keyphrase assignment state-of-the art methods.
Adrien Bougouin, Florian Boudin, Béatrice Daille
COLING2
2016 TermITH-Eval: a French Standard-Based Resource for Keyphrase Extraction Evaluation
Adrien Bougouin, Sabine Barreaux, Laurent Romary, Florian Boudin, Béatrice Daille
LREC4
2015 Concept-based Summarization using Integer Linear Programming: From Concept Pruning to Multiple Optimal Solutions
abstract
In concept-based summarization, sentence selection is modelled as a budgeted maximum coverage problem.As this problem is NP-hard, pruning low-weight concepts is required for the solver to find optimal solutions efficiently.This work shows that reducing the number of concepts in the model leads to lower ROUGE scores, and more importantly to the presence of multiple optimal solutions.We address these issues by extending the model to provide a single optimal solution, and eliminate the need for concept pruning using an approximation algorithm that achieves comparable performance to exact inference.
Florian Boudin, Hugo Mougard, Benoît Favre
EMNLP1
2013 A Comparison of Centrality Measures for Graph-Based Keyphrase Extraction
Florian Boudin
IJCNLP1
2013 TopicRank: Graph-Based Topic Ranking for Keyphrase Extraction
Adrien Bougouin, Florian Boudin, Béatrice Daille
IJCNLP2
2013 Keyphrase Extraction for N-best Reranking in Multi-Sentence Compression
Florian Boudin, Emmanuel Morin
HLT-NAACL1
2012 Using a Medical Thesaurus to Predict Query Difficulty
Florian Boudin, Jian-Yun Nie, Martin Dawes
ECIR1
2010 Improving Medical Information Retrieval with PICO Element Detection
Florian Boudin, Lixin Shi, Jian-Yun Nie
ECIR1
2010 Positional Language Models for Clinical Information Retrieval
Florian Boudin, Jian-Yun Nie, Martin Dawes
EMNLP1
2010 Clinical Information Retrieval using Document and PICO Structure
Florian Boudin, Jian-Yun Nie, Martin Dawes
HLT-NAACL1
2008 Mixing Statistical and Symbolic Approaches for Chemical Names Recognition
Florian Boudin, Juan-Manuel Torres-Moreno, Marc El-Bèze
CICLing1
2007 NEO-CORTEX: A Performant User-Oriented Multi-Document Summarization System
Florian Boudin, Juan-Manuel Torres-Moreno
CICLing1