VLDB 2026 Research / reviewers in the wild / expert
Florian Boudin
dblp:54/2762
· DBLP profile ↗
26ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0001-5849-2261ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 9 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and ChartsabstractWith the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are a core component of scientific work, often presented in varying formats such as tables or charts. Understanding how robust current multimodal large language models (multimodal LLMs) are at verifying scientific claims across different evidence formats remains an important and underexplored challenge. In this paper, we design and conduct a series of experiments to assess the ability of multimodal LLMs to verify scientific claims using both tables and charts as evidence. To enable this evaluation, we adapt two existing datasets of scientific papers by incorporating annotations and structures necessary for a multimodal claim verification task. Using this adapted dataset, we evaluate 12 multimodal LLMs and find that current models perform better with table-based evidence while struggling with chart-based evidence. We further conduct human evaluations and observe that humans maintain strong performance across both formats, unlike the models. Our analysis also reveals that smaller multimodal LLMs (under 8B) show weak correlation in performance between table-based and chart-based tasks, indicating limited cross-modal generalization. These findings highlight a critical gap in current models' multimodal reasoning capabilities. We suggest that future multimodal LLMs should place greater emphasis on improving chart understanding to better support scientific claim verification. Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa |
AAAI | 4 |
| 2026 | SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, André Greiner-Petter, Akiko Aizawa |
LREC | 5 |
| 2026 | Evaluating the Homogeneity of Keyphrase Prediction Models
Maël Houbre, Florian Boudin, Béatrice Daille |
LREC | 2 |
| 2025 | Identifying Reliable Evaluation Metrics for Scientific Text RevisionabstractInternational audience Léane Jourdan, Nicolas Hernandez, Florian Boudin, Richard Dufour |
ACL (1) | 3 |
| 2025 | ACL-rlg: A Dataset for Reading List GenerationabstractFamiliarizing oneself with a new scientific field and its existing literature can be daunting due to the large amount of available articles. Curated lists of academic references, or reading lists, compiled by experts, offer a structured way to gain a comprehensive overview of a domain or a specific scientific challenge. In this work, we introduce ACL-rlg, the largest open expert-annotated reading list dataset. We also provide multiple baselines for evaluating reading list generation and formally define it as a retrieval task. Our qualitative study highlights that traditional scholarly search engines and indexing methods perform poorly on this task, and GPT-4o, despite showing better results, exhibits signs of potential data contamination. Julien Aubert-Béduchaud, Florian Boudin, Béatrice Daille, Richard Dufour |
COLING | 2 |
| 2025 | Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
Kon Woo Kim, Rezarta Islamaj Dogan, Jin-Dong Kim, Florian Boudin, Akiko Aizawa |
NLDB (2) | 4 |
| 2024 | CASIMIR: A Corpus of Scientific Articles Enhanced with Multiple Author-Integrated RevisionsabstractWriting a scientific article is a challenging task as it is a highly codified and specific genre, consequently proficiency in written communication is essential for effectively conveying research findings and ideas. In this article, we propose an original textual resource on the revision step of the writing process of scientific articles. This new dataset, called CASIMIR, contains the multiple revised versions of 15,646 scientific articles from OpenReview, along with their peer reviews. Pairs of consecutive versions of an article are aligned at sentence-level while keeping paragraph location information as metadata for supporting future revision studies at the discourse level. Each pair of revised sentences is enriched with automatically extracted edits and associated revision intention. To assess the initial quality on the dataset, we conducted a qualitative study of several state-of-the-art text revision approaches and compared various evaluation metrics. Our experiments led us to question the relevance of the current evaluation methods for the text revision task. Léane Jourdan, Florian Boudin, Nicolas Hernandez, Richard Dufour |
LREC/COLING | 2 |
| 2022 | From Fundamentals to Recent Advances: A Tutorial on Keyphrasification
Debanjan Mahata, Florian Boudin |
ECIR (2) | 3 |
| 2022 | Cross-lingual and Cross-domain Transfer Learning for Automatic Term Extraction from Low Resource DataabstractAutomatic Term Extraction (ATE) is a key component for domain knowledge understanding and an important basis for further natural language processing applications. Even with persistent improvements, ATE still exhibits weak results exacerbated by small training data inherent to specialized domain corpora. Recently, transformers-based deep neural models, such as BERT, have proven to be efficient in many downstream NLP tasks. However, no systematic evaluation of ATE has been conducted so far. In this paper, we run an extensive study on fine-tuning pre-trained BERT models for ATE. We propose strategies that empirically show BERT’s effectiveness using cross-lingual and cross-domain transfer learning to extract single and multi-word terms. Experiments have been conducted on four specialized domains in three languages. The obtained results suggest that BERT can capture cross-domain and cross-lingual terminologically-marked contexts shared by terms, opening a new design-pattern for ATE. Amir Hazem, Mérième Bouhandi, Florian Boudin, Béatrice Daille |
LREC | 3 |
| 2022 | Extraction and evaluation of formulaic expressions used in scholarly papersabstractFormulaic expressions, such as ‘in this paper we propose’, are helpful for authors of scholarly papers because they convey communicative functions; in the above, it is ‘showing the aim of this paper’. Thus, resources of formulaic expressions, such as a dictionary, that could be looked up easily would be useful. However, forms of formulaic expressions can often vary to a great extent. For example, ‘in this paper we propose’, ‘in this study we propose’ and ‘in this paper we propose a new method to’ are all regarded as formulaic expressions. Such a diversity of spans and forms causes problems in both extraction and evaluation of formulaic expressions. In this paper, we propose a new approach that is robust to variation of spans and forms of formulaic expressions. Our approach regards a sentence as consisting of a formulaic part and non-formulaic part. Then, instead of trying to extract formulaic expressions from a whole corpus, by extracting them from each sentence, different forms can be dealt with at once. Based on this formulation, to avoid the diversity problem, we propose evaluating extraction methods by how much they convey specific communicative functions rather than by comparing extracted expressions to an existing lexicon. We also propose a new extraction method that utilises named entities and dependency structures to remove the non-formulaic part from a sentence. Experimental results show that the proposed extraction method achieved the best performance compared to other existing methods. Kenichi Iwatsuki, Florian Boudin, Akiko Aizawa |
Expert Syst. Appl. | 2 |
| 2021 | Redefining Absent Keyphrases and their Effect on Retrieval EffectivenessabstractNeural keyphrase generation models have recently attracted much interest due to their ability to output absent keyphrases, that is, keyphrases that do not appear in the source text.In this paper, we discuss the usefulness of absent keyphrases from an Information Retrieval (IR) perspective, and show that the commonly drawn distinction between present and absent keyphrases is not made explicit enough.We introduce a finer-grained categorization scheme that sheds more light on the impact of absent keyphrases on scientific document retrieval.Under this scheme, we find that only a fraction (around 20%) of the words that make up keyphrases actually serves as document expansion, but that this small fraction of words is behind much of the gains observed in retrieval effectiveness.We also discuss how the proposed scheme can offer a new angle to evaluate the output of neural keyphrase generation models. Florian Boudin, Ygor Gallina |
NAACL-HLT | 1 |
| 2020 | Keyphrase Generation for Scientific Document RetrievalabstractSequence-to-sequence models have lead to significant progress in keyphrase generation, but it remains unknown whether they are reliable enough to be beneficial for document retrieval.This study provides empirical evidence that such models can significantly improve retrieval performance, and introduces a new extrinsic evaluation framework that allows for a better understanding of the limitations of keyphrase generation models.Using this framework, we point out and discuss the difficulties encountered with supplementing documents with -not present in textkeyphrases, and generalizing models across domains.Our code is available at https:// github.com/boudinfl/ir-using-kg. Florian Boudin, Ygor Gallina, Akiko Aizawa |
ACL | 1 |
| 2020 | An Evaluation Dataset for Identifying Communicative Functions of Sentences in English Scholarly PapersabstractFormulaic expressions, such as ‘in this paper we propose’, are used by authors of scholarly papers to perform communicative functions; the communicative function of the present example is ‘stating the aim of the paper’. Collecting such expressions and pairing them with their communicative functions would be highly valuable for various tasks, particularly for writing assistance. However, such collection and paring in a principled and automated manner would require high-quality annotated data, which are not available. In this study, we address this shortcoming by creating a manually annotated dataset for detecting communicative functions in sentences. Starting from a seed list of labelled formulaic expressions, we retrieved new sentences from scholarly papers in the ACL Anthology and asked multiple human evaluators to label communicative functions. To show the usefulness of our dataset, we conducted a series of experiments that determined to what extent sentence representations acquired by recent models, such as word2vec and BERT, can be employed to detect communicative functions in sentences. Kenichi Iwatsuki, Florian Boudin, Akiko Aizawa |
LREC | 2 |
| 2019 | KPTimes: A Large-Scale Dataset for Keyphrase Generation on News DocumentsabstractKeyphrase generation is the task of predicting a set of lexical units that conveys the main content of a source text.Existing datasets for keyphrase generation are only readily available for the scholarly domain and include nonexpert annotations.In this paper we present KPTimes, a large-scale dataset of news texts paired with editor-curated keyphrases.Exploring the dataset, we show how editors tag documents, and how their annotations differ from those found in existing datasets.We also train and evaluate state-of-the-art neural keyphrase generation models on KPTimes to gain insights on how well they perform on the news domain. Ygor Gallina, Florian Boudin, Béatrice Daille |
INLG | 2 |
| 2016 | Keyphrase Annotation with Graph Co-RankingabstractKeyphrase annotation is the task of identifying textual units that represent the main content of a document. Keyphrase annotation is either carried out by extracting the most important phrases from a document, keyphrase extraction, or by assigning entries from a controlled domain-specific vocabulary, keyphrase assignment. Assignment methods are generally more reliable. They provide better-formed keyphrases, as well as keyphrases that do not occur in the document. But they are often silent on the contrary of extraction methods that do not depend on manually built resources. This paper proposes a new method to perform both keyphrase extraction and keyphrase assignment in an integrated and mutual reinforcing manner. Experiments have been carried out on datasets covering different domains of humanities and social sciences. They show statistically significant improvements compared to both keyphrase extraction and keyphrase assignment state-of-the art methods. Adrien Bougouin, Florian Boudin, Béatrice Daille |
COLING | 2 |
| 2016 | TermITH-Eval: a French Standard-Based Resource for Keyphrase Extraction Evaluation
Adrien Bougouin, Sabine Barreaux, Laurent Romary, Florian Boudin, Béatrice Daille |
LREC | 4 |
| 2015 | Concept-based Summarization using Integer Linear Programming: From Concept Pruning to Multiple Optimal SolutionsabstractIn concept-based summarization, sentence selection is modelled as a budgeted maximum coverage problem.As this problem is NP-hard, pruning low-weight concepts is required for the solver to find optimal solutions efficiently.This work shows that reducing the number of concepts in the model leads to lower ROUGE scores, and more importantly to the presence of multiple optimal solutions.We address these issues by extending the model to provide a single optimal solution, and eliminate the need for concept pruning using an approximation algorithm that achieves comparable performance to exact inference. Florian Boudin, Hugo Mougard, Benoît Favre |
EMNLP | 1 |
| 2013 | A Comparison of Centrality Measures for Graph-Based Keyphrase Extraction
Florian Boudin |
IJCNLP | 1 |
| 2013 | TopicRank: Graph-Based Topic Ranking for Keyphrase Extraction
Adrien Bougouin, Florian Boudin, Béatrice Daille |
IJCNLP | 2 |
| 2013 | Keyphrase Extraction for N-best Reranking in Multi-Sentence Compression
Florian Boudin, Emmanuel Morin |
HLT-NAACL | 1 |
| 2012 | Using a Medical Thesaurus to Predict Query Difficulty
Florian Boudin, Jian-Yun Nie, Martin Dawes |
ECIR | 1 |
| 2010 | Improving Medical Information Retrieval with PICO Element Detection
Florian Boudin, Lixin Shi, Jian-Yun Nie |
ECIR | 1 |
| 2010 | Positional Language Models for Clinical Information Retrieval
Florian Boudin, Jian-Yun Nie, Martin Dawes |
EMNLP | 1 |
| 2010 | Clinical Information Retrieval using Document and PICO Structure
Florian Boudin, Jian-Yun Nie, Martin Dawes |
HLT-NAACL | 1 |
| 2008 | Mixing Statistical and Symbolic Approaches for Chemical Names Recognition
Florian Boudin, Juan-Manuel Torres-Moreno, Marc El-Bèze |
CICLing | 1 |
| 2007 | NEO-CORTEX: A Performant User-Oriented Multi-Document Summarization System
Florian Boudin, Juan-Manuel Torres-Moreno |
CICLing | 1 |