VLDB 2026 Research / reviewers in the wild / expert
Thomas François
dblp:70/4215
· DBLP profile ↗
20ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for FrenchabstractAbstract In Automated Essay Scoring (AES), benchmarking practices have fostered minimalist evaluation practices, in contrast with the broader-view recommendations of evaluation frameworks, such as the argument-based validation framework (ABV), which argued in favor of a multidimensional assessment of systems, especially in the context of high-stakes language tests. In this paper, we introduce an enhanced and more practical version of the ABV framework, incorporating fairness analysis, correlations with linguistic features, prediction error evaluation, and model agreement compared with human raters. Applying this framework to French AES, we compare 8 model architectures on a corpus of 27k exam essays (2 raters each) and a generalization corpus of 961 essays (at least nine raters each). Our analyses illustrate the benefits of applying the ABV framework to better understand the capabilities and pitfalls of AES models, while also advancing the state-of-the-art for French AES. Rodrigo Wilkens, Rémi Cardon, Vincent Folny, Thomas François |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Assessing French Readability for Adults with Low Literacy: A Global and Local PerspectiveabstractThis study presents a novel approach to assessing French text readability for adults with low literacy skills, addressing both global (full-text) and local (segment-level) difficulty.We also introduce a dataset of 461 texts annotated using a difficulty scale developed specifically for this population.Using this corpus, we conducted a systematic comparison of key readability modeling approaches, including machine learning techniques based on linguistic variables, fine-tuning of CamemBERT, a hybrid approach combining CamemBERT with linguistic variables, and the use of generative language models (LLMs) to carry out readability assessment at global and local level. Wafa Aissa, Thibault Bañeras-Roux, Elodie Vanzeveren, Lingyun Gao, Rodrigo Wilkens, Thomas François |
EMNLP | 6 |
| 2025 | UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency AssessmentabstractJoseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds 0001, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi |
EMNLP | 17 |
| 2024 | Contribution of Move Structure to Automatic Genre Identification: An Annotated Corpus of French Tourism WebsitesabstractThe present work studies the contribution of move structure to automatic genre identification. This concept - well known in other branches of genre analysis - seems to have little application in natural language processing. We describe how we collect a corpus of websites in French related to tourism and annotate it with move structure. We conduct experiments on automatic genre identification with our corpus. Our results show that our approach for informing a model with move structure can increase its performance for automatic genre identification, and reduce the need for annotated data and computational power. Rémi Cardon, Trang Tran Hanh Pham, Julien Zakhia Doueihi, Thomas François |
LREC/COLING | 4 |
| 2024 | Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized DomainsabstractPretrained Language Models (PLMs) are the de facto backbone of most state-of-the-art NLP systems. In this paper, we introduce a family of domain-specific pretrained PLMs for French, focusing on three important domains: transcribed speech, medicine, and law. We use a transformer architecture based on efficient methods (LinFormer) to maximise their utility, since these domains often involve processing long documents. We evaluate and compare our models to state-of-the-art models on a diverse set of tasks and datasets, some of which are introduced in this paper. We gather the datasets into a new French-language evaluation benchmark for these three domains. We also compare various training configurations: continued pretraining, pretraining from scratch, as well as single- and multi-domain pretraining. Extensive domain-specific experiments show that it is possible to attain competitive downstream performance even when pre-training with the approximative LinFormer attention mechanism. For full reproducibility, we release the models and pretraining data, as well as contributed datasets. Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Audibert, Cécile Macaire, Adrien Pupier, Yongxin Zhou 0004, Mathilde Aguiar, Felix Herron, Magali Norré, Massih-Reza Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab |
LREC/COLING | 16 |
| 2023 | TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay ScoringabstractAutomated Essay Scoring (AES) aims to automatically assess the quality of essays.Automation enables large-scale assessment, improvements in consistency, reliability, and standardization.Those characteristics are of particular relevance in the context of language certification exams.However, a major bottleneck in the development of AES systems is the availability of corpora, which, unfortunately, are scarce, especially for languages other than English.In this paper, we aim to foster the development of AES for French by providing the TCFLE-8 corpus, a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF -French Knowledge Test) certification exam.We report the strict quality procedure that led to the scoring of each essay by at least two raters according to the levels of the Common European Framework of Reference for Languages (CEFR) and to the creation of a balanced corpus.In addition, we describe how linguistic properties of the essays relate to the learners' proficiency in TCFLE-8.We also advance the state-of-the-art performance for the AES task in French by experimenting with two strong baselines (i.e., RoBERTa and featurebased).Finally, we discuss the challenges of AES using TCFLE-8. 1 Rodrigo Wilkens, Alice Pintard, David Alfter, Vincent Folny, Thomas François |
EMNLP | 5 |
| 2022 | Is Attention Explanation? An Introduction to the DebateabstractAdrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin |
ACL (1) | 6 |
| 2022 | Linguistic Corpus Annotation for Automatic Text Simplification EvaluationabstractRémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter, Magali Norré, Adeline Müller, Watrin Patrick, Thomas François. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Rémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter, Magali Norré, Adeline Müller, Patrick Watrin, Thomas François |
EMNLP | 8 |
| 2022 | HECTOR: A Hybrid TExt SimplifiCation TOol for Raw Texts in FrenchabstractReducing the complexity of texts by applying an Automatic Text Simplification (ATS) system has been sparking interest inthe area of Natural Language Processing (NLP) for several years and a number of methods and evaluation campaigns haveemerged targeting lexical and syntactic transformations. In recent years, several studies exploit deep learning techniques basedon very large comparable corpora. Yet the lack of large amounts of corpora (original-simplified) for French has been hinderingthe development of an ATS tool for this language. In this paper, we present our system, which is based on a combination ofmethods relying on word embeddings for lexical simplification and rule-based strategies for syntax and discourse adaptations. We present an evaluation of the lexical, syntactic and discourse-level simplifications according to automatic and humanevaluations. We discuss the performances of our system at the lexical, syntactic, and discourse levels Amalia Todirascu, Rodrigo Wilkens, Eva Rolin, Thomas François, Delphine Bernhard, Núria Gala |
LREC | 4 |
| 2022 | FABRA: French Aggregator-Based Readability Assessment toolkitabstractIn this paper, we present the FABRA: readability toolkit based on the aggregation of a large number of readability predictor variables. The toolkit is implemented as a service-oriented architecture, which obviates the need for installation, and simplifies its integration into other projects. We also perform a set of experiments to show which features are most predictive on two different corpora, and how the use of aggregators improves performance over standard feature-based readability prediction. Our experiments show that, for the explored corpora, the most important predictors for native texts are measures of lexical diversity, dependency counts and text coherence, while the most important predictors for foreign texts are syntactic variables illustrating language development, as well as features linked to lexical sophistication. FABRA: have the potential to support new research on readability assessment for French. Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard, Anaïs Tack, Kevin P. Yancey, Thomas François |
LREC | 7 |
| 2020 | Alector: A Parallel Corpus of Simplified French Texts with Alignments of Misreadings by Poor and Dyslexic ReadersabstractIn this paper, we present a new parallel corpus addressed to researchers, teachers, and speech therapists interested in text simplification as a means of alleviating difficulties in children learning to read. The corpus is composed of excerpts drawn from 79 authentic literary (tales, stories) and scientific (documentary) texts commonly used in French schools for children aged between 7 to 9 years old. The excerpts were manually simplified at the lexical, morpho-syntactic, and discourse levels in order to propose a parallel corpus for reading tests and for the development of automatic text simplification tools. A sample of 21 poor-reading and dyslexic children with an average reading delay of 2.5 years read a portion of the corpus. The transcripts of readings errors were integrated into the corpus with the goal of identifying lexical difficulty in the target population. By means of statistical testing, we provide evidence that the manual simplifications significantly reduced reading errors, highlighting that the words targeted for simplification were not only well-chosen but also substituted with substantially easier alternatives. The entire corpus is available for consultation through a web interface and available on demand for research purposes. Núria Gala, Anaïs Tack, Ludivine Javourey-Drevet, Thomas François, Johannes C. Ziegler |
LREC | 4 |
| 2018 | ReSyf: a French lexicon with ranked synonymsabstractIn this article, we present ReSyf, a lexical resource of monolingual synonyms ranked according to their difficulty to be read and understood by native learners of French. The synonyms come from an existing lexical network and they have been semantically disambiguated and refined. A ranking algorithm, based on a wide range of linguistic features and validated through an evaluation campaign with human annotators, automatically sorts the synonyms corresponding to a given word sense by reading difficulty. ReSyf is freely available and will be integrated into a web platform for reading assistance. It can also be applied to perform lexical simplification of French texts. Mokhtar Boumedyen Billami, Thomas François, Núria Gala |
COLING | 2 |
| 2018 | EFLLex: A Graded Lexical Resource for Learners of English as a Foreign Language
Luise Dürlich, Thomas François |
LREC | 2 |
| 2016 | Are Cohesive Features Relevant for Text Readability Evaluation?abstractThis paper investigates the effectiveness of 65 cohesion-based variables that are commonly used in the literature as predictive features to assess text readability. We evaluate the efficiency of these variables across narrative and informative texts intended for an audience of L2 French learners. In our experiments, we use a French corpus that has been both manually and automatically annotated as regards to co-reference and anaphoric chains. The efficiency of the 65 variables for readability is analyzed through a correlational analysis and some modelling experiments. Amalia Todirascu, Thomas François, Delphine Bernhard, Núria Gala, Anne-Laure Ligozat |
COLING | 2 |
| 2016 | Combining Manual and Automatic Prosodic Annotation for Expressive Speech Synthesis
Sandrine Brognaux, Thomas François, Marco Saerens |
LREC | 2 |
| 2016 | SVALex: a CEFR-graded Lexical Resource for Swedish Foreign and Second Language Learners
Thomas François, Elena Volodina, Ildikó Pilán, Anaïs Tack |
LREC | 1 |
| 2016 | Evaluating Lexical Simplification and Vocabulary Knowledge for Learners of French: Possibilities of Using the FLELex Resource
Anaïs Tack, Thomas François, Anne-Laure Ligozat, Cédrick Fairon |
LREC | 2 |
| 2014 | FLELex: a graded lexical resource for French foreign learners
Thomas François, Núria Gala, Patrick Watrin, Cédrick Fairon |
LREC | 1 |
| 2014 | Multiple Choice Question Corpus Analysis for Distractor Characterization
Van-Minh Pho, Thibault André, Anne-Laure Ligozat, Brigitte Grau, Gabriel Illouz, Thomas François |
LREC | 6 |
| 2012 | An "AI readability" Formula for French as a Foreign Language
Thomas François, Cédrick Fairon |
EMNLP-CoNLL | 1 |