Rodrigo Wilkens

dblp:16/8694 · also Rodrigo Souza Wilkens · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0003-4366-1215ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Figurative Language in Alzheimer's Discourse: Linguistic and Neural Alignment in Clinical Narratives
Diana Kylymnyk, Vitória Hilgert Tomasel, Helena de Medeiros Caseli, Edward Watkins, Aline Villavicencio, Rodrigo Wilkens
LREC6
2026 Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
abstract
The recent trend towards utilisation of reasoning models has improved the performance of Large Language Models (LLMs) across many tasks which involve logical steps. One linguistic task that could benefit from this framing is idiomaticity detection, as a potentially idiomatic expression must first be understood in relation to the context before it can be disambiguated. In this paper, we explore how reasoning capabilities in LLMs affect idiomaticity detection performance and examine the effect of model size. We evaluate, as open source representative models, the suite of DeepSeek-R1 distillation models ranging from 1.5B to 70B parameters across four idiomaticity detection datasets. We find the effect of reasoning to be smaller and more varied than expected. For smaller models, producing chain-of-thought (CoT) reasoning increases performance from Math-tuned intermediate models, but not to the levels of the base models, whereas larger models (14B, 32B, and 70B) show modest improvements. Our in-depth analyses reveal that larger models demonstrate good understanding of idiomaticity, successfully producing accurate definitions of expressions, while smaller models often fail to output the actual meaning. For this reason, we also experiment with providing definitions in the prompts of smaller models, which we show can improve performance in some cases.
Dylan Phelps, Rodrigo Wilkens, Edward Gow-Smith, Thomas Pickard, Maggie Mi, Marco Idiart, Aline Villavicencio
LREC2
2026 A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
abstract
Potentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
Dilara Torunoglu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 0017, Doruk Eryigit, Thomas Pickard, Adriana S. Pagano, Aline Villavicencio, Gülsen Eryigit, Ágnes Abuczki, Aida Cardoso, Alesia Lazarenka, Dina Almassova, Amália Mendes, Anna Kanellopoulou, Antoni Brosa-Rodríguez, Baiba Valkovska, Beata Wojtowicz, Bolette Pedersen, Carlos Manuel Hidalgo-Ternero, Chaya Liebeskind, Danka Jokic, Diego Alves, Eleni Triantafyllidi, Erik Velldal, Fred Philippy, Giedre Valunaite Oleskeviciene, Ieva Rizgeliene, Inguna Skadina, Irina Lobzhanidze, Isabell Stinessen Haugen, Jauza Akbar Krito, Jelena M. Markovic, Johanna Monti, Josue Alejandro Sauca, Kaja Dobrovoljc, Kingsley O. Ugwuanyi, Laura Rituma, Lilja Øvrelid, Maha Tufail Agro, Manzura Abjalova, Maria Chatzigrigoriou, María del Mar Sánchez Ramos, Marija Pendevska, Masoumeh Seyyedrezaei, Mehrnoush Shamsfard, Momina Ahsan, Muhammad Ahsan Riaz Khan, Nathalie Carmen Hau Norman, Nilay Erdem Ayyildiz, Nina Hosseini-Kivanani, Noémi Ligeti-Nagy, Numaan Naeem, Olha Kanishcheva, Olha Yatsyshyna, Daniil Orel, Petra Giommarelli, Petya Osenova, Radovan Garabík, Regina E. Semou, Rozane Rebechi, Salsabila Zahirah Pranida, Samia Touileb, Sanni Nimb, Sarvinoz Sharipova, Shahar Golan, Shaoxiong Ji, Sopuruchi Christian Aboh, Srdjan Sucur, Stella Markantonatou, Sussi Olsen, Vahideh Tajalli, Veronika Lipp, Voula Giouli, Yelda Yesildal Eraydin, Zahra Saaberi, Zhuohan Xie
LREC3
2026 Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French
abstract
Abstract In Automated Essay Scoring (AES), benchmarking practices have fostered minimalist evaluation practices, in contrast with the broader-view recommendations of evaluation frameworks, such as the argument-based validation framework (ABV), which argued in favor of a multidimensional assessment of systems, especially in the context of high-stakes language tests. In this paper, we introduce an enhanced and more practical version of the ABV framework, incorporating fairness analysis, correlations with linguistic features, prediction error evaluation, and model agreement compared with human raters. Applying this framework to French AES, we compare 8 model architectures on a corpus of 27k exam essays (2 raters each) and a generalization corpus of 961 essays (at least nine raters each). Our analyses illustrate the benefits of applying the ABV framework to better understand the capabilities and pitfalls of AES models, while also advancing the state-of-the-art for French AES.
Rodrigo Wilkens, Rémi Cardon, Vincent Folny, Thomas François
Trans. Assoc. Comput. Linguistics1
2025 Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective
abstract
This study presents a novel approach to assessing French text readability for adults with low literacy skills, addressing both global (full-text) and local (segment-level) difficulty.We also introduce a dataset of 461 texts annotated using a difficulty scale developed specifically for this population.Using this corpus, we conducted a systematic comparison of key readability modeling approaches, including machine learning techniques based on linguistic variables, fine-tuning of CamemBERT, a hybrid approach combining CamemBERT with linguistic variables, and the use of generative language models (LLMs) to carry out readability assessment at global and local level.
Wafa Aissa, Thibault Bañeras-Roux, Elodie Vanzeveren, Lingyun Gao, Rodrigo Wilkens, Thomas François
EMNLP5
2025 UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
abstract
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds 0001, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
EMNLP4
2023 TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring
abstract
Automated Essay Scoring (AES) aims to automatically assess the quality of essays.Automation enables large-scale assessment, improvements in consistency, reliability, and standardization.Those characteristics are of particular relevance in the context of language certification exams.However, a major bottleneck in the development of AES systems is the availability of corpora, which, unfortunately, are scarce, especially for languages other than English.In this paper, we aim to foster the development of AES for French by providing the TCFLE-8 corpus, a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF -French Knowledge Test) certification exam.We report the strict quality procedure that led to the scoring of each essay by at least two raters according to the levels of the Common European Framework of Reference for Languages (CEFR) and to the creation of a balanced corpus.In addition, we describe how linguistic properties of the essays relate to the learners' proficiency in TCFLE-8.We also advance the state-of-the-art performance for the AES task in French by experimenting with two strong baselines (i.e., RoBERTa and featurebased).Finally, we discuss the challenges of AES using TCFLE-8. 1
Rodrigo Wilkens, Alice Pintard, David Alfter, Vincent Folny, Thomas François
EMNLP1
2023 Statistical Methods for Annotation Analysis
abstract
A common task in Natural Language Processing (NLP) is the development of datasets/corpora. It is the crucial initial step for initiatives aiming to train and evaluate Machine Learning and AI systems, for example. Often, these resources must be annotated with additional information (e.g., part-of-speech and named entities), which leads to the question of how to obtain these values. One of the most natural and widely used approaches is to ask for people (e.g., from untrained annotators to domain experts) to identify this information in a given text or document and possibly for more than one annotator per item. However, this is an incomplete solution. It is still necessary to obtain a final annotation per item and to measure agreement among the different annotators (or coders). Presenting a survey on this topic, Ron Artstein and Massimo Poesio published an article (“Inter-coder Agreement for Computational Linguistics”) in 2008 that addressed the mathematics and underlying assumptions of agreement coefficients (e.g., Krippendorff’s α, Scott’s π, and Cohen’s κ) and the use of coefficients in several annotation tasks. However, it left open questions, such as the interpretability of coefficients of agreement, and it did not cover topics that nowadays are important (e.g., the research within the field of statistical methods for annotation analysis, such as latent models of agreement or probabilistic annotation models). In 2022, Silviu Paun, Ron Artstein, and Massimo Poesio published a book addressing primarily the NLP community but also including other communities, such as Data Science. They intended to offer an introduction to latent models of agreement, probabilistic models of aggregation, and learning directly from multiple coders. They also reintroduced the topics presented in 2008, making an incremental contextualization. Although the reliability (agreement between coders) and the validity (the “correctness” of the annotations) are present in the entire book, it is divided into two parts. The first part covers the development of labeling scheme coefficients of agreement such as π, κ, and their variants used in NLP and AI. The second part includes methods developed to analyze the output of annotators (e.g., the most likely label for an item among those provided).Chapter 2 recaps the content presented by Artstein and Poesio (2008), updating the discussion to include recent progress. It mainly targets the coefficients of agreement and their purpose of reliability, which is a prerequisite for demonstrating the validity of a coding scheme. Also, the term “reliability” can be used in different ways: intercoder agreement (or test stability), measuring reproducibility, and accuracy. After presenting a brief context and the notation used in the book, the question “why are custom coefficients to measure agreement necessary?” is explored by reviewing chance-adjusted measures, the percentage agreement by chance, and the specific agreement coefficients. Then, the authors address some of the most common design questions in any annotation project. They start by addressing missing data caused by the coders failing to classify items (for whatever reason). In this discussion, they provide possible actions, raising pros and cons. Then, they discuss the identification of units in tasks where the coder is also required to identify the item boundaries (e.g., beginning and end of a named entity). Finally, they address severely skewed annotated data, discussing the bias problem and the prevalence problem. They finish Chapter 2 presenting proofs for the theorems presented.Chapter 3 presents agreement measures for computational linguistic annotation tasks, dividing them into three main points: methodology, choice of coefficients, and interpretation of coefficients. They describe the challenges of different annotation tasks (e.g., part-of-speech tagging, dialogue act tagging, and named entities), as well as of labeling with and without a predefined set of categories. This description of methodological choices of various studies goes along with an observation that even if a work may report agreement, it may not necessarily follow a methodology as rigorous as that envisaged by Krippendorff (2004). Concerning the choice of coefficients, they discuss the most basic and common form of coding in computational linguistics (i.e., text segment labeling with a limited number of categories), then present coding schemes with hierarchical tagsets and coding schemes with set-valued interpretations (e.g., anaphora and summarization). The discussion about coefficient interpretation looks at the range of values and different authors’ positions. They also discuss the use of weighted coefficients, arguably more appropriate for some annotation tasks, and the challenges in their interpretability.In Chapter 4 the authors present studies on how to interpret the results of reliability by rephrasing the problem as one of confidence estimation of a particular label given the behavior of the coders. The chapter starts by raising an important topic concerning the annotation: The coders easily agree about some items while other items seem more difficult to agree on. This leads to the concept of item difficulty. The items might also be viewed as latent classes, which can be modeled as the likelihood of a coder assigning a given label to an item given that item’s latent class. Considering this reformulation, the authors discuss how to measure and model the agreement (including different probability distributions) and the coders’ stability. This chapter ends Part 1 by moving the reader from an annotation task carried out by experts, which can accurately identify the labels, to a richer formulation where the interaction between both the label and the annotator may be considered. Therefore, it moves away from the simple majority choice, which ignores the accuracy and biases of coders as well the characteristics of the items.Chapter 5 focuses on the probabilistic models of annotation. It begins with a simple annotation model introducing the terminology and some key assumptions frequently made. Next, this model is extended to cover the annotation pattern of the coders. After this introduction, the authors address the issue of item difficulty and how it can affect coders’ annotation. They also discuss hierarchical priors for the annotators (which can be used to estimate annotators’ behavior when the data is scarce), how to model the characteristics of the items to discriminate between the labels, and how to have a richer model of annotator ability. Moving on, they present models where the items have inter-dependent labels (e.g., named entity recognition or information extraction tasks) and where the labels are not predefined classes (e.g., anaphoric annotations). Then, by the chapter’s end, the authors shift from encoding assumptions about the annotation process when inferring the ground-truth labels to neural networks to aggregate the annotations, using a variational autoencoder. Afterwards, they present notes on modeling other types of annotation data.Chapter 6 addresses a different source of disagreement from that presented in Chapter 5—namely the item’s difficulty, which can come from ambiguity, for example. This chapter covers methods for learning from multi-annotated corpora starting from covering the use of soft labels and the coders’ individual labels. Later, the chapter moves to distill the labels dealing with noise and pooling coder confusion. Finally, the authors finish the chapter and the book with recommendations about when to apply each method depending on the characteristics of the datasets the models are to be trained on. This chapter finishes with a summary of the lessons learned including also topics like (1) the decision aggregate or keep all the annotations, (2) crowdsourced labels versus gold labels, and (3) mixed results and what did not work for them.In summary, this book provides a complete perspective of statistical methods for annotation analysis in NLP, covering meaningful references and contextualizing them critically and historically at the same time, while also putting forth the assumptions behind the different coefficients. Moreover, the book provides several practical examples of annotation designs and how to measure their agreement. Thus, it provides an insightful perspective on what the agreement measures can and cannot do, which is present throughout the entire book. The content is suitable for both those who want to carry out research on the subject and for those who are interested in assessing reliability. From the perspective of someone who has an annotated corpus, some sections may be less interesting (i.e., specialized in different tasks), but the coverage of the various NLP tasks makes this book also a good guide for assessing reliability and validity.
Rodrigo Wilkens
Comput. Linguistics1
2022 Is Attention Explanation? An Introduction to the Debate
abstract
Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin
ACL (1)4
2022 Linguistic Corpus Annotation for Automatic Text Simplification Evaluation
abstract
Rémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter, Magali Norré, Adeline Müller, Watrin Patrick, Thomas François. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Rémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter, Magali Norré, Adeline Müller, Patrick Watrin, Thomas François
EMNLP3
2022 HECTOR: A Hybrid TExt SimplifiCation TOol for Raw Texts in French
abstract
Reducing the complexity of texts by applying an Automatic Text Simplification (ATS) system has been sparking interest inthe area of Natural Language Processing (NLP) for several years and a number of methods and evaluation campaigns haveemerged targeting lexical and syntactic transformations. In recent years, several studies exploit deep learning techniques basedon very large comparable corpora. Yet the lack of large amounts of corpora (original-simplified) for French has been hinderingthe development of an ATS tool for this language. In this paper, we present our system, which is based on a combination ofmethods relying on word embeddings for lexical simplification and rule-based strategies for syntax and discourse adaptations. We present an evaluation of the lexical, syntactic and discourse-level simplifications according to automatic and humanevaluations. We discuss the performances of our system at the lexical, syntactic, and discourse levels
Amalia Todirascu, Rodrigo Wilkens, Eva Rolin, Thomas François, Delphine Bernhard, Núria Gala
LREC2
2022 FABRA: French Aggregator-Based Readability Assessment toolkit
abstract
In this paper, we present the FABRA: readability toolkit based on the aggregation of a large number of readability predictor variables. The toolkit is implemented as a service-oriented architecture, which obviates the need for installation, and simplifies its integration into other projects. We also perform a set of experiments to show which features are most predictive on two different corpora, and how the use of aggregators improves performance over standard feature-based readability prediction. Our experiments show that, for the explored corpora, the most important predictors for native texts are measures of lexical diversity, dependency counts and text coherence, while the most important predictors for foreign texts are syntactic variables illustrating language development, as well as features linked to lexical sophistication. FABRA: have the potential to support new research on readability assessment for French.
Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard, Anaïs Tack, Kevin P. Yancey, Thomas François
LREC1
2020 French Coreference for Spoken and Written Language
abstract
Coreference resolution aims at identifying and grouping all mentions referring to the same entity. In French, most systems run different setups, making their comparison difficult. In this paper, we present an extensive comparison of several coreference resolution systems for French. The systems have been trained on two corpora (ANCOR for spoken language and Democrat for written language) annotated with coreference chains, and augmented with syntactic and semantic information. The models are compared with different configurations (e.g. with and without singletons). In addition, we evaluate mention detection and coreference resolution apart. We present a full-stack model that outperforms other approaches. This model allows us to study the impact of mention detection errors on coreference resolution. Our analysis shows that mention detection can be improved by focusing on boundary identification while advances in the pronoun-noun relation detection can help the coreference task. Another contribution of this work is the first end-to-end neural French coreference resolution model trained on Democrat (written texts), which compares to the state-of-the-art systems for oral French.
Rodrigo Wilkens, Bruno Oberle, Frédéric Landragin, Amalia Todirascu
LREC1
2020 Simplifying Coreference Chains for Dyslexic Children
abstract
We present a work aiming to generate adapted content for dyslexic children for French, in the context of the ALECTOR project. Thus, we developed a system to transform the texts at the discourse level. This system modifies the coreference chains, which are markers of text cohesion, by using rules. These rules were designed following a careful study of coreference chains in both original texts and its simplified versions. Moreover, in order to define reliable transformation rules, we analysed several coreference properties as well as the concurrent simplification operations in the aligned texts. This information is coded together with a coreference resolution system and a text rewritten tool in the proposed system, which comprise a coreference module specialised in written text and seven text transformation operations. The evaluation of the system firstly focused on check the simplification by manual validation of three judges. These errors were grouped into five classes that combined can explain 93% of the errors. The second evaluation step consisted of measuring the simplification perception by 23 judges, which allow us to measure the simplification impact of the proposed rules.
Rodrigo Wilkens, Amalia Todirascu
LREC1
2018 Investigating Productive and Receptive Knowledge: A Profile for Second Language Learning
abstract
The literature frequently addresses the differences in receptive and productive vocabulary, but grammar is often left unacknowledged in second language acquisition studies. In this paper, we used two corpora to investigate the divergences in the behavior of pedagogically relevant grammatical structures in reception and production texts. We further improved the divergence scores observed in this investigation by setting a polarity to them that indicates whether there is overuse or underuse of a grammatical structure by language learners. This led to the compilation of a language profile that was later combined with vocabulary and readability features for classifying reception and production texts in three classes: beginner, intermediate, and advanced. The results of the automatic classification task in both production (0.872 of F-measure) and reception (0.942 of F-measure) were comparable to the current state of the art. We also attempted to automatically attribute a score to texts produced by learners, and the correlation results were encouraging, but there is still a good amount of room for improvement in this task. The developed language profile will serve as input for a system that helps language learners to activate more of their passive knowledge in writing texts.
Leonardo Zilio, Rodrigo Wilkens, Cédrick Fairon
COLING2
2018 Document Ranking Applied to Second Language Learning
Rodrigo Wilkens, Leonardo Zilio, Cédrick Fairon
ECIR1
2018 The brWaC Corpus: A New Open Resource for Brazilian Portuguese
Jorge A. Wagner Filho, Rodrigo Wilkens, Marco Idiart, Aline Villavicencio
LREC2
2018 SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning
Rodrigo Wilkens, Leonardo Zilio, Cédrick Fairon
LREC1
2018 An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures
Leonardo Zilio, Rodrigo Wilkens, Cédrick Fairon
LREC2
2016 Multiword Expressions in Child Language
Rodrigo Wilkens, Marco Idiart, Aline Villavicencio
LREC1
2016 B2SG: a TOEFL-like Task for Portuguese
Rodrigo Wilkens, Leonardo Zilio, Aline Villavicencio
LREC1