Karën Fort

dblp:35/8154 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-0723-8850ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 7 first-author · 11 since 2021
YearPublicationVenuePosition
2026 COCOA: Creation and Exploratory Investigation of a COrpus of Claims frOm NLP Articles
abstract
International audience
Clementine Bleuze, Fanny Ducel, Maxime Amblard, Karën Fort
LREC4
2026 Code-switching as a Bias Indicator in LLMs: "the Consequences Are Not the Same Para Nosotros"
abstract
International audience
Fanny Ducel, Aurélie Névéol, Vidit Khazanchi, Loïc Leclere, Arthur Pedrini, Léa Bouchet, Benjamin Caissial, Karën Fort
LREC8
2024 Beyond Model Performance: Can Link Prediction Enrich French Lexical Graphs?
abstract
This paper presents a resource-centric study of link prediction approaches over French lexical-semantic graphs. Our study incorporates two graphs, RezoJDM16k and RL-fr, and we evaluated seven link prediction models, with CompGCN-ConvE emerging as the best performer. We also conducted a qualitative analysis of the predictions using manual annotations. Based on this, we found that predictions with higher confidence scores were more valid for inclusion. Our findings highlight different benefits for the dense graph compared to the sparser graph RL-fr. While the addition of new triples to RezoJDM16k offers limited advantages, RL-fr can benefit substantially from our approach.
Hee-Soo Choi, Priyansh Trivedi, Matthieu Constant, Karën Fort, Bruno Guillaume
LREC/COLING4
2024 Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts
abstract
Warning: This paper contains explicit statements of offensive stereotypes which may be upsetting The study of bias, fairness and social impact in Natural Language Processing (NLP) lacks resources in languages other than English. Our objective is to support the evaluation of bias in language models in a multilingual setting. We use stereotypes across nine types of biases to build a corpus containing contrasting sentence pairs, one sentence that presents a stereotype concerning an underadvantaged group and another minimally changed sentence, concerning a matching advantaged group. We build on the French CrowS-Pairs corpus and guidelines to provide translations of the existing material into seven additional languages. In total, we produce 11,139 new sentence pairs that cover stereotypes dealing with nine types of biases in seven cultural contexts. We use the final resource for the evaluation of relevant monolingual and multilingual masked language models. We find that language models in all languages favor sentences that express stereotypes in most bias categories. The process of creating a resource that covers a wide range of language types and cultural settings highlights the difficulty of bias evaluation, in particular comparability across languages and contexts.
Karën Fort, Laura Alonso Alemany, Luciana Benotti, Julien Bezançon, Claudia Borg, Marthese Borg, Yongjian Chen, Fanny Ducel, Yoann Dupont, Guido Ivetta, Margot Mieskes, Marco Naguib, Yuyan Qian, Matteo Radaelli, Wolfgang Schmeisser-Nieto, Emma Raimundo Schulz, Thiziri Saci, Sarah Saidi, Javier Torroba Marchante, Shilin Xie, Sergio E. Zanotto, Aurélie Névéol
LREC/COLING1
2024 Unveiling Strengths and Weaknesses of NLP Systems Based on a Rich Evaluation Corpus: The Case of NER in French
abstract
Named Entity Recognition (NER) is an applicative task for which annotation schemes vary. To compare the performance of systems which tagsets differ in precision and coverage, it is necessary to assess (i) the comparability of their annotation schemes and (ii) the individual adequacy of the latter to a common annotation scheme. What is more, and given the lack of robustness of some tools towards textual variation, we cannot expect an evaluation led on an homogeneous corpus with low-coverage to provide a reliable prediction of the actual tools performance. To tackle both these limitations in evaluation, we provide a gold corpus for French covering 6 textual genres and annotated with a rich tagset that enables comparison with multiple annotation schemes. We use the flexibility of this gold corpus to provide both: (i) an individual evaluation of four heterogeneous NER systems on their target tagsets, (ii) a comparison of their performance on a common scheme. This rich evaluation framework enables a fair comparison of NER systems across textual genres and annotation schemes.
Alice Millour, Yoann Dupont, Karën Fort, Liam Duignan
LREC/COLING3
2023 The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research
abstract
Mohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mohamed Abdalla 0001, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif M. Mohammad, Karën Fort
ACL (1)7
2023 Can Synthetic Text Help Clinical Named Entity Recognition? A Study of Electronic Health Records in French
abstract
In sensitive domains, the sharing of corpora is restricted due to confidentiality, copyrights, or trade secrets.Automatic text generation can help alleviate these issues by producing synthetic texts that mimic the linguistic properties of real documents while preserving confidentiality.In this study, we assess the usability of synthetic corpus as a substitute training corpus for clinical information extraction.Our goal is to automatically produce a clinical case corpus annotated with clinical entities and to evaluate it for a named entity recognition (NER) task.We use two auto-regressive neural models partially or fully trained on generic French texts and fine-tuned on clinical cases to produce a corpus of synthetic clinical cases.We study variants of the generation process: (i) fine-tuning on annotated vs. plain text (in that case, annotations are obtained a posteriori) and (ii) selection of generated texts based on models' parameters and filtering criteria.We then train NER models with the resulting synthetic text and evaluate them on a gold standard clinical corpus.Our experiments suggest that synthetic text is useful for clinical NER.
Nicolas Hiebel, Olivier Ferret, Karën Fort, Aurélie Névéol
EACL3
2022 French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English
abstract
Warning: This paper contains explicit statements of offensive stereotypes which may be upsetting Much work on biases in natural language processing has addressed biases linked to the social and cultural experience of English speaking individuals in the United States.We seek to widen the scope of bias studies by creating material to measure social bias in language models (LMs) against specific demographic groups in France.We build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages while also characterizing biases that are specific to each country and language.We introduce 1,677 sentence pairs in French that cover stereotypes in ten types of bias like gender and age.1,467 sentence pairs are translated from CrowS-pairs and 210 are newly crowdsourced and translated back into English.The sentence pairs contrast stereotypes concerning underadvantaged groups with the same sentence concerning advantaged groups.We find that four widely used language models (three French, one multilingual) favor sentences that express stereotypes in most bias categories.We report on the translation process, which led to a characterization of stereotypes in CrowS-pairs including the identification of US-centric cultural traits.We offer guidelines to further extend the dataset to other languages and cultural environments.
Aurélie Névéol, Yoann Dupont, Julien Bezançon, Karën Fort
ACL (1)4
2022 Quantification Annotation in ISO 24617-12, Second Draft
abstract
This paper describes the continuation of a project that aims at establishing an interoperable annotation schema for quantification phenomena as part of the ISO suite of standards for semantic annotation, known as the Semantic Annotation Framework. After a break, caused by the Covid-19 pandemic, the project was relaunched in early 2022 with a second working draft of an annotation scheme, which is discussed in this paper. Keywords: semantic annotation, quantification, interoperability, annotation schema, ISO standard
Harry Bunt, Maxime Amblard, Johan Bos, Karën Fort, Bruno Guillaume, Philippe de Groote, Chuyuan Li, Pierre Ludmann, Michel Musiol, Siyana Pavlova, Guy Perrier, Sylvain Pogodalla
LREC4
2022 Do we Name the Languages we Study? The #BenderRule in LREC and ACL articles
abstract
This article studies the application of the #BenderRule in Natural Language Processing (NLP) articles according to two dimensions. Firstly, in a contrastive manner, by considering two major international conferences, LREC and ACL, and secondly, in a diachronic manner, by inspecting nearly 14,000 articles over a period of time ranging from 2000 to 2020 for LREC and from 1979 to 2020 for ACL. For this purpose, we created a corpus from LREC and ACL articles from the above-mentioned periods, from which we manually annotated nearly 1,000. We then developed two classifiers to automatically annotate the rest of the corpus. Our results show that LREC articles tend to respect the #BenderRule (80 to 90% of them respect it), whereas 30 to 40% of ACL articles do not. Interestingly, over the considered periods, the results appear to be stable for the two conferences, even though a rebound in ACL 2020 could be a sign of the influence of the blog post about the #BenderRule.
Fanny Ducel, Karën Fort, Gaël Lejeune, Yves Lepage
LREC2
2022 CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives
abstract
Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluating models. Such resources are scarce, especially for specialized domains in languages other than English. In particular, there are very few resources for semantic similarity in the clinical domain in French. This can be useful for many biomedical natural language processing applications, including text generation. We introduce a definition of similarity that is guided by clinical facts and apply it to the development of a new French corpus of 1,000 sentence pairs manually annotated according to similarity scores. This new sentence similarity corpus is made freely available to the community. We further evaluate the corpus through experiments of automatic similarity measurement. We show that a model of sentence embeddings can capture similarity with state-of-the-art performance on the DEFT STS shared task evaluation data set (Spearman=0.8343). We also show that the corpus is complementary to DEFT STS.
Nicolas Hiebel, Olivier Ferret, Karën Fort, Aurélie Névéol
LREC3
2020 Rigor Mortis: Annotating MWEs with a Gamified Platform
abstract
We present here Rigor Mortis, a gamified crowdsourcing platform designed to evaluate the intuition of the speakers, then train them to annotate multi-word expressions (MWEs) in French corpora. We previously showed that the speakers’ intuition is reasonably good (65% in recall on non-fixed MWE). We detail here the annotation results, after a training phase using some of the tests developed in the PARSEME-FR project.
Karën Fort, Bruno Guillaume, Yann-Alan Pilatte, Matthieu Constant, Nicolas Lefebvre
LREC1
2020 Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning
abstract
We introduce in this paper a generic approach to combine implicit crowdsourcing and language learning in order to mass-produce language resources (LRs) for any language for which a crowd of language learners can be involved. We present the approach by explaining its core paradigm that consists in pairing specific types of LRs with specific exercises, by detailing both its strengths and challenges, and by discussing how much these challenges have been addressed at present. Accordingly, we also report on on-going proof-of-concept efforts aiming at developing the first prototypical implementation of the approach in order to correct and extend an LR called ConceptNet based on the input crowdsourced from language learners. We then present an international network called the European Network for Combining Language Learning with Crowdsourcing Techniques (enetCollect) that provides the context to accelerate the implementation of this generic approach. Finally, we exemplify how it can be used in several language learning scenarios to produce a multitude of NLP resources and how it can therefore alleviate the long-standing NLP issue of the lack of LRs.
Lionel Nicolas, Verena Lyding, Claudia Borg, Corina Forascu, Karën Fort, Katerina Zdravkova, Iztok Kosem, Jaka Cibej, Spela Arhar Holdt, Alice Millour, Christos T. Rodosthenous, Federico Sangati, Umair ul Hassan, Anisia Katinskaia, Anabela Barreiro, Lavinia Aparaschivei, Yaakov HaCohen-Kerner
LREC5
2018 Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing
Alice Millour, Karën Fort
LREC2
2016 Crowdsourcing Complex Language Resources: Playing to Annotate Dependency Syntax
abstract
This article presents the results we obtained on a complex annotation task (that of dependency syntax) using a specifically designed Game with a Purpose, ZombiLingo. We show that with suitable mechanisms (decomposition of the task, training of the players and regular control of the annotation quality during the game), it is possible to obtain annotations whose quality is significantly higher than that obtainable with a parser, provided that enough players participate. The source code of the game and the resulting annotated corpora (for French) are freely available.
Bruno Guillaume, Karën Fort, Nicolas Lefebvre
COLING2
2016 Yes, We Care! Results of the Ethics and Natural Language Processing Surveys
Karën Fort, Alain Couillault
LREC1
2014 Deep Syntax Annotation of the Sequoia French Treebank
Marie Candito, Guy Perrier, Bruno Guillaume, Corentin Ribeyre, Karën Fort, Djamé Seddah, Éric Villemonte de la Clergerie
LREC5
2014 Evaluating corpora documentation with regards to the Ethics and Big Data Charter
Alain Couillault, Karën Fort, Gilles Adda, Hugues de Mazancourt
LREC2
2014 Mapping the Lexique des Verbes du Français (Lexicon of French Verbs) to a NLP lexicon using examples
Bruno Guillaume, Karën Fort, Guy Perrier, Paul Bédaride
LREC2
2014 Propa-L: a semantic filtering service from a lexical network created using Games With A Purpose
Mathieu Lafourcade, Karën Fort
LREC2
2012 Modeling the Complexity of Manual Annotation Tasks: a Grid of Analysis
Karën Fort, Adeline Nazarenko, Sophie Rosset
COLING1
2012 Annotating Football Matches: Influence of the Source Medium on Manual Annotation
Karën Fort, Vincent Claveau
LREC1
2012 Analyzing the Impact of Prevalence on the Evaluation of a Manual Annotation Campaign
Karën Fort, Claire François, Olivier Galibert, Maha Ghribi
LREC1
2011 Amazon Mechanical Turk: Gold Mine or Coal Mine?
abstract
Recently heard at a tutorial in our field: "It cost me less than one hundred bucks to annotate this using Amazon Mechanical Turk!" Assertions like this are increasingly common, but we believe they should not be stated so proudly; they ignore the ethical consequences of using MTurk (Amazon Mechanical Turk) as a source of labour.Manually annotating corpora or manually developing any other linguistic resource, such as a set of judgments about system outputs, represents such a high cost that many researchers are looking for alternative solutions to the standard approach.MTurk is becoming a popular one.However, as in any scientific endeavor involving humans, there is an unspoken ethical dimension involved in resource construction and system evaluation, and this is especially true of MTurk.We would like here to raise some questions about the use of MTurk.To do so, we will define precisely what MTurk is and what it is not, highlighting the issues raised by the system.We hope that this will point out opportunities for our community to deliberately value ethics above cost savings.
Karën Fort, Gilles Adda, Kevin Cohen 0001
Comput. Linguistics1
2010 FastKwic, an "Intelligent" Concordancer Using FASTR
Véronika Lux-Pogodalla, Dominique Besagni, Karën Fort
LREC3