Benoît Sagot

dblp:66/1016 · DBLP profile ↗
← Back
70ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0002-0107-8526ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 11 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
abstract
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa 0001, Nadia Ghezaiel Hammouda, Verrah Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Kranti Chalamalasetti, Carol Muchemi, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada A. Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob van der Goot, Lanwenn Ar C'horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin Rice, Azril Hafizi Amirudin, Jesujoba O. Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata A, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Çelikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger
ACL (1)94
2026 When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
abstract
User-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis. This makes evaluating UGC translation challenging: what counts as a "good" translation depends on the desired standardness level of the output. To explore this, we examine the human translation guidelines of four UGC datasets, and derive a taxonomy of twelve non-standard phenomena and five translation actions (NORMALISE, COPY, TRANSFER, OMIT, CENSOR). Our analysis reveals notable differences in how UGC is treated, resulting in a spectrum of standardness in reference translations. We show that translation scores of large language models are highly sensitive to prompts with explicit UGC translation instructions, and that they improve when they align with the dataset guidelines. We argue that fair evaluation requires both models and metrics to be aware of translation guidelines. Finally, we call for clear guidelines during dataset creation and for the development of controllable, guideline-aware evaluation frameworks for UGC translation.
Lydia Nishimwe, Benoît Sagot, Rachel Bawden
EAMT (1)2
2026 CoMMA, a Large-scale Corpus of Multilingual Medieval Archives
abstract
International audience
Thibault Clérice, Simon Gabay, Malamatenia Vlachou-Efstathiou, Ariane Pinche, Benoît Sagot
LREC5
2026 A Parallel Corpus of the Parable of the Prodigal Son: Building a Resource for Documenting Language Varieties in Mainland France
abstract
International audience
Lucence Ing, Juliette Janes, Sven Ködel, Benoît Sagot
LREC4
2026 Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
abstract
International audience
Malik Marmonier, Benoît Sagot, Rachel Bawden
LREC2
2026 ForumOccitania: A Corpus of User-Generated Content for Multiple Occitan Varieties
abstract
Accepted at LREC 2026
Oriane Nédey, Juliette Janes, Rachel Bawden, Thibault Clérice, Benoît Sagot
LREC5
2025 Explicit Learning and the LLM in Machine Translation
abstract
This study explores an LLM's ability to learn new languages using explanations found in a grammar book-a process we term "explicit learning."To rigorously assess this ability, we design controlled translation experiments between English and constructed languages generated-through specific cryptographic means-from Latin or French.Contrary to previous studies, our results demonstrate that LLMs do possess a measurable capacity for explicit learning.This ability, however, diminishes as the complexity of the linguistic phenomena to be learned increases.Supervised fine-tuning on ad hoc chains of thought significantly enhances LLM performance but struggles to generalize to typologically novel or more complex linguistic features.These findings point to the need for more diverse training sets and alternative fine-tuning strategies to further improve explicit learning by LLMs, benefiting low-resource languages typically described in grammar books but lacking extensive corpora.
Malik Marmonier, Rachel Bawden, Benoît Sagot
EMNLP3
2024 From Text to Source: Results in Detecting Large Language Model-Generated Content
abstract
The widespread use of Large Language Models (LLMs), celebrated for their ability to generate human-like text, has raised concerns about misinformation and ethical implications. Addressing these concerns necessitates the development of robust methods to detect and attribute text generated by LLMs. This paper investigates “Cross-Model Detection,” by evaluating whether a classifier trained to distinguish between source LLM-generated and human-written text can also detect text from a target LLM without further training. The study comprehensively explores various LLM sizes and families and assesses the impact of conversational fine-tuning techniques, quantization, and watermarking on classifier generalization. The research also explores Model Attribution, encompassing source model identification, model family, and model size classification, in addition to quantization and watermarking detection. Our results reveal several key findings: a clear inverse relationship between classifier effectiveness and model size, with larger LLMs being more challenging to detect, especially when the classifier is trained on data from smaller models. Training on data from similarly sized LLMs can improve detection performance from larger models but may lead to decreased performance when dealing with smaller models. Additionally, model attribution experiments show promising results in identifying source models and model families, highlighting detectable signatures in LLM-generated text, with particularly remarkable outcomes in watermarking detection, while no detectable signatures of quantization were observed. Overall, our study contributes valuable insights into the interplay of model size, family, and training data in LLM detection and attribution.
Wissam Antoun, Benoît Sagot, Djamé Seddah
LREC/COLING2
2024 When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
abstract
Most existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful for low-resource languages (LRLs), where large datasets are not available. Often we are interested in building bilingual resources for LRLs against related high-resource languages (HRLs), resulting in severely imbalanced data settings for BLI. We first show that state-of-the-art BLI methods in the literature exhibit near-zero performance for severely data-imbalanced language pairs, indicating that these settings require more robust techniques. We then present a new method for unsupervised BLI between a related LRL and HRL that only requires inference on a masked language model of the HRL, and demonstrate its effectiveness on truly low-resource languages Bhojpuri and Magahi (with <5M monolingual tokens each), against Hindi. We further present experiments on (mid-resource) Marathi and Nepali to compare approach performances by resource range, and release our resulting lexicons for five low-resource Indic languages: Bhojpuri, Magahi, Awadhi, Braj, and Maithili, against Hindi.
Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, Rachel Bawden
LREC/COLING4
2024 On the Scaling Laws of Geographical Representation in Language Models
abstract
Language models have long been shown to embed geographical information in their hidden representations. This line of work has recently been revisited by extending this result to Large Language Models (LLMs). In this paper, we propose to fill the gap between well-established and recent literature by observing how geographical knowledge evolves when scaling language models. We show that geographical knowledge is observable even for tiny models, and that it scales consistently as we increase the model size. Notably, we observe that larger language models cannot mitigate the geographical bias that is inherent to the training data.
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
LREC/COLING3
2024 Making Sentence Embeddings Robust to User-Generated Content
abstract
NLP models have been known to perform poorly on user-generated content (UGC), mainly because it presents a lot of lexical variations and deviates from the standard texts on which most of these models were trained. In this work, we focus on the robustness of LASER, a sentence embedding model, to UGC data. We evaluate this robustness by LASER’s ability to represent non-standard sentences and their standard counterparts close to each other in the embedding space. Inspired by previous works extending LASER to other languages and modalities, we propose RoLASER, a robust English encoder trained using a teacher-student approach to reduce the distances between the representations of standard and UGC sentences. We show that with training only on standard and synthetic UGC-like data, RoLASER significantly improves LASER’s robustness to both natural and artificial UGC data by achieving up to 2x and 11x better scores. We also perform a fine-grained analysis on artificial UGC data and find that our model greatly outperforms LASER on its most challenging UGC phenomena such as keyboard typos and social media abbreviations. Evaluation on downstream tasks shows that RoLASER performs comparably to or better than LASER on standard data, while consistently outperforming it on UGC data.
Lydia Nishimwe, Benoît Sagot, Rachel Bawden
LREC/COLING2
2024 Anisotropy Is Inherent to Self-Attention in Transformers
abstract
The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers.In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity).Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens.We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences.We also show that the anisotropy problem extends to Transformers trained on other modalities.Our observations suggest that anisotropy is actually inherent to Transformers-based models.
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
EACL (1)3
2024 Tree of Problems: Improving structured problem solving with compositionality
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across multiple tasks through in-context learning.For complex reasoning tasks that require step-by-step thinking, Chain-of-Thought (CoT) prompting has given impressive results, especially when combined with self-consistency.Nonetheless, some tasks remain particularly difficult for LLMs to solve.Tree of Thoughts (ToT) and Graph of Thoughts (GoT) emerged as alternatives, dividing the complex problem into paths of subproblems.In this paper, we propose Tree of Problems (ToP), a simpler version of ToT, which we hypothesise can work better for complex tasks that can be divided into identical subtasks.Our empirical results show that our approach outperforms ToT and GoT, and in addition performs better than CoT on complex reasoning tasks.All code for this paper is publicly available here: https://github. com/ArmelRandy/tree-of-problems.
Armel Zebaze, Benoît Sagot, Rachel Bawden
EMNLP2
2024 Headless Language Models: Learning without Predicting with Contrastive Weight Tying
abstract
Self-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative method that shifts away from probability prediction and instead focuses on reconstructing input embeddings in a contrastive fashion via Constrastive Weight Tying (CWT). We apply this approach to pretrain Headless Language Models in both monolingual and multilingual contexts. Our method offers practical advantages, substantially reducing training computational requirements by up to 20 times, while simultaneously enhancing downstream performance and data efficiency. We observe a significant +1.6 GLUE score increase and a notable +2.7 LAMBADA accuracy improvement compared to classical LMs within similar compute budgets.
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
ICLR3
2024 PatentEval: Understanding Errors in Patent Generation
abstract
You Zuo, Kim Gerdes, Éric Clergerie, Benoît Sagot. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
You Zuo, Kim Gerdes, Éric Villemonte de la Clergerie, Benoît Sagot
NAACL-HLT4
2023 SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
abstract
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Paul-Ambroise Duquenne, Hongyu Gong, Jingfei Du, Ann Lee 0001, Vedanuj Goswami, Changhan Wang, Juan Pino 0001, Benoît Sagot, Holger Schwenk
ACL (1)9
2023 Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation
abstract
One of the major challenges of machine translation (MT) is ambiguity, which can in some cases be resolved by accompanying context such as images.However, recent work in multimodal MT (MMT) has shown that obtaining improvements from images is challenging, limited not only by the difficulty of building effective cross-modal representations, but also by the lack of specific evaluation and training data.We present a new MMT approach based on a strong text-only MT model, which uses neural adapters, a novel guided self-attention mechanism and which is jointly trained on both visually-conditioned masking and MMT.We also introduce CoMMuTE, a Contrastive Multilingual Multimodal Translation Evaluation set of ambiguous sentences and their possible translations, accompanied by disambiguating images corresponding to each translation.Our approach obtains competitive results compared to strong text-only models on standard English→French, English→German and English→Czech benchmarks and outperforms baselines and state-of-the-art MMT systems by a large margin on our contrastive test set.Our code 1 and CoMMuTE 2 are freely available.8 mBART is pretrained on CC25 (Wenzek et al., 2020).
Matthieu Futeral, Cordelia Schmid, Ivan Laptev, Benoît Sagot, Rachel Bawden
ACL (1)4
2023 Generative Spoken Language Model based on continuous word-sized audio tokens
abstract
In NLP, text language models based on words or subwords are known to outperform their character-based counterparts.Yet, in the speech community, the standard input of spoken LMs are 20ms or 40ms-long discrete units (shorter than a phoneme).Taking inspiration from wordbased LM, we introduce a Generative Spoken Language Model (GSLM) based on word-size continuous-valued audio embeddings that can generate diverse and expressive language output.This is obtained by replacing lookup table for lexical types with a Lexical Embedding function, the cross entropy loss by a contrastive loss, and multinomial sampling by k-NN sampling.The resulting model is the first generative language model based on word-size continuous embeddings.Its performance is on par with discrete unit GSLMs regarding generation quality as measured by automatic metrics and subjective human judgements.Moreover, it is five times more memory efficient thanks to its large 200ms units.In addition, the embeddings before and after the Lexical Embedder are phonetically and semantically interpretable. 1
Robin Algayres, Yossi Adi, Tu Anh Nguyen, Jade Copet, Gabriel Synnaeve, Benoît Sagot, Emmanuel Dupoux
EMNLP6
2023 Neural Agents Struggle to Take Turns in Bidirectional Emergent Communication
Valentin Taillandier, Dieuwke Hupkes, Benoît Sagot, Emmanuel Dupoux, Paul Michel
ICLR3
2023 Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer
abstract
International audience
Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot
INTERSPEECH3
2023 Generative Spoken Dialogue Language Modeling
abstract
Abstract We introduce dGSLM, the first “textless” model able to generate audio samples of naturalistic spoken dialogues. It uses recent work on unsupervised spoken unit discovery coupled with a dual-tower transformer architecture with cross-attention trained on 2000 hours of two-channel raw conversational audio (Fisher dataset) without any text or labels. We show that our model is able to generate speech, laughter, and other paralinguistic signals in the two channels simultaneously and reproduces more naturalistic and fluid turn taking compared to a text-based cascaded model.1,2
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoît Sagot, Abdel-rahman Mohamed, Emmanuel Dupoux
Trans. Assoc. Comput. Linguistics9
2022 T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation
abstract
We present a new approach to perform zeroshot cross-modal transfer between speech and text for translation tasks.Multilingual speech and text are encoded in a joint fixed-size representation space.Then, we compare different approaches to decode these multimodal and multilingual fixed-size representations, enabling zero-shot translation between languages and modalities.All our models are trained without the need of cross-modal labeled translation data.Despite a fixed-size representation, we achieve very competitive results on several text and speech translation tasks.In particular, we outperform the state of the art for zero-shot speech translation on Must-C.We also introduce the first results for zero-shot direct speechto-speech and text-to-speech translation.
Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk
EMNLP3
2022 Speech Sequence Embeddings using Nearest Neighbors Contrastive Learning
abstract
International audience
Robin Algayres, Adel Nabli, Benoît Sagot, Emmanuel Dupoux
INTERSPEECH3
2022 Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
abstract
The need for large corpora raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to manually curate the amount of data necessary to train large language models, the main way to obtain this data is still through automatic web crawling. In this paper we take the existing multilingual web corpus OSCAR and its pipeline Ungoliant that extracts and classifies data from Common Crawl at the line level, and propose a set of improvements and automatic annotations in order to produce a new document-oriented version of OSCAR that could prove more suitable to pre-train large generative language models as well as hopefully other applications in Natural Language Processing and Digital Humanities.
Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, Benoît Sagot
LREC4
2022 Automatic Normalisation of Early Modern French
abstract
Spelling normalisation is a useful step in the study and analysis of historical language texts, whether it is manual analysis by experts or automatic analysis using downstream natural language processing (NLP) tools. Not only does it help to homogenise the variable spelling that often exists in historical texts, but it also facilitates the use of off-the-shelf contemporary NLP tools, if contemporary spelling conventions are used for normalisation. We present FREEMnorm, a new benchmark for the normalisation of Early Modern French (from the 17th century) into contemporary French and provide a thorough comparison of three different normalisation methods: ABA, an alignment-based approach and MT-approaches, (both statistical and neural), including extensive parameter searching, which is often missing in the normalisation literature.
Rachel Bawden, Jonathan Poinhos, Eleni Kogkitsidou, Philippe Gambette, Benoît Sagot, Simon Gabay
LREC5
2022 Complex Labelling and Similarity Prediction in Legal Texts: Automatic Analysis of France's Court of Cassation Rulings
abstract
Detecting divergences in the applications of the law (where the same legal text is applied differently by two rulings) is an important task. It is the mission of the French Cour de Cassation. The first step in the detection of divergences is to detect similar cases, which is currently done manually by experts. They rely on summarised versions of the rulings (syntheses and keyword sequences), which are currently produced manually and are not available for all rulings. There is also a high degree of variability in the keyword choices and the level of granularity used. In this article, we therefore aim to provide automatic tools to facilitate the search for similar rulings. We do this by (i) providing automatic keyword sequence generation models, which can be used to improve the coverage of the analysis, and (ii) providing measures of similarity based on the available texts and augmented with predicted keyword sequences. Our experiments show that the predictions improve correlations of automatically obtained similarities against our specially colelcted human judgments of similarity.
Thibault Charmet, Inès Cherichi, Matthieu Allain, Urszula Czerwinska, Amaury Fouret, Benoît Sagot, Rachel Bawden
LREC6
2022 From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French
abstract
anguage models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historical states are at the same time more complex to process and more scarce in the corpora available, this paper presents recent efforts to overcome this difficult situation. These efforts include producing a corpus, creating the model, and evaluating it with an NLP task currently used by scholars in other ongoing projects.
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
LREC7
2022 BERTrade: Using Contextual Embeddings to Parse Old French
abstract
The successes of contextual word embeddings learned by training large-scale language models, while remarkable, have mostly occurred for languages where significant amounts of raw texts are available and where annotated data in downstream tasks have a relatively regular spelling. Conversely, it is not yet completely clear if these models are also well suited for lesser-resourced and more irregular languages. We study the case of Old French, which is in the interesting position of having relatively limited amount of available raw text, but enough annotated resources to assess the relevance of contextual word embedding models for downstream NLP tasks. In particular, we use POS-tagging and dependency parsing to evaluate the quality of such models in a large array of configurations, including models trained from scratch from small amounts of raw text and models pre-trained on other languages but fine-tuned on Medieval French data.
Loïc Grobol, Mathilde Regnault, Pedro Ortiz Suarez, Benoît Sagot, Laurent Romary, Benoît Crabbé
LREC4
2022 MUSS: Multilingual Unsupervised Sentence Simplification by Mining Paraphrases
abstract
Progress in sentence simplification has been hindered by a lack of labeled parallel simplification data, particularly in languages other than English. We introduce MUSS, a Multilingual Unsupervised Sentence Simplification system that does not require labeled simplification data. MUSS uses a novel approach to sentence simplification that trains strong models using sentence-level paraphrase data instead of proper simplification data. These models leverage unsupervised pretraining and controllable generation mechanisms to flexibly adjust attributes such as length and lexical complexity at inference time. We further present a method to mine such paraphrase data in any language from Common Crawl using semantic sentence embeddings, thus removing the need for labeled data. We evaluate our approach on English, French, and Spanish simplification benchmarks and closely match or outperform the previous best supervised results, despite not using any labeled simplification data. We push the state of the art further by incorporating labeled simplification data.
Louis Martin, Angela Fan, Éric Villemonte de la Clergerie, Antoine Bordes, Benoît Sagot
LREC5
2022 DP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon
abstract
Abstract Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a ‘space’ delimiter between words. Popular Bayesian non-parametric models for text segmentation (Goldwater et al., 2006, 2009) use a Dirichlet process to jointly segment sentences and build a lexicon of word types. We introduce DP-Parse, which uses similar principles but only relies on an instance lexicon of word tokens, avoiding the clustering errors that arise with a lexicon of word types. On the Zero Resource Speech Benchmark 2017, our model sets a new speech segmentation state-of-the-art in 5 languages. The algorithm monotonically improves with better input representations, achieving yet higher scores when fed with weakly supervised inputs. Despite lacking a type lexicon, DP-Parse can be pipelined to a language model and learn semantic and syntactic representations as assessed by a new spoken word embedding benchmark. 1
Robin Algayres, Tristan Ricoul, Julien Karadayi, Hugo Laurençon, Mohamed Salah Zaïem, Abdel-rahman Mohamed, Benoît Sagot, Emmanuel Dupoux
Trans. Assoc. Comput. Linguistics7
2022 Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
abstract
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases.
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi
Trans. Assoc. Comput. Linguistics14
2021 First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT
abstract
Multilingual pretrained language models have demonstrated remarkable zero-shot crosslingual transfer capabilities.Such transfer emerges by fine-tuning on a task of interest in one language and evaluating on a distinct language, not seen during the fine-tuning.Despite promising results, we still lack a proper understanding of the source of this transfer.Using a novel layer ablation technique and analyses of the model's internal representations, we show that multilingual BERT, a popular multilingual language model, can be viewed as the stacking of two sub-networks: a multilingual encoder followed by a taskspecific language-agnostic predictor.While the encoder is crucial for cross-lingual transfer and remains mostly unchanged during finetuning, the task predictor has little importance on the transfer and can be reinitialized during fine-tuning.We present extensive experiments with three distinct tasks, seventeen typologically diverse languages and multiple domains to support our hypothesis.RANDOM-INIT of layers SRC-TRG REF ∆1-2 ∆3-4 ∆5-6 ∆7-8 ∆9-10 ∆11-12 Parsing EN -EN 88.98 -0.96 -0.66 -0.93 -0.55 0.04 -0.09RU -RU 85.15 -0.82 -1.38 -1.51 -0.86 -0.29 0.18 AR -AR 59.54 -0.78 -2.14 -1.20 -0.67 -0.27 0.08 EN -X 53.23 -15.77 -6.51 -3.39 -1.47 0.29 1.00 RU -X 55.41 -7.69 -3.71 -3.13 -1.70 0.92 0.94 AR -X 27.97 -4.91 -3.17 -1.48 -1.68 -0.36 -0.14 POS EN -EN 96.51 -0.30 -0.25 -0.40 -0.00 0.05 0.02 RU -RU 96.90 -0.52 -0.55 -0.40 -0.07 0.02 -0.03 AR -AR 79.28 -0.35 -0.49-0.36 -0.19 -0.05 -0.00 EN -X 79.37 -8.94 -2.49-1.66 -0.88 0.20 -0.14 RU -X 79.25 -10.08 -2.83 -1.65 -2.74 0.01 -0.45 AR -X 64.81 -6.73 -3.50 -1.63 -1.56 -0.73 -1.29 NER EN -EN 83.30 -2.66 -2.14 -1.43 -0.63 -0.23 -0.12 RU -RU 88.20 -2.08 -2.13 -1.52 -0.64 -0.33 -0.13 AR -AR 87.97 -2.37 -2.11 -0.96 -0.39 -0.15 0.21 EN -X 64.17 -8.28 -5.09 -3.07 -0.79 -0.47 -0.13 RU -X 62.13 -15.85 -9.36 -5.50 -2.44 -1.16 -0.06 AR -X 65.59 -16.10 -8.42 -3.
Benjamin Muller, Yanai Elazar, Benoît Sagot, Djamé Seddah
EACL3
2021 Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering
abstract
Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on the Question Answering task. However, most of those datasets are in English, and the performances of state-of-the-art multilingual models are significantly lower when evaluated on non-English data. Due to high data collection costs, it is not realistic to obtain annotated data for each language one desires to support. We propose a method to improve the Cross-lingual Question Answering performance without requiring additional annotated data, leveraging Question Generation models to produce synthetic samples in a cross-lingual fashion. We show that the proposed method allows to significantly outperform the baselines trained on English data only. We report a new state-of-the-art on four multilingual datasets: MLQA, XQuAD, SQuAD-it and PIAF (fr).
Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, Jacopo Staiano
EMNLP (1)4
2021 When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models
abstract
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah
NAACL-HLT3
2020 ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting Transformations
abstract
In order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e.replacing complex words or phrases by simpler synonyms), reorder components, and/or delete information deemed unnecessary.Despite these varied range of possible text alterations, current models for automatic sentence simplification are evaluated using datasets that are focused on a single transformation, such as lexical paraphrasing or splitting.This makes it impossible to understand the ability of simplification models in more realistic settings.To alleviate this limitation, this paper introduces ASSET, a new dataset for assessing sentence simplification in English.ASSET is a crowdsourced multi-reference corpus where each simplification was produced by executing several rewriting transformations.Through quantitative and qualitative experiments, we show that simplifications in ASSET are better at capturing characteristics of simplicity when compared to other standard evaluation datasets for the task.Furthermore, we motivate the need for developing better methods for automatic evaluation using ASSET, since we show that current popular metrics may not be suitable when multiple simplification transformations are performed.
Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, Lucia Specia
ACL5
2020 CamemBERT: a Tasty French Language Model
abstract
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Louis Martin, Benjamin Muller, Pedro Ortiz Suarez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, Benoît Sagot
ACL8
2020 Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
abstract
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Ortiz Suarez, Benoît Sagot
ACL7
2020 A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
abstract
We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages.We then compare the performance of OSCARbased and Wikipedia-based ELMo embeddings for these languages on the part-ofspeech tagging and parsing tasks.We show that, despite the noise in the Common-Crawlbased OSCAR data, embeddings trained on OSCAR perform much better than monolingual embeddings trained on Wikipedia.They actually equal or improve the current state of the art in tagging and parsing for all five languages.In particular, they also improve over multilingual Wikipedia-based contextual embeddings (multilingual BERT), which almost always constitutes the previous state of the art, thereby showing that the benefit of a larger, more diverse corpus surpasses the crosslingual benefit of multilingual embedding architectures.
Pedro Ortiz Suarez, Laurent Romary, Benoît Sagot
ACL3
2020 Evaluating the Reliability of Acoustic Speech Embeddings
abstract
International audience
Robin Algayres, Mohamed Salah Zaïem, Benoît Sagot, Emmanuel Dupoux
INTERSPEECH3
2020 Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0
abstract
Diachronic lexical information is not only important in the field of historical linguistics, but is also increasingly used in NLP, most recently for machine translation of low resource languages. Therefore, there is a need for fine-grained, large-coverage and accurate etymological lexical resources. In this paper, we propose a set of guidelines to generate such resources, for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation. To illustrate the guidelines, we introduce EtymDB 2.0, an etymological database automatically generated from the Wiktionary, which contains 1.8 million lexemes, linked by more than 700,000 fine-grained etymological relations, across 2,536 living and dead languages. We also introduce use cases for which EtymDB 2.0 could represent a key resource, such as phylogenetic tree generation, low resource machine translation or medieval languages study.
Clémentine Fourrier, Benoît Sagot
LREC2
2020 OFrLex: A Computational Morphological and Syntactic Lexicon for Old French
abstract
In this paper we describe our work on the development and enrichment of OFrLex, a freely available, large-coverage morphological and syntactic Old French lexicon. We rely on several heterogeneous language resources to extract structured and exploitable information. The extraction follows a semi-automatic procedure with substantial manual steps to respond to difficulties encountered while aligning lexical entries from distinct language resources. OFrLex aims at improving natural language processing tasks on Old French such as part-of-speech tagging and dependency parsing. We provide quantitative information on OFrLex and discuss its reliability. We also describe and evaluate a semi-automatic, word-embedding-based lexical enrichment process aimed at increasing the accuracy of the resource. Results of this extension technique will be manually validated in the near future, a step that will take advantage of OFrLex’s viewing, searching and editing interface, which is already accessible online.
Gaël Guibon, Benoît Sagot
LREC2
2020 Controllable Sentence Simplification
abstract
Text simplification aims at making a text easier to read and understand by simplifying grammar and structure while keeping the underlying information identical. It is often considered an all-purpose generic task where the same simplification is suitable for all; however multiple audiences can benefit from simplified text in different ways. We adapt a discrete parametrization mechanism that provides explicit control on simplification systems based on Sequence-to-Sequence models. As a result, users can condition the simplifications returned by a model on attributes such as length, amount of paraphrasing, lexical complexity and syntactic complexity. We also show that carefully chosen values of these attributes allow out-of-the-box Sequence-to-Sequence models to outperform their standard counterparts on simplification benchmarks. Our model, which we call ACCESS (as shorthand for AudienCe-CEntric Sentence Simplification), establishes the state of the art at 41.87 SARI on the WikiLarge test set, a +1.42 improvement over the best previously reported score.
Louis Martin, Éric Villemonte de la Clergerie, Benoît Sagot, Antoine Bordes
LREC3
2020 Establishing a New State-of-the-Art for French Named Entity Recognition
abstract
The French TreeBank developed at the University Paris 7 is the main source of morphosyntactic and syntactic annotations for French. However, it does not include explicit information related to named entities, which are among the most useful information for several natural language processing tasks and applications. Moreover, no large-scale French corpus with named entity annotations contain referential information, which complement the type and the span of each mention with an indication of the entity it refers to. We have manually annotated the French TreeBank with such information, after an automatic pre-annotation step. We sketch the underlying annotation guidelines and we provide a few figures about the resulting annotations.
Pedro Ortiz Suarez, Yoann Dupont, Benjamin Muller, Laurent Romary, Benoît Sagot
LREC5
2019 What Does BERT Learn about the Structure of Language?
abstract
BERT is a recent language representation model that has surprisingly performed well in diverse language understanding benchmarks.This result indicates the possibility that BERT networks capture structural information about language.In this work, we provide novel support for this claim by performing a series of experiments to unpack the elements of English language structure learned by BERT.We first show that BERT's phrasal representation captures phrase-level information in the lower layers.We also show that BERT's intermediate layers encode a rich hierarchy of linguistic information, with surface features at the bottom, syntactic features in the middle and semantic features at the top.BERT turns out to require deeper layers when long-distance dependency information is required, e.g. to track subjectverb agreement.Finally, we show that BERT representations capture linguistic information in a compositional way that mimics classical, tree-like structures.
Ganesh Jawahar, Benoît Sagot, Djamé Seddah
ACL (1)2
2018 CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing
Amir More, Özlem Çetinoglu, Çagri Çöltekin, Nizar Habash, Benoît Sagot, Djamé Seddah, Dima Taji, Reut Tsarfaty
LREC5
2018 A multilingual collection of CoNLL-U-compatible morphological lexicons
Benoît Sagot
LREC1
2018 Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer
Djamé Seddah, Éric Villemonte de la Clergerie, Benoît Sagot, Héctor Martínez Alonso, Marie Candito
LREC3
2014 A Language-independent Approach to Extracting Derivational Relations from an Inflectional Lexicon
Marion Baranes, Benoît Sagot
LREC2
2014 Developing a French FrameNet: Methodology and First results
Marie Candito, Pascal Amsili, Lucie Barque, Farah Benamara, Gaël de Chalendar, Marianne Djemaa, Pauline Haas, Richard Huyghe, Yvette Yannick Mathieu, Philippe Muller, Benoît Sagot, Laure Vieu
LREC11
2014 An Open-Source Heavily Multilingual Translation Graph Extracted from Wiktionaries and Parallel Corpora
Valérie Hanoka, Benoît Sagot
LREC2
2014 DeLex, a freely-avaible, large-scale and linguistically grounded morphological lexicon for German
Benoît Sagot
LREC1
2014 A language-independent and fully unsupervised approach to lexicon induction and part-of-speech tagging for closely related languages
Yves Scherrer, Benoît Sagot
LREC2
2013 Enforcing Subcategorization Constraints in a Parser Using Sub-parses Recombining
Seyed Abolghasem Mirroshandel, Alexis Nasr, Benoît Sagot
HLT-NAACL3
2012 The French Social Media Bank: a Treebank of Noisy User Generated Content
Djamé Seddah, Benoît Sagot, Marie Candito, Virginie Mouilleron, Vanessa Combet
COLING2
2012 Applying cross-lingual WSD to wordnet development
Marianna Apidianaki, Benoît Sagot
LREC2
2012 Boosting the Coverage of a Semantic Lexicon by Automatically Extracted Event Nominalizations
Kata Gábor, Marianna Apidianaki, Benoît Sagot, Éric Villemonte de la Clergerie
LREC3
2012 Wordnet extension made simple: A multilingual lexicon-based approach using wiki resources
Valérie Hanoka, Benoît Sagot
LREC2
2012 Cleaning noisy wordnets
Benoît Sagot, Darja Fiser
LREC1
2012 Aleda, a free large-scale entity database for French
Benoît Sagot, Rosa Stern
LREC1
2012 Evaluating and improving syntactic lexica by plugging them within a parser
Elsa Tolone, Benoît Sagot, Éric Villemonte de la Clergerie
LREC2
2010 Optimal Rank Reduction for Linear Context-Free Rewriting Systems with Fan-Out Two
Benoît Sagot, Giorgio Satta
ACL1
2010 The Lefff, a Freely Available and Large-coverage Morphological and Syntactic Lexicon for French
Benoît Sagot
LREC1
2010 A Lexicon of French Quotation Verbs for Automatic Quotation Extraction
Benoît Sagot, Laurence Danlos, Rosa Stern
LREC1
2010 A Morphological Lexicon for the Persian Language
Benoît Sagot, Géraldine Walther
LREC1
2009 Coupling an Annotated Corpus and a Morphosyntactic Lexicon for State-of-the-Art POS Tagging with Less Human Effort
Pascal Denis, Benoît Sagot
PACLIC2
2008 Computer Aided Correction and Extension of a Syntactic Wide-Coverage Lexicon
Lionel Nicolas, Benoît Sagot, Miguel A. Molinero, Jacques Farré, Éric Villemonte de la Clergerie
COLING2
2006 Error Mining in Parsing Results
abstract
We introduce an error mining technique for automatically detecting errors in resources that are used in parsing systems. We applied this technique on parsing results produced on several million words by two distinct parsing systems, which share the syntactic lexicon and the pre-parsing processing chain. We were thus able to identify missing and erroneous information in these resources.
Benoît Sagot, Éric Villemonte de la Clergerie
ACL1
2006 Deep non-probabilistic parsing of large corpora
Benoît Sagot, Pierre Boullier
LREC1
2006 The Lefff 2 syntactic lexicon for French: architecture, acquisition, use
Benoît Sagot, Lionel Clément, Éric Villemonte de la Clergerie, Pierre Boullier
LREC1
2004 Morphology Based Automatic Acquisition of Large-coverage Lexica
Lionel Clément, Benoît Sagot, Bernard Lang
LREC2