EDBT 2026 Demo / reviewers in the wild / expert
Benoît Sagot
dblp:66/1016
· DBLP profile ↗
70ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0002-0107-8526ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 11 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataabstractPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa 0001, Nadia Ghezaiel Hammouda, Verrah Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Kranti Chalamalasetti, Carol Muchemi, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada A. Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob van der Goot, Lanwenn Ar C'horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin Rice, Azril Hafizi Amirudin, Jesujoba O. Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, Akshata A, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Çelikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger |
ACL (1) | 94 |
| 2026 | When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated ContentabstractUser-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis. This makes evaluating UGC translation challenging: what counts as a "good" translation depends on the desired standardness level of the output. To explore this, we examine the human translation guidelines of four UGC datasets, and derive a taxonomy of twelve non-standard phenomena and five translation actions (NORMALISE, COPY, TRANSFER, OMIT, CENSOR). Our analysis reveals notable differences in how UGC is treated, resulting in a spectrum of standardness in reference translations. We show that translation scores of large language models are highly sensitive to prompts with explicit UGC translation instructions, and that they improve when they align with the dataset guidelines. We argue that fair evaluation requires both models and metrics to be aware of translation guidelines. Finally, we call for clear guidelines during dataset creation and for the development of controllable, guideline-aware evaluation frameworks for UGC translation. Lydia Nishimwe, Benoît Sagot, Rachel Bawden |
EAMT (1) | 2 |
| 2026 | CoMMA, a Large-scale Corpus of Multilingual Medieval ArchivesabstractInternational audience Thibault Clérice, Simon Gabay, Malamatenia Vlachou-Efstathiou, Ariane Pinche, Benoît Sagot |
LREC | 5 |
| 2026 | A Parallel Corpus of the Parable of the Prodigal Son: Building a Resource for Documenting Language Varieties in Mainland FranceabstractInternational audience Lucence Ing, Juliette Janes, Sven Ködel, Benoît Sagot |
LREC | 4 |
| 2026 | Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine TranslationabstractInternational audience Malik Marmonier, Benoît Sagot, Rachel Bawden |
LREC | 2 |
| 2026 | ForumOccitania: A Corpus of User-Generated Content for Multiple Occitan VarietiesabstractAccepted at LREC 2026 Oriane Nédey, Juliette Janes, Rachel Bawden, Thibault Clérice, Benoît Sagot |
LREC | 5 |
| 2025 | Explicit Learning and the LLM in Machine TranslationabstractThis study explores an LLM's ability to learn new languages using explanations found in a grammar book-a process we term "explicit learning."To rigorously assess this ability, we design controlled translation experiments between English and constructed languages generated-through specific cryptographic means-from Latin or French.Contrary to previous studies, our results demonstrate that LLMs do possess a measurable capacity for explicit learning.This ability, however, diminishes as the complexity of the linguistic phenomena to be learned increases.Supervised fine-tuning on ad hoc chains of thought significantly enhances LLM performance but struggles to generalize to typologically novel or more complex linguistic features.These findings point to the need for more diverse training sets and alternative fine-tuning strategies to further improve explicit learning by LLMs, benefiting low-resource languages typically described in grammar books but lacking extensive corpora. Malik Marmonier, Rachel Bawden, Benoît Sagot |
EMNLP | 3 |
| 2024 | From Text to Source: Results in Detecting Large Language Model-Generated ContentabstractThe widespread use of Large Language Models (LLMs), celebrated for their ability to generate human-like text, has raised concerns about misinformation and ethical implications. Addressing these concerns necessitates the development of robust methods to detect and attribute text generated by LLMs. This paper investigates “Cross-Model Detection,” by evaluating whether a classifier trained to distinguish between source LLM-generated and human-written text can also detect text from a target LLM without further training. The study comprehensively explores various LLM sizes and families and assesses the impact of conversational fine-tuning techniques, quantization, and watermarking on classifier generalization. The research also explores Model Attribution, encompassing source model identification, model family, and model size classification, in addition to quantization and watermarking detection. Our results reveal several key findings: a clear inverse relationship between classifier effectiveness and model size, with larger LLMs being more challenging to detect, especially when the classifier is trained on data from smaller models. Training on data from similarly sized LLMs can improve detection performance from larger models but may lead to decreased performance when dealing with smaller models. Additionally, model attribution experiments show promising results in identifying source models and model families, highlighting detectable signatures in LLM-generated text, with particularly remarkable outcomes in watermarking detection, while no detectable signatures of quantization were observed. Overall, our study contributes valuable insights into the interplay of model size, family, and training data in LLM detection and attribution. Wissam Antoun, Benoît Sagot, Djamé Seddah |
LREC/COLING | 2 |
| 2024 | When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced LanguagesabstractMost existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful for low-resource languages (LRLs), where large datasets are not available. Often we are interested in building bilingual resources for LRLs against related high-resource languages (HRLs), resulting in severely imbalanced data settings for BLI. We first show that state-of-the-art BLI methods in the literature exhibit near-zero performance for severely data-imbalanced language pairs, indicating that these settings require more robust techniques. We then present a new method for unsupervised BLI between a related LRL and HRL that only requires inference on a masked language model of the HRL, and demonstrate its effectiveness on truly low-resource languages Bhojpuri and Magahi (with <5M monolingual tokens each), against Hindi. We further present experiments on (mid-resource) Marathi and Nepali to compare approach performances by resource range, and release our resulting lexicons for five low-resource Indic languages: Bhojpuri, Magahi, Awadhi, Braj, and Maithili, against Hindi. Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, Rachel Bawden |
LREC/COLING | 4 |
| 2024 | On the Scaling Laws of Geographical Representation in Language ModelsabstractLanguage models have long been shown to embed geographical information in their hidden representations. This line of work has recently been revisited by extending this result to Large Language Models (LLMs). In this paper, we propose to fill the gap between well-established and recent literature by observing how geographical knowledge evolves when scaling language models. We show that geographical knowledge is observable even for tiny models, and that it scales consistently as we increase the model size. Notably, we observe that larger language models cannot mitigate the geographical bias that is inherent to the training data. Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot |
LREC/COLING | 3 |
| 2024 | Making Sentence Embeddings Robust to User-Generated ContentabstractNLP models have been known to perform poorly on user-generated content (UGC), mainly because it presents a lot of lexical variations and deviates from the standard texts on which most of these models were trained. In this work, we focus on the robustness of LASER, a sentence embedding model, to UGC data. We evaluate this robustness by LASER’s ability to represent non-standard sentences and their standard counterparts close to each other in the embedding space. Inspired by previous works extending LASER to other languages and modalities, we propose RoLASER, a robust English encoder trained using a teacher-student approach to reduce the distances between the representations of standard and UGC sentences. We show that with training only on standard and synthetic UGC-like data, RoLASER significantly improves LASER’s robustness to both natural and artificial UGC data by achieving up to 2x and 11x better scores. We also perform a fine-grained analysis on artificial UGC data and find that our model greatly outperforms LASER on its most challenging UGC phenomena such as keyboard typos and social media abbreviations. Evaluation on downstream tasks shows that RoLASER performs comparably to or better than LASER on standard data, while consistently outperforming it on UGC data. Lydia Nishimwe, Benoît Sagot, Rachel Bawden |
LREC/COLING | 2 |
| 2024 | Anisotropy Is Inherent to Self-Attention in TransformersabstractThe representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers.In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity).Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens.We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences.We also show that the anisotropy problem extends to Transformers trained on other modalities.Our observations suggest that anisotropy is actually inherent to Transformers-based models. Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot |
EACL (1) | 3 |
| 2024 | Tree of Problems: Improving structured problem solving with compositionalityabstractLarge Language Models (LLMs) have demonstrated remarkable performance across multiple tasks through in-context learning.For complex reasoning tasks that require step-by-step thinking, Chain-of-Thought (CoT) prompting has given impressive results, especially when combined with self-consistency.Nonetheless, some tasks remain particularly difficult for LLMs to solve.Tree of Thoughts (ToT) and Graph of Thoughts (GoT) emerged as alternatives, dividing the complex problem into paths of subproblems.In this paper, we propose Tree of Problems (ToP), a simpler version of ToT, which we hypothesise can work better for complex tasks that can be divided into identical subtasks.Our empirical results show that our approach outperforms ToT and GoT, and in addition performs better than CoT on complex reasoning tasks.All code for this paper is publicly available here: https://github. com/ArmelRandy/tree-of-problems. Armel Zebaze, Benoît Sagot, Rachel Bawden |
EMNLP | 2 |
| 2024 | Headless Language Models: Learning without Predicting with Contrastive Weight TyingabstractSelf-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative method that shifts away from probability prediction and instead focuses on reconstructing input embeddings in a contrastive fashion via Constrastive Weight Tying (CWT). We apply this approach to pretrain Headless Language Models in both monolingual and multilingual contexts. Our method offers practical advantages, substantially reducing training computational requirements by up to 20 times, while simultaneously enhancing downstream performance and data efficiency. We observe a significant +1.6 GLUE score increase and a notable +2.7 LAMBADA accuracy improvement compared to classical LMs within similar compute budgets. Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot |
ICLR | 3 |
| 2024 | PatentEval: Understanding Errors in Patent GenerationabstractYou Zuo, Kim Gerdes, Éric Clergerie, Benoît Sagot. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. You Zuo, Kim Gerdes, Éric Villemonte de la Clergerie, Benoît Sagot |
NAACL-HLT | 4 |
| 2023 | SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech TranslationsabstractPaul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Paul-Ambroise Duquenne, Hongyu Gong, Jingfei Du, Ann Lee 0001, Vedanuj Goswami, Changhan Wang, Juan Pino 0001, Benoît Sagot, Holger Schwenk |
ACL (1) | 9 |
| 2023 | Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive EvaluationabstractOne of the major challenges of machine translation (MT) is ambiguity, which can in some cases be resolved by accompanying context such as images.However, recent work in multimodal MT (MMT) has shown that obtaining improvements from images is challenging, limited not only by the difficulty of building effective cross-modal representations, but also by the lack of specific evaluation and training data.We present a new MMT approach based on a strong text-only MT model, which uses neural adapters, a novel guided self-attention mechanism and which is jointly trained on both visually-conditioned masking and MMT.We also introduce CoMMuTE, a Contrastive Multilingual Multimodal Translation Evaluation set of ambiguous sentences and their possible translations, accompanied by disambiguating images corresponding to each translation.Our approach obtains competitive results compared to strong text-only models on standard English→French, English→German and English→Czech benchmarks and outperforms baselines and state-of-the-art MMT systems by a large margin on our contrastive test set.Our code 1 and CoMMuTE 2 are freely available.8 mBART is pretrained on CC25 (Wenzek et al., 2020). Matthieu Futeral, Cordelia Schmid, Ivan Laptev, Benoît Sagot, Rachel Bawden |
ACL (1) | 4 |
| 2023 | Generative Spoken Language Model based on continuous word-sized audio tokensabstractIn NLP, text language models based on words or subwords are known to outperform their character-based counterparts.Yet, in the speech community, the standard input of spoken LMs are 20ms or 40ms-long discrete units (shorter than a phoneme).Taking inspiration from wordbased LM, we introduce a Generative Spoken Language Model (GSLM) based on word-size continuous-valued audio embeddings that can generate diverse and expressive language output.This is obtained by replacing lookup table for lexical types with a Lexical Embedding function, the cross entropy loss by a contrastive loss, and multinomial sampling by k-NN sampling.The resulting model is the first generative language model based on word-size continuous embeddings.Its performance is on par with discrete unit GSLMs regarding generation quality as measured by automatic metrics and subjective human judgements.Moreover, it is five times more memory efficient thanks to its large 200ms units.In addition, the embeddings before and after the Lexical Embedder are phonetically and semantically interpretable. 1 Robin Algayres, Yossi Adi, Tu Anh Nguyen, Jade Copet, Gabriel Synnaeve, Benoît Sagot, Emmanuel Dupoux |
EMNLP | 6 |
| 2023 | Neural Agents Struggle to Take Turns in Bidirectional Emergent Communication
Valentin Taillandier, Dieuwke Hupkes, Benoît Sagot, Emmanuel Dupoux, Paul Michel |
ICLR | 3 |
| 2023 | Modular Speech-to-Text Translation for Zero-Shot Cross-Modal TransferabstractInternational audience Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot |
INTERSPEECH | 3 |
| 2023 | Generative Spoken Dialogue Language ModelingabstractAbstract We introduce dGSLM, the first “textless” model able to generate audio samples of naturalistic spoken dialogues. It uses recent work on unsupervised spoken unit discovery coupled with a dual-tower transformer architecture with cross-attention trained on 2000 hours of two-channel raw conversational audio (Fisher dataset) without any text or labels. We show that our model is able to generate speech, laughter, and other paralinguistic signals in the two channels simultaneously and reproduces more naturalistic and fluid turn taking compared to a text-based cascaded model.1,2 Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoît Sagot, Abdel-rahman Mohamed, Emmanuel Dupoux |
Trans. Assoc. Comput. Linguistics | 9 |
| 2022 | T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine TranslationabstractWe present a new approach to perform zeroshot cross-modal transfer between speech and text for translation tasks.Multilingual speech and text are encoded in a joint fixed-size representation space.Then, we compare different approaches to decode these multimodal and multilingual fixed-size representations, enabling zero-shot translation between languages and modalities.All our models are trained without the need of cross-modal labeled translation data.Despite a fixed-size representation, we achieve very competitive results on several text and speech translation tasks.In particular, we outperform the state of the art for zero-shot speech translation on Must-C.We also introduce the first results for zero-shot direct speechto-speech and text-to-speech translation. Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk |
EMNLP | 3 |
| 2022 | Speech Sequence Embeddings using Nearest Neighbors Contrastive LearningabstractInternational audience Robin Algayres, Adel Nabli, Benoît Sagot, Emmanuel Dupoux |
INTERSPEECH | 3 |
| 2022 | Towards a Cleaner Document-Oriented Multilingual Crawled CorpusabstractThe need for large corpora raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to manually curate the amount of data necessary to train large language models, the main way to obtain this data is still through automatic web crawling. In this paper we take the existing multilingual web corpus OSCAR and its pipeline Ungoliant that extracts and classifies data from Common Crawl at the line level, and propose a set of improvements and automatic annotations in order to produce a new document-oriented version of OSCAR that could prove more suitable to pre-train large generative language models as well as hopefully other applications in Natural Language Processing and Digital Humanities. Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, Benoît Sagot |
LREC | 4 |
| 2022 | Automatic Normalisation of Early Modern FrenchabstractSpelling normalisation is a useful step in the study and analysis of historical language texts, whether it is manual analysis by experts or automatic analysis using downstream natural language processing (NLP) tools. Not only does it help to homogenise the variable spelling that often exists in historical texts, but it also facilitates the use of off-the-shelf contemporary NLP tools, if contemporary spelling conventions are used for normalisation. We present FREEMnorm, a new benchmark for the normalisation of Early Modern French (from the 17th century) into contemporary French and provide a thorough comparison of three different normalisation methods: ABA, an alignment-based approach and MT-approaches, (both statistical and neural), including extensive parameter searching, which is often missing in the normalisation literature. Rachel Bawden, Jonathan Poinhos, Eleni Kogkitsidou, Philippe Gambette, Benoît Sagot, Simon Gabay |
LREC | 5 |
| 2022 | Complex Labelling and Similarity Prediction in Legal Texts: Automatic Analysis of France's Court of Cassation RulingsabstractDetecting divergences in the applications of the law (where the same legal text is applied differently by two rulings) is an important task. It is the mission of the French Cour de Cassation. The first step in the detection of divergences is to detect similar cases, which is currently done manually by experts. They rely on summarised versions of the rulings (syntheses and keyword sequences), which are currently produced manually and are not available for all rulings. There is also a high degree of variability in the keyword choices and the level of granularity used. In this article, we therefore aim to provide automatic tools to facilitate the search for similar rulings. We do this by (i) providing automatic keyword sequence generation models, which can be used to improve the coverage of the analysis, and (ii) providing measures of similarity based on the available texts and augmented with predicted keyword sequences. Our experiments show that the predictions improve correlations of automatically obtained similarities against our specially colelcted human judgments of similarity. Thibault Charmet, Inès Cherichi, Matthieu Allain, Urszula Czerwinska, Amaury Fouret, Benoît Sagot, Rachel Bawden |
LREC | 6 |
| 2022 | From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern Frenchabstractanguage models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historical states are at the same time more complex to process and more scarce in the corpora available, this paper presents recent efforts to overcome this difficult situation. These efforts include producing a corpus, creating the model, and evaluating it with an NLP task currently used by scholars in other ongoing projects. Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot |
LREC | 7 |
| 2022 | BERTrade: Using Contextual Embeddings to Parse Old FrenchabstractThe successes of contextual word embeddings learned by training large-scale language models, while remarkable, have mostly occurred for languages where significant amounts of raw texts are available and where annotated data in downstream tasks have a relatively regular spelling. Conversely, it is not yet completely clear if these models are also well suited for lesser-resourced and more irregular languages. We study the case of Old French, which is in the interesting position of having relatively limited amount of available raw text, but enough annotated resources to assess the relevance of contextual word embedding models for downstream NLP tasks. In particular, we use POS-tagging and dependency parsing to evaluate the quality of such models in a large array of configurations, including models trained from scratch from small amounts of raw text and models pre-trained on other languages but fine-tuned on Medieval French data. Loïc Grobol, Mathilde Regnault, Pedro Ortiz Suarez, Benoît Sagot, Laurent Romary, Benoît Crabbé |
LREC | 4 |
| 2022 | MUSS: Multilingual Unsupervised Sentence Simplification by Mining ParaphrasesabstractProgress in sentence simplification has been hindered by a lack of labeled parallel simplification data, particularly in languages other than English. We introduce MUSS, a Multilingual Unsupervised Sentence Simplification system that does not require labeled simplification data. MUSS uses a novel approach to sentence simplification that trains strong models using sentence-level paraphrase data instead of proper simplification data. These models leverage unsupervised pretraining and controllable generation mechanisms to flexibly adjust attributes such as length and lexical complexity at inference time. We further present a method to mine such paraphrase data in any language from Common Crawl using semantic sentence embeddings, thus removing the need for labeled data. We evaluate our approach on English, French, and Spanish simplification benchmarks and closely match or outperform the previous best supervised results, despite not using any labeled simplification data. We push the state of the art further by incorporating labeled simplification data. Louis Martin, Angela Fan, Éric Villemonte de la Clergerie, Antoine Bordes, Benoît Sagot |
LREC | 5 |
| 2022 | DP-Parse: Finding Word Boundaries from Raw Speech with an Instance LexiconabstractAbstract Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a ‘space’ delimiter between words. Popular Bayesian non-parametric models for text segmentation (Goldwater et al., 2006, 2009) use a Dirichlet process to jointly segment sentences and build a lexicon of word types. We introduce DP-Parse, which uses similar principles but only relies on an instance lexicon of word tokens, avoiding the clustering errors that arise with a lexicon of word types. On the Zero Resource Speech Benchmark 2017, our model sets a new speech segmentation state-of-the-art in 5 languages. The algorithm monotonically improves with better input representations, achieving yet higher scores when fed with weakly supervised inputs. Despite lacking a type lexicon, DP-Parse can be pipelined to a language model and learn semantic and syntactic representations as assessed by a new spoken word embedding benchmark. 1 Robin Algayres, Tristan Ricoul, Julien Karadayi, Hugo Laurençon, Mohamed Salah Zaïem, Abdel-rahman Mohamed, Benoît Sagot, Emmanuel Dupoux |
Trans. Assoc. Comput. Linguistics | 7 |
| 2022 | Quality at a Glance: An Audit of Web-Crawled Multilingual DatasetsabstractAbstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases. Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi |
Trans. Assoc. Comput. Linguistics | 14 |
| 2021 | First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERTabstractMultilingual pretrained language models have demonstrated remarkable zero-shot crosslingual transfer capabilities.Such transfer emerges by fine-tuning on a task of interest in one language and evaluating on a distinct language, not seen during the fine-tuning.Despite promising results, we still lack a proper understanding of the source of this transfer.Using a novel layer ablation technique and analyses of the model's internal representations, we show that multilingual BERT, a popular multilingual language model, can be viewed as the stacking of two sub-networks: a multilingual encoder followed by a taskspecific language-agnostic predictor.While the encoder is crucial for cross-lingual transfer and remains mostly unchanged during finetuning, the task predictor has little importance on the transfer and can be reinitialized during fine-tuning.We present extensive experiments with three distinct tasks, seventeen typologically diverse languages and multiple domains to support our hypothesis.RANDOM-INIT of layers SRC-TRG REF ∆1-2 ∆3-4 ∆5-6 ∆7-8 ∆9-10 ∆11-12 Parsing EN -EN 88.98 -0.96 -0.66 -0.93 -0.55 0.04 -0.09RU -RU 85.15 -0.82 -1.38 -1.51 -0.86 -0.29 0.18 AR -AR 59.54 -0.78 -2.14 -1.20 -0.67 -0.27 0.08 EN -X 53.23 -15.77 -6.51 -3.39 -1.47 0.29 1.00 RU -X 55.41 -7.69 -3.71 -3.13 -1.70 0.92 0.94 AR -X 27.97 -4.91 -3.17 -1.48 -1.68 -0.36 -0.14 POS EN -EN 96.51 -0.30 -0.25 -0.40 -0.00 0.05 0.02 RU -RU 96.90 -0.52 -0.55 -0.40 -0.07 0.02 -0.03 AR -AR 79.28 -0.35 -0.49-0.36 -0.19 -0.05 -0.00 EN -X 79.37 -8.94 -2.49-1.66 -0.88 0.20 -0.14 RU -X 79.25 -10.08 -2.83 -1.65 -2.74 0.01 -0.45 AR -X 64.81 -6.73 -3.50 -1.63 -1.56 -0.73 -1.29 NER EN -EN 83.30 -2.66 -2.14 -1.43 -0.63 -0.23 -0.12 RU -RU 88.20 -2.08 -2.13 -1.52 -0.64 -0.33 -0.13 AR -AR 87.97 -2.37 -2.11 -0.96 -0.39 -0.15 0.21 EN -X 64.17 -8.28 -5.09 -3.07 -0.79 -0.47 -0.13 RU -X 62.13 -15.85 -9.36 -5.50 -2.44 -1.16 -0.06 AR -X 65.59 -16.10 -8.42 -3. Benjamin Muller, Yanai Elazar, Benoît Sagot, Djamé Seddah |
EACL | 3 |
| 2021 | Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question AnsweringabstractCoupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on the Question Answering task. However, most of those datasets are in English, and the performances of state-of-the-art multilingual models are significantly lower when evaluated on non-English data. Due to high data collection costs, it is not realistic to obtain annotated data for each language one desires to support. We propose a method to improve the Cross-lingual Question Answering performance without requiring additional annotated data, leveraging Question Generation models to produce synthetic samples in a cross-lingual fashion. We show that the proposed method allows to significantly outperform the baselines trained on English data only. We report a new state-of-the-art on four multilingual datasets: MLQA, XQuAD, SQuAD-it and PIAF (fr). Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, Jacopo Staiano |
EMNLP (1) | 4 |
| 2021 | When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language ModelsabstractBenjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah |
NAACL-HLT | 3 |
| 2020 | ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting TransformationsabstractIn order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e.replacing complex words or phrases by simpler synonyms), reorder components, and/or delete information deemed unnecessary.Despite these varied range of possible text alterations, current models for automatic sentence simplification are evaluated using datasets that are focused on a single transformation, such as lexical paraphrasing or splitting.This makes it impossible to understand the ability of simplification models in more realistic settings.To alleviate this limitation, this paper introduces ASSET, a new dataset for assessing sentence simplification in English.ASSET is a crowdsourced multi-reference corpus where each simplification was produced by executing several rewriting transformations.Through quantitative and qualitative experiments, we show that simplifications in ASSET are better at capturing characteristics of simplicity when compared to other standard evaluation datasets for the task.Furthermore, we motivate the need for developing better methods for automatic evaluation using ASSET, since we show that current popular metrics may not be suitable when multiple simplification transformations are performed. Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, Lucia Specia |
ACL | 5 |
| 2020 | CamemBERT: a Tasty French Language ModelabstractLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Louis Martin, Benjamin Muller, Pedro Ortiz Suarez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, Benoît Sagot |
ACL | 8 |
| 2020 | Building a User-Generated Content North-African Arabizi Treebank: Tackling HellabstractDjamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Ortiz Suarez, Benoît Sagot |
ACL | 7 |
| 2020 | A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesabstractWe use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages.We then compare the performance of OSCARbased and Wikipedia-based ELMo embeddings for these languages on the part-ofspeech tagging and parsing tasks.We show that, despite the noise in the Common-Crawlbased OSCAR data, embeddings trained on OSCAR perform much better than monolingual embeddings trained on Wikipedia.They actually equal or improve the current state of the art in tagging and parsing for all five languages.In particular, they also improve over multilingual Wikipedia-based contextual embeddings (multilingual BERT), which almost always constitutes the previous state of the art, thereby showing that the benefit of a larger, more diverse corpus surpasses the crosslingual benefit of multilingual embedding architectures. Pedro Ortiz Suarez, Laurent Romary, Benoît Sagot |
ACL | 3 |
| 2020 | Evaluating the Reliability of Acoustic Speech EmbeddingsabstractInternational audience Robin Algayres, Mohamed Salah Zaïem, Benoît Sagot, Emmanuel Dupoux |
INTERSPEECH | 3 |
| 2020 | Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0abstractDiachronic lexical information is not only important in the field of historical linguistics, but is also increasingly used in NLP, most recently for machine translation of low resource languages. Therefore, there is a need for fine-grained, large-coverage and accurate etymological lexical resources. In this paper, we propose a set of guidelines to generate such resources, for each step of the life-cycle of an etymological lexicon: creation, update, evaluation, dissemination, and exploitation. To illustrate the guidelines, we introduce EtymDB 2.0, an etymological database automatically generated from the Wiktionary, which contains 1.8 million lexemes, linked by more than 700,000 fine-grained etymological relations, across 2,536 living and dead languages. We also introduce use cases for which EtymDB 2.0 could represent a key resource, such as phylogenetic tree generation, low resource machine translation or medieval languages study. Clémentine Fourrier, Benoît Sagot |
LREC | 2 |
| 2020 | OFrLex: A Computational Morphological and Syntactic Lexicon for Old FrenchabstractIn this paper we describe our work on the development and enrichment of OFrLex, a freely available, large-coverage morphological and syntactic Old French lexicon. We rely on several heterogeneous language resources to extract structured and exploitable information. The extraction follows a semi-automatic procedure with substantial manual steps to respond to difficulties encountered while aligning lexical entries from distinct language resources. OFrLex aims at improving natural language processing tasks on Old French such as part-of-speech tagging and dependency parsing. We provide quantitative information on OFrLex and discuss its reliability. We also describe and evaluate a semi-automatic, word-embedding-based lexical enrichment process aimed at increasing the accuracy of the resource. Results of this extension technique will be manually validated in the near future, a step that will take advantage of OFrLex’s viewing, searching and editing interface, which is already accessible online. Gaël Guibon, Benoît Sagot |
LREC | 2 |
| 2020 | Controllable Sentence SimplificationabstractText simplification aims at making a text easier to read and understand by simplifying grammar and structure while keeping the underlying information identical. It is often considered an all-purpose generic task where the same simplification is suitable for all; however multiple audiences can benefit from simplified text in different ways. We adapt a discrete parametrization mechanism that provides explicit control on simplification systems based on Sequence-to-Sequence models. As a result, users can condition the simplifications returned by a model on attributes such as length, amount of paraphrasing, lexical complexity and syntactic complexity. We also show that carefully chosen values of these attributes allow out-of-the-box Sequence-to-Sequence models to outperform their standard counterparts on simplification benchmarks. Our model, which we call ACCESS (as shorthand for AudienCe-CEntric Sentence Simplification), establishes the state of the art at 41.87 SARI on the WikiLarge test set, a +1.42 improvement over the best previously reported score. Louis Martin, Éric Villemonte de la Clergerie, Benoît Sagot, Antoine Bordes |
LREC | 3 |
| 2020 | Establishing a New State-of-the-Art for French Named Entity RecognitionabstractThe French TreeBank developed at the University Paris 7 is the main source of morphosyntactic and syntactic annotations for French. However, it does not include explicit information related to named entities, which are among the most useful information for several natural language processing tasks and applications. Moreover, no large-scale French corpus with named entity annotations contain referential information, which complement the type and the span of each mention with an indication of the entity it refers to. We have manually annotated the French TreeBank with such information, after an automatic pre-annotation step. We sketch the underlying annotation guidelines and we provide a few figures about the resulting annotations. Pedro Ortiz Suarez, Yoann Dupont, Benjamin Muller, Laurent Romary, Benoît Sagot |
LREC | 5 |
| 2019 | What Does BERT Learn about the Structure of Language?abstractBERT is a recent language representation model that has surprisingly performed well in diverse language understanding benchmarks.This result indicates the possibility that BERT networks capture structural information about language.In this work, we provide novel support for this claim by performing a series of experiments to unpack the elements of English language structure learned by BERT.We first show that BERT's phrasal representation captures phrase-level information in the lower layers.We also show that BERT's intermediate layers encode a rich hierarchy of linguistic information, with surface features at the bottom, syntactic features in the middle and semantic features at the top.BERT turns out to require deeper layers when long-distance dependency information is required, e.g. to track subjectverb agreement.Finally, we show that BERT representations capture linguistic information in a compositional way that mimics classical, tree-like structures. Ganesh Jawahar, Benoît Sagot, Djamé Seddah |
ACL (1) | 2 |
| 2018 | CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing
Amir More, Özlem Çetinoglu, Çagri Çöltekin, Nizar Habash, Benoît Sagot, Djamé Seddah, Dima Taji, Reut Tsarfaty |
LREC | 5 |
| 2018 | A multilingual collection of CoNLL-U-compatible morphological lexicons
Benoît Sagot |
LREC | 1 |
| 2018 | Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer
Djamé Seddah, Éric Villemonte de la Clergerie, Benoît Sagot, Héctor Martínez Alonso, Marie Candito |
LREC | 3 |
| 2014 | A Language-independent Approach to Extracting Derivational Relations from an Inflectional Lexicon
Marion Baranes, Benoît Sagot |
LREC | 2 |
| 2014 | Developing a French FrameNet: Methodology and First results
Marie Candito, Pascal Amsili, Lucie Barque, Farah Benamara, Gaël de Chalendar, Marianne Djemaa, Pauline Haas, Richard Huyghe, Yvette Yannick Mathieu, Philippe Muller, Benoît Sagot, Laure Vieu |
LREC | 11 |
| 2014 | An Open-Source Heavily Multilingual Translation Graph Extracted from Wiktionaries and Parallel Corpora
Valérie Hanoka, Benoît Sagot |
LREC | 2 |
| 2014 | DeLex, a freely-avaible, large-scale and linguistically grounded morphological lexicon for German
Benoît Sagot |
LREC | 1 |
| 2014 | A language-independent and fully unsupervised approach to lexicon induction and part-of-speech tagging for closely related languages
Yves Scherrer, Benoît Sagot |
LREC | 2 |
| 2013 | Enforcing Subcategorization Constraints in a Parser Using Sub-parses Recombining
Seyed Abolghasem Mirroshandel, Alexis Nasr, Benoît Sagot |
HLT-NAACL | 3 |
| 2012 | The French Social Media Bank: a Treebank of Noisy User Generated Content
Djamé Seddah, Benoît Sagot, Marie Candito, Virginie Mouilleron, Vanessa Combet |
COLING | 2 |
| 2012 | Applying cross-lingual WSD to wordnet development
Marianna Apidianaki, Benoît Sagot |
LREC | 2 |
| 2012 | Boosting the Coverage of a Semantic Lexicon by Automatically Extracted Event Nominalizations
Kata Gábor, Marianna Apidianaki, Benoît Sagot, Éric Villemonte de la Clergerie |
LREC | 3 |
| 2012 | Wordnet extension made simple: A multilingual lexicon-based approach using wiki resources
Valérie Hanoka, Benoît Sagot |
LREC | 2 |
| 2012 | Cleaning noisy wordnets
Benoît Sagot, Darja Fiser |
LREC | 1 |
| 2012 | Aleda, a free large-scale entity database for French
Benoît Sagot, Rosa Stern |
LREC | 1 |
| 2012 | Evaluating and improving syntactic lexica by plugging them within a parser
Elsa Tolone, Benoît Sagot, Éric Villemonte de la Clergerie |
LREC | 2 |
| 2010 | Optimal Rank Reduction for Linear Context-Free Rewriting Systems with Fan-Out Two
Benoît Sagot, Giorgio Satta |
ACL | 1 |
| 2010 | The Lefff, a Freely Available and Large-coverage Morphological and Syntactic Lexicon for French
Benoît Sagot |
LREC | 1 |
| 2010 | A Lexicon of French Quotation Verbs for Automatic Quotation Extraction
Benoît Sagot, Laurence Danlos, Rosa Stern |
LREC | 1 |
| 2010 | A Morphological Lexicon for the Persian Language
Benoît Sagot, Géraldine Walther |
LREC | 1 |
| 2009 | Coupling an Annotated Corpus and a Morphosyntactic Lexicon for State-of-the-Art POS Tagging with Less Human Effort
Pascal Denis, Benoît Sagot |
PACLIC | 2 |
| 2008 | Computer Aided Correction and Extension of a Syntactic Wide-Coverage Lexicon
Lionel Nicolas, Benoît Sagot, Miguel A. Molinero, Jacques Farré, Éric Villemonte de la Clergerie |
COLING | 2 |
| 2006 | Error Mining in Parsing ResultsabstractWe introduce an error mining technique for automatically detecting errors in resources that are used in parsing systems. We applied this technique on parsing results produced on several million words by two distinct parsing systems, which share the syntactic lexicon and the pre-parsing processing chain. We were thus able to identify missing and erroneous information in these resources. Benoît Sagot, Éric Villemonte de la Clergerie |
ACL | 1 |
| 2006 | Deep non-probabilistic parsing of large corpora
Benoît Sagot, Pierre Boullier |
LREC | 1 |
| 2006 | The Lefff 2 syntactic lexicon for French: architecture, acquisition, use
Benoît Sagot, Lionel Clément, Éric Villemonte de la Clergerie, Pierre Boullier |
LREC | 1 |
| 2004 | Morphology Based Automatic Acquisition of Large-coverage Lexica
Lionel Clément, Benoît Sagot, Bernard Lang |
LREC | 2 |