VLDB 2026 Research / reviewers in the wild / expert
Maximin Coavoux
dblp:184/3728
· DBLP profile ↗
17ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-4089-4558ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Radio Haiti-Inter: A Large-Scale Annotated Corpus of Spoken Haitian CreoleabstractInternational audience William Havard, Rayan Ziane, Mélissa Menclé, Maximin Coavoux, Benjamin Lecouteux, Emmanuel Schang |
LREC | 4 |
| 2026 | Pantagruel: Unified Self-Supervised Encoders for French Text and SpeechabstractInternational audience Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane, Mohammed Ghennai, Qianwen Guan, Kirill Milintsevich, Salima Mdhaffar, Aidan Mannion, Nils Defauw, Shuyue Gu, Alexandre Audibert, Marco Dinarelli, Yannick Estève, Lorraine Goeuriot, Steffen Lalande, Nicolas Hervé, Maximin Coavoux, François Portet, Étienne Ollion, Marie Candito, Maxime Peyrard, Solange Rossato, Benjamin Lecouteux, Aurélie Nardy, Gilles Sérasset, Vincent Segonne, Solène Evain, Diandra Fabre, Didier Schwab |
LREC | 18 |
| 2024 | Limitations of Human Identification of Automatically Generated TextabstractNeural text generation is receiving broad attention with the publication of new tools such as ChatGPT. The main reason for that is that the achieved quality of the generated text may be attributed to a human writer by the naked eye of a human evaluator. In this paper, we propose a new corpus in French and English for the task of recognising automatically generated texts and we conduct a study of how humans perceive the text. Our results show, as previous work before the ChatGPT era, that the generated texts by tools such as ChatGPT share some common characteristics but they are not clearly identifiable which generates different perceptions of these texts. Nadège Alavoine, Maximin Coavoux, Emmanuelle Esperança-Rodier, Romane Gallienne, Carlos E. González-Gallardo, Jérôme Goulian, José G. Moreno 0001, Aurélie Névéol, Didier Schwab, Vincent Segonne, Johanna Simoens |
LREC/COLING | 2 |
| 2024 | What Has LeBenchmark Learnt about French Syntax?abstractThe paper reports on a series of experiments aiming at probing LeBenchmark, a pretrained acoustic model trained on 7k hours of spoken French, for syntactic information. Pretrained acoustic models are increasingly used for downstream speech tasks such as automatic speech recognition, speech translation, spoken language understanding or speech parsing. They are trained on very low level information (the raw speech signal), and do not have explicit lexical knowledge. Despite that, they obtained reasonable results on tasks that requires higher level linguistic knowledge. As a result, an emerging question is whether these models encode syntactic information. We probe each representation layer of LeBenchmark for syntax, using the Orféo treebank, and observe that it has learnt some syntactic information. Our results show that syntactic information is more easily extractable from the middle layers of the network, after which a very sharp decrease is observed. Zdravko Dugonjic, Adrien Pupier, Benjamin Lecouteux, Maximin Coavoux |
LREC/COLING | 4 |
| 2024 | Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized DomainsabstractPretrained Language Models (PLMs) are the de facto backbone of most state-of-the-art NLP systems. In this paper, we introduce a family of domain-specific pretrained PLMs for French, focusing on three important domains: transcribed speech, medicine, and law. We use a transformer architecture based on efficient methods (LinFormer) to maximise their utility, since these domains often involve processing long documents. We evaluate and compare our models to state-of-the-art models on a diverse set of tasks and datasets, some of which are introduced in this paper. We gather the datasets into a new French-language evaluation benchmark for these three domains. We also compare various training configurations: continued pretraining, pretraining from scratch, as well as single- and multi-domain pretraining. Extensive domain-specific experiments show that it is possible to attain competitive downstream performance even when pre-training with the approximative LinFormer attention mechanism. For full reproducibility, we release the models and pretraining data, as well as contributed datasets. Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Audibert, Cécile Macaire, Adrien Pupier, Yongxin Zhou 0004, Mathilde Aguiar, Felix Herron, Magali Norré, Massih-Reza Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab |
LREC/COLING | 24 |
| 2024 | LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
Titouan Parcollet, Solène Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le 0001, Sina Alisamir, Natalia A. Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Estève, Mickael Rouvier, Jérôme Goulian, Benjamin Lecouteux, François Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier |
Comput. Speech Lang. | 13 |
| 2023 | BERT Is Not The Count: Learning to Match Mathematical Statements with ProofsabstractWe introduce a task consisting in matching a proof to a given mathematical statement.The task fits well within current research on Mathematical Information Retrieval and, more generally, mathematical article analysis (Mathematical Sciences, 2014).We present a dataset for the task (the MATCH dataset) consisting of over 180k statement-proof pairs extracted from modern mathematical research articles.1 We find this dataset highly representative of our task, as it consists of relatively new findings useful to mathematicians.We propose a bilinear similarity model and two decoding methods to match statements to proofs effectively.While the first decoding method matches a proof to a statement without being aware of other statements or proofs, the second method treats the task as a global matching problem.Through a symbol replacement procedure, we analyze the "insights" that pre-trained language models have in such mathematical article analysis and show that while these models perform well on this task with the best performing mean reciprocal rank of 73.7, they follow a relatively shallow symbolic analysis and matching to achieve that performance.2 * Work mostly done at the University of Edinburgh. 1 Our dataset and code are available at https:// github.com/waylonli/MATcH.2 Like Bert, The Count (or Count von Count; ) is a character from the television show Sesame Street.The Count likes counting, and his main role in the show is to teach this skill to children. Weixian Waylon Li, Yftah Ziser, Maximin Coavoux, Shay B. Cohen |
EACL | 3 |
| 2023 | PROPICTO: Developing Speech-to-Pictograph Translation Systems to Enhance Communication AccessibilityabstractPROPICTO is a project funded by the French National Research Agency and the Swiss National Science Foundation, that aims at creating Speech-to-Pictograph translation systems, with a special focus on French as an input language. By developing such technologies, we intend to enhance communication access for non-French speaking patients and people with cognitive impairments. Lucia Ormaechea Grijalba, Pierrette Bouillon, Maximin Coavoux, Emmanuelle Esperança-Rodier, Johanna Gerlach, Jérôme Goulian, Benjamin Lecouteux, Cécile Macaire, Jonathan Mutal, Magali Norré, Adrien Pupier, Didier Schwab |
EAMT | 3 |
| 2023 | On Detecting Policy-Related Political Ads: An Exploratory Analysis of Meta Ads in 2022 French ElectionabstractOnline political advertising has become the cornerstone of political campaigns. The budget spent solely on political advertising in the U.S. has increased by more than 100% from $ 700 million during the 2017-2018 U.S. election cycle to $ 1.6 billion during the 2020 U.S. presidential elections. Naturally, the capacity offered by online platforms to micro-target ads with political content has been worrying lawmakers, journalists, and online platforms, especially after the 2016 U.S. presidential election, where Cambridge Analytica has targeted voters with political ads congruent with their personality. Vera Sosnovik, Romaissa Kessi, Maximin Coavoux, Oana Goga |
WWW | 3 |
| 2022 | End-to-End Dependency Parsing of Spoken FrenchabstractInternational audience Adrien Pupier, Maximin Coavoux, Benjamin Lecouteux, Jérôme Goulian |
INTERSPEECH | 2 |
| 2021 | Self-Supervised and Controlled Multi-Document Opinion SummarizationabstractWe address the problem of unsupervised abstractive summarization of collections of user generated reviews through self-supervision and control.We propose a self-supervised setup that considers an individual document as a target summary for a set of similar documents.This setting makes training simpler than previous approaches by relying only on standard log-likelihood loss and mainstream models.We address the problem of hallucinations through the use of control codes, to steer the generation towards more coherent and relevant summaries.Our benchmarks on two English datasets against graph-based and recent neural abstractive unsupervised models show that our proposed method generates summaries with a superior quality and relevance, as well as a high sentiment and topic alignment with the input reviews.This is confirmed in our human evaluation which focuses explicitly on the faithfulness of generated summaries.We also provide an ablation study showing the importance of the control setup in controlling hallucinations. Hady ElSahar, Maximin Coavoux, Jos Rozen, Matthias Gallé |
EACL | 2 |
| 2020 | FlauBERT: Unsupervised Language Model Pre-training for FrenchabstractLanguage models have become a key step to achieve state-of-the art results in many different Natural Language Processing (NLP) tasks. Leveraging the huge amount of unlabeled texts nowadays available, they provide an efficient way to pre-train continuous word representations that can be fine-tuned for a downstream task, along with their contextualization at the sentence level. This has been widely demonstrated for English using contextualized representations (Dai and Le, 2015; Peters et al., 2018; Howard and Ruder, 2018; Radford et al., 2018; Devlin et al., 2019; Yang et al., 2019b). In this paper, we introduce and share FlauBERT, a model learned on a very large and heterogeneous French corpus. Models of different sizes are trained using the new CNRS (French National Centre for Scientific Research) Jean Zay supercomputer. We apply our French language models to diverse NLP tasks (text classification, paraphrasing, natural language inference, parsing, word sense disambiguation) and show that most of the time they outperform other pre-training approaches. Different versions of FlauBERT as well as a unified evaluation protocol for the downstream tasks, called FLUE (French Language Understanding Evaluation), are shared to the research community for further reproducible experiments in French NLP. Hang Le 0001, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, Didier Schwab |
LREC | 5 |
| 2019 | Unlexicalized Transition-based Discontinuous Constituency ParsingabstractAbstract Lexicalized parsing models are based on the assumptions that (i) constituents are organized around a lexical head and (ii) bilexical statistics are crucial to solve ambiguities. In this paper, we introduce an unlexicalized transition-based parser for discontinuous constituency structures, based on a structure-label transition system and a bi-LSTM scoring system. We compare it with lexicalized parsing models in order to address the question of lexicalization in the context of discontinuous constituency parsing. Our experiments show that unlexicalized models systematically achieve higher results than lexicalized models, and provide additional empirical evidence that lexicalization is not necessary to achieve strong parsing results. Our best unlexicalized model sets a new state of the art on English and German discontinuous constituency treebanks. We further provide a per-phenomenon analysis of its errors on discontinuous constituents. Maximin Coavoux, Benoît Crabbé, Shay B. Cohen |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | Privacy-preserving Neural Representations of TextabstractThis article deals with adversarial attacks towards deep learning systems for Natural Language Processing (NLP), in the context of privacy protection.We study a specific type of attack: an attacker eavesdrops on the hidden representations of a neural text classifier and tries to recover information about the input text.Such scenario may arise in situations when the computation of a neural network is shared across multiple devices, e.g.some hidden representation is computed by a user's device and sent to a cloud-based model.We measure the privacy of a hidden representation by the ability of an attacker to predict accurately specific private information from it and characterize the tradeoff between the privacy and the utility of neural representations.Finally, we propose several defense methods based on modified training objectives and show that they improve the privacy of neural representations. Maximin Coavoux, Shashi Narayan, Shay B. Cohen |
EMNLP | 1 |
| 2017 | Incremental Discontinuous Phrase Structure Parsing with the GAP TransitionabstractThis article introduces a novel transition system for discontinuous lexicalized constituent parsing called SR-GAP.It is an extension of the shift-reduce algorithm with an additional gap transition.Evaluation on two German treebanks shows that SR-GAP outperforms the previous best transitionbased discontinuous parser (Maier, 2015) by a large margin (it is notably twice as accurate on the prediction of discontinuous constituents), and is competitive with the state of the art (Fernández-González and Martins, 2015).As a side contribution, we adapt span features (Hall et al., 2014) to discontinuous parsing. Maximin Coavoux, Benoît Crabbé |
EACL (1) | 1 |
| 2017 | Cross-lingual RST Discourse ParsingabstractDiscourse parsing is an integral part of understanding information flow and argumentative structure in documents.Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank.However, discourse treebanks for other languages exist, including Spanish, German, Basque, Dutch and Brazilian Portuguese.The treebanks share the same underlying linguistic theory, but differ slightly in the way documents are annotated.In this paper, we present (a) a new discourse parser which is simpler, yet competitive (significantly better on 2/3 metrics) to state of the art for English, (b) a harmonization of discourse treebanks across languages, enabling us to present (c) what to the best of our knowledge are the first experiments on crosslingual discourse parsing. Chloé Braud, Maximin Coavoux, Anders Søgaard |
EACL (1) | 2 |
| 2016 | Neural Greedy Constituent Parsing with Dynamic OraclesabstractDynamic oracle training has shown substantial improvements for dependency parsing in various settings, but has not been explored for constituent parsing.The present article introduces a dynamic oracle for transition-based constituent parsing.Experiments on the 9 languages of the SPMRL dataset show that a neural greedy parser with morphological features, trained with a dynamic oracle, leads to accuracies comparable with the best non-reranking and non-ensemble parsers. Maximin Coavoux, Benoît Crabbé |
ACL (1) | 1 |