VLDB 2026 Research / reviewers in the wild / expert
Éric Villemonte de la Clergerie
dblp:54/5373 · also Éric de la Clergerie
· DBLP profile ↗
33ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0001-6428-9219ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 first-authorTheory of computation · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 61% Transfer learning and domain adaptation · 20% Language models and text generation · 15% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 9 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 1 | 2024 | Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.8 | 1 | 2024 | Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation › structured regularization
weight tying |
0.8 | 1 | 2024 | Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.4 | 1 | 2020 | CamemBERT: a Tasty French Language Model · ACL 2020 |
Natural language and speech › Language models and text generation
masked language modeling |
0.1 | 1 | 2020 | CamemBERT: a Tasty French Language Model · ACL 2020 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 2 | 2006 | Error Mining in Parsing Results · ACL 2006 Guided Parsing of Range Concatenation Languages · ACL 2001 |
Programming languages and type systems
logic programming |
0.0 | 1 | 1993 | Layer Sharing: An Improved Structure-Sharing Framework · POPL 1993 |
Compilers and program optimization
structure sharing |
0.0 | 1 | 1993 | Layer Sharing: An Improved Structure-Sharing Framework · POPL 1993 |
Methods — techniques the papers use, named apart from their topics
weight tying · 0.8contrastive learning · 0.8masked language modeling · 0.4shared derivation forest · 0.1guided parsing · 0.1error mining · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On the Scaling Laws of Geographical Representation in Language ModelsabstractLanguage models have long been shown to embed geographical information in their hidden representations. This line of work has recently been revisited by extending this result to Large Language Models (LLMs). In this paper, we propose to fill the gap between well-established and recent literature by observing how geographical knowledge evolves when scaling language models. We show that geographical knowledge is observable even for tiny models, and that it scales consistently as we increase the model size. Notably, we observe that larger language models cannot mitigate the geographical bias that is inherent to the training data. Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot |
LREC/COLING | 2 |
| 2024 | CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical DataabstractClinical data in hospitals are increasingly accessible for research through clinical data warehouses. However these documents are unstructured and it is therefore necessary to extract information from medical reports to conduct clinical studies. Transfer learning with BERT-like models such as CamemBERT has allowed major advances for French, especially for named entity recognition. However, these models are trained for plain language and are less efficient on biomedical data. Addressing this gap, we introduce CamemBERT-bio, a dedicated French biomedical model derived from a new public French biomedical dataset. Through continual pre-training of the original CamemBERT, CamemBERT-bio achieves an improvement of 2.54 points of F1-score on average across various biomedical named entity recognition tasks, reinforcing the potential of continual pre-training as an equally proficient yet less computationally intensive alternative to training from scratch. Additionally, we highlight the importance of using a standard evaluation protocol that provides a clear view of the current state-of-the-art for French biomedical models. Rian Touchent, Éric Villemonte de la Clergerie |
LREC/COLING | 2 |
| 2024 | Anisotropy Is Inherent to Self-Attention in TransformersabstractThe representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers.In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity).Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens.We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences.We also show that the anisotropy problem extends to Transformers trained on other modalities.Our observations suggest that anisotropy is actually inherent to Transformers-based models. Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot |
EACL (1) | 2 |
| 2024 | Translate your Own: a Post-Editing Experiment in the NLP domainabstractThe improvements in neural machine translation make translation and post-editing pipelines ever more effective for a wider range of applications. In this paper, we evaluate the effectiveness of such a pipeline for the translation of scientific documents (limited here to article abstracts). Using a dedicated interface, we collect, then analyse the post-edits of approximately 350 abstracts (English→French) in the Natural Language Processing domain for two groups of post-editors: domain experts (academics encouraged to post-edit their own articles) on the one hand and trained translators on the other. Our results confirm that such pipelines can be effective, at least for high-resource language pairs. They also highlight the difference in the post-editing strategy of the two subgroups. Finally, they suggest that working on term translation is the most pressing issue to improve fully automatic translations, but that in a post-editing setup, other error types can be equally annoying for post-editors. Rachel Bawden, Ziqian Peng, Maud Bénard, Éric Villemonte de la Clergerie, Raphaël Esamotunu, Mathilde Huguin, Natalie Kübler, Alexandra Mestivier, Mona Michelot, Laurent Romary, Lichao Zhu, François Yvon |
EAMT (1) | 4 |
| 2024 | Headless Language Models: Learning without Predicting with Contrastive Weight TyingabstractSelf-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative method that shifts away from probability prediction and instead focuses on reconstructing input embeddings in a contrastive fashion via Constrastive Weight Tying (CWT). We apply this approach to pretrain Headless Language Models in both monolingual and multilingual contexts. Our method offers practical advantages, substantially reducing training computational requirements by up to 20 times, while simultaneously enhancing downstream performance and data efficiency. We observe a significant +1.6 GLUE score increase and a notable +2.7 LAMBADA accuracy improvement compared to classical LMs within similar compute budgets. Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot |
ICLR | 2 |
| 2024 | PatentEval: Understanding Errors in Patent GenerationabstractYou Zuo, Kim Gerdes, Éric Clergerie, Benoît Sagot. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. You Zuo, Kim Gerdes, Éric Villemonte de la Clergerie, Benoît Sagot |
NAACL-HLT | 3 |
| 2022 | MUSS: Multilingual Unsupervised Sentence Simplification by Mining ParaphrasesabstractProgress in sentence simplification has been hindered by a lack of labeled parallel simplification data, particularly in languages other than English. We introduce MUSS, a Multilingual Unsupervised Sentence Simplification system that does not require labeled simplification data. MUSS uses a novel approach to sentence simplification that trains strong models using sentence-level paraphrase data instead of proper simplification data. These models leverage unsupervised pretraining and controllable generation mechanisms to flexibly adjust attributes such as length and lexical complexity at inference time. We further present a method to mine such paraphrase data in any language from Common Crawl using semantic sentence embeddings, thus removing the need for labeled data. We evaluate our approach on English, French, and Spanish simplification benchmarks and closely match or outperform the previous best supervised results, despite not using any labeled simplification data. We push the state of the art further by incorporating labeled simplification data. Louis Martin, Angela Fan, Éric Villemonte de la Clergerie, Antoine Bordes, Benoît Sagot |
LREC | 3 |
| 2020 | CamemBERT: a Tasty French Language ModelabstractLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Louis Martin, Benjamin Muller, Pedro Ortiz Suarez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, Benoît Sagot |
ACL | 6 |
| 2020 | Clustering-based Automatic Construction of Legal Entity Knowledge Base from ContractsabstractIn contract analysis and contract automation, a Knowledge Base (KB) of legal entities is fundamental for performing tasks such as contract verification, contract generation and contract analytic. However, such a knowledge base does not always exist nor can be produced in a short time. In this paper, we propose a clustering-based approach to automatically generate a reliable knowledge base of legal entities from given contracts without any supplemental references. The proposed method is robust to different types of errors produced by preprocessing such as Optical Character Recognition (OCR) and Named Entity Recognition (NER), as well as editing errors such as typos. We evaluate our method on a dataset that consists of 800 real contracts with various qualities from 15 clients. Compared to the collected ground-truth data, our method is able to recall 84% of the knowledge. Fuqi Song, Éric Villemonte de la Clergerie |
IEEE BigData | 2 |
| 2020 | Controllable Sentence SimplificationabstractText simplification aims at making a text easier to read and understand by simplifying grammar and structure while keeping the underlying information identical. It is often considered an all-purpose generic task where the same simplification is suitable for all; however multiple audiences can benefit from simplified text in different ways. We adapt a discrete parametrization mechanism that provides explicit control on simplification systems based on Sequence-to-Sequence models. As a result, users can condition the simplifications returned by a model on attributes such as length, amount of paraphrasing, lexical complexity and syntactic complexity. We also show that carefully chosen values of these attributes allow out-of-the-box Sequence-to-Sequence models to outperform their standard counterparts on simplification benchmarks. Our model, which we call ACCESS (as shorthand for AudienCe-CEntric Sentence Simplification), establishes the state of the art at 41.87 SARI on the WikiLarge test set, a +1.42 improvement over the best previously reported score. Louis Martin, Éric Villemonte de la Clergerie, Benoît Sagot, Antoine Bordes |
LREC | 2 |
| 2018 | ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations
Loïc Grobol, Isabelle Tellier, Éric Villemonte de la Clergerie, Marco Dinarelli, Frédéric Landragin |
LREC | 3 |
| 2018 | Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer
Djamé Seddah, Éric Villemonte de la Clergerie, Benoît Sagot, Héctor Martínez Alonso, Marie Candito |
LREC | 2 |
| 2016 | Accurate Deep Syntactic Parsing of Graphs: The Case of French
Corentin Ribeyre, Éric Villemonte de la Clergerie, Djamé Seddah |
LREC | 2 |
| 2015 | Because Syntax Does Matter: Improving Predicate-Argument Structures Parsing with Syntactic FeaturesabstractCorentin Ribeyre, Eric Villemonte de la Clergerie, Djamé Seddah. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Corentin Ribeyre, Éric Villemonte de la Clergerie, Djamé Seddah |
HLT-NAACL | 2 |
| 2014 | Deep Syntax Annotation of the Sequoia French Treebank
Marie Candito, Guy Perrier, Bruno Guillaume, Corentin Ribeyre, Karën Fort, Djamé Seddah, Éric Villemonte de la Clergerie |
LREC | 7 |
| 2014 | Towards an environment for the production and the validation of lexical semantic resources
Mikaël Morardo, Éric Villemonte de la Clergerie |
LREC | 2 |
| 2012 | Boosting the Coverage of a Semantic Lexicon by Automatically Extracted Event Nominalizations
Kata Gábor, Marianna Apidianaki, Benoît Sagot, Éric Villemonte de la Clergerie |
LREC | 4 |
| 2012 | Evaluating and improving syntactic lexica by plugging them within a parser
Elsa Tolone, Benoît Sagot, Éric Villemonte de la Clergerie |
LREC | 3 |
| 2010 | PASSAGE Syntactic Representation: a Minimal Common Ground for Evaluation
Anne Vilnat, Patrick Paroubek, Éric Villemonte de la Clergerie, Gil Francopoulo, Marie-Laure Guénot |
LREC | 3 |
| 2008 | Computer Aided Correction and Extension of a Syntactic Wide-Coverage Lexicon
Lionel Nicolas, Benoît Sagot, Miguel A. Molinero, Jacques Farré, Éric Villemonte de la Clergerie |
COLING | 5 |
| 2008 | PASSAGE: from French Parser Evaluation to Large Sized Treebank
Éric Villemonte de la Clergerie, Olivier Hamon, Djamel Mostefa, Christelle Ayache, Patrick Paroubek, Anne Vilnat |
LREC | 1 |
| 2007 | Large-Scale Knowledge Acquisition from Botanical Texts
François Role, Milagros Fernández Gavilanes, Éric Villemonte de la Clergerie |
NLDB | 3 |
| 2006 | Error Mining in Parsing ResultsabstractWe introduce an error mining technique for automatically detecting errors in resources that are used in parsing systems. We applied this technique on parsing results produced on several million words by two distinct parsing systems, which share the syntactic lexicon and the pre-parsing processing chain. We were thus able to identify missing and erroneous information in these resources. Benoît Sagot, Éric Villemonte de la Clergerie |
ACL | 2 |
| 2006 | The Lefff 2 syntactic lexicon for French: architecture, acquisition, use
Benoît Sagot, Lionel Clément, Éric Villemonte de la Clergerie, Pierre Boullier |
LREC | 3 |
| 2004 | Towards an International Standard on Feature Structure Representation
Kiyong Lee, Lou Burnard, Laurent Romary, Éric Villemonte de la Clergerie, Thierry Declerck, Syd Bauman, Harry Bunt, Lionel Clément, Tomaz Erjavec, Azim Roussanaly, Claude Roux |
LREC | 4 |
| 2002 | Parsing Mildly Context-Sensitive Languages with Thread Automata
Éric Villemonte de la Clergerie |
COLING | 1 |
| 2001 | Guided Parsing of Range Concatenation LanguagesabstractThe theoretical study of the range concatenation grammar [RCG] formalism has revealed many attractive properties which may be used in NLP. In particular, range concatenation languages [RCL] can be parsed in polynomial time and many classical grammatical formalisms can be translated into equivalent RCGs without increasing their worst-case parsing time complexity. For example, after translation into an equivalent RCG, any tree adjoining grammar can be parsed in O(n6) time. In this paper, we study a parsing technique whose purpose is to improve the practical efficiency of RCL parsers. The non-deterministic parsing choices of the main parser for a language L are directed by a guide which uses the shared derivation forest output by a prior RCL parser for a suitable superset of L. The results of a practical evaluation of this method on a wide coverage English grammar are given. François Barthélemy, Pierre Boullier, Philippe Deschamp, Éric Villemonte de la Clergerie |
ACL | 4 |
| 2001 | Natural Language Tabular Parsing
Éric Villemonte de la Clergerie |
ICLP | 1 |
| 2001 | Refining Tabular Parsers for TAGs
Éric Villemonte de la Clergerie |
NAACL | 1 |
| 1999 | Tabular Algorithms for TAG Parsing
Miguel A. Alonso 0001, David Cabrero Souto, Éric Villemonte de la Clergerie, Manuel Vilares Ferro |
EACL | 3 |
| 1998 | Information Flow in Tabular Interpretations for Generalized Push-Down Automata
Éric Villemonte de la Clergerie, François Barthélemy |
Theor. Comput. Sci. | 1 |
| 1994 | LPDA: Another look at Tabulation in Logic Programming
Éric Villemonte de la Clergerie, Bernard Lang |
ICLP | 1 |
| 1993 | Layer Sharing: An Improved Structure-Sharing FrameworkabstractWe present in this paper a structure–sharing framework originally developed for a Dynamic Programming interpreter of Logic programs called DyALog. This mechanism should be of interest for alternative execution models of PROLOG which maintain multiple computation branches and reuse sub-computations in various contexts (computation sharing). This category includes, besides our Dynamic Programming model, the tabular models (OLDT, SLDAL, XWAM), the “magic-set” models, and the independent AND and OR parallelism with solution sharing models. These models raise the problem of storing vast amount of data, motivating us to discard copying mechanisms in favor of structure-sharing mechanisms. Unfortunately, computation sharing requires joining computation branches and possibly renaming some variables, which generally leads to complex structure-sharing mechanisms. The proposed “layer-sharing” framework succeeds however in remaining understandable and easy to implement. Éric Villemonte de la Clergerie |
POPL | 1 |