Éric Villemonte de la Clergerie

dblp:54/5373 · also Éric de la Clergerie · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0001-6428-9219ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 first-authorTheory of computation · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Representation and self-supervised learning · 61% Transfer learning and domain adaptation · 20% Language models and text generation · 15%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 9 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024
Machine learning › Representation and self-supervised learning
pre-training
0.812024
Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.812024
Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024
Machine learning › Transfer learning and domain adaptation › structured regularization
weight tying
0.812024
Headless Language Models: Learning without Predicting with Contrastive Weight Tying · ICLR 2024
Natural language and speech › Language models and text generation
pre-trained language model
0.412020
CamemBERT: a Tasty French Language Model · ACL 2020
Natural language and speech › Language models and text generation
masked language modeling
0.112020
CamemBERT: a Tasty French Language Model · ACL 2020
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.122006
Error Mining in Parsing Results · ACL 2006
Guided Parsing of Range Concatenation Languages · ACL 2001
Programming languages and type systems
logic programming
0.011993
Layer Sharing: An Improved Structure-Sharing Framework · POPL 1993
Compilers and program optimization
structure sharing
0.011993
Layer Sharing: An Improved Structure-Sharing Framework · POPL 1993

Methods — techniques the papers use, named apart from their topics

weight tying · 0.8contrastive learning · 0.8masked language modeling · 0.4shared derivation forest · 0.1guided parsing · 0.1error mining · 0.1
YearPublicationVenuePosition
2024 On the Scaling Laws of Geographical Representation in Language Models
abstract
Language models have long been shown to embed geographical information in their hidden representations. This line of work has recently been revisited by extending this result to Large Language Models (LLMs). In this paper, we propose to fill the gap between well-established and recent literature by observing how geographical knowledge evolves when scaling language models. We show that geographical knowledge is observable even for tiny models, and that it scales consistently as we increase the model size. Notably, we observe that larger language models cannot mitigate the geographical bias that is inherent to the training data.
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
LREC/COLING2
2024 CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
abstract
Clinical data in hospitals are increasingly accessible for research through clinical data warehouses. However these documents are unstructured and it is therefore necessary to extract information from medical reports to conduct clinical studies. Transfer learning with BERT-like models such as CamemBERT has allowed major advances for French, especially for named entity recognition. However, these models are trained for plain language and are less efficient on biomedical data. Addressing this gap, we introduce CamemBERT-bio, a dedicated French biomedical model derived from a new public French biomedical dataset. Through continual pre-training of the original CamemBERT, CamemBERT-bio achieves an improvement of 2.54 points of F1-score on average across various biomedical named entity recognition tasks, reinforcing the potential of continual pre-training as an equally proficient yet less computationally intensive alternative to training from scratch. Additionally, we highlight the importance of using a standard evaluation protocol that provides a clear view of the current state-of-the-art for French biomedical models.
Rian Touchent, Éric Villemonte de la Clergerie
LREC/COLING2
2024 Anisotropy Is Inherent to Self-Attention in Transformers
abstract
The representation degeneration problem is a phenomenon that is widely observed among self-supervised learning methods based on Transformers.In NLP, it takes the form of anisotropy, a singular property of hidden representations which makes them unexpectedly close to each other in terms of angular distance (cosine-similarity).Some recent works tend to show that anisotropy is a consequence of optimizing the cross-entropy loss on long-tailed distributions of tokens.We show in this paper that anisotropy can also be observed empirically in language models with specific objectives that should not suffer directly from the same consequences.We also show that the anisotropy problem extends to Transformers trained on other modalities.Our observations suggest that anisotropy is actually inherent to Transformers-based models.
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
EACL (1)2
2024 Translate your Own: a Post-Editing Experiment in the NLP domain
abstract
The improvements in neural machine translation make translation and post-editing pipelines ever more effective for a wider range of applications. In this paper, we evaluate the effectiveness of such a pipeline for the translation of scientific documents (limited here to article abstracts). Using a dedicated interface, we collect, then analyse the post-edits of approximately 350 abstracts (English→French) in the Natural Language Processing domain for two groups of post-editors: domain experts (academics encouraged to post-edit their own articles) on the one hand and trained translators on the other. Our results confirm that such pipelines can be effective, at least for high-resource language pairs. They also highlight the difference in the post-editing strategy of the two subgroups. Finally, they suggest that working on term translation is the most pressing issue to improve fully automatic translations, but that in a post-editing setup, other error types can be equally annoying for post-editors.
Rachel Bawden, Ziqian Peng, Maud Bénard, Éric Villemonte de la Clergerie, Raphaël Esamotunu, Mathilde Huguin, Natalie Kübler, Alexandra Mestivier, Mona Michelot, Laurent Romary, Lichao Zhu, François Yvon
EAMT (1)4
2024 Headless Language Models: Learning without Predicting with Contrastive Weight Tying
abstract
Self-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative method that shifts away from probability prediction and instead focuses on reconstructing input embeddings in a contrastive fashion via Constrastive Weight Tying (CWT). We apply this approach to pretrain Headless Language Models in both monolingual and multilingual contexts. Our method offers practical advantages, substantially reducing training computational requirements by up to 20 times, while simultaneously enhancing downstream performance and data efficiency. We observe a significant +1.6 GLUE score increase and a notable +2.7 LAMBADA accuracy improvement compared to classical LMs within similar compute budgets.
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
ICLR2
2024 PatentEval: Understanding Errors in Patent Generation
abstract
You Zuo, Kim Gerdes, Éric Clergerie, Benoît Sagot. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
You Zuo, Kim Gerdes, Éric Villemonte de la Clergerie, Benoît Sagot
NAACL-HLT3
2022 MUSS: Multilingual Unsupervised Sentence Simplification by Mining Paraphrases
abstract
Progress in sentence simplification has been hindered by a lack of labeled parallel simplification data, particularly in languages other than English. We introduce MUSS, a Multilingual Unsupervised Sentence Simplification system that does not require labeled simplification data. MUSS uses a novel approach to sentence simplification that trains strong models using sentence-level paraphrase data instead of proper simplification data. These models leverage unsupervised pretraining and controllable generation mechanisms to flexibly adjust attributes such as length and lexical complexity at inference time. We further present a method to mine such paraphrase data in any language from Common Crawl using semantic sentence embeddings, thus removing the need for labeled data. We evaluate our approach on English, French, and Spanish simplification benchmarks and closely match or outperform the previous best supervised results, despite not using any labeled simplification data. We push the state of the art further by incorporating labeled simplification data.
Louis Martin, Angela Fan, Éric Villemonte de la Clergerie, Antoine Bordes, Benoît Sagot
LREC3
2020 CamemBERT: a Tasty French Language Model
abstract
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Louis Martin, Benjamin Muller, Pedro Ortiz Suarez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, Benoît Sagot
ACL6
2020 Clustering-based Automatic Construction of Legal Entity Knowledge Base from Contracts
abstract
In contract analysis and contract automation, a Knowledge Base (KB) of legal entities is fundamental for performing tasks such as contract verification, contract generation and contract analytic. However, such a knowledge base does not always exist nor can be produced in a short time. In this paper, we propose a clustering-based approach to automatically generate a reliable knowledge base of legal entities from given contracts without any supplemental references. The proposed method is robust to different types of errors produced by preprocessing such as Optical Character Recognition (OCR) and Named Entity Recognition (NER), as well as editing errors such as typos. We evaluate our method on a dataset that consists of 800 real contracts with various qualities from 15 clients. Compared to the collected ground-truth data, our method is able to recall 84% of the knowledge.
Fuqi Song, Éric Villemonte de la Clergerie
IEEE BigData2
2020 Controllable Sentence Simplification
abstract
Text simplification aims at making a text easier to read and understand by simplifying grammar and structure while keeping the underlying information identical. It is often considered an all-purpose generic task where the same simplification is suitable for all; however multiple audiences can benefit from simplified text in different ways. We adapt a discrete parametrization mechanism that provides explicit control on simplification systems based on Sequence-to-Sequence models. As a result, users can condition the simplifications returned by a model on attributes such as length, amount of paraphrasing, lexical complexity and syntactic complexity. We also show that carefully chosen values of these attributes allow out-of-the-box Sequence-to-Sequence models to outperform their standard counterparts on simplification benchmarks. Our model, which we call ACCESS (as shorthand for AudienCe-CEntric Sentence Simplification), establishes the state of the art at 41.87 SARI on the WikiLarge test set, a +1.42 improvement over the best previously reported score.
Louis Martin, Éric Villemonte de la Clergerie, Benoît Sagot, Antoine Bordes
LREC2
2018 ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations
Loïc Grobol, Isabelle Tellier, Éric Villemonte de la Clergerie, Marco Dinarelli, Frédéric Landragin
LREC3
2018 Cheating a Parser to Death: Data-driven Cross-Treebank Annotation Transfer
Djamé Seddah, Éric Villemonte de la Clergerie, Benoît Sagot, Héctor Martínez Alonso, Marie Candito
LREC2
2016 Accurate Deep Syntactic Parsing of Graphs: The Case of French
Corentin Ribeyre, Éric Villemonte de la Clergerie, Djamé Seddah
LREC2
2015 Because Syntax Does Matter: Improving Predicate-Argument Structures Parsing with Syntactic Features
abstract
Corentin Ribeyre, Eric Villemonte de la Clergerie, Djamé Seddah. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Corentin Ribeyre, Éric Villemonte de la Clergerie, Djamé Seddah
HLT-NAACL2
2014 Deep Syntax Annotation of the Sequoia French Treebank
Marie Candito, Guy Perrier, Bruno Guillaume, Corentin Ribeyre, Karën Fort, Djamé Seddah, Éric Villemonte de la Clergerie
LREC7
2014 Towards an environment for the production and the validation of lexical semantic resources
Mikaël Morardo, Éric Villemonte de la Clergerie
LREC2
2012 Boosting the Coverage of a Semantic Lexicon by Automatically Extracted Event Nominalizations
Kata Gábor, Marianna Apidianaki, Benoît Sagot, Éric Villemonte de la Clergerie
LREC4
2012 Evaluating and improving syntactic lexica by plugging them within a parser
Elsa Tolone, Benoît Sagot, Éric Villemonte de la Clergerie
LREC3
2010 PASSAGE Syntactic Representation: a Minimal Common Ground for Evaluation
Anne Vilnat, Patrick Paroubek, Éric Villemonte de la Clergerie, Gil Francopoulo, Marie-Laure Guénot
LREC3
2008 Computer Aided Correction and Extension of a Syntactic Wide-Coverage Lexicon
Lionel Nicolas, Benoît Sagot, Miguel A. Molinero, Jacques Farré, Éric Villemonte de la Clergerie
COLING5
2008 PASSAGE: from French Parser Evaluation to Large Sized Treebank
Éric Villemonte de la Clergerie, Olivier Hamon, Djamel Mostefa, Christelle Ayache, Patrick Paroubek, Anne Vilnat
LREC1
2007 Large-Scale Knowledge Acquisition from Botanical Texts
François Role, Milagros Fernández Gavilanes, Éric Villemonte de la Clergerie
NLDB3
2006 Error Mining in Parsing Results
abstract
We introduce an error mining technique for automatically detecting errors in resources that are used in parsing systems. We applied this technique on parsing results produced on several million words by two distinct parsing systems, which share the syntactic lexicon and the pre-parsing processing chain. We were thus able to identify missing and erroneous information in these resources.
Benoît Sagot, Éric Villemonte de la Clergerie
ACL2
2006 The Lefff 2 syntactic lexicon for French: architecture, acquisition, use
Benoît Sagot, Lionel Clément, Éric Villemonte de la Clergerie, Pierre Boullier
LREC3
2004 Towards an International Standard on Feature Structure Representation
Kiyong Lee, Lou Burnard, Laurent Romary, Éric Villemonte de la Clergerie, Thierry Declerck, Syd Bauman, Harry Bunt, Lionel Clément, Tomaz Erjavec, Azim Roussanaly, Claude Roux
LREC4
2002 Parsing Mildly Context-Sensitive Languages with Thread Automata
Éric Villemonte de la Clergerie
COLING1
2001 Guided Parsing of Range Concatenation Languages
abstract
The theoretical study of the range concatenation grammar [RCG] formalism has revealed many attractive properties which may be used in NLP. In particular, range concatenation languages [RCL] can be parsed in polynomial time and many classical grammatical formalisms can be translated into equivalent RCGs without increasing their worst-case parsing time complexity. For example, after translation into an equivalent RCG, any tree adjoining grammar can be parsed in O(n6) time. In this paper, we study a parsing technique whose purpose is to improve the practical efficiency of RCL parsers. The non-deterministic parsing choices of the main parser for a language L are directed by a guide which uses the shared derivation forest output by a prior RCL parser for a suitable superset of L. The results of a practical evaluation of this method on a wide coverage English grammar are given.
François Barthélemy, Pierre Boullier, Philippe Deschamp, Éric Villemonte de la Clergerie
ACL4
2001 Natural Language Tabular Parsing
Éric Villemonte de la Clergerie
ICLP1
2001 Refining Tabular Parsers for TAGs
Éric Villemonte de la Clergerie
NAACL1
1999 Tabular Algorithms for TAG Parsing
Miguel A. Alonso 0001, David Cabrero Souto, Éric Villemonte de la Clergerie, Manuel Vilares Ferro
EACL3
1998 Information Flow in Tabular Interpretations for Generalized Push-Down Automata
Éric Villemonte de la Clergerie, François Barthélemy
Theor. Comput. Sci.1
1994 LPDA: Another look at Tabulation in Logic Programming
Éric Villemonte de la Clergerie, Bernard Lang
ICLP1
1993 Layer Sharing: An Improved Structure-Sharing Framework
abstract
We present in this paper a structure–sharing framework originally developed for a Dynamic Programming interpreter of Logic programs called DyALog. This mechanism should be of interest for alternative execution models of PROLOG which maintain multiple computation branches and reuse sub-computations in various contexts (computation sharing). This category includes, besides our Dynamic Programming model, the tabular models (OLDT, SLDAL, XWAM), the “magic-set” models, and the independent AND and OR parallelism with solution sharing models. These models raise the problem of storing vast amount of data, motivating us to discard copying mechanisms in favor of structure-sharing mechanisms. Unfortunately, computation sharing requires joining computation branches and possibly renaming some variables, which generally leads to complex structure-sharing mechanisms. The proposed “layer-sharing” framework succeeds however in remaining understandable and easy to implement.
Éric Villemonte de la Clergerie
POPL1