EDBT 2026 Demo / reviewers in the wild / expert
Yves Lepage
dblp:51/130
· DBLP profile ↗
45ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0002-3059-4271ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 11 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalizing Analogical Inference from Boolean to Continuous DomainsabstractAnalogical reasoning is a powerful inductive mechanism, widely used in human cognition and increasingly applied in artificial intelligence. Formal frameworks for analogical inference have been developed for Boolean domains, where inference is provably sound for affine functions and approximately correct for functions close to affine. These results have informed the design of analogy-based classifiers. However, they do not extend to regression tasks or continuous domains. In this paper, we revisit analogical inference from a foundational perspective. We first present a counterexample showing that existing generalization bounds fail even in the Boolean setting. We then introduce a unified framework for analogical reasoning in real-valued domains based on parameterized analogies defined via generalized means. This model subsumes both Boolean classification and regression, and supports analogical inference over continuous functions. We characterize the class of analogy-preserving functions in this setting and derive both worst-case and average-case error bounds under smoothness assumptions. Our results offer a general theory of analogical inference across discrete and continuous domains. Francisco Cunha, Yves Lepage, Miguel Couceiro, Zied Bouraoui |
AAAI | 2 |
| 2025 | ALF: A Fine-Grained French Analogical Dataset for Evaluating Lexical Knowledge of Large Language ModelsabstractThe undeniable revolution brought forth by Large Language Models (LLMs) stems from the amazing fluency of the texts they generate, mastering language with seemingly human-like finesse. This fluency raises a key scientific question: How much lexical knowledge do LLMs actually capture in order to produce such fluent language? To address this, we present ALF, a freely available, analogical dataset endowed with rich lexicographic information grounded in Meaning-Text Theory for the French language. It comprises 2600 fine-grained lexical analogies with which we evaluate the lexical ability of five off-the-shelf LLMs, namely ChatGPT-4o mini, Llama3.0-8B, Llama3.1-8B, Qwen2.5-14B, and Mistral7B. Their performance spans from 45% for Mistral, through about 55% for the ChatGPT and Llama models, and up to nearly 60% for Qwen2.5-14B, thus qualifying ALF as a challenging dataset. Experimenting with larger models (OpenAI o1, Llama3.0/3.1-70B, and Qwen2.5-32B) yields rather limited returns considering the drastic increase in computational cost. We further identify certain types of analogies and prompting methods that reveal performance disparities. Alexander Petrov, Antoine Venant, François Lareau, Yves Lepage, Philippe Langlais |
ECAI | 4 |
| 2025 | AnaScore: Understanding Semantic Parallelism in Proportional AnalogiesabstractLiyan Wang, Haotong Wang, Yves Lepage. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haotong Wang, Yves Lepage |
NAACL (Long Papers) | 3 |
| 2025 | Eliciting analogical reasoning from language models in retrieval-augmented translation under low-resource scenariosabstractRetrieval-Augmented Neural Machine Translation (RANMT), which augments translation models with relevant examples fetched by a similarity retriever, is proficient in well-resourced translations. However, its inherent advantages are not fully realized in low-resource contexts, where the sparsity of data can often result in less relevant or less useful information for translation. Recent literature indicates that nearest-neighbor examples from small training data have the unfortunate effect of impairing RANMT performance. Our examination of 16 low-resource tasks reveals a sharp deterioration in performance of a multilingual language model when conditioned on retrieved examples compared to direct translation. To address this problem, we explore a framework based on analogical reasoning, aiming to enhance the capacity of language models to infer translations from parallel examples in limited data settings. This framework mimics a cognitive process of human translation by structuring examples in analogy patterns. We propose a multi-objective learning strategy that augments vanilla training for conditional translation to learn latent knowledge from examples. We also investigate different retrieval methods for selecting translation examples based on lexical similarity, semantic relatedness, or a combination of both. The results show that our approach is effective in optimizing RANMT in low-resource settings, delivering notable improvements across all retrieval settings. In particular, augmented training akin to reasoning with analogies in two directions, contributes significantly to deriving benefits from examples, even when their relevance is limited. Moreover, our approach demonstrates superior performance in low-resource translation tasks compared to prompting large language models in few-shot contexts. It also proves to be competitive with models that have been extensively trained using substantial amounts of supervised data. Bartholomäus Wloka, Yves Lepage |
Neurocomputing | 3 |
| 2025 | Mixup Helps Translation, But Do the Coefficients and the Selection Strategy Influence Translation Quality?abstractMixup, an interpolation-based method that implicitly generates synthetic examples for training, has shown effectiveness in tasks such as image and text classification. Standard mixup randomly interpolates two samples of images and their labels. In this article, we apply mixup to low-resource machine translation tasks by interpolating in the hidden space. We investigate the impact of different mixing coefficients on this technique. We also explore whether semantically related or unrelated samples provide more benefits for interpolation compared to random selection. To investigate this, we extend the standard mixup approach by selecting samples based on distance and experimenting with different sampling settings. Our experiments are conducted across several low-resource language pairs, including Lower Sorbian and Upper Sorbian, Lower Sorbian and German, and Upper Sorbian and German. Through systematic experiments on multiple language pairs, we evaluate the effectiveness of mixup data augmentation in improving low-resource machine translation performance. Our findings indicate that the standard mixup technique enhances the quality of machine translation, resulting in an average increase of 1.9 BLEU points over the baseline Transformer model. The choice of mixing coefficients has minimal impact on translation quality, which suggests that fine-tuning these coefficients is not essential to benefit from mixup. In addition, the standard mixup performs robustly, as selecting either the most similar or most dissimilar samples for mixing does not provide a significant improvement over it. Yves Lepage |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | Continued Pre-training on Sentence Analogies for Translation with Small DataabstractThis paper introduces Continued Pre-training on Analogies (CPoA) to incorporate pre-trained language models with analogical abilities, aiming at improving performance in low-resource translations without data augmentation. We continue training the models on sentence analogies retrieved from a translation corpus. Considering the sparsity of analogy in corpora, especially in low-resource scenarios, we propose exploring approximate analogies between sentences. We attempt to find sentence analogies that might not conform to formal criteria for entire sentences but partial pieces. When training the models, we introduce a weighting scalar pertaining to the quality of analogies to adjust the influence: emphasizing closer analogies while diminishing the impact of far ones. We evaluate our approach on a low-resource translation task: German-Upper Sorbian. The results show that CPoA using 10 times fewer instances can effectively attain gains of +1.4 and +1.3 BLEU points over the original model in two translation directions. This improvement is more pronounced when there are fewer parallel examples. Haotong Wang, Yves Lepage |
LREC/COLING | 3 |
| 2024 | Leveraging Knowledge from Translation Memory for Globally and Locally Guiding Neural Machine Translation
Ruibo Hou, Hengjie Liu, Yves Lepage |
PACLIC | 3 |
| 2024 | Organising lexica into analogical grids: a study of a holistic approach for morphological generation under various sizes of data in various languagesabstractMorphological generation is a task where given a lemma and a morphosyntactic description of the target form, we are asked to generate the target form. Knowing that the syntactic and semantic relations to other forms are reflected by the word form itself, we show how to exploit these relations between word forms, holistically, that is, as a whole, to derive the target form without even breaking them into morphemes. Experimental results show that by organising the lexica into analogical grids we are able to improve the accuracy of morphological generation by up to 8% in low data scenarios. Our holistic approach always performs better than a morpheme-based baseline. We also enquire possible improvements by using data augmentation for neural approaches, especially in low data scenarios. However, our system seems not to gain any advantage from having more data after some point in time. Rashel Fam, Yves Lepage |
J. Exp. Theor. Artif. Intell. | 2 |
| 2023 | A Dual Reinforcement Method for Data Augmentation using Middle Sentences for Machine TranslationabstractThis paper presents an approach to enhance the quality of machine translation by leveraging middle sentences as pivot points and employing dual reinforcement learning. Conventional methods for generating parallel sentence pairs for machine translation rely on parallel corpora, which may be scarce, resulting in limitations in translation quality. In contrast, our proposed method entails training two machine translation models in opposite directions, utilizing the middle sentence as a bridge for a virtuous feedback loop between the two models. This feedback loop resembles reinforcement learning, facilitating the models to make informed decisions based on mutual feedback. Experimental results substantiate that our proposed method significantly improves machine translation quality. Wenyi Tang, Yves Lepage |
MTSummit (1) | 2 |
| 2022 | WAPITI - Web-based Assignment Preparation and Instruction Tool for Interpreters
Bartholomäus Wloka, Yves Lepage, Werner Winiwarter |
iiWAS | 2 |
| 2022 | Do we Name the Languages we Study? The #BenderRule in LREC and ACL articlesabstractThis article studies the application of the #BenderRule in Natural Language Processing (NLP) articles according to two dimensions. Firstly, in a contrastive manner, by considering two major international conferences, LREC and ACL, and secondly, in a diachronic manner, by inspecting nearly 14,000 articles over a period of time ranging from 2000 to 2020 for LREC and from 1979 to 2020 for ACL. For this purpose, we created a corpus from LREC and ACL articles from the above-mentioned periods, from which we manually annotated nearly 1,000. We then developed two classifiers to automatically annotate the rest of the corpus. Our results show that LREC articles tend to respect the #BenderRule (80 to 90% of them respect it), whereas 30 to 40% of ACL articles do not. Interestingly, over the considered periods, the results appear to be stable for the two conferences, even though a rebound in ACL 2020 could be a sign of the influence of the blog post about the #BenderRule. Fanny Ducel, Karën Fort, Gaël Lejeune, Yves Lepage |
LREC | 4 |
| 2022 | A Study of Re-generating Sentences Given Similar Sentences that Cover Them on the Level of Form and Meaning
Hsuan-Wei Lo, Rashel Fam, Yves Lepage |
PACLIC | 4 |
| 2022 | Can the Translation Memory Principle Benefit Neural Machine Translation? A Series of Extensive Experiments with Input Sentence Annotation
Yaling Wang, Yves Lepage |
PACLIC | 2 |
| 2021 | Covering a sentence in form and meaning with fewer retrieved sentences
Yves Lepage |
PACLIC | 2 |
| 2020 | The French Correction: When Retrieval Is Harder to Specify than Adaptation
Yves Lepage, Jean Lieber, Isabelle Mornard, Emmanuel Nauer, Julien Romary, Reynault Sies |
ICCBR | 1 |
| 2019 | An Approach to Case-Based Reasoning Based on Local Enrichment of the Case Base
Yves Lepage, Jean Lieber |
ICCBR | 1 |
| 2019 | Neural Morphological Segmentation Model for MongolianabstractMorphological segmentation is useful for processing Mongolian. In this paper, we manually build a morphological segmentation data set for Mongolian. We then present a character-based encoder-decoder model with attention mechanism to perform the morphological segmentation task. We further investigate the influence of analogy features extracted from scratch and improve the performance of our model using multi languages setting. Experimental results show that our encoder-decoder model with attention mechanism provides a strong baseline for Mongolian morphological segmentation. The analogy features provide useful information to the model and improve the performance of the system. The use of multi languages data set shows the capability of our model to acquire knowledge through different languages and delivers the best result. Weihua Wang 0006, Rashel Fam, Feilong Bao, Yves Lepage, Guanglai Gao |
IJCNN | 4 |
| 2018 | Production of Large Analogical Clusters from Smaller Example Seed Clusters Using Word Embeddings
Yuzhong Hong, Yves Lepage |
ICCBR | 2 |
| 2018 | Case-Based Translation: First Steps from a Knowledge-Light Approach Based on Analogy to a Knowledge-Intensive One
Yves Lepage, Jean Lieber |
ICCBR | 1 |
| 2018 | Tools for The Production of Analogical Grids and a Resource of N-gram Analogical Grids in 11 Languages
Rashel Fam, Yves Lepage |
LREC | 2 |
| 2018 | Korean L2 Vocabulary Prediction: Can a Large Annotated Corpus be Used to Train Better Models for Predicting Unknown Words?
Kevin P. Yancey, Yves Lepage |
LREC | 2 |
| 2018 | Context Encoder for Analogies on Strings
Tianjing Zhao, Yves Lepage |
PACLIC | 2 |
| 2017 | Unsupervised Bilingual Segmentation using MDL for Machine Translation
Bin Shan, Yves Lepage |
PACLIC | 3 |
| 2017 | BTG-based Machine Translation with Simple Reordering Model using Structured Perceptron
Yves Lepage |
PACLIC | 2 |
| 2016 | Yet Another Symmetrical and Real-time Word Alignment Method: Hierarchical Sub-sentential Alignment using F-measure
Yves Lepage |
PACLIC | 2 |
| 2016 | HSSA tree structures for BTG-based preordering in machine translation
Yves Lepage |
PACLIC | 3 |
| 2015 | Translation of Unseen Bigrams by Analogy Using an SVM Classifier
Lu Lyu, Yves Lepage |
PACLIC | 3 |
| 2015 | Chinese Word Segmentation based on analogy and majority voting
Zongrong Zheng, Yves Lepage |
PACLIC | 3 |
| 2014 | Production of Phrase Tables in 11 European Languages using an Improved Sub-sentential Aligner
Juan Luo, Yves Lepage |
LREC | 2 |
| 2013 | Exploiting Parallel Corpus for Handling Out-of-Vocabulary Words
Juan Luo, John Tinsley, Yves Lepage |
PACLIC | 3 |
| 2013 | Generalizing sampling-based multilingual alignment
Adrien Lardilleux, François Yvon, Yves Lepage |
Mach. Transl. | 3 |
| 2012 | Hierarchical Sub-sentential Alignment with Anymalign
Adrien Lardilleux, François Yvon, Yves Lepage |
EAMT | 3 |
| 2012 | Can Word Segmentation be Considered Harmful for Statistical Machine Translation Tasks between Japanese and Chinese?
Yves Lepage |
PACLIC | 2 |
| 2011 | Marker-based Chunking for Analogy-based Translation of Chunks
Kota Takeya, Yves Lepage |
MTSummit | 2 |
| 2011 | Improving Sampling-based Alignment by Investigating the Distribution of N-grams in Phrase Translation Tables
Juan Luo, Adrien Lardilleux, Yves Lepage |
PACLIC | 3 |
| 2011 | Fully-Automatic Marker-based Chunking in 11 European Languages and Counts of the Number of Analogies between Chunks
Kota Takeya, Yves Lepage |
PACLIC | 2 |
| 2010 | Bilingual Lexicon Induction: Effortless Evaluation of Word Alignment Tools and Production of Resources for Improbable Language Pairs
Adrien Lardilleux, Julien Gosme, Yves Lepage |
LREC | 3 |
| 2005 | Purest ever example-based machine translation: Detailed presentation and assessment
Yves Lepage, Etienne Denoual |
Mach. Transl. | 1 |
| 2004 | Lower and higher estimates of the number of "true analogies" between sentences contained in a large multilingual corpus
Yves Lepage |
COLING | 1 |
| 2004 | Using Paradigm Tables to Generate New Utterances Similar to those Existing in Linguistic Resources
Yves Lepage, Guilhem Peralta |
LREC | 1 |
| 2000 | Languages Of Analogical Strings
Yves Lepage |
COLING | 1 |
| 2000 | A tool to build a treebank for conversational Chinese
Yves Lepage, Nicolas Auclerc, Satoshi Shirai |
INTERSPEECH | 1 |
| 1996 | Saussurian analogy: a theoretical account and its application
Yves Lepage, Ando Shin-ichi |
COLING | 1 |
| 1994 | Non-directionality and Self-Assessment in an Example-based System Using Genetic Algorithms
Yves Lepage |
COLING | 1 |
| 1986 | A Language for Transcriptions
Yves Lepage |
COLING | 1 |