EDBT 2026 Demo / reviewers in the wild / expert
Philippe Langlais
dblp:66/1102
· DBLP profile ↗
74ranked-venue papers
17as first author
19since 2021 · last 2025
0000-0002-7319-1595ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 17 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Part-Of-Speech Sensitivity of Routers in Mixture of Experts ModelsabstractThis study investigates the behavior of model-integrated routers in Mixture of Experts (MoE) models, focusing on how tokens are routed based on their linguistic features, specifically Part-of-Speech (POS) tags. The goal is to explore across different MoE architectures whether experts specialize in processing tokens with similar linguistic traits. By analyzing token trajectories across experts and layers, we aim to uncover how MoE models handle linguistic information. Findings from six popular MoE models reveal expert specialization for specific POS categories, with routing paths showing high predictive accuracy for POS, highlighting the value of routing paths in characterizing tokens. Elie Antoine, Frédéric Béchet, Philippe Langlais |
COLING | 3 |
| 2025 | On Evaluation Protocols for Data Augmentation in a Limited Data ScenarioabstractTextual data augmentation (DA) is a prolific field of study where novel techniques to create artificial data are regularly proposed, and that has demonstrated great efficiency on small data settings, at least for text classification tasks. In this paper, we challenge those results, showing that classical data augmentation (which modify sentences) is simply a way of performing better fine-tuning, and that spending more time doing so before applying data augmentation negates its effect. This is a significant contribution as it answers several questions that were left open in recent years, namely : which DA technique performs best (all of them as long as they generate data close enough to the training set, as to not impair training) and why did DA show positive results (facilitates training of network). We further show that zero- and few-shot DA via conversational agents such as ChatGPT or LLama2 can increase performances, confirming that this form of data augmentation is preferable to classical methods. Frédéric Piedboeuf, Philippe Langlais |
COLING | 2 |
| 2025 | An Interpretable Quantum-Inspired Model for Multi-Task Natural Language UnderstandingabstractMulti-task learning has demonstrated remarkable success across a broad spectrum of natural language processing tasks, particularly with neural network-based methods. Despite these advances, a fundamental gap remains in explaining the relationship between task-relatedness and model effectiveness. To address this issue, we propose a novel approach for implicitly modeling task-relatedness by leveraging a quantum physical mathematical framework. In this paper, we introduce a complex-valued neural network designed to encapsulate and analyze task-relatedness. Within this framework, sentences originating from diverse tasks are encoded as mixed quantum systems, represented on a meticulously defined Semantic Hilbert Space. This allows the network to interpret inter-task relationships through the explicit physical semantics of well-constrained components grounded in quantum probability theory. By adhering to these rigorous principles, our model not only establishes a robust method for quantifying task-relatedness but also fosters a deeper, self-explanatory understanding of the underlying processes. To validate the efficacy of our approach, we conducted extensive experiments across five benchmark text classification tasks. The results demonstrate both the superior performance and the interpretability of the proposed model, highlighting its potential as a self-explanatory system for multi-task learning in NLP. Peng Lu 0006, Jerry Huang, Xinyu Wang 0061, Philippe Langlais |
ECAI | 4 |
| 2025 | ALF: A Fine-Grained French Analogical Dataset for Evaluating Lexical Knowledge of Large Language ModelsabstractThe undeniable revolution brought forth by Large Language Models (LLMs) stems from the amazing fluency of the texts they generate, mastering language with seemingly human-like finesse. This fluency raises a key scientific question: How much lexical knowledge do LLMs actually capture in order to produce such fluent language? To address this, we present ALF, a freely available, analogical dataset endowed with rich lexicographic information grounded in Meaning-Text Theory for the French language. It comprises 2600 fine-grained lexical analogies with which we evaluate the lexical ability of five off-the-shelf LLMs, namely ChatGPT-4o mini, Llama3.0-8B, Llama3.1-8B, Qwen2.5-14B, and Mistral7B. Their performance spans from 45% for Mistral, through about 55% for the ChatGPT and Llama models, and up to nearly 60% for Qwen2.5-14B, thus qualifying ALF as a challenging dataset. Experimenting with larger models (OpenAI o1, Llama3.0/3.1-70B, and Qwen2.5-32B) yields rather limited returns considering the drastic increase in computational cost. We further identify certain types of analogies and prompting methods that reveal performance disparities. Alexander Petrov, Antoine Venant, François Lareau, Yves Lepage, Philippe Langlais |
ECAI | 5 |
| 2025 | ReGLA: Refining Gated Linear AttentionabstractPeng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Peng Lu 0006, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais |
NAACL (Long Papers) | 5 |
| 2025 | Mamba Modulation: On the Length Generalization of Mamba ModelsabstractThe quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged as a leading architecture, achieving state-of-the-art results across a range of language modeling tasks. However, Mamba’s performance significantly deteriorates when applied to contexts longer than those seen during pre-training, revealing a sharp sensitivity to context length extension. Through detailed analysis, we attribute this limitation to the out-of-distribution behavior of its state-space dynamics, particularly within the parameterization of the state transition matrix $A$. Unlike recent works which attribute this sensitivity to the vanished accumulation of discretization time steps, $\exp(-\sum_{t=1}^N{\Delta}_t)$, we establish a connection between state convergence behavior as the input length approaches infinity and the spectrum of the transition matrix $A$, offering a well-founded explanation of its role in length extension. Next, to overcome this challenge, we propose an approach that applies spectrum scaling to pre-trained Mamba models to enable robust long-context generalization by selectively modulating the spectrum of $A$ matrices in each layer. We show that this can significantly improve performance in settings where simply modulating ${\Delta}_t$ fails, validating our insights and providing avenues for better length generalization of state-space models with structured transition matrices. Peng Lu 0006, Jerry Huang, Qiuhao Zeng, Xinyu Wang 0061, Boxing Chen, Philippe Langlais, Yufei Cui |
NeurIPS | 6 |
| 2024 | EUROPA: A Legal Multilingual Keyphrase Generation DatasetabstractOlivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso-Hermelo, Philippe Langlais. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Olivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso-Hermelo, Philippe Langlais |
ACL (1) | 5 |
| 2024 | A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasksabstractWe introduce an evaluation methodology for reading comprehension tasks based on the intuition that certain examples, by the virtue of their linguistic complexity, consistently yield lower scores regardless of model size or architecture.We capitalize on semantic frame annotation for characterizing this complexity, and study seven complexity factors that may account for model's difficulty.We first deploy this methodology on a carefully annotated French reading comprehension benchmark showing that two of those complexity factors are indeed good predictors of models' failure, while others are less so.We further deploy our methodology on a well studied English benchmark by using Chat-GPT as a proxy for semantic annotation.Our study reveals that fine-grained linguisticallymotivated automatic evaluation of a reading comprehension task is not only possible, but helps understand models' abilities to handle specific linguistic characteristics of input examples.It also shows that current state-of-the-art models fail with some for those characteristics which suggests that adequately handling them requires more than merely increasing model size. Elie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe Langlais |
EMNLP | 4 |
| 2023 | On the utility of enhancing BERT syntactic bias with Token Reordering PretrainingabstractYassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Phillippe Langlais, Prasanna Parthasarathi. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. Yassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Philippe Langlais, Prasanna Parthasarathi |
CoNLL | 5 |
| 2023 | An analysis of entity normalization evaluation biases in specialized domainsabstractBACKGROUND: Entity normalization is an important information extraction task which has recently gained attention, particularly in the clinical/biomedical and life science domains. On several datasets, state-of-the-art methods perform rather well on popular benchmarks. Yet, we argue that the task is far from resolved. RESULTS: We have selected two gold standard corpora and two state-of-the-art methods to highlight some evaluation biases. We present non-exhaustive initial findings on the existence of evaluation problems of the entity normalization task. CONCLUSIONS: Our analysis suggests better evaluation practices to support the methodological research in this field. Arnaud Ferré, Philippe Langlais |
BMC Bioinform. | 2 |
| 2022 | CILDA: Contrastive Data Augmentation Using Intermediate Layer Knowledge DistillationabstractKnowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by leveraging Contrastive Learning, Intermediate Layer Distillation, Data Augmentation, and Adversarial Training. In this work, we propose a learning-based data augmentation technique tailored for knowledge distillation, called CILDA. To the best of our knowledge, this is the first time that intermediate layer representations of the main task are used in improving the quality of augmented samples. More precisely, we introduce an augmentation technique for KD based on intermediate layer matching using contrastive loss to improve masked adversarial data augmentation. CILDA outperforms existing state-of-the-art KD approaches on the GLUE benchmark, as well as in an out-of-domain evaluation. Md. Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar, Khalil Bibi, Philippe Langlais, Pascal Poupart |
COLING | 5 |
| 2022 | Effective Data Augmentation for Sentence Classification Using One VAE per ClassabstractIn recent years, data augmentation has become an important field of machine learning. While images can use simple techniques such as cropping or rotating, textual data augmentation needs more complex manipulations to ensure that the generated examples are useful. Variational auto-encoders (VAE) and its conditional variant the Conditional-VAE (CVAE) are often used to generate new textual data, both relying on a good enough training of the generator so that it doesn’t create examples of the wrong class. In this paper, we explore a simpler way to use VAE for data augmentation: the training of one VAE per class. We show on several dataset sizes, as well as on four different binary classification tasks, that it systematically outperforms other generative data augmentation techniques. Frédéric Piedboeuf, Philippe Langlais |
COLING | 2 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 14 |
| 2022 | Why Do Tenants Sue Their Landlords? Answers from a Topic ModelabstractTopic modeling is widely used in various domains for extracting latent topics underlying large corpora, including judicial texts. In the latter, topics tend to be made by and for domain experts, but remain unintelligible for laymen. In the framework of housing law court decisions in French which mixes abstract legal terminology with real-life situations described in common language, similarly to [1], we aim at identifying different situations that can cause a tenant to prosecute their landlord in court with the application of topic models. Upon quantitative evaluation, LDA and BERTopic deliver the best results, but a closer manual analysis reveals that the second embedding-based approach is much better at producing and even uncovering topics that describe a tenant’s real-life issues and situations. Olivier Salaün, Fabrizio Gotti, Philippe Langlais, Karim Benyekhlef |
JURIX | 3 |
| 2022 | Conditional Abstractive Summarization of Court Decisions for Laymen and Insights from Human EvaluationabstractLegal text summarization is generally formalized as an extractive text summarization task applied to court decisions from which the most relevant sentences are identified and returned as a gist meant to be read by legal experts. However, such summaries are not suitable for laymen seeking intelligible legal information. In the scope of the JusticeBot, a question-answering system in French that provides information about housing law, we intend to generate summaries of court decisions that are, on the one hand, conditioned by a question-answer-decision triplet, and on the other hand, intelligible for ordinary citizens not familiar with legal documents. So far, our best model, a further pre-trained BARThez, achieves an average ROUGE-1 score of 37.7 and a deepened manual evaluation of summaries reveals that there is still room for improvement. Olivier Salaün, Aurore Clément Troussel, Sylvain Longhais, Hannes Westermann, Philippe Langlais, Karim Benyekhlef |
JURIX | 5 |
| 2022 | A Methodology for Building a Diachronic Dataset of Semantic Shifts and its Application to QC-FR-Diac-V1.0, a Free Reference for FrenchabstractDifferent algorithms have been proposed to detect semantic shifts (changes in a word meaning over time) in a diachronic corpus. Yet, and somehow surprisingly, no reference corpus has been designed so far to evaluate them, leaving researchers to fallback to troublesome evaluation strategies. In this work, we introduce a methodology for the construction of a reference dataset for the evaluation of semantic shift detection, that is, a list of words where we know for sure whether they present a word meaning change over a period of interest. We leverage a state-of-the-art word-sense disambiguation model to associate a date of first appearance to all the senses of a word. Significant changes in sense distributions as well as clear stability are detected and the resulting words are inspected by experts using a dedicated interface before populating a reference dataset. As a proof of concept, we apply this methodology to a corpus of newspapers from Quebec covering the whole 20th century. We manually verified a subset of candidates, leading to QC-FR-Diac-V1.0, a corpus of 151 words allowing one to evaluate the identification of semantic shifts in French between 1910 and 1990. David Kletz, Philippe Langlais, François Lareau, Patrick Drouin |
LREC | 2 |
| 2022 | A new dataset for multilingual keyphrase generationabstractKeyphrases are an important tool for efficiently dealing with the ever-increasing amount of information present on the internet. While there are many recent papers on English keyphrase generation, keyphrase generation for other languages remains vastly understudied, mostly due to the absence of datasets. To address this, we present a novel dataset called Papyrus, composed of 16427 pairs of abstracts and keyphrases. We release four versions of this dataset, corresponding to different subtasks. Papyrus-e considers only English keyphrases, Papyrus-f considers French keyphrases, Papyrus-m considers keyphrase generation in any language (mostly French and English), and Papyrus-a considers keyphrase generation in several languages. We train a state-of-the-art model on all four tasks and show that they lead to better results for non-English languages, with an average improvement of 14.2\% on keyphrase extraction and 2.0\% on generation. We also show an improvement of 0.4\% on extraction and 0.7\% on generation over English state-of-the-art results by concatenating Papyrus-e with the Kp20K training set. Frédéric Piedboeuf, Philippe Langlais |
NeurIPS | 2 |
| 2021 | Labels distribution matters in performance achieved in legal judgment prediction tasksabstractIn recent years, transformer [4] and BERT models [1] have been widely used in plain NLP tasks with the assumption that models first pretrained on massive corpora then fine-tuned on the dataset of a given task may suffice to achieve significant improvements. At the intersection of machine learning and law, legal judgment prediction (LJP) is a task that aims at predicting the outcome of a lawsuit based on a representation of the case. Such task is usually formalized in NLP as a text classification with different classes or labels corresponding to the verdicts. One specificity of court rulings is that their decisions are based on the application of legal articles to the facts described by the two parties (applicant and defendant). Olivier Salaün, Philippe Langlais, Karim Benyekhlef |
ICAIL | 2 |
| 2021 | Context-aware Adversarial Training for Name Regularity Bias in Named Entity RecognitionabstractAbstract In this work, we examine the ability of NER models to use contextual information when predicting the type of an ambiguous entity. We introduce NRB, a new testbed carefully designed to diagnose Name Regularity Bias of NER models. Our results indicate that all state-of-the-art models we tested show such a bias; BERT fine-tuned models significantly outperforming feature-based (LSTM-CRF) ones on NRB, despite having comparable (sometimes lower) performance on standard benchmarks. To mitigate this bias, we propose a novel model-agnostic training method that adds learnable adversarial noise to some entity mentions, thus enforcing models to focus more strongly on the contextual signal, leading to significant gains on NRB. Combining it with two other training strategies, data augmentation and parameter freezing, leads to further gains. Abbas Ghaddar, Philippe Langlais, Ahmad Rashid, Mehdi Rezagholizadeh |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Human or Neural Translation?abstractShivendra Bhardwaj, David Alfonso Hermelo, Phillippe Langlais, Gabriel Bernier-Colborne, Cyril Goutte, Michel Simard. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Shivendra Bhardwaj, David Alfonso-Hermelo, Philippe Langlais, Gabriel Bernier-Colborne, Cyril Goutte, Michel Simard |
COLING | 3 |
| 2020 | Data Selection for Bilingual Lexicon Induction from Specialized Comparable CorporaabstractNarrow specialized comparable corpora are often small in size.This particularity makes it difficult to build efficient models to acquire translation equivalents, especially for less frequent and rare words.One way to overcome this issue is to enrich the specialized corpora with out-ofdomain resources.Although some recent studies have shown improvements using data augmentation, the enrichment method was roughly conducted by adding out-of-domain data with no particular attention given to how to enrich words and how to do it optimally.In this paper, we contrast several data selection techniques to improve bilingual lexicon induction from specialized comparable corpora.We first apply two well-established data selection techniques often used in machine translation that is: Tf-Idf and cross entropy.Then, we propose to exploit BERT for data selection.Overall, all the proposed techniques improve the quality of the extracted bilingual lexicons by a large margin.The best performing model is the cross entropy, obtaining a gain of about 4 points in MAP while decreasing computation time by a factor of 10. Martin Laville, Amir Hazem, Emmanuel Morin, Philippe Langlais |
COLING | 4 |
| 2020 | Predicting S&P500 Monthly Direction with Informed Machine Learning
David Romain Djoumbissie, Philippe Langlais |
IPMU (3) | 2 |
| 2020 | HardEval: Focusing on Challenging Tokens to Assess Robustness of NERabstractTo assess the robustness of NER systems, we propose an evaluation method that focuses on subsets of tokens that represent specific sources of errors: unknown words and label shift or ambiguity. These subsets provide a system-agnostic basis for evaluating specific sources of NER errors and assessing room for improvement in terms of robustness. We analyze these subsets of challenging tokens in two widely-used NER benchmarks, then exploit them to evaluate NER systems in both in-domain and out-of-domain settings. Results show that these challenging tokens explain the majority of errors made by modern NER systems, although they represent only a small fraction of test tokens. They also indicate that label shift is harder to deal with than unknown words, and that there is much more room for improvement than the standard NER evaluation procedure would suggest. We hope this work will encourage NLP researchers to adopt rigorous and meaningful evaluation methods, and will help them develop more robust models. Gabriel Bernier-Colborne, Philippe Langlais |
LREC | 2 |
| 2020 | SEDAR: a Large Scale French-English Financial Domain Parallel CorpusabstractThis paper describes the acquisition, preprocessing and characteristics of SEDAR, a large scale English-French parallel corpus for the financial domain. Our extensive experiments on machine translation show that SEDAR is essential to obtain good performance on finance. We observe a large gain in the performance of machine translation systems trained on SEDAR when tested on finance, which makes SEDAR suitable to study domain adaptation for neural machine translation. The first release of the corpus comprises 8.6 million high quality sentence pairs that are publicly available for research at https://github.com/autorite/sedar-bitext. Abbas Ghaddar, Philippe Langlais |
LREC | 2 |
| 2020 | Analysis and Multilabel Classification of Quebec Court Decisions in the Domain of Housing Law
Olivier Salaün, Philippe Langlais, Andrés Lou, Hannes Westermann, Karim Benyekhlef |
NLDB | 2 |
| 2018 | Robust Lexical Features for Improved Neural Network Named-Entity RecognitionabstractNeural network approaches to Named-Entity Recognition reduce the need for carefully hand-crafted features. While some features do remain in state-of-the-art systems, lexical features have been mostly discarded, with the exception of gazetteers. In this work, we show that this is unfair: lexical features are actually quite useful. We propose to embed words and entity types into a low-dimensional vector space we train from annotated data produced by distant supervision thanks to Wikipedia. From this, we compute — offline — a feature vector representing each word. When used with a vanilla recurrent neural network model, this representation yields substantial improvements. We establish a new state-of-the-art F1 score of 87.95 on ONTONOTES 5.0, while matching state-of-the-art performance with a F1 score of 91.73 on the over-studied CONLL-2003 dataset. Abbas Ghaddar, Philippe Langlais |
COLING | 2 |
| 2018 | Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine TranslationabstractParallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose a bidirectional recurrent neural network based approach to extract parallel sentences from collections of multilingual texts. Our experiments with noisy parallel corpora show that we can achieve promising results against a competitive baseline by removing the need of specific feature engineering or additional external resources. To justify the utility of our approach, we extract sentence pairs from Wikipedia articles to train machine translation systems and show significant improvements in translation performance. Francis Grégoire, Philippe Langlais |
COLING | 2 |
| 2018 | Experiments in Learning to Solve Formal Analogical Equations
Rafik Rhouma, Philippe Langlais |
ICCBR | 2 |
| 2018 | Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus
Abbas Ghaddar, Philippe Langlais |
LREC | 2 |
| 2018 | Revisiting the Task of Scoring Open IE Relations
William Léchelle, Philippe Langlais |
LREC | 2 |
| 2018 | From French Wikipedia to Erudit: A test case for cross-domain open information extractionabstractAbstract In this paper, we describe an open information extraction pipeline based on ReVerb for extracting knowledge from French text. We put it to the test by using the information triples extracted to build an entity classifier, ie, a system able to label a given instance with its type (for instance, Michel Foucault is a philosopher). The classifier requires little supervision. One novel aspect of this study is that we show how general domain information triples (extracted from French Wikipedia) can be used for deriving new knowledge from domain‐specific documents unrelated to Wikipedia, in our case scholarly articles focusing on the humanities. We believe that the present study is the first that focuses on such a cross‐domain, recall‐oriented approach in open information extraction. While our system's performance shows room for improvement, manual assessments show that the task is quite hard, even for a human, in part because of the cross‐domain aspect of the problem we tackle. Fabrizio Gotti, Philippe Langlais |
Comput. Intell. | 2 |
| 2017 | WiNER: A Wikipedia Annotated Corpus for Named Entity RecognitionabstractWe revisit the idea of mining Wikipedia in order to generate named-entity annotations. We propose a new methodology that we applied to English Wikipedia to build WiNER, a large, high quality, annotated corpus. We evaluate its usefulness on 6 NER tasks, comparing 4 popular state-of-the art approaches. We show that LSTM-CRF is the approach that benefits the most from our corpus. We report impressive gains with this model when using a small portion of WiNER on top of the CONLL training material. Last, we propose a simple but efficient method for exploiting the full range of WiNER, leading to further improvements. Abbas Ghaddar, Philippe Langlais |
IJCNLP(1) | 2 |
| 2016 | An Informativeness Approach to Open IE Evaluation
William Léchelle, Philippe Langlais |
CICLing (2) | 2 |
| 2016 | Coreference in Wikipedia: Main Concept ResolutionabstractWikipedia is a resource of choice exploited in many NLP applications, yet we are not aware of recent attempts to adapt coreference resolution to this resource. In this work, we revisit a seldom studied task which consists in identifying in a Wikipedia article all the mentions of the main concept being described. We show that by exploiting the Wikipedia markup of a document, as well as links to external knowledge bases such as Freebase, we can acquire useful information on entities that helps to classify mentions as coreferent or not. We designed a classifier which drastically outperforms fair baselines built on top of state-of-the-art coreference resolution systems. We also measure the benefits of this classifier in a full coreference resolution pipeline applied to Wikipedia texts. Abbas Ghaddar, Philippe Langlais |
CoNLL | 2 |
| 2016 | WikiCoref: An English Coreference-annotated Corpus of Wikipedia Articles
Abbas Ghaddar, Philippe Langlais |
LREC | 2 |
| 2014 | Fourteen Light Tasks for comparing Analogical and Phrase-based Machine Translation
Rafik Rhouma, Philippe Langlais |
COLING | 2 |
| 2014 | Hashtag Occurrences, Layout and Translation: A Corpus-driven Analysis of Tweets Published by the Canadian Government
Fabrizio Gotti, Philippe Langlais, Anna Farzindar |
LREC | 2 |
| 2014 | An Iterative Approach for Mining Parallel Sentences in a Comparable Corpus
Lise Rebout, Philippe Langlais |
LREC | 2 |
| 2014 | Designing a machine translation system for Canadian weather warnings: A case studyabstractIn this paper we describe the many steps involved in building a production quality Machine Translation system for translating weather warnings between French and English. Although in principle this task may seem straightforward, the details, especially corpus preparation and final text presentation, involve many difficult aspects that are often glossed over in the literature. On top of the classic Statistical Machine Translation evaluation metric results, four manual evaluations have been performed to assess and improve translation quality. We also show the usefulness of the integration of out-of-domain information sources in a Statistical Machine Translation system to produce high quality translated text. Fabrizio Gotti, Philippe Langlais, Guy Lapalme |
Nat. Lang. Eng. | 2 |
| 2013 | Yet Another Fast, Robust and Open Source Sentence Aligner. Time toReconsider Sentence Alignment?
Fethi Lamraoui, Philippe Langlais |
MTSummit | 2 |
| 2012 | Texto4Science: a Quebec French Database of Annotated Short Text Messages
Philippe Langlais, Patrick Drouin, Amélie Paulus, Eugénie Rompré Brodeur, Florent Cottin |
LREC | 1 |
| 2011 | Reducing Overdetections in a French Symbolic Grammar Checker by Classification
Fabrizio Gotti, Philippe Langlais, Guy Lapalme, Simon Charest, Éric Brunelle |
CICLing (2) | 2 |
| 2011 | How Good is Your Comment? A Study of Comments in Java ProgramsabstractComments are very useful to developers during maintenance tasks and are useful as well to help structuring a code at development time. They convey useful information about the system functionalities as well as the state of mind of a developer. Comments in code have been the focus of several studies, but none of them was targeted at analyzing commenting habits precisely. In this paper, we present an empirical study which analyzes existing comments in different open source Java projects. We study comments from both a quantitative and a qualitative point of view. We propose a taxonomy of comments that we used for conducting our analysis. Dorsaf Haouari, Houari Sahraoui, Philippe Langlais |
ESEM | 3 |
| 2011 | Going Beyond Word Cooccurrences in Global Lexical Selection for Statistical Machine Translation using a Multilayer Perceptron
Alexandre Patry, Philippe Langlais |
IJCNLP | 2 |
| 2010 | Revisiting Context-based Projection Methods for Term-Translation Spotting in Comparable Corpora
Audrey Laroche, Philippe Langlais |
COLING | 2 |
| 2010 | TransSearch: from a bilingual concordancer to a translation finder
Julien Bourdaillet, Stéphane Huet, Philippe Langlais, Guy Lapalme |
Mach. Transl. | 3 |
| 2009 | Improvements in Analogical Learning: Application to Translating Multi-Terms of the Medical Domain
Philippe Langlais, François Yvon, Pierre Zweigenbaum |
EACL | 1 |
| 2009 | TS3: an Improved Version of the Bilingual Concordancer TransSearch
Stéphane Huet, Julien Bourdaillet, Philippe Langlais |
EAMT | 3 |
| 2009 | Prediction of Words in Statistical Machine Translation using a Multilayer Perceptron
Alexandre Patry, Philippe Langlais |
MTSummit | 2 |
| 2008 | Explorations in using grammatical dependencies for contextual phrase translation disambiguation
Aurélien Max, Rafik Makhloufi, Philippe Langlais |
EAMT | 3 |
| 2008 | MISTRAL: a Statistical Machine Translation Decoder for Speech Recognition Lattices
Alexandre Patry, Philippe Langlais |
LREC | 2 |
| 2007 | Translating Unknown Words by Analogical Learning
Philippe Langlais, Alexandre Patry |
EMNLP-CoNLL | 1 |
| 2006 | MOOD: A Modular Object-Oriented Decoder for Statistical Machine Translation
Alexandre Patry, Fabrizio Gotti, Philippe Langlais |
LREC | 3 |
| 2006 | EBMT by tree-phrasing
Philippe Langlais, Fabrizio Gotti |
Mach. Transl. | 1 |
| 2005 | From the real world to real words: the METEO case
Philippe Langlais, Thomas Leplus, Simona Gandrabur, Guy Lapalme |
EAMT | 1 |
| 2005 | The Long-Term Forecast for Weather Bulletin Translation
Philippe Langlais, Simona Gandrabur, Thomas Leplus, Guy Lapalme |
Mach. Transl. | 1 |
| 2004 | Adaptive Language and Translation Models for Interactive Machine Translation
Laurent Nepveu, Guy Lapalme, Philippe Langlais, George F. Foster |
EMNLP | 3 |
| 2004 | Evaluating Variants of the Lesk Approach for Disambiguating Words
Florentina Armaselu, Philippe Langlais, Guy Lapalme |
LREC | 2 |
| 2003 | Statistical machine translation: rapid development with limited resourcesabstractWe describe an experiment in rapid development of a statistical machine translation (SMT) system from scratch, using limited resources: under this heading we include not only training data, but also computing power, linguistic knowledge, programming effort, and absolute time. George F. Foster, Simona Gandrabur, Philippe Langlais, Pierre Plamondon, Graham Russell, Michel Simard |
MTSummit | 3 |
| 2002 | User-Friendly Text Prediction For TranslatorsabstractText prediction is a form of interactive machine translation that is well suited to skilled translators. In principle it can assist in the production of a target text with minimal disruption to a translator's normal routine. However, recent evaluations of a prototype prediction system showed that it significantly decreased the productivity of most translators who used it. In this paper, we analyze the reasons for this and propose a solution which consists in seeking predictions that maximize the expected benefit to the translator, rather than just trying to anticipate some amount of upcoming text. Using a model of a "typical translator" constructed from data collected in the evaluations of the prediction prototype, we show that this approach has the potential to turn text prediction into a help rather than a hindrance to a translator. George F. Foster, Philippe Langlais, Guy Lapalme |
EMNLP | 2 |
| 2002 | Translators at work with TRANSTYPE: Resource and Evaluation
Philippe Langlais, Marie Loranger, Guy Lapalme |
LREC | 1 |
| 2002 | Opening Statistical Translation Engines to Terminological Resources
Philippe Langlais |
NLDB | 1 |
| 2002 | Trans Type: Development-Evaluation Cycles to Boost Translator's Productivity
Philippe Langlais, Guy Lapalme |
Mach. Transl. | 1 |
| 2001 | Integrating bilingual lexicons in a probabilistic translation assistantabstractIn this paper, we present a way to integrate bilingual lexicons into an operational probabilistic translation assistant (TransType). These lexicons could be any resource available to the translator (e.g. terminological lexicons) or any resource statistically derived from training material. We describe a bilingual lexicon acquisition process that we developped and we evaluate from a theoretical point of view its benefits to a translation completion task. Philippe Langlais, George F. Foster, Guy Lapalme |
MTSummit | 1 |
| 2001 | Sub-sentential exploitation of translation memoriesabstractTranslation memory systems (TMS) are a family of computer tools whose purpose is to facilitate and encourage the re-use of existing translations. By searching a database of past translations, these systems can retrieve the translation of whole segments of text and propose them to the translator for re-use. However, the usefulness of existing TMS’s is limited by the nature of the text segments that that they are able to put in correspondence, generally whole sentences. This article examines the potential of a type of system that is able to recuperate the translation of sub-sentential sequences of words. Michel Simard, Philippe Langlais |
MTSummit | 2 |
| 2000 | Evaluation of TRANSTYPE, a Computer-aided Translation Typing System: A Comparison of a Theoretical- and a User-oriented Evaluation Procedures
Philippe Langlais, Sébastien Sauvé, George F. Foster, Elliott Macklovitch, Guy Lapalme |
LREC | 1 |
| 2000 | TransSearch: A Free Translation Memory on the World Wide Web
Elliott Macklovitch, Michel Simard, Philippe Langlais |
LREC | 3 |
| 2000 | Unit Completion for a Computer-aided Translation Typing System
Philippe Langlais, George F. Foster, Guy Lapalme |
Mach. Transl. | 1 |
| 1998 | Phonetic-level mispronunciation detection in non-native Swedish speechabstractThis contribution presents part of the work initiated at the CTT for the development of speech technology to assist non-native speakers learn Swedish. This study focuses mainly on the automatic location of mispronunciations at a phonetic level. We first describe the database we created for this work and then report on the reliability of several phonetic scores to automatically locate segmental problems in student utterances. 1. INTRODUCTION Over the last decade, advancements in speech technology have opened up new possibilities for interactive language teaching systems [1,7,10,11]. And more recently, a growing number of studies have addressed the problem of automatically rating nonnative speakers by providing measurements that can be correlated with human judgment [3,5]. These studies show that rating a speaker on a 5-point scale can be achieved with performance inversely proportional to the size of the speechunit rated. In other words, speech technology is mature enough to grade a s... Philippe Langlais, Anne-Marie Öster, Björn Granström |
ICSLP | 1 |
| 1998 | ARCADE: a cooperative research project on parallel text alignment evaluation
Philippe Langlais, Marc Simard, Jean Véronis, Susan Armstrong, Pierre Bonhomme, Fathi Debili, Isabelle P. Oswald, Emna Soussi, P. Therón |
LREC | 1 |
| 1997 | Estimating prosodic weights in a syntactic-rhythmical prediction systemabstractThis paper concerns the study of information derived from the melodic, temporal and intensity characteristics of the material to be recognized in a speech recognition system, in French. More precisely, it describes experiments we achieved at the suprasegmental levels with a system that outperform automatic correlation between prosodic labels and linguistic organization of a message to decode. Firstly an overview of the system is described along with the results of experiments carried out to determine which prosodic indexes are bestsuited for syntactic and rhythmycal prediction. 1 INTRODUCTION It is well known that prosodic structure and syntactic structure are not identical; neither are they unrelated. Knowing when and how the two correspond could help in the disambiguation of competing syntactic hypotheses in a speech understanding system. This practical reason explains the renewed interest in the use of prosody in ASR [8, 2, 4, 6]. This paper reports our experiments trying to ans... Philippe Langlais |
EUROSPEECH | 1 |
| 1995 | Microprosodic study of isolated French word corporaabstractThe present contribution aims at inventoring the microprosodic phenomena already abundantly described in the past for languages such as English and French (see [1] for a commented review of these studies) . We propose to measure their usefulness in an automatic processing and also to estimate the expected confidency of microprosodic "corrections" --- at least in an automatic way --- by studying each parameter separately and by presenting statistical distributions of the different phenomena measured on isolated french word corpora uttered by several speakers. Philippe Langlais |
EUROSPEECH | 1 |
| 1993 | Integration of a prosodic component in an automatic speech recognition system
Philippe Langlais, Henri Meloni |
EUROSPEECH | 1 |
| 1991 | Analytical strategy for speaker identification
Jean-François Bonastre, Henri Meloni, Philippe Langlais |
EUROSPEECH | 3 |