Philippe Langlais

dblp:66/1102 · DBLP profile ↗
← Back
74ranked-venue papers
17as first author
19since 2021 · last 2025
0000-0002-7319-1595ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 17 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
abstract
This study investigates the behavior of model-integrated routers in Mixture of Experts (MoE) models, focusing on how tokens are routed based on their linguistic features, specifically Part-of-Speech (POS) tags. The goal is to explore across different MoE architectures whether experts specialize in processing tokens with similar linguistic traits. By analyzing token trajectories across experts and layers, we aim to uncover how MoE models handle linguistic information. Findings from six popular MoE models reveal expert specialization for specific POS categories, with routing paths showing high predictive accuracy for POS, highlighting the value of routing paths in characterizing tokens.
Elie Antoine, Frédéric Béchet, Philippe Langlais
COLING3
2025 On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
abstract
Textual data augmentation (DA) is a prolific field of study where novel techniques to create artificial data are regularly proposed, and that has demonstrated great efficiency on small data settings, at least for text classification tasks. In this paper, we challenge those results, showing that classical data augmentation (which modify sentences) is simply a way of performing better fine-tuning, and that spending more time doing so before applying data augmentation negates its effect. This is a significant contribution as it answers several questions that were left open in recent years, namely : which DA technique performs best (all of them as long as they generate data close enough to the training set, as to not impair training) and why did DA show positive results (facilitates training of network). We further show that zero- and few-shot DA via conversational agents such as ChatGPT or LLama2 can increase performances, confirming that this form of data augmentation is preferable to classical methods.
Frédéric Piedboeuf, Philippe Langlais
COLING2
2025 An Interpretable Quantum-Inspired Model for Multi-Task Natural Language Understanding
abstract
Multi-task learning has demonstrated remarkable success across a broad spectrum of natural language processing tasks, particularly with neural network-based methods. Despite these advances, a fundamental gap remains in explaining the relationship between task-relatedness and model effectiveness. To address this issue, we propose a novel approach for implicitly modeling task-relatedness by leveraging a quantum physical mathematical framework. In this paper, we introduce a complex-valued neural network designed to encapsulate and analyze task-relatedness. Within this framework, sentences originating from diverse tasks are encoded as mixed quantum systems, represented on a meticulously defined Semantic Hilbert Space. This allows the network to interpret inter-task relationships through the explicit physical semantics of well-constrained components grounded in quantum probability theory. By adhering to these rigorous principles, our model not only establishes a robust method for quantifying task-relatedness but also fosters a deeper, self-explanatory understanding of the underlying processes. To validate the efficacy of our approach, we conducted extensive experiments across five benchmark text classification tasks. The results demonstrate both the superior performance and the interpretability of the proposed model, highlighting its potential as a self-explanatory system for multi-task learning in NLP.
Peng Lu 0006, Jerry Huang, Xinyu Wang 0061, Philippe Langlais
ECAI4
2025 ALF: A Fine-Grained French Analogical Dataset for Evaluating Lexical Knowledge of Large Language Models
abstract
The undeniable revolution brought forth by Large Language Models (LLMs) stems from the amazing fluency of the texts they generate, mastering language with seemingly human-like finesse. This fluency raises a key scientific question: How much lexical knowledge do LLMs actually capture in order to produce such fluent language? To address this, we present ALF, a freely available, analogical dataset endowed with rich lexicographic information grounded in Meaning-Text Theory for the French language. It comprises 2600 fine-grained lexical analogies with which we evaluate the lexical ability of five off-the-shelf LLMs, namely ChatGPT-4o mini, Llama3.0-8B, Llama3.1-8B, Qwen2.5-14B, and Mistral7B. Their performance spans from 45% for Mistral, through about 55% for the ChatGPT and Llama models, and up to nearly 60% for Qwen2.5-14B, thus qualifying ALF as a challenging dataset. Experimenting with larger models (OpenAI o1, Llama3.0/3.1-70B, and Qwen2.5-32B) yields rather limited returns considering the drastic increase in computational cost. We further identify certain types of analogies and prompting methods that reveal performance disparities.
Alexander Petrov, Antoine Venant, François Lareau, Yves Lepage, Philippe Langlais
ECAI5
2025 ReGLA: Refining Gated Linear Attention
abstract
Peng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Peng Lu 0006, Ivan Kobyzev, Mehdi Rezagholizadeh, Boxing Chen, Philippe Langlais
NAACL (Long Papers)5
2025 Mamba Modulation: On the Length Generalization of Mamba Models
abstract
The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged as a leading architecture, achieving state-of-the-art results across a range of language modeling tasks. However, Mamba’s performance significantly deteriorates when applied to contexts longer than those seen during pre-training, revealing a sharp sensitivity to context length extension. Through detailed analysis, we attribute this limitation to the out-of-distribution behavior of its state-space dynamics, particularly within the parameterization of the state transition matrix $A$. Unlike recent works which attribute this sensitivity to the vanished accumulation of discretization time steps, $\exp(-\sum_{t=1}^N{\Delta}_t)$, we establish a connection between state convergence behavior as the input length approaches infinity and the spectrum of the transition matrix $A$, offering a well-founded explanation of its role in length extension. Next, to overcome this challenge, we propose an approach that applies spectrum scaling to pre-trained Mamba models to enable robust long-context generalization by selectively modulating the spectrum of $A$ matrices in each layer. We show that this can significantly improve performance in settings where simply modulating ${\Delta}_t$ fails, validating our insights and providing avenues for better length generalization of state-space models with structured transition matrices.
Peng Lu 0006, Jerry Huang, Qiuhao Zeng, Xinyu Wang 0061, Boxing Chen, Philippe Langlais, Yufei Cui
NeurIPS6
2024 EUROPA: A Legal Multilingual Keyphrase Generation Dataset
abstract
Olivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso-Hermelo, Philippe Langlais. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Olivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso-Hermelo, Philippe Langlais
ACL (1)5
2024 A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
abstract
We introduce an evaluation methodology for reading comprehension tasks based on the intuition that certain examples, by the virtue of their linguistic complexity, consistently yield lower scores regardless of model size or architecture.We capitalize on semantic frame annotation for characterizing this complexity, and study seven complexity factors that may account for model's difficulty.We first deploy this methodology on a carefully annotated French reading comprehension benchmark showing that two of those complexity factors are indeed good predictors of models' failure, while others are less so.We further deploy our methodology on a well studied English benchmark by using Chat-GPT as a proxy for semantic annotation.Our study reveals that fine-grained linguisticallymotivated automatic evaluation of a reading comprehension task is not only possible, but helps understand models' abilities to handle specific linguistic characteristics of input examples.It also shows that current state-of-the-art models fail with some for those characteristics which suggests that adequately handling them requires more than merely increasing model size.
Elie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe Langlais
EMNLP4
2023 On the utility of enhancing BERT syntactic bias with Token Reordering Pretraining
abstract
Yassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Phillippe Langlais, Prasanna Parthasarathi. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023.
Yassir El Mesbahi, Atif Mahmud, Abbas Ghaddar, Mehdi Rezagholizadeh, Philippe Langlais, Prasanna Parthasarathi
CoNLL5
2023 An analysis of entity normalization evaluation biases in specialized domains
abstract
BACKGROUND: Entity normalization is an important information extraction task which has recently gained attention, particularly in the clinical/biomedical and life science domains. On several datasets, state-of-the-art methods perform rather well on popular benchmarks. Yet, we argue that the task is far from resolved. RESULTS: We have selected two gold standard corpora and two state-of-the-art methods to highlight some evaluation biases. We present non-exhaustive initial findings on the existence of evaluation problems of the entity normalization task. CONCLUSIONS: Our analysis suggests better evaluation practices to support the methodological research in this field.
Arnaud Ferré, Philippe Langlais
BMC Bioinform.2
2022 CILDA: Contrastive Data Augmentation Using Intermediate Layer Knowledge Distillation
abstract
Knowledge distillation (KD) is an efficient framework for compressing large-scale pre-trained language models. Recent years have seen a surge of research aiming to improve KD by leveraging Contrastive Learning, Intermediate Layer Distillation, Data Augmentation, and Adversarial Training. In this work, we propose a learning-based data augmentation technique tailored for knowledge distillation, called CILDA. To the best of our knowledge, this is the first time that intermediate layer representations of the main task are used in improving the quality of augmented samples. More precisely, we introduce an augmentation technique for KD based on intermediate layer matching using contrastive loss to improve masked adversarial data augmentation. CILDA outperforms existing state-of-the-art KD approaches on the GLUE benchmark, as well as in an out-of-domain evaluation.
Md. Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar, Khalil Bibi, Philippe Langlais, Pascal Poupart
COLING5
2022 Effective Data Augmentation for Sentence Classification Using One VAE per Class
abstract
In recent years, data augmentation has become an important field of machine learning. While images can use simple techniques such as cropping or rotating, textual data augmentation needs more complex manipulations to ensure that the generated examples are useful. Variational auto-encoders (VAE) and its conditional variant the Conditional-VAE (CVAE) are often used to generate new textual data, both relying on a good enough training of the generator so that it doesn’t create examples of the wrong class. In this paper, we explore a simpler way to use VAE for data augmentation: the training of one VAE per class. We show on several dataset sizes, as well as on four different binary classification tasks, that it systematically outperforms other generative data augmentation techniques.
Frédéric Piedboeuf, Philippe Langlais
COLING2
2022 Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing
abstract
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais
EMNLP14
2022 Why Do Tenants Sue Their Landlords? Answers from a Topic Model
abstract
Topic modeling is widely used in various domains for extracting latent topics underlying large corpora, including judicial texts. In the latter, topics tend to be made by and for domain experts, but remain unintelligible for laymen. In the framework of housing law court decisions in French which mixes abstract legal terminology with real-life situations described in common language, similarly to [1], we aim at identifying different situations that can cause a tenant to prosecute their landlord in court with the application of topic models. Upon quantitative evaluation, LDA and BERTopic deliver the best results, but a closer manual analysis reveals that the second embedding-based approach is much better at producing and even uncovering topics that describe a tenant’s real-life issues and situations.
Olivier Salaün, Fabrizio Gotti, Philippe Langlais, Karim Benyekhlef
JURIX3
2022 Conditional Abstractive Summarization of Court Decisions for Laymen and Insights from Human Evaluation
abstract
Legal text summarization is generally formalized as an extractive text summarization task applied to court decisions from which the most relevant sentences are identified and returned as a gist meant to be read by legal experts. However, such summaries are not suitable for laymen seeking intelligible legal information. In the scope of the JusticeBot, a question-answering system in French that provides information about housing law, we intend to generate summaries of court decisions that are, on the one hand, conditioned by a question-answer-decision triplet, and on the other hand, intelligible for ordinary citizens not familiar with legal documents. So far, our best model, a further pre-trained BARThez, achieves an average ROUGE-1 score of 37.7 and a deepened manual evaluation of summaries reveals that there is still room for improvement.
Olivier Salaün, Aurore Clément Troussel, Sylvain Longhais, Hannes Westermann, Philippe Langlais, Karim Benyekhlef
JURIX5
2022 A Methodology for Building a Diachronic Dataset of Semantic Shifts and its Application to QC-FR-Diac-V1.0, a Free Reference for French
abstract
Different algorithms have been proposed to detect semantic shifts (changes in a word meaning over time) in a diachronic corpus. Yet, and somehow surprisingly, no reference corpus has been designed so far to evaluate them, leaving researchers to fallback to troublesome evaluation strategies. In this work, we introduce a methodology for the construction of a reference dataset for the evaluation of semantic shift detection, that is, a list of words where we know for sure whether they present a word meaning change over a period of interest. We leverage a state-of-the-art word-sense disambiguation model to associate a date of first appearance to all the senses of a word. Significant changes in sense distributions as well as clear stability are detected and the resulting words are inspected by experts using a dedicated interface before populating a reference dataset. As a proof of concept, we apply this methodology to a corpus of newspapers from Quebec covering the whole 20th century. We manually verified a subset of candidates, leading to QC-FR-Diac-V1.0, a corpus of 151 words allowing one to evaluate the identification of semantic shifts in French between 1910 and 1990.
David Kletz, Philippe Langlais, François Lareau, Patrick Drouin
LREC2
2022 A new dataset for multilingual keyphrase generation
abstract
Keyphrases are an important tool for efficiently dealing with the ever-increasing amount of information present on the internet. While there are many recent papers on English keyphrase generation, keyphrase generation for other languages remains vastly understudied, mostly due to the absence of datasets. To address this, we present a novel dataset called Papyrus, composed of 16427 pairs of abstracts and keyphrases. We release four versions of this dataset, corresponding to different subtasks. Papyrus-e considers only English keyphrases, Papyrus-f considers French keyphrases, Papyrus-m considers keyphrase generation in any language (mostly French and English), and Papyrus-a considers keyphrase generation in several languages. We train a state-of-the-art model on all four tasks and show that they lead to better results for non-English languages, with an average improvement of 14.2\% on keyphrase extraction and 2.0\% on generation. We also show an improvement of 0.4\% on extraction and 0.7\% on generation over English state-of-the-art results by concatenating Papyrus-e with the Kp20K training set.
Frédéric Piedboeuf, Philippe Langlais
NeurIPS2
2021 Labels distribution matters in performance achieved in legal judgment prediction tasks
abstract
In recent years, transformer [4] and BERT models [1] have been widely used in plain NLP tasks with the assumption that models first pretrained on massive corpora then fine-tuned on the dataset of a given task may suffice to achieve significant improvements. At the intersection of machine learning and law, legal judgment prediction (LJP) is a task that aims at predicting the outcome of a lawsuit based on a representation of the case. Such task is usually formalized in NLP as a text classification with different classes or labels corresponding to the verdicts. One specificity of court rulings is that their decisions are based on the application of legal articles to the facts described by the two parties (applicant and defendant).
Olivier Salaün, Philippe Langlais, Karim Benyekhlef
ICAIL2
2021 Context-aware Adversarial Training for Name Regularity Bias in Named Entity Recognition
abstract
Abstract In this work, we examine the ability of NER models to use contextual information when predicting the type of an ambiguous entity. We introduce NRB, a new testbed carefully designed to diagnose Name Regularity Bias of NER models. Our results indicate that all state-of-the-art models we tested show such a bias; BERT fine-tuned models significantly outperforming feature-based (LSTM-CRF) ones on NRB, despite having comparable (sometimes lower) performance on standard benchmarks. To mitigate this bias, we propose a novel model-agnostic training method that adds learnable adversarial noise to some entity mentions, thus enforcing models to focus more strongly on the contextual signal, leading to significant gains on NRB. Combining it with two other training strategies, data augmentation and parameter freezing, leads to further gains.
Abbas Ghaddar, Philippe Langlais, Ahmad Rashid, Mehdi Rezagholizadeh
Trans. Assoc. Comput. Linguistics2
2020 Human or Neural Translation?
abstract
Shivendra Bhardwaj, David Alfonso Hermelo, Phillippe Langlais, Gabriel Bernier-Colborne, Cyril Goutte, Michel Simard. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Shivendra Bhardwaj, David Alfonso-Hermelo, Philippe Langlais, Gabriel Bernier-Colborne, Cyril Goutte, Michel Simard
COLING3
2020 Data Selection for Bilingual Lexicon Induction from Specialized Comparable Corpora
abstract
Narrow specialized comparable corpora are often small in size.This particularity makes it difficult to build efficient models to acquire translation equivalents, especially for less frequent and rare words.One way to overcome this issue is to enrich the specialized corpora with out-ofdomain resources.Although some recent studies have shown improvements using data augmentation, the enrichment method was roughly conducted by adding out-of-domain data with no particular attention given to how to enrich words and how to do it optimally.In this paper, we contrast several data selection techniques to improve bilingual lexicon induction from specialized comparable corpora.We first apply two well-established data selection techniques often used in machine translation that is: Tf-Idf and cross entropy.Then, we propose to exploit BERT for data selection.Overall, all the proposed techniques improve the quality of the extracted bilingual lexicons by a large margin.The best performing model is the cross entropy, obtaining a gain of about 4 points in MAP while decreasing computation time by a factor of 10.
Martin Laville, Amir Hazem, Emmanuel Morin, Philippe Langlais
COLING4
2020 Predicting S&P500 Monthly Direction with Informed Machine Learning
David Romain Djoumbissie, Philippe Langlais
IPMU (3)2
2020 HardEval: Focusing on Challenging Tokens to Assess Robustness of NER
abstract
To assess the robustness of NER systems, we propose an evaluation method that focuses on subsets of tokens that represent specific sources of errors: unknown words and label shift or ambiguity. These subsets provide a system-agnostic basis for evaluating specific sources of NER errors and assessing room for improvement in terms of robustness. We analyze these subsets of challenging tokens in two widely-used NER benchmarks, then exploit them to evaluate NER systems in both in-domain and out-of-domain settings. Results show that these challenging tokens explain the majority of errors made by modern NER systems, although they represent only a small fraction of test tokens. They also indicate that label shift is harder to deal with than unknown words, and that there is much more room for improvement than the standard NER evaluation procedure would suggest. We hope this work will encourage NLP researchers to adopt rigorous and meaningful evaluation methods, and will help them develop more robust models.
Gabriel Bernier-Colborne, Philippe Langlais
LREC2
2020 SEDAR: a Large Scale French-English Financial Domain Parallel Corpus
abstract
This paper describes the acquisition, preprocessing and characteristics of SEDAR, a large scale English-French parallel corpus for the financial domain. Our extensive experiments on machine translation show that SEDAR is essential to obtain good performance on finance. We observe a large gain in the performance of machine translation systems trained on SEDAR when tested on finance, which makes SEDAR suitable to study domain adaptation for neural machine translation. The first release of the corpus comprises 8.6 million high quality sentence pairs that are publicly available for research at https://github.com/autorite/sedar-bitext.
Abbas Ghaddar, Philippe Langlais
LREC2
2020 Analysis and Multilabel Classification of Quebec Court Decisions in the Domain of Housing Law
Olivier Salaün, Philippe Langlais, Andrés Lou, Hannes Westermann, Karim Benyekhlef
NLDB2
2018 Robust Lexical Features for Improved Neural Network Named-Entity Recognition
abstract
Neural network approaches to Named-Entity Recognition reduce the need for carefully hand-crafted features. While some features do remain in state-of-the-art systems, lexical features have been mostly discarded, with the exception of gazetteers. In this work, we show that this is unfair: lexical features are actually quite useful. We propose to embed words and entity types into a low-dimensional vector space we train from annotated data produced by distant supervision thanks to Wikipedia. From this, we compute — offline — a feature vector representing each word. When used with a vanilla recurrent neural network model, this representation yields substantial improvements. We establish a new state-of-the-art F1 score of 87.95 on ONTONOTES 5.0, while matching state-of-the-art performance with a F1 score of 91.73 on the over-studied CONLL-2003 dataset.
Abbas Ghaddar, Philippe Langlais
COLING2
2018 Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation
abstract
Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. We propose a bidirectional recurrent neural network based approach to extract parallel sentences from collections of multilingual texts. Our experiments with noisy parallel corpora show that we can achieve promising results against a competitive baseline by removing the need of specific feature engineering or additional external resources. To justify the utility of our approach, we extract sentence pairs from Wikipedia articles to train machine translation systems and show significant improvements in translation performance.
Francis Grégoire, Philippe Langlais
COLING2
2018 Experiments in Learning to Solve Formal Analogical Equations
Rafik Rhouma, Philippe Langlais
ICCBR2
2018 Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus
Abbas Ghaddar, Philippe Langlais
LREC2
2018 Revisiting the Task of Scoring Open IE Relations
William Léchelle, Philippe Langlais
LREC2
2018 From French Wikipedia to Erudit: A test case for cross-domain open information extraction
abstract
Abstract In this paper, we describe an open information extraction pipeline based on ReVerb for extracting knowledge from French text. We put it to the test by using the information triples extracted to build an entity classifier, ie, a system able to label a given instance with its type (for instance, Michel Foucault is a philosopher). The classifier requires little supervision. One novel aspect of this study is that we show how general domain information triples (extracted from French Wikipedia) can be used for deriving new knowledge from domain‐specific documents unrelated to Wikipedia, in our case scholarly articles focusing on the humanities. We believe that the present study is the first that focuses on such a cross‐domain, recall‐oriented approach in open information extraction. While our system's performance shows room for improvement, manual assessments show that the task is quite hard, even for a human, in part because of the cross‐domain aspect of the problem we tackle.
Fabrizio Gotti, Philippe Langlais
Comput. Intell.2
2017 WiNER: A Wikipedia Annotated Corpus for Named Entity Recognition
abstract
We revisit the idea of mining Wikipedia in order to generate named-entity annotations. We propose a new methodology that we applied to English Wikipedia to build WiNER, a large, high quality, annotated corpus. We evaluate its usefulness on 6 NER tasks, comparing 4 popular state-of-the art approaches. We show that LSTM-CRF is the approach that benefits the most from our corpus. We report impressive gains with this model when using a small portion of WiNER on top of the CONLL training material. Last, we propose a simple but efficient method for exploiting the full range of WiNER, leading to further improvements.
Abbas Ghaddar, Philippe Langlais
IJCNLP(1)2
2016 An Informativeness Approach to Open IE Evaluation
William Léchelle, Philippe Langlais
CICLing (2)2
2016 Coreference in Wikipedia: Main Concept Resolution
abstract
Wikipedia is a resource of choice exploited in many NLP applications, yet we are not aware of recent attempts to adapt coreference resolution to this resource. In this work, we revisit a seldom studied task which consists in identifying in a Wikipedia article all the mentions of the main concept being described. We show that by exploiting the Wikipedia markup of a document, as well as links to external knowledge bases such as Freebase, we can acquire useful information on entities that helps to classify mentions as coreferent or not. We designed a classifier which drastically outperforms fair baselines built on top of state-of-the-art coreference resolution systems. We also measure the benefits of this classifier in a full coreference resolution pipeline applied to Wikipedia texts.
Abbas Ghaddar, Philippe Langlais
CoNLL2
2016 WikiCoref: An English Coreference-annotated Corpus of Wikipedia Articles
Abbas Ghaddar, Philippe Langlais
LREC2
2014 Fourteen Light Tasks for comparing Analogical and Phrase-based Machine Translation
Rafik Rhouma, Philippe Langlais
COLING2
2014 Hashtag Occurrences, Layout and Translation: A Corpus-driven Analysis of Tweets Published by the Canadian Government
Fabrizio Gotti, Philippe Langlais, Anna Farzindar
LREC2
2014 An Iterative Approach for Mining Parallel Sentences in a Comparable Corpus
Lise Rebout, Philippe Langlais
LREC2
2014 Designing a machine translation system for Canadian weather warnings: A case study
abstract
In this paper we describe the many steps involved in building a production quality Machine Translation system for translating weather warnings between French and English. Although in principle this task may seem straightforward, the details, especially corpus preparation and final text presentation, involve many difficult aspects that are often glossed over in the literature. On top of the classic Statistical Machine Translation evaluation metric results, four manual evaluations have been performed to assess and improve translation quality. We also show the usefulness of the integration of out-of-domain information sources in a Statistical Machine Translation system to produce high quality translated text.
Fabrizio Gotti, Philippe Langlais, Guy Lapalme
Nat. Lang. Eng.2
2013 Yet Another Fast, Robust and Open Source Sentence Aligner. Time toReconsider Sentence Alignment?
Fethi Lamraoui, Philippe Langlais
MTSummit2
2012 Texto4Science: a Quebec French Database of Annotated Short Text Messages
Philippe Langlais, Patrick Drouin, Amélie Paulus, Eugénie Rompré Brodeur, Florent Cottin
LREC1
2011 Reducing Overdetections in a French Symbolic Grammar Checker by Classification
Fabrizio Gotti, Philippe Langlais, Guy Lapalme, Simon Charest, Éric Brunelle
CICLing (2)2
2011 How Good is Your Comment? A Study of Comments in Java Programs
abstract
Comments are very useful to developers during maintenance tasks and are useful as well to help structuring a code at development time. They convey useful information about the system functionalities as well as the state of mind of a developer. Comments in code have been the focus of several studies, but none of them was targeted at analyzing commenting habits precisely. In this paper, we present an empirical study which analyzes existing comments in different open source Java projects. We study comments from both a quantitative and a qualitative point of view. We propose a taxonomy of comments that we used for conducting our analysis.
Dorsaf Haouari, Houari Sahraoui, Philippe Langlais
ESEM3
2011 Going Beyond Word Cooccurrences in Global Lexical Selection for Statistical Machine Translation using a Multilayer Perceptron
Alexandre Patry, Philippe Langlais
IJCNLP2
2010 Revisiting Context-based Projection Methods for Term-Translation Spotting in Comparable Corpora
Audrey Laroche, Philippe Langlais
COLING2
2010 TransSearch: from a bilingual concordancer to a translation finder
Julien Bourdaillet, Stéphane Huet, Philippe Langlais, Guy Lapalme
Mach. Transl.3
2009 Improvements in Analogical Learning: Application to Translating Multi-Terms of the Medical Domain
Philippe Langlais, François Yvon, Pierre Zweigenbaum
EACL1
2009 TS3: an Improved Version of the Bilingual Concordancer TransSearch
Stéphane Huet, Julien Bourdaillet, Philippe Langlais
EAMT3
2009 Prediction of Words in Statistical Machine Translation using a Multilayer Perceptron
Alexandre Patry, Philippe Langlais
MTSummit2
2008 Explorations in using grammatical dependencies for contextual phrase translation disambiguation
Aurélien Max, Rafik Makhloufi, Philippe Langlais
EAMT3
2008 MISTRAL: a Statistical Machine Translation Decoder for Speech Recognition Lattices
Alexandre Patry, Philippe Langlais
LREC2
2007 Translating Unknown Words by Analogical Learning
Philippe Langlais, Alexandre Patry
EMNLP-CoNLL1
2006 MOOD: A Modular Object-Oriented Decoder for Statistical Machine Translation
Alexandre Patry, Fabrizio Gotti, Philippe Langlais
LREC3
2006 EBMT by tree-phrasing
Philippe Langlais, Fabrizio Gotti
Mach. Transl.1
2005 From the real world to real words: the METEO case
Philippe Langlais, Thomas Leplus, Simona Gandrabur, Guy Lapalme
EAMT1
2005 The Long-Term Forecast for Weather Bulletin Translation
Philippe Langlais, Simona Gandrabur, Thomas Leplus, Guy Lapalme
Mach. Transl.1
2004 Adaptive Language and Translation Models for Interactive Machine Translation
Laurent Nepveu, Guy Lapalme, Philippe Langlais, George F. Foster
EMNLP3
2004 Evaluating Variants of the Lesk Approach for Disambiguating Words
Florentina Armaselu, Philippe Langlais, Guy Lapalme
LREC2
2003 Statistical machine translation: rapid development with limited resources
abstract
We describe an experiment in rapid development of a statistical machine translation (SMT) system from scratch, using limited resources: under this heading we include not only training data, but also computing power, linguistic knowledge, programming effort, and absolute time.
George F. Foster, Simona Gandrabur, Philippe Langlais, Pierre Plamondon, Graham Russell, Michel Simard
MTSummit3
2002 User-Friendly Text Prediction For Translators
abstract
Text prediction is a form of interactive machine translation that is well suited to skilled translators. In principle it can assist in the production of a target text with minimal disruption to a translator's normal routine. However, recent evaluations of a prototype prediction system showed that it significantly decreased the productivity of most translators who used it. In this paper, we analyze the reasons for this and propose a solution which consists in seeking predictions that maximize the expected benefit to the translator, rather than just trying to anticipate some amount of upcoming text. Using a model of a "typical translator" constructed from data collected in the evaluations of the prediction prototype, we show that this approach has the potential to turn text prediction into a help rather than a hindrance to a translator.
George F. Foster, Philippe Langlais, Guy Lapalme
EMNLP2
2002 Translators at work with TRANSTYPE: Resource and Evaluation
Philippe Langlais, Marie Loranger, Guy Lapalme
LREC1
2002 Opening Statistical Translation Engines to Terminological Resources
Philippe Langlais
NLDB1
2002 Trans Type: Development-Evaluation Cycles to Boost Translator's Productivity
Philippe Langlais, Guy Lapalme
Mach. Transl.1
2001 Integrating bilingual lexicons in a probabilistic translation assistant
abstract
In this paper, we present a way to integrate bilingual lexicons into an operational probabilistic translation assistant (TransType). These lexicons could be any resource available to the translator (e.g. terminological lexicons) or any resource statistically derived from training material. We describe a bilingual lexicon acquisition process that we developped and we evaluate from a theoretical point of view its benefits to a translation completion task.
Philippe Langlais, George F. Foster, Guy Lapalme
MTSummit1
2001 Sub-sentential exploitation of translation memories
abstract
Translation memory systems (TMS) are a family of computer tools whose purpose is to facilitate and encourage the re-use of existing translations. By searching a database of past translations, these systems can retrieve the translation of whole segments of text and propose them to the translator for re-use. However, the usefulness of existing TMS’s is limited by the nature of the text segments that that they are able to put in correspondence, generally whole sentences. This article examines the potential of a type of system that is able to recuperate the translation of sub-sentential sequences of words.
Michel Simard, Philippe Langlais
MTSummit2
2000 Evaluation of TRANSTYPE, a Computer-aided Translation Typing System: A Comparison of a Theoretical- and a User-oriented Evaluation Procedures
Philippe Langlais, Sébastien Sauvé, George F. Foster, Elliott Macklovitch, Guy Lapalme
LREC1
2000 TransSearch: A Free Translation Memory on the World Wide Web
Elliott Macklovitch, Michel Simard, Philippe Langlais
LREC3
2000 Unit Completion for a Computer-aided Translation Typing System
Philippe Langlais, George F. Foster, Guy Lapalme
Mach. Transl.1
1998 Phonetic-level mispronunciation detection in non-native Swedish speech
abstract
This contribution presents part of the work initiated at the CTT for the development of speech technology to assist non-native speakers learn Swedish. This study focuses mainly on the automatic location of mispronunciations at a phonetic level. We first describe the database we created for this work and then report on the reliability of several phonetic scores to automatically locate segmental problems in student utterances. 1. INTRODUCTION Over the last decade, advancements in speech technology have opened up new possibilities for interactive language teaching systems [1,7,10,11]. And more recently, a growing number of studies have addressed the problem of automatically rating nonnative speakers by providing measurements that can be correlated with human judgment [3,5]. These studies show that rating a speaker on a 5-point scale can be achieved with performance inversely proportional to the size of the speechunit rated. In other words, speech technology is mature enough to grade a s...
Philippe Langlais, Anne-Marie Öster, Björn Granström
ICSLP1
1998 ARCADE: a cooperative research project on parallel text alignment evaluation
Philippe Langlais, Marc Simard, Jean Véronis, Susan Armstrong, Pierre Bonhomme, Fathi Debili, Isabelle P. Oswald, Emna Soussi, P. Therón
LREC1
1997 Estimating prosodic weights in a syntactic-rhythmical prediction system
abstract
This paper concerns the study of information derived from the melodic, temporal and intensity characteristics of the material to be recognized in a speech recognition system, in French. More precisely, it describes experiments we achieved at the suprasegmental levels with a system that outperform automatic correlation between prosodic labels and linguistic organization of a message to decode. Firstly an overview of the system is described along with the results of experiments carried out to determine which prosodic indexes are bestsuited for syntactic and rhythmycal prediction. 1 INTRODUCTION It is well known that prosodic structure and syntactic structure are not identical; neither are they unrelated. Knowing when and how the two correspond could help in the disambiguation of competing syntactic hypotheses in a speech understanding system. This practical reason explains the renewed interest in the use of prosody in ASR [8, 2, 4, 6]. This paper reports our experiments trying to ans...
Philippe Langlais
EUROSPEECH1
1995 Microprosodic study of isolated French word corpora
abstract
The present contribution aims at inventoring the microprosodic phenomena already abundantly described in the past for languages such as English and French (see [1] for a commented review of these studies) . We propose to measure their usefulness in an automatic processing and also to estimate the expected confidency of microprosodic "corrections" --- at least in an automatic way --- by studying each parameter separately and by presenting statistical distributions of the different phenomena measured on isolated french word corpora uttered by several speakers.
Philippe Langlais
EUROSPEECH1
1993 Integration of a prosodic component in an automatic speech recognition system
Philippe Langlais, Henri Meloni
EUROSPEECH1
1991 Analytical strategy for speaker identification
Jean-François Bonastre, Henri Meloni, Philippe Langlais
EUROSPEECH3