François Yvon

dblp:05/2701 · DBLP profile ↗
← Back
108ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-7972-7442ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 6 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
abstract
Pierre-Antoine Lequeu, Léo Labat, Laurène Cave, Gaël Lejeune, François Yvon, Benjamin Piwowarski. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pierre-Antoine Lequeu, Léo Labat, Laurène Cave, Gaël Lejeune, François Yvon, Benjamin Piwowarski
ACL (1)5
2026 Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
abstract
Conventional retrieval-augmented neural machine translation (RANMT) systems leverage bilingual corpora, e.g., translation memories (TMs). Yet, in many settings, monolingual corpora in the target language are often available. This work explores ways to take advantage of such resources by directly retrieving relevant target language segments, based on a source-side query. For this, we design improved cross-lingual retrieval systems, trained with both sentence level and word-level matching objectives. In our experiments with two RANMT architectures, we assess of such cross-lingual objectives in a controlled setting, reaching performances that match those of standard TM-based models. We also showcase our method on real-world settings, using much larger monolingual corpora, and observe strong improvements over both the baseline setting, and general-purpose cross-lingual retrievers.
Maxime Bouthors, Josep Maria Crego, Dakun Zhang, François Yvon
EAMT (1)4
2026 MetaDocEval: A Contrastive Framework for Evaluating Machine Translation Metrics at the Document-Level
abstract
Recent advances in neural machine translation (MT) have spurred increased interest in evaluating translations beyond the sentence level, making it possible to assess discourse-level phenomena related to coherence and consistency. While existing metrics can be applied to multi-sentence spans, it remains unclear whether their scores truly capture document-level quality. We introduce MetaDocEval, an automatic contrastive test set for evaluating MT metrics across three language pairs (en–fr, en–es, en–de) when applied at the document-level. It targets a range of discourse-level phenomena and potential problems linked to translation at the document level. To evaluate how metrics behave as a function of context size, we apply them under a sliding-window protocol, varying the input from single sentences up to full documents. Our experiments show that no current metric genuinely captures document-level coherence: reference-based metrics overfit lexical overlap, reference+source metrics gain little from added context, reference-free encoders show brief context sensitivity before degrading on longer spans, and LLM-based scorers collapse beyond short inputs. A key finding is that reference access can be actively harmful for detecting discourse-level errors. Using short windows (≈ 3 sentences) offers the best trade-off between discourse error detection and score dilution.
Nicolas Dahan, Rachel Bawden, François Yvon
EAMT (1)3
2026 The MaTOS Pipeline for the Translation of Scientific Abstracts on the HAL Platform
abstract
English dominates scientific publishing, which disadvantages researchers who are not native English speakers, especially those in the earlier stages of their careers. Being able to write and engage with scientific content written in their own language would clearly facilitate scientific production. The MaTOS project (Machine Translation for Open Science) seeks to reduce these barriers by developing machine translation tools for scientific documents in English and French. This article presents the design of the MaTOS pipeline for the HAL platform to automatically translate article abstracts, with author validation, to increase the number of bilingual abstracts on the platform. We also report preliminary experiments comparing translation of sentence, three-sentence chunks, and whole abstracts, evaluated using quality estimation metrics.
Panagiotis Tsolakis, Ziqian Peng, Laurent Romary, François Yvon, Rachel Bawden
EAMT (2)4
2026 Assessing the Political Fairness of Multilingual LLMs: A Case Study Based on a 21-Way Multiparallel EuroParl Dataset
abstract
The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fairness in multilingual translation. We systematically compare the translation quality of speeches in the European Parliament (EP), observing systematic differences with majority parties from left and right being better translated than outsider parties. This study is made possible by a new, 21-way multiparallel version of EuroParl, the parliamentary proceedings of the EP, which includes the political affiliations of each speaker. The dataset consists of 1.5M sentences for a total of 40M words and 249M characters. It covers three years, 1000+ speakers, 7 countries, 12 EU parties, 25 EU committees, and hundreds of national parties.
Paul Lerner, François Yvon
LREC2
2026 Biases in Translation: Assessing Opinion Distortion in Machine Translated Texts
Nazanin Shafiabadi, François Yvon
LREC2
2026 GlotWeb: Web Indexing for Minority Languages
abstract
International audience
Abdullah Al Sefat, Amir Hossein Kargaran, François Yvon, Hinrich Schütze
WWW3
2025 MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
abstract
International audience
Dávid Javorský, Ondrej Bojar, François Yvon
ACL (1)3
2025 Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu
abstract
In-context machine translation (MT) with large language models (LLMs) is a promising approach for low-resource MT, as it can readily take advantage of linguistic resources such as grammar books and dictionaries.Such resources are usually selectively integrated into the prompt so that LLMs can directly perform translation without any specific training, via their in-context learning capability (ICL).However, the relative importance of each type of resource, e.g., dictionary, grammar book, and retrieved parallel examples, is not entirely clear.To address this gap, this study systematically investigates how each resource and its quality affect the translation performance, with the Manchu language as our case study. To remove any prior knowledge of Manchu encoded in the LLM parameters and single out the effect of ICL, we also experiment with an enciphered version of Manchu texts.Our results indicate that high-quality dictionaries and good parallel examples are very helpful, while grammars hardly help.In a follow-up study, we showcase a promising application of in-context MT: parallel data augmentation as a way to bootstrap a conventional MT model. When monolingual data abound, generating synthetic parallel data through in-context MT offers a pathway to mitigate data scarcity and build effective and efficient low-resource neural MT systems.
Renhao Pei, Yihong Liu 0001, Peiqin Lin, François Yvon, Hinrich Schütze
ACL (1)4
2025 Towards the Machine Translation of Scientific Neologisms
abstract
Scientific research continually discovers and invents new concepts, which are then referred to by new terms, neologisms, or neonyms in this context. As the vast majority of publications are written in English, disseminating this new knowledge to the general public often requires translating these terms. However, by definition, no parallel data exist to provide such translations. Therefore, we propose to leverage term definitions as a useful source of information for the translation process. As we discuss, Large Language Models are well suited for this task and can benefit from in-context learning with co-hyponyms and terms sharing the same derivation paradigm. These models, however, are sensitive to the superficial and morphological similarity between source and target terms. Their predictions are also impacted by subword tokenization, especially for prefixed terms.
Paul Lerner, François Yvon
COLING2
2025 Unlike "Likely", "Unlike" is Unlikely: BPE-based Segmentation hurts Morphological Derivations in LLMs
abstract
Large Language Models (LLMs) rely on subword vocabularies to process and generate text. However, because subwords are marked as initial- or intra-word, we find that LLMs perform poorly at handling some types of affixations, which hinders their ability to generate novel (unobserved) word forms. The largest models trained on enough data can mitigate this tendency because their initial- and intra-word embeddings are aligned; in-context learning also helps when all examples are selected in a consistent way; but only morphological segmentation can achieve a near-perfect accuracy.
Paul Lerner, François Yvon
COLING2
2025 How Transliterations Improve Crosslingual Alignment
abstract
Recent studies have shown that post-aligning multilingual pretrained language models (mPLMs) using alignment objectives on both original and transliterated data can improve crosslingual alignment. This improvement further leads to better crosslingual transfer performance. However, it remains unclear how and why a better crosslingual alignment is achieved, as this technique only involves transliterations, and does not use any parallel data. This paper attempts to explicitly evaluate the crosslingual alignment and identify the key elements in transliteration-based approaches that contribute to better performance. For this, we train multiple models under varying setups for two pairs of related languages: (1) Polish and Ukrainian and (2) Hindi and Urdu. To assess alignment, we define four types of similarities based on sentence representations. Our experimental results show that adding transliterations alone improves the overall similarities, even for random sentence pairs. With the help of auxiliary transliteration-based alignment objectives, especially the contrastive objective, the model learns to distinguish matched from random pairs, leading to better crosslingual alignment. However, we also show that better alignment does not always yield better downstream performance, suggesting that further research is needed to clarify the connection between alignment and performance. The code implementation is based on https://github.com/cisnlp/Transliteration-PPA.
Yihong Liu 0001, Mingyang Wang 0003, Amir Hossein Kargaran, Ayyoob Imani, Orgest Xhelili, Haotian Ye, Chunlan Ma, François Yvon, Hinrich Schütze
COLING8
2025 An Interdisciplinary Approach to Human-Centered Machine Translation
abstract
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon
EMNLP20
2025 On Relation-Specific Neurons in Large Language Models
abstract
Yihong Liu, Runsheng Chen, Lea Hirlimann, Ahmad Dawar Hakimi, Mingyang Wang, Amir Hossein Kargaran, Sascha Rothe, François Yvon, Hinrich Schuetze. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yihong Liu 0001, Runsheng Chen, Lea Hirlimann, Ahmad Dawar Hakimi, Mingyang Wang 0003, Amir Hossein Kargaran, Sascha Rothe, François Yvon, Hinrich Schütze
EMNLP8
2025 MaTOS: Machine Translation for Open Science
abstract
This paper is a short presentation of MaTOS, a project focusing on the automatic translation of scholarly documents. Its main aims are threefold: (a) to develop resources (term lists and corpora) for high-quality machine translation; (b) to study methods for translating complete, structured documents in a cohesive and consistent manner; (c) to propose novel metrics to evaluate machine translation in technical domains. Publications and resources are available on the project web site: https://anr-matos.gihub.io.
Rachel Bawden, Maud Bénard, José Cornejo Cárcamo, Nicolas Dahan, Manon Delorme, Mathilde Huguin, Natalie Kübler, Paul Lerner, Alexandra Mestivier, Joachim Minder, Jean-François Nominé, Ziqian Peng, Laurent Romary, Panagiotis Tsolakis, Lichao Zhu, François Yvon
MTSummit (2)16
2025 Investigating Length Issues in Document-level Machine Translation
abstract
Transformer architectures are increasingly effective at processing and generating very long chunks of texts, opening new perspectives for document-level machine translation (MT). In this work, we challenge the ability of MT systems to handle texts comprising up to several thousands of tokens. We design and implement a new approach designed to precisely measure the effect of length increments on MT outputs. Our experiments with two representative architectures unambiguously show that (a) translation performance decreases with the length of the input text; (b) the position of sentences within the document matters and translation quality is higher for sentences occurring earlier in a document. We further show that manipulating the distribution of document lengths and of positional embeddings only marginally mitigates such problems. Our results suggest that even though document-level MT is computationally feasible, it does not yet match the performance of sentence-based MT.
Ziqian Peng, Rachel Bawden, François Yvon
MTSummit (1)3
2024 GlotScript: A Resource and Tool for Low Resource Writing System Identification
abstract
We present GlotScript, an open resource and tool for low resource writing system identification. GlotScript-R is a resource that provides the attested writing systems for more than 7,000 languages. It is compiled by aggregating information from existing writing system resources. GlotScript-T is a writing system identification tool that covers all 161 Unicode 15.0 scripts. For an input text, it returns its script distribution where scripts are identified by ISO 15924 codes. We also present two use cases for GlotScript. First, we demonstrate that GlotScript can help cleaning multilingual corpora such as mC4 and OSCAR. Second, we analyze the tokenization of a number of language models such as GPT-4 using GlotScript and provide insights on the coverage of low resource scripts and languages by each language model. We hope that GlotScript will become a useful resource for work on low resource languages in the NLP community. GlotScript-R and GlotScript-T are available at https://github.com/cisnlp/GlotScript.
Amir Hossein Kargaran, François Yvon, Hinrich Schütze
LREC/COLING2
2024 Translate your Own: a Post-Editing Experiment in the NLP domain
abstract
The improvements in neural machine translation make translation and post-editing pipelines ever more effective for a wider range of applications. In this paper, we evaluate the effectiveness of such a pipeline for the translation of scientific documents (limited here to article abstracts). Using a dedicated interface, we collect, then analyse the post-edits of approximately 350 abstracts (English→French) in the Natural Language Processing domain for two groups of post-editors: domain experts (academics encouraged to post-edit their own articles) on the one hand and trained translators on the other. Our results confirm that such pipelines can be effective, at least for high-resource language pairs. They also highlight the difference in the post-editing strategy of the two subgroups. Finally, they suggest that working on term translation is the most pressing issue to improve fully automatic translations, but that in a post-editing setup, other error types can be equally annoying for post-editors.
Rachel Bawden, Ziqian Peng, Maud Bénard, Éric Villemonte de la Clergerie, Raphaël Esamotunu, Mathilde Huguin, Natalie Kübler, Alexandra Mestivier, Mona Michelot, Laurent Romary, Lichao Zhu, François Yvon
EAMT (1)12
2024 GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
abstract
The need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models. Most available corpora have sufficient data only for languages with large dominant communities. However, there is no corpus available that (i) covers a wide range of minority languages; (ii) is generated by an open-source reproducible pipeline; and (iii) is rigorously cleaned from noise, making it trustworthy to use. We present GlotCC, a clean, document-level, 2TB general domain corpus derived from CommonCrawl, covering more than 1000 languages. We make GlotCC and the system used to generate it— including the pipeline, language identification model, and filters—available to the research community.Corpus v. 1.0 https://huggingface.co/datasets/cis-lmu/GlotCC-v1Pipeline v. 3.0 https://github.com/cisnlp/GlotCC
Amir Hossein Kargaran, François Yvon, Hinrich Schütze
NeurIPS2
2024 Translating scientific abstracts in the bio-medical domain with structure-aware models
Sadaf Abdul-Rauf, François Yvon
Comput. Speech Lang.2
2023 Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
abstract
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André Martins, François Yvon, Hinrich Schütze. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André F. T. Martins, François Yvon, Hinrich Schütze
ACL (1)10
2023 Integrating Translation Memories into Non-Autoregressive Machine Translation
abstract
Non-autoregressive machine translation (NAT) has recently made great progress.However, most works to date have focused on standard translation tasks, even though some edit-based NAT models, such as the Levenshtein Transformer (LevT), seem well suited to translate with a Translation Memory (TM).This is the scenario considered here.We first analyze the vanilla LevT model and explain why it does not do well in this setting.We then propose a new variant, TM-LevT, and show how to effectively train this model.By modifying the data presentation and introducing an extra deletion operation, we obtain performance that are on par with an autoregressive approach, while reducing the decoding load.We also show that incorporating TMs during training dispenses to use knowledge distillation, a well-known trick used to mitigate the multimodality issue.
Jitao Xu 0001, Josep Maria Crego, François Yvon
EACL3
2023 Investigating the Translation Performance of a Large Multilingual Language Model: the Case of BLOOM
abstract
The NLP community recently saw the release of a new large open-access multilingual language model, BLOOM (BigScience et al., 2022) covering 46 languages. We focus on BLOOM’s multilingual ability by evaluating its machine translation performance across several datasets (WMT, Flores-101 and DiaBLa) and language pairs (high- and low-resourced). Our results show that 0-shot performance suffers from overgeneration and generating in the wrong language, but this is greatly improved in the few-shot setting, with very good results for a number of language pairs. We study several aspects including prompt design, model sizes, cross-lingual transfer and the use of discursive context.
Rachel Bawden, François Yvon
EAMT2
2023 Towards Example-Based NMT with Multi-Levenshtein Transformers
abstract
Retrieval-Augmented Machine Translation (RAMT) is attracting growing attention.This is because RAMT not only improves translation metrics, but is also assumed to implement some form of domain adaptation.In this contribution, we study another salient trait of RAMT, its ability to make translation decisions more transparent by allowing users to go back to examples that contributed to these decisions.For this, we propose a novel architecture aiming to increase this transparency.This model adapts a retrieval-augmented version of the Levenshtein Transformer and makes it amenable to simultaneously edit multiple fuzzy matches found in memory.We discuss how to perform training and inference in this model, based on multiway alignment algorithms and imitation learning.Our experiments show that editing several examples positively impacts translation scores, notably increasing the number of target spans that are copied from existing instances.
Maxime Bouthors, Josep Maria Crego, François Yvon
EMNLP3
2023 Structural generalization in COGS: Supertagging is (almost) all you need
abstract
In many Natural Language Processing applications, neural networks have been found to fail to generalize on out-of-distribution examples.In particular, several recent semantic parsing datasets have put forward important limitations of neural networks in cases where compositional generalization is required.In this work, we extend a neural graph-based semantic parsing framework in several ways to alleviate this issue.Notably, we propose: (1) the introduction of a supertagging step with valency constraints, expressed as an integer linear program;(2) a reduction of the graph prediction problem to the maximum matching problem; (3) the design of an incremental early-stopping training strategy to prevent overfitting.Experimentally, our approach significantly improves results on examples that require structural generalization in the COGS dataset, a known challenging benchmark for compositional generalization.Overall, our results confirm that structural constraints are important for generalization in semantic parsing.
Alban Petit, Caio F. Corro, François Yvon
EMNLP3
2022 Weakly Supervised Word Segmentation for Computational Language Documentation
abstract
Word and morpheme segmentation are fundamental steps of language documentation as they allow to discover lexical units in a language for which the lexicon is unknown.However, in most language documentation scenarios, linguists do not start from a blank page: they may already have a pre-existing dictionary or have initiated manual segmentation of a small part of their data.This paper studies how such a weak supervision can be taken advantage of in Bayesian non-parametric models of segmentation.Our experiments on two very low resource languages (Mboshi and Japhug), whose documentation is still in progress, show that weak supervision can be beneficial to the segmentation quality.In addition, we investigate an incremental learning scenario where manual segmentations are provided in a sequential manner.This work opens the way for interactive annotation tools for documentary linguists.
Shu Okabe, Laurent Besacier, François Yvon
ACL (1)3
2022 Multi-Domain Adaptation in Neural Machine Translation with Dynamic Sampling Strategies
abstract
Building effective Neural Machine Translation models often implies accommodating diverse sets of heterogeneous data so as to optimize performance for the domain(s) of interest. Such multi-source / multi-domain adaptation problems are typically approached through instance selection or reweighting strategies, based on a static assessment of the relevance of training instances with respect to the task at hand. In this paper, we study dynamic data selection strategies that are able to automatically re-evaluate the usefulness of data samples and to evolve a data selection policy in the course of training. Based on the results of multiple experiments, we show that such methods constitute a generic framework to automatically and effectively handle a variety of real-world situations, from multi-source domain adaptation to multi-domain learning and unsupervised domain adaptation.
Minh Quang Pham, Josep Maria Crego, François Yvon
EAMT3
2022 Bilingual Synchronization: Restoring Translational Relationships with Editing Operations
abstract
Machine Translation (MT) is usually viewed as a one-shot process that generates the target language equivalent of some source text from scratch.We consider here a more general setting which assumes an initial target sequence, that must be transformed into a valid translation of the source, thereby restoring parallelism between source and target.For this bilingual synchronization task, we consider several architectures (both autoregressive and non-autoregressive) and training regimes, and experiment with multiple practical settings such as simulated interactive MT, translating with Translation Memory (TM) and TM cleaning.Our results suggest that one single generic edit-based system, once fine-tuned, can compare with, or even outperform, dedicated systems specifically trained for these tasks.
Jitao Xu 0001, Josep Maria Crego, François Yvon
EMNLP3
2022 Graph-Based Multilingual Label Propagation for Low-Resource Part-of-Speech Tagging
abstract
Part-of-Speech (POS) tagging is an important component of the NLP pipeline, but many lowresource languages lack labeled data for training.An established method for training a POS tagger in such a scenario is to create a labeled training set by transferring from high-resource languages.In this paper, we propose a novel method for transferring labels from multiple high-resource source to low-resource target languages.We formalize POS tag projection as graph-based label propagation.Given translations of a sentence in multiple languages, we create a graph with words as nodes and alignment links as edges by aligning words for all language pairs.We then propagate node labels from source to target using a Graph Neural Network augmented with transformer layers.We show that our propagation creates training sets that allow us to train POS taggers for a diverse set of languages.When combined with enhanced contextualized embeddings, our method achieves a new state-ofthe-art for unsupervised POS tagging of low resource languages.
Ayyoob Imani, Silvia Severini, Masoud Jalili Sabet, François Yvon, Hinrich Schütze
EMNLP4
2022 Evaluating Subtitle Segmentation for End-to-end Generation Systems
abstract
Subtitles appear on screen as short pieces of text, segmented based on formal constraints (length) and syntactic/semantic criteria. Subtitle segmentation can be evaluated with sequence segmentation metrics against a human reference. However, standard segmentation metrics cannot be applied when systems generate outputs different than the reference, e.g. with end-to-end subtitling systems. In this paper, we study ways to conduct reference-based evaluations of segmentation accuracy irrespective of the textual content. We first conduct a systematic analysis of existing metrics for evaluating subtitle segmentation. We then introduce Sigma, a Subtitle Segmentation Score derived from an approximate upper-bound of BLEU on segmentation boundaries, which allows us to disentangle the effect of good segmentation from text quality. To compare Sigma with existing metrics, we further propose a boundary projection method from imperfect hypotheses to the true reference. Results show that all metrics are able to reward high quality output but for similar outputs system ranking depends on each metric’s sensitivity to error type. Our thorough analyses suggest Sigma is a promising segmentation candidate but its reliability over other segmentation metrics remains to be validated through correlations with human judgements.
Alina Karakanta, François Buet, Mauro Cettolo, François Yvon
LREC4
2021 One Source, Two Targets: Challenges and Rewards of Dual Decoding
abstract
Machine translation is generally understood as generating one target text from an input source document.In this paper, we consider a stronger requirement: to jointly generate two texts so that each output side effectively depends on the other.As we discuss, such a device serves several practical purposes, from multi-target machine translation to the generation of controlled variations of the target text.We present an analysis of possible implementations of dual decoding, and experiment with four applications.Viewing the problem from multiple angles allows us to better highlight the challenges of dual decoding and to also thoroughly analyze the benefits of generating matched, rather than independent, translations.
Jitao Xu 0001, François Yvon
EMNLP (1)2
2021 Graph Algorithms for Multiparallel Word Alignment
abstract
Ayyoob ImaniGooghari, Masoud Jalili Sabet, Lutfi Kerem Senel, Philipp Dufter, François Yvon, Hinrich Schütze. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Ayyoob Imani, Masoud Jalili Sabet, Lutfi Kerem Senel, Philipp Dufter, François Yvon, Hinrich Schütze
EMNLP (1)5
2021 Toward Genre Adapted Closed Captioning
abstract
International audience
François Buet, François Yvon
Interspeech2
2021 Optimizing Word Alignments with Better Subword Tokenization
abstract
Word alignment identify translational correspondences between words in a parallel sentence pair and are used and for example and to train statistical machine translation and learn bilingual dictionaries or to perform quality estimation. Subword tokenization has become a standard preprocessing step for a large number of applications and notably for state-of-the-art open vocabulary machine translation systems. In this paper and we thoroughly study how this preprocessing step interacts with the word alignment task and propose several tokenization strategies to obtain well-segmented parallel corpora. Using these new techniques and we were able to improve baseline word-based alignment models for six language pairs.
Anh Khoa Ngo Ho, François Yvon
MTSummit (1)2
2021 Revisiting Multi-Domain Machine Translation
abstract
When building machine translation systems, one often needs to make the best out of heterogeneous sets of parallel data in training, and to robustly handle inputs from unexpected domains in testing. This multi-domain scenario has attracted a lot of recent work that fall under the general umbrella of transfer learning. In this study, we revisit multi-domain machine translation, with the aim to formulate the motivations for developing such systems and the associated expectations with respect to performance. Our experiments with a large sample of multi-domain systems show that most of these expectations are hardly met and suggest that further work is needed to better analyze the current behaviour of multi-domain systems and to make them fully hold their promises.
Minh Quang Pham, Josep Maria Crego, François Yvon
Trans. Assoc. Comput. Linguistics3
2020 The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
abstract
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions.
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
LREC47
2019 Quality Estimation for Machine Translation
abstract
Many natural language processing tasks aim to generate a human readable text (or human audible speech) in response to some input: Machine translation (MT) generates a target translation of a source input; document summarization generates a shortened version of its input document(s); text generation converts a formal representation into an utterance or a document, and so on. For such tasks, the automatic evaluation of the system’s performance is often performed by comparison to a reference output, deemed representative of human-level performance.Evaluation of MT is a typical illustration of this approach and relies on metrics such as BLEU (Papineni et al. 2002), Translation Edit Rate (TER) (Snover et al. 2006), or METEOR (Banerjee and Lavie 2005) implementing various string comparison routines between the system output and the corresponding reference(s). This strategy has the merit of making evaluation fully automatic and reproducible. Preparing human translation references is, however, a costly process, which requires highly trained experts; it is also prone to much variability and subjectivity. This implies that the failure to match the reference does not necessarily entail an error of the system. Reference-based evaluations are also considered too crude for many language pairs and tend to only evaluate the system’s ability to reproduce one specific human annotation. Organizers of shared tasks in MT have therefore abandoned reference-based metrics to compare systems and resort to human judgments (Callison-Burch et al. 2008).The book by Specia, Scarton, and Paetzold surveys an alternative approach to automatic evaluation of MT, Quality Estimation (QE). In essence, QE aims to move away from human references and to evaluate a generated text based only on automatically computed features. QE was initially proposed in the context of Automatic Speech Recognition (ASR) systems, an area where much of the foundational work has been performed (Jiang 2005). As explained in the introductory chapter, QE for MT also has many applications and has emerged in the last decade as a very active subfield of MT, with its own evaluation campaigns and metrics. In a nutshell, QE predicts the quality score or quality label of some target fragment, produced in response to some source text. Assuming that texts annotated with their quality level are available, QE is usually cast as a supervised machine learning task. QE for MT needs to simultaneously take two dimensions into account: (a) Is the proposed output appropriate for the input data? and (b) Is the generated text grammatically correct? Dimension (a) is usually associated with the concept of adequacy with respect to the input signal, whereas dimension (b) is associated with the correctness or fluency of the output target fragment. Correctness can be defined at various levels of granularity, depending on the size of the output chunk: Smaller chunks need to contain the right words or to be syntactically correct; larger chunks, in addition, need to contain valid discourse relationships and to display lexical cohesiveness.Specia and her colleagues address all these issues, and many more, in this book, where they survey QE for MT from an increasingly larger perspective: first words and phrases, then sentences, and finally complete documents. The last chapter widens the perspective, surveying QE for related NLP tasks, such as text simplification and generation.After the short introduction, the second chapter covers QE for MT at the subsentential level, adopting a fixed outline that will also be used in the subsequent chapters. It successively presents the main applications, the targeted labels, the features, evaluation metrics, and some selected state-of-the-art approaches. This task nicely illustrates the difference between reference-based evaluation and QE: Evaluating the correctness of a hypothetical isolated word or phrase with respect to a reference translation would only be possible in rare cases of technical terms or named entities. Yet, it is possible to devise useful quality indicators at the word or phrase levels using, for instance, traces of the translation post-editor’s work. Post-edition produces an annotation of which words are “correct” and kept in the revised version, and which are “wrong” and need to be discarded, replaced, or moved. Building QE systems using such labels amounts to training a sequence labeling system, for which many computational architectures, many features, and obvious evaluation metrics (e.g., label accuracy) are readily available. As discussed at length in the last section, more complex views of this task have proven useful to establish state-of-the-art results. For instance, Automatic Post-Edition–based QE predicts quality labels using an automatic post-edition of the output.Chapter 3 is devoted to sentence-level QE, and follows the same outline as the previous chapter. Sentence-level QE is useful for many downstream applications, as evidenced by the large choice of labels and scores that can be targeted. It is still possible to use discrete labels such as “usable/useless sentence,” or richer measures of the post-editing effort. Considering complete sentences makes it possible to also predict continuous values, such as the TER score, or the post-editing time. For each of these settings, off-the-shelf learning tools are available, as well as appropriate evaluation metrics. Note that such tasks are hard, probably as hard as MT itself: Indeed, any system successful for this task could effectively re-rank the output of an MT system and readily yield improved automatic translations. Working at the sentence level also encourages the use of very diverse sets of features, which the authors have taken the burden to list exhaustively: In addition to the obvious syntactic features for evaluating fluency, the authors also describe features that detect translation difficulties in the source (complexity indicators), extract information from the internals of the MT (confidence indicators), and globally evaluate the source and target alignment (translation features).Chapter 4 covers document-level QE, which may provide users with a useful score, both for gisting applications and also in post-edition or translation revision scenarios. This is a new and difficult topic, as acknowledged by the small number of studies on these issues. One of the main challenges concerns the definition of metrics that could be used to define objective training functions: Possible sources of inspiration come from automatic text comprehension metrics or variants of post-edition effort. Document level QE also needs features that measure document-level properties of a text such as discourse-structure appropriateness or lexical cohesion, an area where techniques borrowed from the statistical and neural text-mining literature such as topic models can be very useful. Much, however, remains to be done in this domain.QE for other language generation applications is discussed in Chapter 5, which notably covers Text Simplification, Automatic Text Summarization, Grammatical Error Correction, and Natural Language Generation. With the exception of QE for ASR, which is already well documented, developing QE for these tasks is relatively new, as acknowledged by the small amount of prior work. Considering multiple applications poses new questions regarding quality measures and techniques that can be used to approximate them. For each task, the authors follow the same organization as in the previous chapters, which altogether puts considerable emphasis on the descriptions of features, as many features are useful in more than one task. This chapter is nonetheless very informative regarding the development of QE methods for systems generating a text output.The concluding chapter discusses some directions for future research, in the light of the advent of a new generation of neural machine translation systems: On the one hand, neural machine translation outputs are quite different from statistical machine translation outputs, suggesting that QE methods need to evolve and take this difference into account; on the other hand, state-of–the-art QE systems are using neural components and need to be improved in their ability to train with scarce annotated data and to adapt to new domains and language pairs. This concluding section also plays the role of an appendix, as it includes a very complete list of existing resources and software packages for quality estimation of MT output. This supplementary material, and the bibliography that follows, will be of great help for readers willing to re-implement the systems described in this book or to develop new ideas.In summary, this book discusses problems whose significance (at least for MT) is increasing with the general quality of machine translation output: With MT becoming more widespread and useful, it becomes necessary to inform users with indications of the cases where MT can be relied on—or not. Due to the invaluable expertise of the authors, who have been directly involved in the development of this field, the coverage of the recent literature on QE for MT (and on QE in general) is extremely thorough, which makes this book a must-read for any NLP practitioner willing to develop QE systems. The reading of this quite technical book, however, requires a solid working knowledge of machine translation and previous exposure to both automatic language analysis and to machine learning methodologies. For the sake of accessibility to a larger audience, adding definitions of basic concepts of the domain and a glossary should be considered in the second edition of the book.
François Yvon
Comput. Linguistics1
2018 Unsupervised Learning of Word Segmentation: Does Tone Matter?
Pierre Godard, Kevin Löser, Alexandre Allauzen, Laurent Besacier, François Yvon
CICLing (1)5
2018 Quantifying training challenges of dependency parsers
abstract
Not all dependencies are equal when training a dependency parser: some are straightforward enough to be learned with only a sample of data, others embed more complexity. This work introduces a series of metrics to quantify those differences, and thereby to expose the shortcomings of various parsing algorithms and strategies. Apart from a more thorough comparison of parsing systems, these new tools also prove useful for characterizing the information conveyed by cross-lingual parsers, in a quantitative but still interpretable way.
Lauriane Aufrant, Guillaume Wisniewski, François Yvon
COLING3
2018 Fixing Translation Divergences in Parallel Corpora for Neural MT
abstract
Corpus-based approaches to machine translation rely on the availability of clean parallel corpora.Such resources are scarce, and because of the automatic processes involved in their preparation, they are often noisy.This paper describes an unsupervised method for detecting translation divergences in parallel sentences.We rely on a neural network that computes cross-lingual sentence similarity scores, which are then used to effectively filter out divergent translations.Furthermore, similarity scores predicted by the network are used to identify and fix some partial divergences, yielding additional parallel segments.We evaluate these methods for English-French and English-German machine translation tasks, and show that using filtered/corrected corpora actually improves MT performance.
Minh Quang Pham, Josep Maria Crego, Jean Senellart, François Yvon
EMNLP4
2018 Bayesian Models for Unit Discovery on a Very Low Resource Language
abstract
Developing speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the I-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus.
Lucas Ondel Yang, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Lukás Burget, François Yvon, Sanjeev Khudanpur
ICASSP9
2018 Unsupervised Word Segmentation from Speech with Attention
abstract
International audience
Pierre Godard, Marcely Zanon Boito, Lucas Ondel Yang, Alexandre Berard, François Yvon, Aline Villavicencio, Laurent Besacier
INTERSPEECH5
2018 A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noël Kouarata, Lori Lamel, Hélène Bonneau-Maynard, Markus Müller 0001, Annie Rialland, Sebastian Stüker, François Yvon, Marcely Zanon Boito
LREC13
2018 Reassessing the proper place of man and machine in translation: a pre-translation scenario
Julia Ive, Aurélien Max, François Yvon
Mach. Transl.3
2017 Learning the Structure of Variable-Order CRFs: a finite-state perspective
abstract
The computational complexity of linearchain Conditional Random Fields (CRFs) makes it difficult to deal with very large label sets and long range dependencies.Such situations are not rare and arise when dealing with morphologically rich languages or joint labelling tasks.We extend here recent proposals to consider variable order CRFs.Using an effective finitestate representation of variable-length dependencies, we propose new ways to perform feature selection at large scale and report experimental results where we outperform strong baselines on a tagging task.
Thomas Lavergne, François Yvon
EMNLP2
2017 A comparison of discriminative training criteria for continuous space translation models
Alexandre Allauzen, Quoc-Khanh Do, François Yvon
Mach. Transl.3
2016 Zero-resource Dependency Parsing: Boosting Delexicalized Cross-lingual Transfer with Linguistic Knowledge
abstract
This paper studies cross-lingual transfer for dependency parsing, focusing on very low-resource settings where delexicalized transfer is the only fully automatic option. We show how to boost parsing performance by rewriting the source sentences so as to better match the linguistic regularities of the target language. We contrast a data-driven approach with an approach relying on linguistically motivated rules automatically extracted from the World Atlas of Language Structures. Our findings are backed up by experiments involving 40 languages. They show that both approaches greatly outperform the baseline, the knowledge-driven method yielding the best accuracies, with average improvements of +2.9 UAS, and up to +90 UAS (absolute) on some frequent PoS configurations.
Lauriane Aufrant, Guillaume Wisniewski, François Yvon
COLING3
2016 Parallel Sentence Compression
abstract
Sentence compression is a way to perform text simplification and is usually handled in a monolingual setting. In this paper, we study ways to extend sentence compression in a bilingual context, where the goal is to obtain parallel compressions of parallel sentences. This can be beneficial for a series of multilingual natural language processing (NLP) tasks. We compare two ways to take bilingual information into account when compressing parallel sentences. Their efficiency is contrasted on a parallel corpus of News articles.
Julia Ive, François Yvon
COLING2
2016 Preliminary Experiments on Unsupervised Word Discovery in Mboshi
abstract
International audience
Pierre Godard, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Laurent Besacier, Hélène Bonneau-Maynard, Guy-Noël Kouarata, Kevin Löser, Annie Rialland, François Yvon
INTERSPEECH10
2016 Cross-lingual and Supervised Models for Morphosyntactic Annotation: a Comparison on Romanian
Lauriane Aufrant, Guillaume Wisniewski, François Yvon
LREC3
2016 Novel elicitation and annotation schemes for sentential and sub-sentential alignments of bitexts
Yong Xu 0006, François Yvon
LREC2
2016 Frustratingly Easy Cross-Lingual Transfer for Transition-Based Dependency Parsing
abstract
Ophélie Lacroix, Lauriane Aufrant, Guillaume Wisniewski, François Yvon. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Ophélie Lacroix, Lauriane Aufrant, Guillaume Wisniewski, François Yvon
HLT-NAACL4
2015 A Discriminative Training Procedure for Continuous Translation Models
abstract
Continuous-space translation models have recently emerged as extremely powerful ways to boost the performance of existing translation systems.A simple, yet effective way to integrate such models in inference is to use them in an N -best rescoring step.In this paper, we focus on this scenario and show that the performance gains in rescoring can be greatly increased when the neural network is trained jointly with all the other model parameters, using an appropriate objective function.Our approach is validated on two domains, where it outperforms strong baselines.
Quoc-Khanh Do, Alexandre Allauzen, François Yvon
EMNLP3
2015 Structured prediction for speaker identification in TV series
abstract
International audience
Elena Knyazeva, Guillaume Wisniewski, Hervé Bredin, François Yvon
INTERSPEECH4
2014 Cross-Lingual Part-of-Speech Tagging through Ambiguous Learning
abstract
International audience
Guillaume Wisniewski, Nicolas Pécheux, Souhir Gahbiche-Braham, François Yvon
EMNLP4
2014 Rule-based Reordering Space in Statistical Machine Translation
Nicolas Pécheux, Alexandre Allauzen, François Yvon
LREC3
2014 A Corpus of Machine Translation Errors Extracted from Translation Students Exercises
Guillaume Wisniewski, Natalie Kübler, François Yvon
LREC3
2014 Maximum-entropy word alignment and posterior-based phrase extraction for machine translation
Nadi Tomeh, Alexandre Allauzen, François Yvon
Mach. Transl.3
2013 Discriminative training of a phoneme confusion model for a dynamic lexicon in ASR
abstract
International audience
Panagiota Karanasou, François Yvon, Thomas Lavergne, Lori Lamel
INTERSPEECH2
2013 Structure learning in hidden conditional random fields for grapheme-to-phoneme conversion
abstract
Accurate grapheme-to-phoneme (g2p) conversion is needed for several speech processing applications, such as automatic speech synthesis and recognition.For some languages, notably English, improvements of g2p systems are very slow, due to the intricacy of the associations between letter and sounds.In recent years, several improvements have been obtained either by using variable-length associations in generative models (jointn-grams), or by recasting the problem as a conventional sequence labeling task, enabling to integrate rich dependencies in discriminative models.In this paper, we consider several ways to reconciliate these two approaches.Introducing hidden variable-length alignments through latent variables, our Hidden Conditional Random Field (HCRF) models are able to produce comparative performance compared to strong generative and discriminative models on the CELEX database.
Patrick Lehnen, Alexandre Allauzen, Thomas Lavergne, François Yvon, Stefan Hahn, Hermann Ney
INTERSPEECH4
2013 Design and Analysis of a Large Corpus of Post-Edited Translations: Quality Estimation, Failure Analysis and the Variability of Post-Edition
Guillaume Wisniewski, Anil Kumar Singh 0001, Natalia Segal, François Yvon
MTSummit4
2013 Generalizing sampling-based multilingual alignment
Adrien Lardilleux, François Yvon, Yves Lepage
Mach. Transl.2
2013 Quality estimation for machine translation: some lessons learned
Guillaume Wisniewski, Anil Kumar Singh 0001, François Yvon
Mach. Transl.3
2013 Oracle decoding as a new way to analyze phrase-based machine translation
Guillaume Wisniewski, François Yvon
Mach. Transl.2
2013 Structured Output Layer Neural Network Language Models for Speech Recognition
abstract
This paper extends a novel neural network language model (NNLM) which relies on word clustering to structure the output vocabulary: Structured OUtput Layer (SOUL) NNLM. This model is able to handle arbitrarily-sized vocabularies, hence dispensing with the need for shortlists that are commonly used in NNLMs. Several softmax layers replace the standard output layer in this model. The output structure depends on the word clustering which is based on the continuous word representation determined by the NNLM. Mandarin and Arabic data are used to evaluate the SOUL NNLM accuracy via speech-to-text experiments. Well tuned speech-to-text systems (with error rates around 10%) serve as the baselines. The SOUL model achieves consistent improvements over a classical shortlist NNLM both in terms of perplexity and recognition accuracy for these two languages that are quite different in terms of their internal structure and recognition vocabulary size. An enhanced training scheme is proposed that allows more data to be used at each training iteration of the neural network.
Hai Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon
IEEE Trans. Speech Audio Process.5
2012 Computing Lattice BLEU Oracle Scores for Machine Translation
Artem Sokolov 0001, Guillaume Wisniewski, François Yvon
EACL3
2012 Hierarchical Sub-sentential Alignment with Anymalign
Adrien Lardilleux, François Yvon, Yves Lepage
EAMT2
2012 Joint Segmentation and POS Tagging for Arabic Using a CRF-based Classifier
Souhir Gahbiche-Braham, Hélène Bonneau-Maynard, Thomas Lavergne, François Yvon
LREC4
2012 Continuous Space Translation Models with Neural Networks
Hai Son Le, Alexandre Allauzen, François Yvon
HLT-NAACL3
2011 Minimum Error Rate Training Semiring
Artem Sokolov 0001, François Yvon
EAMT2
2011 Discriminative Weighted Alignment Matrices For Statistical Machine Translation
Nadi Tomeh, Alexandre Allauzen, François Yvon
EAMT3
2011 Structured Output Layer neural network language model
abstract
This paper introduces a new neural network language model (NNLM) based on word clustering to structure the output vocabulary: Structured Output Layer NNLM. This model is able to handle vocabularies of arbitrary size, hence dispensing with the design of short-lists that are commonly used in NNLMs. Several softmax layers replace the standard output layer in this model. The output structure depends on the word clustering which uses the continuous word representation induced by a NNLM. The GALE Mandarin data was used to carry out the speech-to-text experiments and evaluate the NNLMs. On this data the well tuned baseline system has a character error rate under 10%. Our model achieves consistent improvements over the combination of an n-gram model and classical short-list NNLMs both in terms of perplexity and recognition accuracy.
Hai Son Le, Ilya Oparin, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon
ICASSP5
2011 Large Vocabulary SOUL Neural Network Language Models
abstract
International audience
Hai Son Le, Ilya Oparin, Abdelkhalek Messaoudi, Alexandre Allauzen, Jean-Luc Gauvain, François Yvon
INTERSPEECH6
2011 Text segmentation: A topic modeling perspective
Hemant Misra, François Yvon, Olivier Cappé, Joemon M. Jose
Inf. Process. Manag.2
2010 Practical Very Large Scale CRFs
Thomas Lavergne, Olivier Cappé, François Yvon
ACL3
2010 Local lexical adaptation in Machine Translation through triangulation: SMT helping SMT
Josep Maria Crego, Aurélien Max, François Yvon
COLING3
2010 Training Continuous Space Language Models: Some Practical Issues
Hai Son Le, Alexandre Allauzen, Guillaume Wisniewski, François Yvon
EMNLP4
2010 Assessing Phrase-Based Translation Models with Oracle Decoding
Guillaume Wisniewski, Alexandre Allauzen, François Yvon
EMNLP3
2010 Contrastive Lexical Evaluation of Machine Translation
Aurélien Max, Josep Maria Crego, François Yvon
LREC3
2010 Factored bilingual n-gram language models for statistical machine translation
Josep Maria Crego, François Yvon
Mach. Transl.2
2010 Rewriting the orthography of SMS messages
abstract
Abstract Electronic written texts used in computer-mediated interactions (emails, blogs, chats, and the like) contain significant deviations from the norm of the language. This paper presents the detail of a system aiming at normalizing the orthography of French SMS messages: after discussing the linguistic peculiarities of these messages and possible approaches to their automatic normalization, we present, compare, and evaluate various instanciations of a normalization device based on weighted finite-state transducers. These experiments show that using an intermediate phonemic representation and training, our system outperforms an alternative normalization system based on phrase-based statistical machine translation techniques.
François Yvon
Nat. Lang. Eng.1
2009 Text segmentation via topic modeling: an analytical study
abstract
In this paper, the task of text segmentation is approached from a topic modeling perspective. We investigate the use of latent Dirichlet allocation (LDA) topic model to segment a text into semantically coherent segments. A major benefit of the proposed approach is that along with the segment boundaries, it outputs the topic distribution associated with each segment. This information is of potential use in applications like segment retrieval and discourse analysis. The new approach outperforms a standard baseline method and yields significantly better performance than most of the available unsupervised methods on a benchmark dataset.
Hemant Misra, François Yvon, Joemon M. Jose, Olivier Cappé
CIKM2
2009 Improvements in Analogical Learning: Application to Translating Multi-Terms of the Medical Domain
Philippe Langlais, François Yvon, Pierre Zweigenbaum
EACL2
2009 Gappy Translation Units under Left-to-Right SMT Decoding
Josep Maria Crego, François Yvon
EAMT2
2008 Normalizing SMS: are Two Metaphors Better than One ?
Catherine Kobus, François Yvon, Géraldine Damnati
COLING2
2008 Robust Similarity Measures for Named Entities Matching
Erwan Moreau, François Yvon, Olivier Cappé
COLING2
2008 Using LDA to detect semantically incoherent documents
Hemant Misra, Olivier Cappé, François Yvon
CoNLL3
2008 The asymptotics of semi-supervised learning in discriminative probabilistic models
abstract
Semi-supervised learning aims at taking advantage of unlabeled data to improve the efficiency of supervised learning procedures. For discriminative models however, this is a challenging task. In this contribution, we introduce an original methodology for using unlabeled data through the design of a simple semi-supervised objective function. We prove that the corresponding semi-supervised estimator is asymptotically optimal. The practical consequences of this result are discussed for the case of the logistic regression model.
Nataliya Sokolovska, Olivier Cappé, François Yvon
ICML3
2007 Approaches for adaptive database reduction for text-to-speech synthesis
Aleksandra Krul, Géraldine Damnati, François Yvon, Cédric Boidin, Thierry Moudenc
INTERSPEECH3
2007 Optimization on decoding graphs by discriminative training
abstract
Les trois sources principalement utilisées en reconnaissance vocale automatique (Automatic Speech Recognition, ASR) sont les modèles acoustiques, le dictionnaire et le modèle de langage. Elles sont habituellement conçues et optimisées de manière séparée. Notre travail a proposé une méthodologie, à savoir un apprentissage discriminant sur un grand graphe de décodage, pour optimiser conjointement les paramètres de ces différents modèles, en se fondant sur l'intégration des ressources dans un transducteur fini pondéré dont les poids des transitions sont estimés par de manière discriminante. Dans ce cadre d'apprentissage, les paramètres du modèle sont ajustés itérativement de façon à réduire progressivement le nombre d'erreurs de retranscription commises par le système. Nous considérons en particulier dans ce travail de mettre en oeuvre ce cadre d'apprentissage pour une tâche de reconnaissance à grand vocabulaire : la transcription automatique des nouvelles de la radio française. Nous proposons plusieurs techniques pour un accélérer les algorithmes de décodage, afin de rendre ce type d'apprentissage computationnellement faisable. Une série d'expériences conduites sur cette tâche montrent qu'une réduction de 1 point du taux d'erreur de retranscription peut être obtenu, démontrant que cette méthodologie d'apprentissage permet d'améliorer les performances des systèmes de reconnaissance. Diverses extensions de cette méthode seront finalement présentées et discutées.
Shiuan-Sung Lin, François Yvon
INTERSPEECH2
2007 Inference and evaluation of the multinomial mixture model for text clustering
Loïs Rigouste, Olivier Cappé, François Yvon
Inf. Process. Manag.3
2006 Corpus design based on the kullback-leibler divergence for text-to-speech synthesis application
Aleksandra Krul, Géraldine Damnati, François Yvon, Thierry Moudenc
INTERSPEECH3
2005 An Analogical Learner for Morphological Analysis
Nicolas Stroppa, François Yvon
CoNLL2
2005 On the use of morphological constraints in n-gram statistical language model
abstract
State of the art Speech Recognition systems use statistical language modeling and in particular N-gram models to represent the language structure. The Arabic language has a rich morphology, which motivates the introduction of morphological constraints in the language model. Class-based N-gram models have shown satisfactory results, especially for language model adaptation and training from reduced datasets. They were also proven quite effective in their use of memory space. In this paper, we investigate a new morphological classbased language model. Morphological rules are used to derive the different words in a class from their stem. As morphological analyzer, a rule-based stemming method is proposed for the Arabic language. The language model has been evaluated on a database composed of articles from Lebanese newspaper Al-Nahar for the years 1998 and 1999. In addition, a linear interpolation between the N-gram model and the morphological model is also evaluated. Preliminary experiments detailed in this paper show satisfactory results.
A. Ghaoui, François Yvon, Chafic Mokbel, Gérard Chollet
INTERSPEECH2
2005 Discriminative training of finite state decoding graphs
Shiuan-Sung Lin, François Yvon
INTERSPEECH2
2004 Arc minimization in finite-state decoding graphs with cross-word acoustic context
François Yvon, Geoffrey Zweig, George Saon
Comput. Speech Lang.1
2003 Improving Rocchio with Weakly Supervised Clustering
Romain Vinot, François Yvon
ECML2
2003 Proper Names Extraction from Fax Images Combining Textual and Image Features
abstract
In the frame of a unified messaging system, a crucial task of the system is to provide the user with key information on every message received, like keywords reflecting the object of the message, or the name of the sender. However, in the case of facsimiles, this information is not as easy to detect as in the case of e-mails, since no standard headers are defined. The aim of the presented work is to identify and extract specific information (the name of the sender) from a fax cover page. For this purpose, methods based on image document analysis (OCR recognition, physical blocks selection), and text analysis methods (optimized dictionary lookup, local grammar rules), are implemented to work in parallel. The fusion of their results brings a more accurate guess than any of the methods would achieve separately.
Laurence Likforman-Sulem, Pascal Vaillant, François Yvon
ICDAR3
2002 Arc minimization in finite state decoding graphs with cross-word acoustic context
abstract
Recent approaches to large vocabulary decoding with finite state graphs have focused on the use of state minimization algorithms to produce relatively compact graphs. This paper extends the finite state approach by developing complementary arc-minimization techniques. The use of these techniques in concert with state minimization allows us to statically compile decoding graphs in which the acoustic models utilize a full word of cross-word context. This is in significant contrast to typical systems which use only a single phone. We show that the particular arc-minimization problem that arises is in fact an NP-complete combinatorial optimization problem, and describe the reduction from 3-SAT. We present experimental results that illustrate the moderate sizes and runtimes of graphs for the Switchboard task. 1.
Geoffrey Zweig, George Saon, François Yvon
INTERSPEECH3
2002 Using the Web as a Linguistic Resource for Learning Reformulations Automatically
Florence Duclaye, François Yvon, Olivier Collin
LREC2
2001 Integrating contextual phonological rules in a large vocabulary decoder
abstract
International audience
Guillaume Gravier, François Yvon, Bruno Jacob, Frédéric Bimbot
INTERSPEECH2
2000 A French Phonetic Lexicon with Variants for Speech and Language Processing
Philippe Boula de Mareüil, Christophe d'Alessandro, François Yvon, Véronique Aubergé, Jacqueline Vaissière, Angélique Amelot
LREC3
1999 Pronouncing unknown words using multi-dimensional analogies
abstract
In this paper, a model of analogy-based learning is presented, whose main novelty is the crucial ability to produce analogies in multi-dimensional input and output spaces. Evaluations are performed on various word pronunciation tasks, revealing the effectiveness of such joint learning strategies.
François Yvon
EUROSPEECH1
1999 The hidden dimension: a paradigmatic view of data-driven NLP
Vito Pirrelli, François Yvon
J. Exp. Theor. Artif. Intell.2
1998 Evaluation of grapheme-to phoneme conversion for text-to-speech synthesis in French
Philippe Boula de Mareüil, François Yvon, Christophe d'Alessandro, V. Auberg, Michel Bagein, Gérard Bailly, Frédéric Béchet, S. Fonkia, Jean-Philippe Goldman, Eric Keller, Douglas D. O'Shaughnessy, Steve Pagel, F. Sannier, Jean Véronis, Brigitte Zellner Keller
LREC2
1998 Objective evaluation of grapheme to phoneme conversion for text-to-speech synthesis in French
François Yvon, Philippe Boula de Mareüil, Christophe d'Alessandro, Véronique Aubergé, Michel Bagein, Gérard Bailly, Frédéric Béchet, S. Foukia, J.-F. Goldman, Eric Keller, Douglas D. O'Shaughnessy, Vincent Pagel, Fred Sannier, Jean Véronis, Brigitte Zellner
Comput. Speech Lang.1
1997 Paradigmatic Cascades: a Linguistically Sound Model of Pronunciation by Analogy
abstract
We present and experimentally evaluate a new model of prounciation by analogy: the paradigmatic cascades model. Given a pronunciation lexicon, this algorithm first extracts the most productive paradigmatic mappings in the graphemic domain, and pairs them statistically with their correlate(s) in the phonemic domain. These mappings are used to search and retrieve in the lexical database the most promising analog of unseen words. We finally apply to the analogs pronunciation the correlated series of mappings in the phonemic domain to get the desired pronunciation.
François Yvon
ACL1
1995 Variable-length sequence matching for phonetic transcription using joint multigrams
Sabine Deligne, François Yvon, Frédéric Bimbot
EUROSPEECH2