Philipp Koehn

dblp:84/4538 · DBLP profile ↗
← Back
92ranked-venue papers
15as first author
25since 2021 · last 2025
0000-0003-1565-064XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 14 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
abstract
This paper investigates the propagation of information in multilingual large language models (LLMs) and evaluates the efficacy of various unlearning methods.We demonstrate that misinformation, regardless of the language it is in, once introduced into these models through training data, can spread across different languages, compromising the integrity and reliability of the generated content.Our findings reveal that standard unlearning techniques, which typically focus on English data, are insufficient in mitigating the spread of fake content in multilingual contexts and could inadvertently reinforce misinformation across languages.We show that only by addressing misinformative responses in both English and the original language of the fake data we can effectively eliminate it for all languages.This underscores the critical need for comprehensive unlearning strategies that consider the multilingual nature of modern LLMs to enhance their safety and reliability across landscapes.Code and data is accessible here: https
Taiming Lu, Philipp Koehn
EMNLP2
2025 Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
abstract
We present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions.Compared to the baseline method Global Mining (Duquenne et al., 2023a), a variant of speech mining, Speech Vecalign produces longer speech-to-speech alignments.It also demonstrates greater robustness than Local Mining, another speech mining variant, as it produces less noise.We applied Speech Vecalign to 3,000 hours of unlabeled parallel English-German (En-De) speech documents from VoxPopuli, yielding about 1,000 hours of high-quality alignments.We then trained En-De speech-to-speech translation models on the aligned data.Speech Vecalign improves the En-to-De and De-to-En performance over Global Mining by 0.37 and 0.18 ASR-BLEU, respectively.Moreover, our models match or outperform SpeechMatrix model performance, despite using 8 times fewer raw speech documents.1
Chutong Meng, Philipp Koehn
EMNLP2
2025 HadaSmileNet: Hadamard Fusion of Handcrafted and Deep-Learning Features for Enhancing Facial Emotion Recognition of Genuine Smiles
abstract
The distinction between genuine and posed emotions represents a fundamental pattern recognition challenge with significant implications for data mining applications in social sciences, healthcare, and human-computer interaction. While recent multitask learning frameworks have shown promise in combining deep learning architectures with handcrafted D-Marker features for smile facial emotion recognition, these approaches exhibit computational inefficiencies due to auxiliary task supervision and complex loss balancing requirements. This paper introduces HadaSmileNet, a novel feature fusion framework that directly integrates transformer-based representations with physiologically-grounded D-Markers through parameter-free multiplicative interactions. Through systematic evaluation of 15 fusion strategies, we demonstrate that Hadamard multi-plicative fusion achieves optimal performance by enabling direct feature interactions while maintaining computational efficiency. The proposed approach establishes new state-of-the-art results for deep learning methods across four benchmark datasets: UvA-NEMO (88.7%, +0.8%), MMI (99.7%), SPOS (98.5%, +0.7%), and BBC (100%, +5.0%). Comprehensive computational analysis reveals 26% parameter reduction and simplified training compared to multitask alternatives, while feature visualization demonstrates enhanced discriminative power through direct domain knowledge integration. The framework's efficiency and effectiveness make it particularly suitable for practical deployment in multimedia data mining applications that require realtime affective computing capabilities.
Mohammad Junayed Hasan, Nabeel Mohammed, Shafin Rahman, Philipp Koehn
ICDM4
2025 X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale
abstract
Large language models (LLMs) have achieved remarkable success across various NLP tasks with a focus on English due to English-centric pre-training and limited multilingual data. In this work, we focus on the problem of translation, and while some multilingual LLMs claim to support for hundreds of languages, models often fail to provide high-quality responses for mid- and low-resource languages, leading to imbalanced performance heavily skewed in favor of high-resource languages. We introduce **X-ALMA**, a model designed to ensure top-tier performance across 50 diverse languages, regardless of their resource levels. X-ALMA surpasses state-of-the-art open-source multilingual LLMs, such as Aya-101 and Aya-23, in every single translation direction on the FLORES-200 and WMT'23 test datasets according to COMET-22. This is achieved by plug-and-play language-specific module architecture to prevent language conflicts during training and a carefully designed training regimen with novel optimization methods to maximize the translation performance. After the final stage of training regimen, our proposed **A**daptive **R**ejection **P**reference **O**ptimization (**ARPO**) surpasses existing preference optimization methods in translation tasks.
Kenton Murray, Philipp Koehn, Hieu Hoang, Akiko Eriguchi, Huda Khayrallah
ICLR3
2024 Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
abstract
Text generation models are notoriously vulnerable to errors in the training data. With the wide-spread availability of massive amounts of web-crawled data becoming more commonplace, how can we enhance the robustness of models trained on a massive amount of noisy web-crawled text? In our work, we propose Error Norm Truncation (ENT), a robust enhancement method to the standard training objective that truncates noisy data. Compared to methods that only uses the negative log-likelihood loss to estimate data quality, our method provides a more accurate estimation by considering the distribution of non-target tokens, which is often overlooked by previous work. Through comprehensive experiments across language modeling, machine translation, and text summarization, we show that equipping text generation models with ENT improves generation quality over standard training and previous soft and hard truncation methods. Furthermore, we show that our method improves the robustness of models against two of the most detrimental types of noise in machine translation, resulting in an increase of more than 2 BLEU points over the MLE baseline when up to 50\% of noise is added to the data.
Tianjian Li, Philipp Koehn, Daniel Khashabi, Kenton Murray
ICLR3
2024 Where are you from? Geolocating Speech and Applications to Language Identification
abstract
Patrick Foley, Matthew Wiesner, Bismarck Odoom, Leibny Paola Garcia Perera, Kenton Murray, Philipp Koehn. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Patrick Foley, Matthew Wiesner, Bismarck Bamfo Odoom, L. Paola García-Perera, Kenton Murray, Philipp Koehn
NAACL-HLT6
2024 DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
abstract
Non-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate text data. Although NATs generate high-quality outputs and offer faster inference than autoregressive models, they tend to produce incoherent and repetitive results due to complex data distribution (e.g., acoustic and linguistic variations in speech). In this work, we introduce DiffNorm, a diffusion-based normalization strategy that simplifies data distributions for training NAT models. After training with a self-supervised noise estimation objective, DiffNorm constructs normalized target data by denoising synthetically corrupted speech features. Additionally, we propose to regularize NATs with classifier-free guidance, improving model robustness and translation quality by randomly dropping out source information during training. Our strategies result in a notable improvement of about $+7$ ASR-BLEU for English-Spanish (En-Es) translation and $+2$ ASR-BLEU for English-French (En-Fr) on the CVSS benchmark, while attaining over $14\times$ speedup for En-Es and $5 \times$ speedup for En-Fr translations compared to autoregressive baselines.
Weiting Tan, Lingfeng Shen, Daniel Khashabi, Philipp Koehn
NeurIPS5
2023 Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation
abstract
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzmán
ACL (1)6
2023 Multilingual Representation Distillation with Contrastive Learning
abstract
Multilingual sentence representations from large models encode semantic information from two or more languages and can be used for different cross-lingual information retrieval and matching tasks.In this paper, we integrate contrastive learning into multilingual representation distillation and use it for quality estimation of parallel sentences (i.e., find semantically similar sentences that can be used as translations of each other).We validate our approach with multilingual similarity search and corpus filtering tasks.Experiments across different low-resource languages show that our method greatly outperforms previous sentence encoders such as LASER, LASER3, and LaBSE.
Weiting Tan, Kevin Heffernan, Holger Schwenk, Philipp Koehn
EACL4
2023 Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer
abstract
We introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations.We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings.We explore various properties of pixel representations such as parameter sharing within and across scripts to better understand where they lead to positive transfer.We observe that these properties not only enable seamless cross-lingual transfer to unseen scripts, but make pixel representations more data-efficient than alternatives such as vocabulary expansion.We hope this work contributes to more extensible multilingual models for all languages and scripts.
Elizabeth Salesky, Neha Verma 0001, Philipp Koehn, Matt Post
EMNLP3
2023 Condensing Multilingual Knowledge with Lightweight Language-Specific Modules
abstract
Incorporating language-specific (LS) modules or Mixture-of-Experts (MoE) are proven methods to boost performance in multilingual model performance, but the scalability of these approaches to hundreds of languages or experts tends to be hard to manage.We present Language-specific Matrix Synthesis (LMS), a novel method that addresses the issue.LMS utilizes parameter-efficient and lightweight modules, reducing the number of parameters while outperforming existing methods, e.g., +1.73 BLEU over Switch Transformer on OPUS-100 multilingual translation.Additionally, we introduce Fuse Distillation (FD) to condense multilingual knowledge from multiple LS modules into a single shared module, improving model inference and storage efficiency.Our approach demonstrates superior scalability and performance compared to state-of-the-art methods. 1 * Equal contribution computational cost may only come from communication among devices (such as ALLToALL) or gate routing.
Weiting Tan, Shuyue Stella Li, Yunmo Chen, Benjamin Van Durme, Philipp Koehn, Kenton Murray
EMNLP6
2023 Learning from Mistakes: Towards Robust Neural Machine Translation for Disfluent L2 Sentences
abstract
We study the sentences written by second-language (L2) learners to improve the robustness of current neural machine translation (NMT) models on this type of data. Current large datasets used to train NMT systems are mostly Wikipedia or government documents written by highly competent speakers of that language, especially English. However, given that English is the most common second language, it is crucial that machine translation systems are robust against the large number of sentences written by L2 learners of English. By studying the difficulties faced by humans in their L2 acquisition process, we are able to transfer such insights to machine translation systems to recover from source-side fluency variations. In this work, we create additional training data with artificial errors similar to mistakes made by L2 learners of various fluency levels to improve the quality of the machine translation system. We test our method in zero-shot settings on the JFLEG-es (English-Spanish) dataset. The quality of our machine translation system on disfluent sentences outperforms the baseline by 1.8 BLEU scores.
Shuyue Stella Li, Philipp Koehn
MTSummit (1)2
2022 Alternative Input Signals Ease Transfer in Multilingual Machine Translation
abstract
Simeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán
ACL (1)6
2022 Toward the Limitation of Code-Switching in Cross-Lingual Transfer
abstract
Multilingual pretrained models have shown strong cross-lingual transfer ability.Some works used code-switching sentences, which consist of tokens from multiple languages, to enhance the cross-lingual representation further, and have shown success in many zero-shot cross-lingual tasks.However, code-switched tokens are likely to cause grammatical incoherence in newly substituted sentences, and negatively affect the performance on tokensensitive tasks, such as Part-of-Speech (POS) tagging and Named-Entity-Recognition (NER).This paper mitigates the limitation of the codeswitching method by not only making the token replacement but considering the similarity between the context and the switched tokens so that the newly substituted sentences are grammatically consistent during both training and inference.We conduct experiments on crosslingual POS and NER over 30+ languages, and demonstrate the effectiveness of our method by outperforming the mBERT by 0.95 and original code-switching method by 1.67 on F1 scores.How do menschen look at and अनु भव 艺术 ?How do people look at and experience art ?How do people look at and experience art ?How do Menschen look at and अनु भव 艺术 ?(b): Ours (a): Code-Switched Training Sentence SCONJ AUX NOUN VERB ADP CCPNJ VERB NOUN PUNCT
Yukun Feng, Philipp Koehn
EMNLP3
2022 Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal Transport
abstract
Bilingual lexicons form a critical component of various natural language processing applications, including unsupervised and semisupervised machine translation and crosslingual information retrieval.We improve bilingual lexicon induction performance across 40 language pairs with a graph-matching method based on optimal transport.The method is especially strong with low amounts of supervision.
Kelly Marchisio, Ali Saad-Eldin, Kevin Duh, Carey E. Priebe, Philipp Koehn
EMNLP5
2022 IsoVec: Controlling the Relative Isomorphism of Word Embedding Spaces
abstract
The ability to extract high-quality translation dictionaries from monolingual word embedding spaces depends critically on the geometric similarity of the spaces-their degree of "isomorphism."We address the root-cause of faulty cross-lingual mapping: that word embedding training resulted in the underlying spaces being non-isomorphic.We incorporate global measures of isomorphism directly into the Skip-gram loss function, successfully increasing the relative isomorphism of trained word embedding spaces and improving their ability to be mapped to a shared crosslingual space.The result is improved bilingual lexicon induction in general data conditions, under domain mismatch, and with training algorithm dissimilarities.We release IsoVec at https://github.com/ kellymarchisio/isovec.
Kelly Marchisio, Neha Verma 0001, Kevin Duh, Philipp Koehn
EMNLP4
2022 The Importance of Being Parameters: An Intra-Distillation Method for Serious Gains
abstract
Recent model pruning methods have demonstrated the ability to remove redundant parameters without sacrificing model performance.Common methods remove redundant parameters according to the parameter sensitivity, a gradient-based measure reflecting the contribution of the parameters.In this paper, however, we argue that redundant parameters can be trained to make beneficial contributions.We first highlight the large sensitivity (contribution) gap among high-sensitivity and lowsensitivity parameters and show that the model generalization performance can be significantly improved after balancing the contribution of all parameters.Our goal is to balance the sensitivity of all parameters and encourage all of them to contribute equally.We propose a general task-agnostic method, namely intradistillation, appended to the regular training loss to balance parameter sensitivity.Moreover, we also design a novel adaptive learning method to control the strength of intradistillation loss for faster convergence.Our experiments show the strong effectiveness of our methods on machine translation, natural language understanding, and zero-shot crosslingual transfer across up to 48 languages 1 , e.g., a gain of 3.54 BLEU on average across 8 language pairs from the IWSLT'14 dataset.
Philipp Koehn, Kenton Murray
EMNLP2
2022 Contrastive Clustering to Mine Pseudo Parallel Data for Unsupervised Translation
Xuan-Phi Nguyen, Hongyu Gong, Yun Tang 0002, Changhan Wang, Philipp Koehn, Shafiq R. Joty
ICLR5
2021 Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data
abstract
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab
ACL/IJCNLP (1)8
2021 Levenshtein Training for Word-level Quality Estimation
abstract
We propose a novel scheme to use the Levenshtein Transformer to perform the task of word-level quality estimation.A Levenshtein Transformer is a natural fit for this task: trained to perform decoding in an iterative manner, a Levenshtein Transformer can learn to post-edit without explicit supervision.To further minimize the mismatch between the translation task and the word-level QE task, we propose a two-stage transfer learning procedure on both augmented data and human postediting data.We also propose heuristics to construct reference labels that are compatible with subword-level finetuning and inference.Results on WMT 2020 QE shared task dataset show that our proposed method has superior data efficiency under the data-constrained setting and competitive performance under the unconstrained setting.* Shuoyang Ding had a part-time affiliation with Microsoft at the time of this work.
Shuoyang Ding, Marcin Junczys-Dowmunt, Matt Post, Philipp Koehn
EMNLP (1)4
2021 XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment
abstract
Cross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification.While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing.To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data.We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English.We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community.
Ahmed El-Kishky, Adithya Renduchintala, James Cross 0003, Francisco Guzmán, Philipp Koehn
EMNLP (1)5
2021 Streaming Simultaneous Speech Translation with Augmented Memory Transformer
abstract
Transformer-based models have achieved state-of-the-art performance on speech translation tasks. However, the model architecture is not efficient enough for streaming scenarios since self-attention is computed over an entire input sequence and the computational cost grows quadratically with the length of the input sequence. Nevertheless, most of the previous work on simultaneous speech translation, the task of generating translations from partial audio input, ignores the time spent in generating the translation when analyzing the latency. With this assumption, a system may have good latency quality trade-offs but be inapplicable in real-time scenarios. In this paper, we focus on the task of streaming simultaneous speech translation, where the systems are not only capable of translating with partial input but are also able to handle very long or continuous input. We propose an end-to-end transformer-based sequence-to-sequence model, equipped with an augmented memory transformer encoder, which has shown great success on the streaming automatic speech recognition task with hybrid or transducer-based models. We conduct an empirical evaluation of the proposed model on segment, context and memory sizes and we compare our approach to a transformer with a unidirectional mask.1
Xutai Ma, Mohammad Javad Dousti, Philipp Koehn, Juan Pino 0001
ICASSP4
2021 Learning Curricula for Multilingual Neural Machine Translation Training
abstract
Low-resource Multilingual Neural Machine Translation (MNMT) is typically tasked with improving the translation performance on one or more language pairs with the aid of high-resource language pairs. In this paper and we propose two simple search based curricula – orderings of the multilingual training data – which help improve translation performance in conjunction with existing techniques such as fine-tuning. Additionally and we attempt to learn a curriculum for MNMT from scratch jointly with the training of the translation system using contextual multi-arm bandits. We show on the FLORES low-resource translation dataset that these learned curricula can provide better starting points for fine tuning and improve overall performance of the translation system.
Philipp Koehn, Sanjeev Khudanpur
MTSummit (1)2
2021 An Alignment-Based Approach to Semi-Supervised Bilingual Lexicon Induction with Small Parallel Corpora
abstract
Aimed at generating a seed lexicon for use in downstream natural language tasks and unsupervised methods for bilingual lexicon induction have received much attention in the academic literature recently. While interesting and fully unsupervised settings are unrealistic; small amounts of bilingual data are usually available due to the existence of massively multilingual parallel corpora and or linguists can create small amounts of parallel data. In this work and we demonstrate an effective bootstrapping approach for semi-supervised bilingual lexicon induction that capitalizes upon the complementary strengths of two disparate methods for inducing bilingual lexicons. Whereas statistical methods are highly effective at inducing correct translation pairs for words frequently occurring in a parallel corpus and monolingual embedding spaces have the advantage of having been trained on large amounts of data and and therefore may induce accurate translations for words absent from the small corpus. By combining these relative strengths and our method achieves state-of-the-art results on 3 of 4 language pairs in the challenging VecMap test set using minimal amounts of parallel data and without the need for a translation dictionary. We release our implementation at www.blind-review.code.
Kelly Marchisio, Philipp Koehn, Conghao Xiong
MTSummit (1)2
2021 Evaluating Saliency Methods for Neural Language Models
abstract
Saliency methods are widely used to interpret neural network predictions, but different variants of saliency methods often disagree even on the interpretations of the same prediction made by the same model.In these cases, how do we identify when are these interpretations trustworthy enough to be used in analyses?To address this question, we conduct a comprehensive and quantitative evaluation of saliency methods on a fundamental category of NLP models: neural language models.We evaluate the quality of prediction interpretations from two perspectives that each represents a desirable property of these interpretations: plausibility and faithfulness.Our evaluation is conducted on four different datasets constructed from the existing human annotation of syntactic and semantic agreements, on both sentencelevel and document-level.Through our evaluation, we identified various ways saliency methods could yield interpretations of low quality.We recommend that future work deploying such methods to neural language models should carefully validate their interpretations before drawing insights.
Shuoyang Ding, Philipp Koehn
NAACL-HLT2
2020 ParaCrawl: Web-Scale Acquisition of Parallel Corpora
abstract
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza
ACL10
2020 CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs
abstract
Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other.We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English.In addition to curating this massive dataset, we introduce baseline methods that leverage crosslingual representations to identify aligned documents based on their textual content.Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data.Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages.
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn
EMNLP (1)4
2020 Statistical Power and Translationese in Machine Translation Evaluation
abstract
The term translationese has been used to describe features of translated text, and in this paper, we provide detailed analysis of potential adverse effects of translationese on machine translation evaluation.Our analysis shows differences in conclusions drawn from evaluations that include translationese in test data compared to experiments that tested only with text originally composed in that language.For this reason we recommend that reverse-created test data be omitted from future machine translation test sets.In addition, we provide a reevaluation of a past machine translation evaluation claiming human-parity of MT.One important issue not previously considered is statistical power of significance tests applied to comparison of human and machine translation.Since the very aim of past evaluations was the investigation of ties between human and MT systems, power analysis is of particular importance, to avoid, for example, claims of human parity simply corresponding to Type II error resulting from the application of a low powered test.We provide detailed analysis of tests used in such evaluations to provide an indication of a suitable minimum sample size for future studies.
Yvette Graham, Barry Haddow, Philipp Koehn
EMNLP (1)3
2020 Simulated multiple reference training improves low-resource machine translation
abstract
Many valid translations exist for a given sentence, yet machine translation (MT) is trained with a single reference translation, exacerbating data sparsity in low-resource settings. We introduce Simulated Multiple Reference Training (SMRT), a novel MT training method that approximates the full space of possible translations by sampling a paraphrase of the reference sentence from a paraphraser and training the MT model to predict the paraphraser's distribution over possible tokens. We demonstrate the effectiveness of SMRT in low-resource settings when translating to English, with improvements of 1.2 to 7.0 BLEU. We also find SMRT is complementary to back-translation.
Huda Khayrallah, Brian Thompson 0001, Matt Post, Philipp Koehn
EMNLP (1)4
2020 Exploiting Sentence Order in Document Alignment
abstract
We present a simple document alignment method that incorporates sentence order information in both candidate generation and candidate re-scoring. Our method results in 61% relative reduction in error compared to the best previously published result on the WMT16 document alignment shared task. Our method improves downstream MT performance on web-scraped Sinhala--English documents from ParaCrawl, outperforming the document alignment method used in the most recent ParaCrawl release. It also outperforms a comparable corpora method which uses the same multilingual embeddings, demonstrating that exploiting sentence order is beneficial even if the end goal is sentence-level bitext.
Brian Thompson 0001, Philipp Koehn
EMNLP (1)2
2020 Searching the Web for Cross-lingual Parallel Data
abstract
While the World Wide Web provides a large amount of text in many languages, cross-lingual parallel data is more difficult to obtain. Despite its scarcity, this parallel cross-lingual data plays a crucial role in a variety of tasks in natural language processing with applications in machine translation, cross-lingual information retrieval, and document classification, as well as learning cross-lingual representations. Here, we describe the end-to-end process of searching the web for parallel cross-lingual texts. We motivate obtaining parallel text as a retrieval problem whereby the goal is to retrieve cross-lingual parallel text from a large, multilingual web-crawled corpus. We introduce techniques for searching for cross-lingual parallel data based on language, content, and other metadata. We motivate and introduce multilingual sentence embeddings as a core tool and demonstrate techniques and models that leverage them for identifying parallel documents and sentences as well as techniques for retrieving and filtering this data. We describe several large-scale datasets curated using these techniques and show how training on sentences extracted from parallel or comparable documents mined from the Web can improve machine translation models and facilitate cross-lingual NLP.
Ahmed El-Kishky, Philipp Koehn, Holger Schwenk
SIGIR2
2019 The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English
abstract
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino 0001, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc'Aurelio Ranzato
EMNLP/IJCNLP (1)6
2019 Spelling-Aware Construction of Macaronic Texts for Teaching Foreign-Language Vocabulary
abstract
Adithya Renduchintala, Philipp Koehn, Jason Eisner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Adithya Renduchintala, Philipp Koehn, Jason Eisner
EMNLP/IJCNLP (1)2
2019 Vecalign: Improved Sentence Alignment in Linear Time and Space
abstract
Brian Thompson, Philipp Koehn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Brian Thompson 0001, Philipp Koehn
EMNLP/IJCNLP (1)2
2019 HABLex: Human Annotated Bilingual Lexicons for Experiments in Machine Translation
abstract
Brian Thompson, Rebecca Knowles, Xuan Zhang, Huda Khayrallah, Kevin Duh, Philipp Koehn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Brian Thompson 0001, Rebecca Knowles, Xuan Zhang 0008, Huda Khayrallah, Kevin Duh, Philipp Koehn
EMNLP/IJCNLP (1)6
2019 Controlling the Reading Level of Machine Translation Output
Kelly Marchisio, Jialiang Guo, Cheng-I Lai, Philipp Koehn
MTSummit (1)4
2019 Character-Aware Decoder for Translation into Morphologically Rich Languages
Adithya Renduchintala, Pamela Shapiro, Kevin Duh, Philipp Koehn
MTSummit (1)4
2019 Robust Document Representations for Cross-Lingual Information Retrieval in Low-Resource Settings
Mahsa Yarmohammadi, Xutai Ma, Sorami Hisamoto, Muhammad Mahbubur Rahman 0001, Yiming Wang 0006, Hainan Xu, Daniel Povey, Philipp Koehn, Kevin Duh
MTSummit (1)8
2019 A user study of neural interactive translation prediction
abstract
Machine translation (MT) on its own is generally not good enough to produce high-quality translations, so it is common to have humans intervening in the translation process to improve MT output. A typical intervention is post-editing (PE), where a human translator corrects errors in the MT output. Another is interactive translation prediction (ITP), which involves an MT system presenting a translator with translation suggestions they can accept or reject, actions the MT system then uses to present them with new, corrected suggestions. Both Macklovitch ( 2006 ) and Koehn ( 2009 ) found ITP to be an efficient alternative to unassisted translation in terms of processing time. So far, phrase-based statistical ITP has not yet proven to be faster than PE (Koehn 2009 ; Sanchis-Trilles et al. 2014 ; Underwood et al. 2014 ; Green et al. 2014 ; Alves et al. 2016 ; Alabau et al. 2016 ). In this paper we present the results of an empirical study on translation productivity in ITP with an underlying neural MT system (NITP). Our results show that over half of the professional translators in our study translated faster with NITP compared to PE, and most preferred it over PE. We also examine differences between PE and ITP in other translation productivity indicators and translators’ reactions to the technology.
Rebecca Knowles, Marina Sanchez-Torron, Philipp Koehn
Mach. Transl.3
2018 An Analysis of Source Context Dependency in Neural Machine Translation
abstract
The encoder-decoder with attention model has become the state of the art for machine translation. However, more investigations are still needed to understand the internal mechanism of this end-to-end model. In this paper, we focus on how neural machine translation (NMT) models consider source information while decoding. We propose a numerical measurement of source context dependency in the NMT models and analyze the behaviors of the NMT decoder with this measurement under several circumstances. Experimental results show that this measurement is an appropriate estimate for source context dependency and consistent over different domains.
Xutai Ma, Philipp Koehn
EAMT3
2018 Context and Copying in Neural Machine Translation
abstract
Neural machine translation systems with subword vocabularies are capable of translating or copying unknown words.In this work, we show that they learn to copy words based on both the context in which the words appear as well as features of the words themselves.In contexts that are particularly copy-prone, they even copy words that they have already learned they should translate.We examine the influence of context and subword features on this and other types of copying behavior.
Rebecca Knowles, Philipp Koehn
EMNLP2
2017 Knowledge Tracing in Sequential Learning of Inflected Vocabulary
abstract
We present a feature-rich knowledge tracing method that captures a student's acquisition and retention of knowledge during a foreign language phrase learning task.We model the student's behavior as making predictions under a log-linear model, and adopt a neural gating mechanism to model how the student updates their log-linear parameters in response to feedback.The gating mechanism allows the model to learn complex patterns of retention and acquisition for each feature, while the log-linear parameterization results in an interpretable knowledge state.We collect human data and evaluate several versions of the model.
Adithya Renduchintala, Philipp Koehn, Jason Eisner
CoNLL2
2017 Zipporah: a Fast and Scalable Data Cleaning System for Noisy Web-Crawled Parallel Corpora
abstract
We introduce Zipporah, a fast and scalable data cleaning system.We propose a novel type of bag-of-words translation feature, and train logistic regression models to classify good data and synthetic noisy data in the proposed feature space.The trained model is used to score parallel sentences in the data pool for selection.As shown in experiments, Zipporah selects a high-quality parallel corpus from a large, mixed quality data pool.In particular, for one noisy dataset, Zipporah achieves a 2.1 BLEU score improvement with using 1/5 of the data over using the entire corpus.
Hainan Xu, Philipp Koehn
EMNLP2
2017 Syntax-Based Statistical Machine Translation Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn (University of Edinburgh, University of Edinburgh, Johns Hopkins University, Johns Hopkins University), edited by Graeme Hirst, volume 33), 2016, xvii+190 pp; paperback, ISBN 978-1-62705-900-8; ebook, ISBN 978-1-62705-502-4; doi: 10.2200/S00716ED1V04Y201604HLT033, $70
abstract
In its early development, machine translation adopted rule-based approaches, which can include the use of language syntax. The late 1980s and early 1990s saw the inception of the statistical machine translation (SMT) approach, where translation models can be learned automatically from a parallel corpus rather than created manually by humans. Initial SMT models were word-based and phrase-based, without the use of syntactic knowledge. In phrase-based SMT, a source sentence is first segmented into phrases and then translated phrase-by-phrase with some reordering of the translated phrases in the target sentence. This has posed challenges when translating between two syntactically different languages. Syntax-based SMT approaches take advantage of syntactic knowledge within the framework of SMT. This book provides an introduction to syntax-based SMT approaches. It is a valuable resource for those who are interested in syntax-based SMT.The book consists of seven chapters. There is not an introduction chapter in this book, aside from the preface, which can be considered as a brief introduction. Readers are referred to Koehn (2010) for background knowledge. I think an introduction chapter categorized into sections would have been useful, before proceeding to describe the various models. The first two chapters provide principles applicable across various syntax-based SMT approaches. The next three chapters describe syntax-based SMT decoding in detail; this constitutes half of the book. Selected extended topics are provided in the next chapter, which is followed by a concluding chapter.Chapter 1 describes the models and formalisms applicable to syntax-based SMT. The first section describes the phrasal translation units in phrase-based SMT, its limitations, and how tree structures address the limitations of the phrase-based approach. This explanation is useful as translation units are the key difference between the phrase-based and syntax-based SMT approaches. The next two sections describe the grammar formalisms and the statistical models that define syntax-based SMT. The section that covers the grammar formalisms (i.e., synchronous context-free grammar [SCFG] and synchronous tree-substitution grammar [STSG]), would have been clearer if their differences were presented in a side-by-side illustrating example. The remainder of the chapter discusses different categories of syntax-based SMT approaches and the history of these approaches, which include string-to-string, string-to-tree, tree-to-string, and tree-to-tree SMT approaches. Although the syntax-based translation model in Galley et al. (2006) falls under the string-to-tree category, I wonder why hierarchical phrase-based SMT, or Hiero (Chiang, 2007), is not explicitly put under the string-to-string category, since Hiero also uses “unlabeled hierarchical phrases where there is no representation of linguistic categories.”Chapter 2 focuses on how the statistical framework of a syntax-based SMT approach learns its model from a word-aligned and parsed parallel text. The first section explains how phrase pairs are extracted as translation rules from a word-aligned sentence pair in phrase-based SMT (Koehn, Och, and Marcu, 2003), highlighting the definition of a phrase as a sequence of words and the alignment-consistency property of a phrase pair as defined in Och and Ney (2004). The remainder of the chapter introduces three predominant instantiations of syntax-based models: hierarchical phrase-based SMT (Hiero) (Chiang, 2007), which is a non-labeled syntax-based SMT approach arising from the phrase-based approach; syntax-augmented machine translation (SAMT), which introduces the notion of soft labels while keeping the nonlinguistic phrase notion; and GHKM (Galley et al., 2004), which only extracts translation rules consistent with constituency parse subtrees. This chapter is nicely organized and it is easy to follow the gradual evolution from phrase-based SMT to GHKM.Chapter 3 introduces the decoding formalism in the form of a directed hypergraph, defined as a set of vertices and a set of directed hyperedges. The first section introduces the notion of a weighted parse forest represented in a weighted hypergraph, representing alternative parse trees of a sentence. I found it important to pay careful attention to this section, in order to understand the next section and the following chapters. The next section presents various algorithms on a hypergraph to translate a sentence in a hypergraph representation of possible tree derivations. Overall, I found this chapter to contain many technical details. The last section of this chapter provides historical notes on the sources of these concepts. This chapter needs to be read before the next chapter, which assumes understanding of the concepts introduced in Chapter 3.Chapter 4 describes tree decoding—that is, decoding with the constituency parse tree of a source sentence as its input, focusing on the tree-to-string approach. The first two sections highlight decoding with local and non-local features, where non-local features accommodate n-gram language models and are more complex than local features. The next section is devoted to an in-depth description of a beam search algorithm on the parse tree of a source sentence. The description could have been improved if the running example showed the decoding steps. The next two sections present extensions to the concepts introduced in the earlier part of this chapter, by providing references to more efficient hypergraph operations. The content of this section requires readers who are interested in implementing an efficient tree-based algorithm to go through the cited references. Brief historical notes conclude this chapter nicely, by pointing to relevant materials for further reading.Chapter 5 describes string decoding with a source sentence string as its input. The first two sections describe beam search decoding algorithms in a binary SCFG, namely, a maximum of two non-terminal symbols on the right-hand side of each rule, adopted in Hiero and SAMT. The algorithms covered are a basic algorithm and an optimized algorithm. The complexity comparison between the two is nicely presented here, emphasizing the complexity reduction achieved by algorithm optimization. The handling of non-binary rules is described in the following section, illustrated by GHKM rule extraction. A mid-chapter summary section divides this chapter into two parts: beam search decoding and parsing. The second part describes parsing algorithms in the context of shared-category SCFG, assuming the same set of non-terminal symbols for the left-hand and right-hand sides of a rule, followed by a section extending the algorithm to STSG and distinct-category SCFG. The organization of this chapter is excellent. However, I feel that the inclusion of distinct-category SCFG decoding does not fit well into this chapter, as string decoding in string-to-tree SMT requires no knowledge of the source syntax. The historical notes also do not provide any references of prior work on string decoding using distinct-category SCFG.Chapter 6 contains various selected topics on syntax-based SMT. The first section discusses tree transformations, which make translation rule learning more effective. The description of non-context-free models serves as a prelude to the next section on dependency-based SMT, which covers dependency treelet (equivalent to the tree-to-string approach) and string-to-dependency (equivalent to the string-to-tree approach). The next section focuses on the ability of syntax-based SMT to have a more grammatical output compared with phrase-based SMT, although there is still room for improvement, including the use of unification grammars and semantic properties. Finally, the last section of this chapter explains how MT evaluation benefits from syntax-based SMT principles. Overall, this chapter enriches readers' knowledge beyond basic syntax-based SMT in the earlier chapters. I would also suggest the inclusion of phrase-based decoding approaches that use syntax-based features (Cherry, 2008; Chang et al., 2009).Chapter 7 nicely concludes this book by discussing the comparison between phrase-based and syntax-based SMT approaches and proposing possible future developments of syntax-based SMT. The chapter also highlights that syntax-driven MT predates statistical MT, as I mentioned at the beginning of this review.Overall, I found this book to be a useful reference book for those interested in syntax-based SMT. The book is well organized, which makes it easy for readers to refer to specific aspects of syntax-based SMT. An improvement can be made to the presentation of ideas in this book. Throughout the book, there are many technical keywords, resulting from the complexity of syntax-based SMT. It would be useful to highlight these keywords in a side bar to remind readers that they are important keywords. In addition, although examples are given throughout the book, it would be even more useful to use these examples to illustrate how the algorithms work, so that readers can gain a better understanding of the algorithms.
Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn, Graeme Hirst, Christian Hadiwinoto
Comput. Linguistics4
2016 User Modeling in Language Learning with Macaronic Texts
abstract
Foreign language learners can acquire new vocabulary by using cognate and context clues when reading.To measure such incidental comprehension, we devise an experimental framework that involves reading mixed-language "macaronic" sentences.Using data collected via Amazon Mechanical Turk, we train a graphical model to simulate a human subject's comprehension of foreign words, based on cognate clues (edit distance to an English word), context clues (pointwise mutual information), and prior exposure.Our model does a reasonable job at predicting which words a user will be able to understand, which should facilitate the automatic construction of comprehensible text for personalized foreign language education.
Adithya Renduchintala, Rebecca Knowles, Philipp Koehn, Jason Eisner
ACL (1)3
2016 Analyzing Learner Understanding of Novel L2 Vocabulary
abstract
In this work, we explore how learners can infer second-language noun meanings in the context of their native language. Motivated by an interest in building interactive tools for language learning, we collect data on three word-guessing tasks, analyze their difficulty, and explore the types of errors that novice learners make. We train a log-linear model for predicting our subjects’ guesses of word meanings in varying kinds of contexts. The model’s predictions correlate well with subject performance, and we provide quantitative and qualitative analyses of both human and model performance.
Rebecca Knowles, Adithya Renduchintala, Philipp Koehn, Jason Eisner
CoNLL3
2015 The Operation Sequence Model - Combining N-Gram-Based and Phrase-Based Statistical Machine Translation
abstract
In this article, we present a novel machine translation model, the Operation Sequence Model (OSM), which combines the benefits of phrase-based and N-gram-based statistical machine translation (SMT) and remedies their drawbacks. The model represents the translation process as a linear sequence of operations. The sequence includes not only translation operations but also reordering operations. As in N-gram-based SMT, the model is: (i) based on minimal translation units, (ii) takes both source and target information into account, (iii) does not make a phrasal independence assumption, and (iv) avoids the spurious phrasal segmentation problem. As in phrase-based SMT, the model (i) has the ability to memorize lexical reordering triggers, (ii) builds the search graph dynamically, and (iii) decodes with large translation units during search. The unique properties of the model are (i) its strong coupling of reordering and translation where translation and reordering decisions are conditioned on n previous translation and reordering decisions, and (ii) the ability to model local and long-range reorderings consistently. Using BLEU as a metric of translation accuracy, we found that our system performs significantly better than state-of-the-art phrase-based systems (Moses and Phrasal) and N-gram-based systems (Ncode) on standard translation tasks. We compare the reordering component of the OSM to the Moses lexical reordering model by integrating it into Moses. Our results show that OSM outperforms lexicalized reordering on all translation tasks. The translation quality is shown to be improved further by learning generalized representations with a POS-based OSM.
Nadir Durrani, Helmut Schmid, Alexander Fraser 0001, Philipp Koehn, Hinrich Schütze
Comput. Linguistics4
2014 Investigating the Usefulness of Generalized Word Representations in SMT
Nadir Durrani, Philipp Koehn, Helmut Schmid, Alexander Fraser 0001
COLING2
2014 CASMACAT: A Computer-assisted Translation Workbench
abstract
CASMACAT is a modular, web-based translation workbench that offers advanced functionalities for computer-aided translation and the scientific study of human translation: automatic interaction with machine translation (MT) engines and translation memories (TM) to obtain raw translations or close TM matchesn for conventional post-editing; interactive translation prediction based on an MT engine’s search graph, detailed recording and replay of edit actions and translator’s gaze (the latter via eye-tracking), and the support of e-pen as an alternative input device. The system is open source sofware and interfaces with multiple MT systems.
Vicente Alabau, Christian Buck, Michael Carl, Francisco Casacuberta, Mercedes García-Martínez, Ulrich Germann, Jesús González-Rubio, Robin L. Hill, Philipp Koehn, Luis A. Leiva, Bartolomé Mesa-Lao, Daniel Ortiz-Martínez, Herve Saint-Amand, Germán Sanchis-Trilles, Chara Tsoukala
EACL9
2014 Integrating an Unsupervised Transliteration Model into Statistical Machine Translation
abstract
We investigate three methods for integrating an unsupervised transliteration model into an end-to-end SMT system.We induce a transliteration model from parallel data and use it to translate OOV words.Our approach is fully unsupervised and language independent.In the methods to integrate transliterations, we observed improvements from 0.23-0.75(∆ 0.41) BLEU points across 7 language pairs.We also show that our mined transliteration corpora provide better rule coverage and translation quality compared to the gold standard transliteration corpora.
Nadir Durrani, Hassan Sajjad 0001, Hieu Hoang, Philipp Koehn
EACL4
2014 Dynamic Topic Adaptation for Phrase-based MT
abstract
Translating text from diverse sources poses a challenge to current machine translation systems which are rarely adapted to structure beyond corpus level. We explore topic adaptation on a diverse data set and present a new bilingual vari-ant of Latent Dirichlet Allocation to com-pute topic-adapted, probabilistic phrase translation features. We dynamically in-fer document-specific translation proba-bilities for test sets of unknown origin, thereby capturing the effects of document context on phrase translations. We show gains of up to 1.26 BLEU over the base-line and 1.04 over a domain adaptation benchmark. We further provide an anal-ysis of the domain-specific data and show additive gains of our model in combination with other types of topic-adapted features. 1
Eva Hasler, Phil Blunsom, Philipp Koehn, Barry Haddow
EACL3
2014 Improving machine translation via triangulation and transliteration
Nadir Durrani, Philipp Koehn
EAMT2
2014 CASMACAT: cognitive analysis and statistical methods for advanced computer aided translation
Philipp Koehn, Michael Carl, Francisco Casacuberta, Eva Marcos
EAMT1
2014 Interactive translation prediction versus conventional post-editing in practice: a study with the CasMaCat workbench
Germán Sanchis-Trilles, Vicente Alabau, Christian Buck, Michael Carl, Francisco Casacuberta, Mercedes García-Martínez, Ulrich Germann, Jesús González-Rubio, Robin L. Hill, Philipp Koehn, Luis A. Leiva, Bartolomé Mesa-Lao, Daniel Ortiz-Martínez, Herve Saint-Amand, Chara Tsoukala, Enrique Vidal 0001
Mach. Transl.10
2013 Dirt Cheap Web-Scale Parallel Text from the Common Crawl
Jason Smith 0006, Herve Saint-Amand, Magdalena Plamada, Philipp Koehn, Chris Callison-Burch, Adam Lopez
ACL (1)4
2013 Grouping Language Model Boundary Words to Speed K-Best Extraction from Hypergraphs
Kenneth Heafield, Philipp Koehn, Alon Lavie
HLT-NAACL2
2012 Language Model Rest Costs and Space-Efficient Storage
Kenneth Heafield, Philipp Koehn, Alon Lavie
EMNLP-CoNLL2
2012 Semi-supervised discriminative language modeling for Turkish ASR
abstract
We present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results.
Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP19
2012 Hallucinated n-best lists for discriminative language modeling
abstract
This paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods.
Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP16
2012 Continuous space discriminative language modeling
abstract
Discriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well.
Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP16
2012 Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
INTERSPEECH17
2011 Soft Dependency Constraints for Reordering in Hierarchical Phrase-Based Translation
Yang Gao 0005, Philipp Koehn, Alexandra Birch
EMNLP2
2010 Enabling Monolingual Translators: Post-Editing vs. Options
Philipp Koehn
HLT-NAACL1
2010 Monte Carlo techniques for phrase-based translation
Abhishek Arun, Barry Haddow, Philipp Koehn, Adam Lopez, Chris Dyer, Phil Blunsom
Mach. Transl.3
2009 Monte Carlo inference and maximization for phrase-based translation
Abhishek Arun, Chris Dyer, Barry Haddow, Phil Blunsom, Adam Lopez, Philipp Koehn
CoNLL6
2009 Improving Mid-Range Re-Ordering Using Templates of Factors
Hieu Hoang, Philipp Koehn
EACL2
2009 Word Lattices for Multi-Source Translation
Josh Schroeder, Trevor Cohn, Philipp Koehn
EACL3
2009 462 Machine Translation Systems for Europe
Philipp Koehn, Alexandra Birch, Ralf Steinberger
MTSummit1
2009 Interactive Assistance to Human Translators using Statistical Machine Translation Methods
Philipp Koehn, Barry Haddow
MTSummit1
2009 A process study of computer-aided translation
Philipp Koehn
Mach. Transl.1
2009 Review of Cyril Goutte, Nicola Cancedda, Marc Dymetman, and George Foster (eds): Learning machine translation
Philipp Koehn
Mach. Transl.1
2009 Introduction to the Special Issue on Machine Translation of Asian Languages
abstract
No abstract available.
David Chiang 0001, Philipp Koehn
ACM Trans. Asian Lang. Inf. Process.2
2008 Enriching Morphologically Poor Languages for Statistical Machine Translation
Eleftherios Avramidis, Philipp Koehn
ACL2
2008 Predicting Success in Machine Translation
Alexandra Birch, Miles Osborne, Philipp Koehn
EMNLP3
2008 Large and Diverse Language Models for Statistical Machine Translation
Holger Schwenk, Philipp Koehn
IJCNLP2
2007 Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst
ACL1
2007 Factored Translation Models
Philipp Koehn, Hieu Hoang
EMNLP-CoNLL1
2007 Chinese Syntactic Reordering for Statistical Machine Translation
Chao Wang 0018, Michael Collins 0001, Philipp Koehn
EMNLP-CoNLL3
2007 Online learning methods for discriminative training of phrase based statistical machine translation
Abhishek Arun, Philipp Koehn
MTSummit2
2006 Re-evaluation the Role of Bleu in Machine Translation Research
Chris Callison-Burch, Miles Osborne, Philipp Koehn
EACL3
2006 Improved Statistical Machine Translation Using Paraphrases
Chris Callison-Burch, Philipp Koehn, Miles Osborne
HLT-NAACL2
2005 Clause Restructuring for Statistical Machine Translation
abstract
We describe a method for incorporating syntactic information in statistical machine translation systems. The first step of the method is to parse the source language string that is being translated. The second step is to apply a series of transformations to the parse tree, effectively reordering the surface string on the source language side of the translation system. The goal of this step is to recover an underlying word order that is closer to the target language word-order than the original string. The reordering approach is applied as a pre-processing step in both the training and decoding phases of a phrase-based statistical MT system. We describe experiments on translation from German to English, showing an improvement from 25.2% Bleu score for a baseline system to 26.8% Bleu score for the system with reordering, a statistically significant improvement.
Michael Collins 0001, Philipp Koehn, Ivona Kucerova
ACL2
2005 Europarl: A Parallel Corpus for Statistical Machine Translation
abstract
We collected a corpus of parallel text in 11 languages from the proceedings of the European Parliament, which are published on the web. This corpus has found widespread use in the NLP community. Here, we focus on its acquisition and its application as training data for statistical machine translation (SMT). We trained SMT systems for 110 language pairs, which reveal interesting clues into the challenges ahead.
Philipp Koehn
MTSummit1
2004 Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn
EMNLP1
2003 Feature-Rich Statistical Translation of Noun Phrases
abstract
We define noun phrase translation as a subtask of machine translation. This enables us to build a dedicated noun phrase translation subsystem that improves over the currently best general statistical machine translation methods by incorporating special modeling and special features. We achieved 65.5% translation accuracy in a German-English translation task vs. 53.2% with IBM Model 4.
Philipp Koehn, Kevin Knight
ACL1
2003 Empirical Methods for Compound Splitting
Philipp Koehn, Kevin Knight
EACL1
2003 What's New in Statistical Machine Translation
Kevin Knight, Philipp Koehn
HLT-NAACL2
2003 Statistical Phrase-Based Translation
Philipp Koehn, Franz Josef Och, Daniel Marcu
HLT-NAACL1
2003 Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight
HLT-NAACL11
2002 Translation with Scarce Bilingual Resources
Yaser Al-Onaizan, Ulrich Germann, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Daniel Marcu, Kenji Yamada
Mach. Transl.5
2001 Knowledge Sources for Word-Level Translation Models
Philipp Koehn, Kevin Knight
EMNLP1
2000 Improving intonational phrasing with syntactic information
abstract
The prediction of intonational phrase boundaries from raw text is an important step for a text-to-speech system: locating where to place short pauses enables more natural sounding speech, that can be more easily understood. We improved upon earlier work [Hirschberg and Prieto, 1996] by adding syntactic information gained from a high-accuracy parser [Collins, 1999]. We report significant improvement using various experimental setups. We also show that our improved method comes close to interannotator agreement.
Philipp Koehn, Steven P. Abney, Julia Hirschberg, Michael Collins 0001
ICASSP1