EDBT 2026 Demo / reviewers in the wild / expert
Philipp Koehn
dblp:84/4538
· DBLP profile ↗
92ranked-venue papers
15as first author
25since 2021 · last 2025
0000-0003-1565-064XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 86 · 14 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learn and Unlearn: Addressing Misinformation in Multilingual LLMsabstractThis paper investigates the propagation of information in multilingual large language models (LLMs) and evaluates the efficacy of various unlearning methods.We demonstrate that misinformation, regardless of the language it is in, once introduced into these models through training data, can spread across different languages, compromising the integrity and reliability of the generated content.Our findings reveal that standard unlearning techniques, which typically focus on English data, are insufficient in mitigating the spread of fake content in multilingual contexts and could inadvertently reinforce misinformation across languages.We show that only by addressing misinformative responses in both English and the original language of the fake data we can effectively eliminate it for all languages.This underscores the critical need for comprehensive unlearning strategies that consider the multilingual nature of modern LLMs to enhance their safety and reliability across landscapes.Code and data is accessible here: https Taiming Lu, Philipp Koehn |
EMNLP | 2 |
| 2025 | Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech DocumentsabstractWe present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions.Compared to the baseline method Global Mining (Duquenne et al., 2023a), a variant of speech mining, Speech Vecalign produces longer speech-to-speech alignments.It also demonstrates greater robustness than Local Mining, another speech mining variant, as it produces less noise.We applied Speech Vecalign to 3,000 hours of unlabeled parallel English-German (En-De) speech documents from VoxPopuli, yielding about 1,000 hours of high-quality alignments.We then trained En-De speech-to-speech translation models on the aligned data.Speech Vecalign improves the En-to-De and De-to-En performance over Global Mining by 0.37 and 0.18 ASR-BLEU, respectively.Moreover, our models match or outperform SpeechMatrix model performance, despite using 8 times fewer raw speech documents.1 Chutong Meng, Philipp Koehn |
EMNLP | 2 |
| 2025 | HadaSmileNet: Hadamard Fusion of Handcrafted and Deep-Learning Features for Enhancing Facial Emotion Recognition of Genuine SmilesabstractThe distinction between genuine and posed emotions represents a fundamental pattern recognition challenge with significant implications for data mining applications in social sciences, healthcare, and human-computer interaction. While recent multitask learning frameworks have shown promise in combining deep learning architectures with handcrafted D-Marker features for smile facial emotion recognition, these approaches exhibit computational inefficiencies due to auxiliary task supervision and complex loss balancing requirements. This paper introduces HadaSmileNet, a novel feature fusion framework that directly integrates transformer-based representations with physiologically-grounded D-Markers through parameter-free multiplicative interactions. Through systematic evaluation of 15 fusion strategies, we demonstrate that Hadamard multi-plicative fusion achieves optimal performance by enabling direct feature interactions while maintaining computational efficiency. The proposed approach establishes new state-of-the-art results for deep learning methods across four benchmark datasets: UvA-NEMO (88.7%, +0.8%), MMI (99.7%), SPOS (98.5%, +0.7%), and BBC (100%, +5.0%). Comprehensive computational analysis reveals 26% parameter reduction and simplified training compared to multitask alternatives, while feature visualization demonstrates enhanced discriminative power through direct domain knowledge integration. The framework's efficiency and effectiveness make it particularly suitable for practical deployment in multimedia data mining applications that require realtime affective computing capabilities. Mohammad Junayed Hasan, Nabeel Mohammed, Shafin Rahman, Philipp Koehn |
ICDM | 4 |
| 2025 | X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleabstractLarge language models (LLMs) have achieved remarkable success across various NLP tasks with a focus on English due to English-centric pre-training and limited multilingual data. In this work, we focus on the problem of translation, and
while some multilingual LLMs claim to support for hundreds of languages, models often fail to provide high-quality responses for mid- and low-resource languages, leading to imbalanced performance heavily skewed in favor of high-resource languages. We introduce **X-ALMA**, a model designed to ensure top-tier performance across 50 diverse languages, regardless of their resource levels. X-ALMA surpasses state-of-the-art open-source multilingual LLMs, such as Aya-101 and Aya-23, in every single translation direction on the FLORES-200 and WMT'23 test datasets according to COMET-22. This is achieved by plug-and-play language-specific module architecture to prevent language conflicts during training and a carefully designed training regimen with novel optimization methods to maximize the translation performance. After the final stage of training regimen, our proposed **A**daptive **R**ejection **P**reference **O**ptimization (**ARPO**) surpasses existing preference optimization methods in translation tasks. Kenton Murray, Philipp Koehn, Hieu Hoang, Akiko Eriguchi, Huda Khayrallah |
ICLR | 3 |
| 2024 | Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsabstractText generation models are notoriously vulnerable to errors in the training data. With the wide-spread availability of massive amounts of web-crawled data becoming more commonplace, how can we enhance the robustness of models trained on a massive amount of noisy web-crawled text? In our work, we propose Error Norm Truncation (ENT), a robust enhancement method to the standard training objective that truncates noisy data. Compared to methods that only uses the negative log-likelihood loss to estimate data quality, our method provides a more accurate estimation by considering the distribution of non-target tokens, which is often overlooked by previous work. Through comprehensive experiments across language modeling, machine translation, and text summarization, we show that equipping text generation models with ENT improves generation quality over standard training and previous soft and hard truncation methods. Furthermore, we show that our method improves the robustness of models against two of the most detrimental types of noise in machine translation, resulting in an increase of more than 2 BLEU points over the MLE baseline when up to 50\% of noise is added to the data. Tianjian Li, Philipp Koehn, Daniel Khashabi, Kenton Murray |
ICLR | 3 |
| 2024 | Where are you from? Geolocating Speech and Applications to Language IdentificationabstractPatrick Foley, Matthew Wiesner, Bismarck Odoom, Leibny Paola Garcia Perera, Kenton Murray, Philipp Koehn. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Patrick Foley, Matthew Wiesner, Bismarck Bamfo Odoom, L. Paola García-Perera, Kenton Murray, Philipp Koehn |
NAACL-HLT | 6 |
| 2024 | DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech TranslationabstractNon-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate text data. Although NATs generate high-quality outputs and offer faster inference than autoregressive models, they tend to produce incoherent and repetitive results due to complex data distribution (e.g., acoustic and linguistic variations in speech). In this work, we introduce DiffNorm, a diffusion-based normalization strategy that simplifies data distributions for training NAT models. After training with a self-supervised noise estimation objective, DiffNorm constructs normalized target data by denoising synthetically corrupted speech features. Additionally, we propose to regularize NATs with classifier-free guidance, improving model robustness and translation quality by randomly dropping out source information during training. Our strategies result in a notable improvement of about $+7$ ASR-BLEU for English-Spanish (En-Es) translation and $+2$ ASR-BLEU for English-French (En-Fr) on the CVSS benchmark, while attaining over $14\times$ speedup for En-Es and $5 \times$ speedup for En-Fr translations compared to autoregressive baselines. Weiting Tan, Lingfeng Shen, Daniel Khashabi, Philipp Koehn |
NeurIPS | 5 |
| 2023 | Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationabstractJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzmán |
ACL (1) | 6 |
| 2023 | Multilingual Representation Distillation with Contrastive LearningabstractMultilingual sentence representations from large models encode semantic information from two or more languages and can be used for different cross-lingual information retrieval and matching tasks.In this paper, we integrate contrastive learning into multilingual representation distillation and use it for quality estimation of parallel sentences (i.e., find semantically similar sentences that can be used as translations of each other).We validate our approach with multilingual similarity search and corpus filtering tasks.Experiments across different low-resource languages show that our method greatly outperforms previous sentence encoders such as LASER, LASER3, and LaBSE. Weiting Tan, Kevin Heffernan, Holger Schwenk, Philipp Koehn |
EACL | 4 |
| 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferabstractWe introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations.We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings.We explore various properties of pixel representations such as parameter sharing within and across scripts to better understand where they lead to positive transfer.We observe that these properties not only enable seamless cross-lingual transfer to unseen scripts, but make pixel representations more data-efficient than alternatives such as vocabulary expansion.We hope this work contributes to more extensible multilingual models for all languages and scripts. Elizabeth Salesky, Neha Verma 0001, Philipp Koehn, Matt Post |
EMNLP | 3 |
| 2023 | Condensing Multilingual Knowledge with Lightweight Language-Specific ModulesabstractIncorporating language-specific (LS) modules or Mixture-of-Experts (MoE) are proven methods to boost performance in multilingual model performance, but the scalability of these approaches to hundreds of languages or experts tends to be hard to manage.We present Language-specific Matrix Synthesis (LMS), a novel method that addresses the issue.LMS utilizes parameter-efficient and lightweight modules, reducing the number of parameters while outperforming existing methods, e.g., +1.73 BLEU over Switch Transformer on OPUS-100 multilingual translation.Additionally, we introduce Fuse Distillation (FD) to condense multilingual knowledge from multiple LS modules into a single shared module, improving model inference and storage efficiency.Our approach demonstrates superior scalability and performance compared to state-of-the-art methods. 1 * Equal contribution computational cost may only come from communication among devices (such as ALLToALL) or gate routing. Weiting Tan, Shuyue Stella Li, Yunmo Chen, Benjamin Van Durme, Philipp Koehn, Kenton Murray |
EMNLP | 6 |
| 2023 | Learning from Mistakes: Towards Robust Neural Machine Translation for Disfluent L2 SentencesabstractWe study the sentences written by second-language (L2) learners to improve the robustness of current neural machine translation (NMT) models on this type of data. Current large datasets used to train NMT systems are mostly Wikipedia or government documents written by highly competent speakers of that language, especially English. However, given that English is the most common second language, it is crucial that machine translation systems are robust against the large number of sentences written by L2 learners of English. By studying the difficulties faced by humans in their L2 acquisition process, we are able to transfer such insights to machine translation systems to recover from source-side fluency variations. In this work, we create additional training data with artificial errors similar to mistakes made by L2 learners of various fluency levels to improve the quality of the machine translation system. We test our method in zero-shot settings on the JFLEG-es (English-Spanish) dataset. The quality of our machine translation system on disfluent sentences outperforms the baseline by 1.8 BLEU scores. Shuyue Stella Li, Philipp Koehn |
MTSummit (1) | 2 |
| 2022 | Alternative Input Signals Ease Transfer in Multilingual Machine TranslationabstractSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán |
ACL (1) | 6 |
| 2022 | Toward the Limitation of Code-Switching in Cross-Lingual TransferabstractMultilingual pretrained models have shown strong cross-lingual transfer ability.Some works used code-switching sentences, which consist of tokens from multiple languages, to enhance the cross-lingual representation further, and have shown success in many zero-shot cross-lingual tasks.However, code-switched tokens are likely to cause grammatical incoherence in newly substituted sentences, and negatively affect the performance on tokensensitive tasks, such as Part-of-Speech (POS) tagging and Named-Entity-Recognition (NER).This paper mitigates the limitation of the codeswitching method by not only making the token replacement but considering the similarity between the context and the switched tokens so that the newly substituted sentences are grammatically consistent during both training and inference.We conduct experiments on crosslingual POS and NER over 30+ languages, and demonstrate the effectiveness of our method by outperforming the mBERT by 0.95 and original code-switching method by 1.67 on F1 scores.How do menschen look at and अनु भव 艺术 ?How do people look at and experience art ?How do people look at and experience art ?How do Menschen look at and अनु भव 艺术 ?(b): Ours (a): Code-Switched Training Sentence SCONJ AUX NOUN VERB ADP CCPNJ VERB NOUN PUNCT Yukun Feng, Philipp Koehn |
EMNLP | 3 |
| 2022 | Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal TransportabstractBilingual lexicons form a critical component of various natural language processing applications, including unsupervised and semisupervised machine translation and crosslingual information retrieval.We improve bilingual lexicon induction performance across 40 language pairs with a graph-matching method based on optimal transport.The method is especially strong with low amounts of supervision. Kelly Marchisio, Ali Saad-Eldin, Kevin Duh, Carey E. Priebe, Philipp Koehn |
EMNLP | 5 |
| 2022 | IsoVec: Controlling the Relative Isomorphism of Word Embedding SpacesabstractThe ability to extract high-quality translation dictionaries from monolingual word embedding spaces depends critically on the geometric similarity of the spaces-their degree of "isomorphism."We address the root-cause of faulty cross-lingual mapping: that word embedding training resulted in the underlying spaces being non-isomorphic.We incorporate global measures of isomorphism directly into the Skip-gram loss function, successfully increasing the relative isomorphism of trained word embedding spaces and improving their ability to be mapped to a shared crosslingual space.The result is improved bilingual lexicon induction in general data conditions, under domain mismatch, and with training algorithm dissimilarities.We release IsoVec at https://github.com/ kellymarchisio/isovec. Kelly Marchisio, Neha Verma 0001, Kevin Duh, Philipp Koehn |
EMNLP | 4 |
| 2022 | The Importance of Being Parameters: An Intra-Distillation Method for Serious GainsabstractRecent model pruning methods have demonstrated the ability to remove redundant parameters without sacrificing model performance.Common methods remove redundant parameters according to the parameter sensitivity, a gradient-based measure reflecting the contribution of the parameters.In this paper, however, we argue that redundant parameters can be trained to make beneficial contributions.We first highlight the large sensitivity (contribution) gap among high-sensitivity and lowsensitivity parameters and show that the model generalization performance can be significantly improved after balancing the contribution of all parameters.Our goal is to balance the sensitivity of all parameters and encourage all of them to contribute equally.We propose a general task-agnostic method, namely intradistillation, appended to the regular training loss to balance parameter sensitivity.Moreover, we also design a novel adaptive learning method to control the strength of intradistillation loss for faster convergence.Our experiments show the strong effectiveness of our methods on machine translation, natural language understanding, and zero-shot crosslingual transfer across up to 48 languages 1 , e.g., a gain of 3.54 BLEU on average across 8 language pairs from the IWSLT'14 dataset. Philipp Koehn, Kenton Murray |
EMNLP | 2 |
| 2022 | Contrastive Clustering to Mine Pseudo Parallel Data for Unsupervised Translation
Xuan-Phi Nguyen, Hongyu Gong, Yun Tang 0002, Changhan Wang, Philipp Koehn, Shafiq R. Joty |
ICLR | 5 |
| 2021 | Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataabstractWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab |
ACL/IJCNLP (1) | 8 |
| 2021 | Levenshtein Training for Word-level Quality EstimationabstractWe propose a novel scheme to use the Levenshtein Transformer to perform the task of word-level quality estimation.A Levenshtein Transformer is a natural fit for this task: trained to perform decoding in an iterative manner, a Levenshtein Transformer can learn to post-edit without explicit supervision.To further minimize the mismatch between the translation task and the word-level QE task, we propose a two-stage transfer learning procedure on both augmented data and human postediting data.We also propose heuristics to construct reference labels that are compatible with subword-level finetuning and inference.Results on WMT 2020 QE shared task dataset show that our proposed method has superior data efficiency under the data-constrained setting and competitive performance under the unconstrained setting.* Shuoyang Ding had a part-time affiliation with Microsoft at the time of this work. Shuoyang Ding, Marcin Junczys-Dowmunt, Matt Post, Philipp Koehn |
EMNLP (1) | 4 |
| 2021 | XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word AlignmentabstractCross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification.While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing.To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data.We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English.We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community. Ahmed El-Kishky, Adithya Renduchintala, James Cross 0003, Francisco Guzmán, Philipp Koehn |
EMNLP (1) | 5 |
| 2021 | Streaming Simultaneous Speech Translation with Augmented Memory TransformerabstractTransformer-based models have achieved state-of-the-art performance on speech translation tasks. However, the model architecture is not efficient enough for streaming scenarios since self-attention is computed over an entire input sequence and the computational cost grows quadratically with the length of the input sequence. Nevertheless, most of the previous work on simultaneous speech translation, the task of generating translations from partial audio input, ignores the time spent in generating the translation when analyzing the latency. With this assumption, a system may have good latency quality trade-offs but be inapplicable in real-time scenarios. In this paper, we focus on the task of streaming simultaneous speech translation, where the systems are not only capable of translating with partial input but are also able to handle very long or continuous input. We propose an end-to-end transformer-based sequence-to-sequence model, equipped with an augmented memory transformer encoder, which has shown great success on the streaming automatic speech recognition task with hybrid or transducer-based models. We conduct an empirical evaluation of the proposed model on segment, context and memory sizes and we compare our approach to a transformer with a unidirectional mask.1 Xutai Ma, Mohammad Javad Dousti, Philipp Koehn, Juan Pino 0001 |
ICASSP | 4 |
| 2021 | Learning Curricula for Multilingual Neural Machine Translation TrainingabstractLow-resource Multilingual Neural Machine Translation (MNMT) is typically tasked with improving the translation performance on one or more language pairs with the aid of high-resource language pairs. In this paper and we propose two simple search based curricula – orderings of the multilingual training data – which help improve translation performance in conjunction with existing techniques such as fine-tuning. Additionally and we attempt to learn a curriculum for MNMT from scratch jointly with the training of the translation system using contextual multi-arm bandits. We show on the FLORES low-resource translation dataset that these learned curricula can provide better starting points for fine tuning and improve overall performance of the translation system. Philipp Koehn, Sanjeev Khudanpur |
MTSummit (1) | 2 |
| 2021 | An Alignment-Based Approach to Semi-Supervised Bilingual Lexicon Induction with Small Parallel CorporaabstractAimed at generating a seed lexicon for use in downstream natural language tasks and unsupervised methods for bilingual lexicon induction have received much attention in the academic literature recently. While interesting and fully unsupervised settings are unrealistic; small amounts of bilingual data are usually available due to the existence of massively multilingual parallel corpora and or linguists can create small amounts of parallel data. In this work and we demonstrate an effective bootstrapping approach for semi-supervised bilingual lexicon induction that capitalizes upon the complementary strengths of two disparate methods for inducing bilingual lexicons. Whereas statistical methods are highly effective at inducing correct translation pairs for words frequently occurring in a parallel corpus and monolingual embedding spaces have the advantage of having been trained on large amounts of data and and therefore may induce accurate translations for words absent from the small corpus. By combining these relative strengths and our method achieves state-of-the-art results on 3 of 4 language pairs in the challenging VecMap test set using minimal amounts of parallel data and without the need for a translation dictionary. We release our implementation at www.blind-review.code. Kelly Marchisio, Philipp Koehn, Conghao Xiong |
MTSummit (1) | 2 |
| 2021 | Evaluating Saliency Methods for Neural Language ModelsabstractSaliency methods are widely used to interpret neural network predictions, but different variants of saliency methods often disagree even on the interpretations of the same prediction made by the same model.In these cases, how do we identify when are these interpretations trustworthy enough to be used in analyses?To address this question, we conduct a comprehensive and quantitative evaluation of saliency methods on a fundamental category of NLP models: neural language models.We evaluate the quality of prediction interpretations from two perspectives that each represents a desirable property of these interpretations: plausibility and faithfulness.Our evaluation is conducted on four different datasets constructed from the existing human annotation of syntactic and semantic agreements, on both sentencelevel and document-level.Through our evaluation, we identified various ways saliency methods could yield interpretations of low quality.We recommend that future work deploying such methods to neural language models should carefully validate their interpretations before drawing insights. Shuoyang Ding, Philipp Koehn |
NAACL-HLT | 2 |
| 2020 | ParaCrawl: Web-Scale Acquisition of Parallel CorporaabstractMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza |
ACL | 10 |
| 2020 | CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsabstractCross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other.We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English.In addition to curating this massive dataset, we introduce baseline methods that leverage crosslingual representations to identify aligned documents based on their textual content.Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data.Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages. Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn |
EMNLP (1) | 4 |
| 2020 | Statistical Power and Translationese in Machine Translation EvaluationabstractThe term translationese has been used to describe features of translated text, and in this paper, we provide detailed analysis of potential adverse effects of translationese on machine translation evaluation.Our analysis shows differences in conclusions drawn from evaluations that include translationese in test data compared to experiments that tested only with text originally composed in that language.For this reason we recommend that reverse-created test data be omitted from future machine translation test sets.In addition, we provide a reevaluation of a past machine translation evaluation claiming human-parity of MT.One important issue not previously considered is statistical power of significance tests applied to comparison of human and machine translation.Since the very aim of past evaluations was the investigation of ties between human and MT systems, power analysis is of particular importance, to avoid, for example, claims of human parity simply corresponding to Type II error resulting from the application of a low powered test.We provide detailed analysis of tests used in such evaluations to provide an indication of a suitable minimum sample size for future studies. Yvette Graham, Barry Haddow, Philipp Koehn |
EMNLP (1) | 3 |
| 2020 | Simulated multiple reference training improves low-resource machine translationabstractMany valid translations exist for a given sentence, yet machine translation (MT) is trained with a single reference translation, exacerbating data sparsity in low-resource settings. We introduce Simulated Multiple Reference Training (SMRT), a novel MT training method that approximates the full space of possible translations by sampling a paraphrase of the reference sentence from a paraphraser and training the MT model to predict the paraphraser's distribution over possible tokens. We demonstrate the effectiveness of SMRT in low-resource settings when translating to English, with improvements of 1.2 to 7.0 BLEU. We also find SMRT is complementary to back-translation. Huda Khayrallah, Brian Thompson 0001, Matt Post, Philipp Koehn |
EMNLP (1) | 4 |
| 2020 | Exploiting Sentence Order in Document AlignmentabstractWe present a simple document alignment method that incorporates sentence order information in both candidate generation and candidate re-scoring. Our method results in 61% relative reduction in error compared to the best previously published result on the WMT16 document alignment shared task. Our method improves downstream MT performance on web-scraped Sinhala--English documents from ParaCrawl, outperforming the document alignment method used in the most recent ParaCrawl release. It also outperforms a comparable corpora method which uses the same multilingual embeddings, demonstrating that exploiting sentence order is beneficial even if the end goal is sentence-level bitext. Brian Thompson 0001, Philipp Koehn |
EMNLP (1) | 2 |
| 2020 | Searching the Web for Cross-lingual Parallel DataabstractWhile the World Wide Web provides a large amount of text in many languages, cross-lingual parallel data is more difficult to obtain. Despite its scarcity, this parallel cross-lingual data plays a crucial role in a variety of tasks in natural language processing with applications in machine translation, cross-lingual information retrieval, and document classification, as well as learning cross-lingual representations. Here, we describe the end-to-end process of searching the web for parallel cross-lingual texts. We motivate obtaining parallel text as a retrieval problem whereby the goal is to retrieve cross-lingual parallel text from a large, multilingual web-crawled corpus. We introduce techniques for searching for cross-lingual parallel data based on language, content, and other metadata. We motivate and introduce multilingual sentence embeddings as a core tool and demonstrate techniques and models that leverage them for identifying parallel documents and sentences as well as techniques for retrieving and filtering this data. We describe several large-scale datasets curated using these techniques and show how training on sentences extracted from parallel or comparable documents mined from the Web can improve machine translation models and facilitate cross-lingual NLP. Ahmed El-Kishky, Philipp Koehn, Holger Schwenk |
SIGIR | 2 |
| 2019 | The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-EnglishabstractFrancisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino 0001, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc'Aurelio Ranzato |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Spelling-Aware Construction of Macaronic Texts for Teaching Foreign-Language VocabularyabstractAdithya Renduchintala, Philipp Koehn, Jason Eisner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Adithya Renduchintala, Philipp Koehn, Jason Eisner |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Vecalign: Improved Sentence Alignment in Linear Time and SpaceabstractBrian Thompson, Philipp Koehn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Brian Thompson 0001, Philipp Koehn |
EMNLP/IJCNLP (1) | 2 |
| 2019 | HABLex: Human Annotated Bilingual Lexicons for Experiments in Machine TranslationabstractBrian Thompson, Rebecca Knowles, Xuan Zhang, Huda Khayrallah, Kevin Duh, Philipp Koehn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Brian Thompson 0001, Rebecca Knowles, Xuan Zhang 0008, Huda Khayrallah, Kevin Duh, Philipp Koehn |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Controlling the Reading Level of Machine Translation Output
Kelly Marchisio, Jialiang Guo, Cheng-I Lai, Philipp Koehn |
MTSummit (1) | 4 |
| 2019 | Character-Aware Decoder for Translation into Morphologically Rich Languages
Adithya Renduchintala, Pamela Shapiro, Kevin Duh, Philipp Koehn |
MTSummit (1) | 4 |
| 2019 | Robust Document Representations for Cross-Lingual Information Retrieval in Low-Resource Settings
Mahsa Yarmohammadi, Xutai Ma, Sorami Hisamoto, Muhammad Mahbubur Rahman 0001, Yiming Wang 0006, Hainan Xu, Daniel Povey, Philipp Koehn, Kevin Duh |
MTSummit (1) | 8 |
| 2019 | A user study of neural interactive translation predictionabstractMachine translation (MT) on its own is generally not good enough to produce high-quality translations, so it is common to have humans intervening in the translation process to improve MT output. A typical intervention is post-editing (PE), where a human translator corrects errors in the MT output. Another is interactive translation prediction (ITP), which involves an MT system presenting a translator with translation suggestions they can accept or reject, actions the MT system then uses to present them with new, corrected suggestions. Both Macklovitch ( 2006 ) and Koehn ( 2009 ) found ITP to be an efficient alternative to unassisted translation in terms of processing time. So far, phrase-based statistical ITP has not yet proven to be faster than PE (Koehn 2009 ; Sanchis-Trilles et al. 2014 ; Underwood et al. 2014 ; Green et al. 2014 ; Alves et al. 2016 ; Alabau et al. 2016 ). In this paper we present the results of an empirical study on translation productivity in ITP with an underlying neural MT system (NITP). Our results show that over half of the professional translators in our study translated faster with NITP compared to PE, and most preferred it over PE. We also examine differences between PE and ITP in other translation productivity indicators and translators’ reactions to the technology. Rebecca Knowles, Marina Sanchez-Torron, Philipp Koehn |
Mach. Transl. | 3 |
| 2018 | An Analysis of Source Context Dependency in Neural Machine TranslationabstractThe encoder-decoder with attention model has become the state of the art for machine translation. However, more investigations are still needed to understand the internal mechanism of this end-to-end model. In this paper, we focus on how neural machine translation (NMT) models consider source information while decoding. We propose a numerical measurement of source context dependency in the NMT models and analyze the behaviors of the NMT decoder with this measurement under several circumstances. Experimental results show that this measurement is an appropriate estimate for source context dependency and consistent over different domains. Xutai Ma, Philipp Koehn |
EAMT | 3 |
| 2018 | Context and Copying in Neural Machine TranslationabstractNeural machine translation systems with subword vocabularies are capable of translating or copying unknown words.In this work, we show that they learn to copy words based on both the context in which the words appear as well as features of the words themselves.In contexts that are particularly copy-prone, they even copy words that they have already learned they should translate.We examine the influence of context and subword features on this and other types of copying behavior. Rebecca Knowles, Philipp Koehn |
EMNLP | 2 |
| 2017 | Knowledge Tracing in Sequential Learning of Inflected VocabularyabstractWe present a feature-rich knowledge tracing method that captures a student's acquisition and retention of knowledge during a foreign language phrase learning task.We model the student's behavior as making predictions under a log-linear model, and adopt a neural gating mechanism to model how the student updates their log-linear parameters in response to feedback.The gating mechanism allows the model to learn complex patterns of retention and acquisition for each feature, while the log-linear parameterization results in an interpretable knowledge state.We collect human data and evaluate several versions of the model. Adithya Renduchintala, Philipp Koehn, Jason Eisner |
CoNLL | 2 |
| 2017 | Zipporah: a Fast and Scalable Data Cleaning System for Noisy Web-Crawled Parallel CorporaabstractWe introduce Zipporah, a fast and scalable data cleaning system.We propose a novel type of bag-of-words translation feature, and train logistic regression models to classify good data and synthetic noisy data in the proposed feature space.The trained model is used to score parallel sentences in the data pool for selection.As shown in experiments, Zipporah selects a high-quality parallel corpus from a large, mixed quality data pool.In particular, for one noisy dataset, Zipporah achieves a 2.1 BLEU score improvement with using 1/5 of the data over using the entire corpus. Hainan Xu, Philipp Koehn |
EMNLP | 2 |
| 2017 | Syntax-Based Statistical Machine Translation Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn (University of Edinburgh, University of Edinburgh, Johns Hopkins University, Johns Hopkins University), edited by Graeme Hirst, volume 33), 2016, xvii+190 pp; paperback, ISBN 978-1-62705-900-8; ebook, ISBN 978-1-62705-502-4; doi: 10.2200/S00716ED1V04Y201604HLT033, $70abstractIn its early development, machine translation adopted rule-based approaches, which can include the use of language syntax. The late 1980s and early 1990s saw the inception of the statistical machine translation (SMT) approach, where translation models can be learned automatically from a parallel corpus rather than created manually by humans. Initial SMT models were word-based and phrase-based, without the use of syntactic knowledge. In phrase-based SMT, a source sentence is first segmented into phrases and then translated phrase-by-phrase with some reordering of the translated phrases in the target sentence. This has posed challenges when translating between two syntactically different languages. Syntax-based SMT approaches take advantage of syntactic knowledge within the framework of SMT. This book provides an introduction to syntax-based SMT approaches. It is a valuable resource for those who are interested in syntax-based SMT.The book consists of seven chapters. There is not an introduction chapter in this book, aside from the preface, which can be considered as a brief introduction. Readers are referred to Koehn (2010) for background knowledge. I think an introduction chapter categorized into sections would have been useful, before proceeding to describe the various models. The first two chapters provide principles applicable across various syntax-based SMT approaches. The next three chapters describe syntax-based SMT decoding in detail; this constitutes half of the book. Selected extended topics are provided in the next chapter, which is followed by a concluding chapter.Chapter 1 describes the models and formalisms applicable to syntax-based SMT. The first section describes the phrasal translation units in phrase-based SMT, its limitations, and how tree structures address the limitations of the phrase-based approach. This explanation is useful as translation units are the key difference between the phrase-based and syntax-based SMT approaches. The next two sections describe the grammar formalisms and the statistical models that define syntax-based SMT. The section that covers the grammar formalisms (i.e., synchronous context-free grammar [SCFG] and synchronous tree-substitution grammar [STSG]), would have been clearer if their differences were presented in a side-by-side illustrating example. The remainder of the chapter discusses different categories of syntax-based SMT approaches and the history of these approaches, which include string-to-string, string-to-tree, tree-to-string, and tree-to-tree SMT approaches. Although the syntax-based translation model in Galley et al. (2006) falls under the string-to-tree category, I wonder why hierarchical phrase-based SMT, or Hiero (Chiang, 2007), is not explicitly put under the string-to-string category, since Hiero also uses “unlabeled hierarchical phrases where there is no representation of linguistic categories.”Chapter 2 focuses on how the statistical framework of a syntax-based SMT approach learns its model from a word-aligned and parsed parallel text. The first section explains how phrase pairs are extracted as translation rules from a word-aligned sentence pair in phrase-based SMT (Koehn, Och, and Marcu, 2003), highlighting the definition of a phrase as a sequence of words and the alignment-consistency property of a phrase pair as defined in Och and Ney (2004). The remainder of the chapter introduces three predominant instantiations of syntax-based models: hierarchical phrase-based SMT (Hiero) (Chiang, 2007), which is a non-labeled syntax-based SMT approach arising from the phrase-based approach; syntax-augmented machine translation (SAMT), which introduces the notion of soft labels while keeping the nonlinguistic phrase notion; and GHKM (Galley et al., 2004), which only extracts translation rules consistent with constituency parse subtrees. This chapter is nicely organized and it is easy to follow the gradual evolution from phrase-based SMT to GHKM.Chapter 3 introduces the decoding formalism in the form of a directed hypergraph, defined as a set of vertices and a set of directed hyperedges. The first section introduces the notion of a weighted parse forest represented in a weighted hypergraph, representing alternative parse trees of a sentence. I found it important to pay careful attention to this section, in order to understand the next section and the following chapters. The next section presents various algorithms on a hypergraph to translate a sentence in a hypergraph representation of possible tree derivations. Overall, I found this chapter to contain many technical details. The last section of this chapter provides historical notes on the sources of these concepts. This chapter needs to be read before the next chapter, which assumes understanding of the concepts introduced in Chapter 3.Chapter 4 describes tree decoding—that is, decoding with the constituency parse tree of a source sentence as its input, focusing on the tree-to-string approach. The first two sections highlight decoding with local and non-local features, where non-local features accommodate n-gram language models and are more complex than local features. The next section is devoted to an in-depth description of a beam search algorithm on the parse tree of a source sentence. The description could have been improved if the running example showed the decoding steps. The next two sections present extensions to the concepts introduced in the earlier part of this chapter, by providing references to more efficient hypergraph operations. The content of this section requires readers who are interested in implementing an efficient tree-based algorithm to go through the cited references. Brief historical notes conclude this chapter nicely, by pointing to relevant materials for further reading.Chapter 5 describes string decoding with a source sentence string as its input. The first two sections describe beam search decoding algorithms in a binary SCFG, namely, a maximum of two non-terminal symbols on the right-hand side of each rule, adopted in Hiero and SAMT. The algorithms covered are a basic algorithm and an optimized algorithm. The complexity comparison between the two is nicely presented here, emphasizing the complexity reduction achieved by algorithm optimization. The handling of non-binary rules is described in the following section, illustrated by GHKM rule extraction. A mid-chapter summary section divides this chapter into two parts: beam search decoding and parsing. The second part describes parsing algorithms in the context of shared-category SCFG, assuming the same set of non-terminal symbols for the left-hand and right-hand sides of a rule, followed by a section extending the algorithm to STSG and distinct-category SCFG. The organization of this chapter is excellent. However, I feel that the inclusion of distinct-category SCFG decoding does not fit well into this chapter, as string decoding in string-to-tree SMT requires no knowledge of the source syntax. The historical notes also do not provide any references of prior work on string decoding using distinct-category SCFG.Chapter 6 contains various selected topics on syntax-based SMT. The first section discusses tree transformations, which make translation rule learning more effective. The description of non-context-free models serves as a prelude to the next section on dependency-based SMT, which covers dependency treelet (equivalent to the tree-to-string approach) and string-to-dependency (equivalent to the string-to-tree approach). The next section focuses on the ability of syntax-based SMT to have a more grammatical output compared with phrase-based SMT, although there is still room for improvement, including the use of unification grammars and semantic properties. Finally, the last section of this chapter explains how MT evaluation benefits from syntax-based SMT principles. Overall, this chapter enriches readers' knowledge beyond basic syntax-based SMT in the earlier chapters. I would also suggest the inclusion of phrase-based decoding approaches that use syntax-based features (Cherry, 2008; Chang et al., 2009).Chapter 7 nicely concludes this book by discussing the comparison between phrase-based and syntax-based SMT approaches and proposing possible future developments of syntax-based SMT. The chapter also highlights that syntax-driven MT predates statistical MT, as I mentioned at the beginning of this review.Overall, I found this book to be a useful reference book for those interested in syntax-based SMT. The book is well organized, which makes it easy for readers to refer to specific aspects of syntax-based SMT. An improvement can be made to the presentation of ideas in this book. Throughout the book, there are many technical keywords, resulting from the complexity of syntax-based SMT. It would be useful to highlight these keywords in a side bar to remind readers that they are important keywords. In addition, although examples are given throughout the book, it would be even more useful to use these examples to illustrate how the algorithms work, so that readers can gain a better understanding of the algorithms. Philip Williams, Rico Sennrich, Matt Post, Philipp Koehn, Graeme Hirst, Christian Hadiwinoto |
Comput. Linguistics | 4 |
| 2016 | User Modeling in Language Learning with Macaronic TextsabstractForeign language learners can acquire new vocabulary by using cognate and context clues when reading.To measure such incidental comprehension, we devise an experimental framework that involves reading mixed-language "macaronic" sentences.Using data collected via Amazon Mechanical Turk, we train a graphical model to simulate a human subject's comprehension of foreign words, based on cognate clues (edit distance to an English word), context clues (pointwise mutual information), and prior exposure.Our model does a reasonable job at predicting which words a user will be able to understand, which should facilitate the automatic construction of comprehensible text for personalized foreign language education. Adithya Renduchintala, Rebecca Knowles, Philipp Koehn, Jason Eisner |
ACL (1) | 3 |
| 2016 | Analyzing Learner Understanding of Novel L2 VocabularyabstractIn this work, we explore how learners can infer second-language noun meanings in the context of their native language. Motivated by an interest in building interactive tools for language learning, we collect data on three word-guessing tasks, analyze their difficulty, and explore the types of errors that novice learners make. We train a log-linear model for predicting our subjects’ guesses of word meanings in varying kinds of contexts. The model’s predictions correlate well with subject performance, and we provide quantitative and qualitative analyses of both human and model performance. Rebecca Knowles, Adithya Renduchintala, Philipp Koehn, Jason Eisner |
CoNLL | 3 |
| 2015 | The Operation Sequence Model - Combining N-Gram-Based and Phrase-Based Statistical Machine TranslationabstractIn this article, we present a novel machine translation model, the Operation Sequence Model (OSM), which combines the benefits of phrase-based and N-gram-based statistical machine translation (SMT) and remedies their drawbacks. The model represents the translation process as a linear sequence of operations. The sequence includes not only translation operations but also reordering operations. As in N-gram-based SMT, the model is: (i) based on minimal translation units, (ii) takes both source and target information into account, (iii) does not make a phrasal independence assumption, and (iv) avoids the spurious phrasal segmentation problem. As in phrase-based SMT, the model (i) has the ability to memorize lexical reordering triggers, (ii) builds the search graph dynamically, and (iii) decodes with large translation units during search. The unique properties of the model are (i) its strong coupling of reordering and translation where translation and reordering decisions are conditioned on n previous translation and reordering decisions, and (ii) the ability to model local and long-range reorderings consistently. Using BLEU as a metric of translation accuracy, we found that our system performs significantly better than state-of-the-art phrase-based systems (Moses and Phrasal) and N-gram-based systems (Ncode) on standard translation tasks. We compare the reordering component of the OSM to the Moses lexical reordering model by integrating it into Moses. Our results show that OSM outperforms lexicalized reordering on all translation tasks. The translation quality is shown to be improved further by learning generalized representations with a POS-based OSM. Nadir Durrani, Helmut Schmid, Alexander Fraser 0001, Philipp Koehn, Hinrich Schütze |
Comput. Linguistics | 4 |
| 2014 | Investigating the Usefulness of Generalized Word Representations in SMT
Nadir Durrani, Philipp Koehn, Helmut Schmid, Alexander Fraser 0001 |
COLING | 2 |
| 2014 | CASMACAT: A Computer-assisted Translation WorkbenchabstractCASMACAT is a modular, web-based translation workbench that offers advanced functionalities for computer-aided translation and the scientific study of human translation: automatic interaction with machine translation (MT) engines and translation memories (TM) to obtain raw translations or close TM matchesn for conventional post-editing; interactive translation prediction based on an MT engine’s search graph, detailed recording and replay of edit actions and translator’s gaze (the latter via eye-tracking), and the support of e-pen as an alternative input device. The system is open source sofware and interfaces with multiple MT systems. Vicente Alabau, Christian Buck, Michael Carl, Francisco Casacuberta, Mercedes García-Martínez, Ulrich Germann, Jesús González-Rubio, Robin L. Hill, Philipp Koehn, Luis A. Leiva, Bartolomé Mesa-Lao, Daniel Ortiz-Martínez, Herve Saint-Amand, Germán Sanchis-Trilles, Chara Tsoukala |
EACL | 9 |
| 2014 | Integrating an Unsupervised Transliteration Model into Statistical Machine TranslationabstractWe investigate three methods for integrating an unsupervised transliteration model into an end-to-end SMT system.We induce a transliteration model from parallel data and use it to translate OOV words.Our approach is fully unsupervised and language independent.In the methods to integrate transliterations, we observed improvements from 0.23-0.75(∆ 0.41) BLEU points across 7 language pairs.We also show that our mined transliteration corpora provide better rule coverage and translation quality compared to the gold standard transliteration corpora. Nadir Durrani, Hassan Sajjad 0001, Hieu Hoang, Philipp Koehn |
EACL | 4 |
| 2014 | Dynamic Topic Adaptation for Phrase-based MTabstractTranslating text from diverse sources poses a challenge to current machine translation systems which are rarely adapted to structure beyond corpus level. We explore topic adaptation on a diverse data set and present a new bilingual vari-ant of Latent Dirichlet Allocation to com-pute topic-adapted, probabilistic phrase translation features. We dynamically in-fer document-specific translation proba-bilities for test sets of unknown origin, thereby capturing the effects of document context on phrase translations. We show gains of up to 1.26 BLEU over the base-line and 1.04 over a domain adaptation benchmark. We further provide an anal-ysis of the domain-specific data and show additive gains of our model in combination with other types of topic-adapted features. 1 Eva Hasler, Phil Blunsom, Philipp Koehn, Barry Haddow |
EACL | 3 |
| 2014 | Improving machine translation via triangulation and transliteration
Nadir Durrani, Philipp Koehn |
EAMT | 2 |
| 2014 | CASMACAT: cognitive analysis and statistical methods for advanced computer aided translation
Philipp Koehn, Michael Carl, Francisco Casacuberta, Eva Marcos |
EAMT | 1 |
| 2014 | Interactive translation prediction versus conventional post-editing in practice: a study with the CasMaCat workbench
Germán Sanchis-Trilles, Vicente Alabau, Christian Buck, Michael Carl, Francisco Casacuberta, Mercedes García-Martínez, Ulrich Germann, Jesús González-Rubio, Robin L. Hill, Philipp Koehn, Luis A. Leiva, Bartolomé Mesa-Lao, Daniel Ortiz-Martínez, Herve Saint-Amand, Chara Tsoukala, Enrique Vidal 0001 |
Mach. Transl. | 10 |
| 2013 | Dirt Cheap Web-Scale Parallel Text from the Common Crawl
Jason Smith 0006, Herve Saint-Amand, Magdalena Plamada, Philipp Koehn, Chris Callison-Burch, Adam Lopez |
ACL (1) | 4 |
| 2013 | Grouping Language Model Boundary Words to Speed K-Best Extraction from Hypergraphs
Kenneth Heafield, Philipp Koehn, Alon Lavie |
HLT-NAACL | 2 |
| 2012 | Language Model Rest Costs and Space-Efficient Storage
Kenneth Heafield, Philipp Koehn, Alon Lavie |
EMNLP-CoNLL | 2 |
| 2012 | Semi-supervised discriminative language modeling for Turkish ASRabstractWe present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results. Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 19 |
| 2012 | Hallucinated n-best lists for discriminative language modelingabstractThis paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods. Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 16 |
| 2012 | Continuous space discriminative language modelingabstractDiscriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well. Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
ICASSP | 16 |
| 2012 | Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley |
INTERSPEECH | 17 |
| 2011 | Soft Dependency Constraints for Reordering in Hierarchical Phrase-Based Translation
Yang Gao 0005, Philipp Koehn, Alexandra Birch |
EMNLP | 2 |
| 2010 | Enabling Monolingual Translators: Post-Editing vs. Options
Philipp Koehn |
HLT-NAACL | 1 |
| 2010 | Monte Carlo techniques for phrase-based translation
Abhishek Arun, Barry Haddow, Philipp Koehn, Adam Lopez, Chris Dyer, Phil Blunsom |
Mach. Transl. | 3 |
| 2009 | Monte Carlo inference and maximization for phrase-based translation
Abhishek Arun, Chris Dyer, Barry Haddow, Phil Blunsom, Adam Lopez, Philipp Koehn |
CoNLL | 6 |
| 2009 | Improving Mid-Range Re-Ordering Using Templates of Factors
Hieu Hoang, Philipp Koehn |
EACL | 2 |
| 2009 | Word Lattices for Multi-Source Translation
Josh Schroeder, Trevor Cohn, Philipp Koehn |
EACL | 3 |
| 2009 | 462 Machine Translation Systems for Europe
Philipp Koehn, Alexandra Birch, Ralf Steinberger |
MTSummit | 1 |
| 2009 | Interactive Assistance to Human Translators using Statistical Machine Translation Methods
Philipp Koehn, Barry Haddow |
MTSummit | 1 |
| 2009 | A process study of computer-aided translation
Philipp Koehn |
Mach. Transl. | 1 |
| 2009 | Review of Cyril Goutte, Nicola Cancedda, Marc Dymetman, and George Foster (eds): Learning machine translation
Philipp Koehn |
Mach. Transl. | 1 |
| 2009 | Introduction to the Special Issue on Machine Translation of Asian LanguagesabstractNo abstract available. David Chiang 0001, Philipp Koehn |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | Enriching Morphologically Poor Languages for Statistical Machine Translation
Eleftherios Avramidis, Philipp Koehn |
ACL | 2 |
| 2008 | Predicting Success in Machine Translation
Alexandra Birch, Miles Osborne, Philipp Koehn |
EMNLP | 3 |
| 2008 | Large and Diverse Language Models for Statistical Machine Translation
Holger Schwenk, Philipp Koehn |
IJCNLP | 2 |
| 2007 | Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst |
ACL | 1 |
| 2007 | Factored Translation Models
Philipp Koehn, Hieu Hoang |
EMNLP-CoNLL | 1 |
| 2007 | Chinese Syntactic Reordering for Statistical Machine Translation
Chao Wang 0018, Michael Collins 0001, Philipp Koehn |
EMNLP-CoNLL | 3 |
| 2007 | Online learning methods for discriminative training of phrase based statistical machine translation
Abhishek Arun, Philipp Koehn |
MTSummit | 2 |
| 2006 | Re-evaluation the Role of Bleu in Machine Translation Research
Chris Callison-Burch, Miles Osborne, Philipp Koehn |
EACL | 3 |
| 2006 | Improved Statistical Machine Translation Using Paraphrases
Chris Callison-Burch, Philipp Koehn, Miles Osborne |
HLT-NAACL | 2 |
| 2005 | Clause Restructuring for Statistical Machine TranslationabstractWe describe a method for incorporating syntactic information in statistical machine translation systems. The first step of the method is to parse the source language string that is being translated. The second step is to apply a series of transformations to the parse tree, effectively reordering the surface string on the source language side of the translation system. The goal of this step is to recover an underlying word order that is closer to the target language word-order than the original string. The reordering approach is applied as a pre-processing step in both the training and decoding phases of a phrase-based statistical MT system. We describe experiments on translation from German to English, showing an improvement from 25.2% Bleu score for a baseline system to 26.8% Bleu score for the system with reordering, a statistically significant improvement. Michael Collins 0001, Philipp Koehn, Ivona Kucerova |
ACL | 2 |
| 2005 | Europarl: A Parallel Corpus for Statistical Machine TranslationabstractWe collected a corpus of parallel text in 11 languages from the proceedings of the European Parliament, which are published on the web. This corpus has found widespread use in the NLP community. Here, we focus on its acquisition and its application as training data for statistical machine translation (SMT). We trained SMT systems for 110 language pairs, which reveal interesting clues into the challenges ahead. Philipp Koehn |
MTSummit | 1 |
| 2004 | Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn |
EMNLP | 1 |
| 2003 | Feature-Rich Statistical Translation of Noun PhrasesabstractWe define noun phrase translation as a subtask of machine translation. This enables us to build a dedicated noun phrase translation subsystem that improves over the currently best general statistical machine translation methods by incorporating special modeling and special features. We achieved 65.5% translation accuracy in a German-English translation task vs. 53.2% with IBM Model 4. Philipp Koehn, Kevin Knight |
ACL | 1 |
| 2003 | Empirical Methods for Compound Splitting
Philipp Koehn, Kevin Knight |
EACL | 1 |
| 2003 | What's New in Statistical Machine Translation
Kevin Knight, Philipp Koehn |
HLT-NAACL | 2 |
| 2003 | Statistical Phrase-Based Translation
Philipp Koehn, Franz Josef Och, Daniel Marcu |
HLT-NAACL | 1 |
| 2003 | Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight |
HLT-NAACL | 11 |
| 2002 | Translation with Scarce Bilingual Resources
Yaser Al-Onaizan, Ulrich Germann, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Daniel Marcu, Kenji Yamada |
Mach. Transl. | 5 |
| 2001 | Knowledge Sources for Word-Level Translation Models
Philipp Koehn, Kevin Knight |
EMNLP | 1 |
| 2000 | Improving intonational phrasing with syntactic informationabstractThe prediction of intonational phrase boundaries from raw text is an important step for a text-to-speech system: locating where to place short pauses enables more natural sounding speech, that can be more easily understood. We improved upon earlier work [Hirschberg and Prieto, 1996] by adding syntactic information gained from a high-accuracy parser [Collins, 1999]. We report significant improvement using various experimental setups. We also show that our improved method comes close to interannotator agreement. Philipp Koehn, Steven P. Abney, Julia Hirschberg, Michael Collins 0001 |
ICASSP | 1 |