George F. Foster

dblp:02/1712 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
10since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Importance-Aware Data Augmentation for Document-Level Neural Machine Translation
abstract
Minghao Wu, Yufei Wang, George Foster, Lizhen Qu, Gholamreza Haffari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Minghao Wu, Yufei Wang 0003, George F. Foster, Lizhen Qu, Gholamreza Haffari
EACL (1)3
2024 Finding Replicable Human Evaluations via Stable Ranking Probability
abstract
Parker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, Markus Freitag. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Parker Riley, Daniel Deutsch, George F. Foster, Viresh Ratnakar, Ali Dabirmoghaddam, Markus Freitag
NAACL-HLT3
2024 To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation
abstract
Abstract We conduct a large-scale fine-grained comparative analysis of machine translations (MTs) against human translations (HTs) through the lens of morphosyntactic divergence. Across three language pairs and two types of divergence defined as the structural difference between the source and the target, MT is consistently more conservative than HT, with less morphosyntactic diversity, more convergent patterns, and more one-to-one alignments. Through analysis on different decoding algorithms, we attribute this discrepancy to the use of beam search that biases MT towards more convergent patterns. This bias is most amplified when the convergent pattern appears around 50% of the time in training data. Lastly, we show that for a majority of morphosyntactic divergences, their presence in HT is correlated with decreased MT performance, presenting a greater challenge for MT systems.
Jiaming Luo, Colin Cherry, George F. Foster
Trans. Assoc. Comput. Linguistics3
2023 Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability
abstract
Large, multilingual language models exhibit surprisingly good zero-or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural translation systems.We investigate the role of incidental bilingualism-the unintentional consumption of bilingual signals, including translation examples-in explaining the translation capabilities of large language models, taking the Pathways Language Model (PaLM) as a case study.We introduce a mixed-method approach to measure and understand incidental bilingualism at scale.We show that PaLM is exposed to over 30 million translation pairs across at least 44 languages.Furthermore, the amount of incidental bilingual content is highly correlated with the amount of monolingual in-language content for non-English languages.We relate incidental bilingual content to zero-shot prompts and show that it can be used to mine new prompts to improve PaLM's out-of-English zero-shot translation quality.Finally, in a series of small-scale ablations, we show that its presence has a substantial impact on translation capabilities, although this impact diminishes with model scale.
Eleftheria Briakou, Colin Cherry, George F. Foster
ACL (1)3
2023 Prompting PaLM for Translation: Assessing Strategies and Performance
abstract
David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, George Foster. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, George F. Foster
ACL (1)6
2023 Document Flattening: Beyond Concatenating Context for Document-Level Neural Machine Translation
abstract
Existing work in document-level neural machine translation commonly concatenates several consecutive sentences as a pseudodocument, and then learns inter-sentential dependencies.This strategy limits the model's ability to leverage information from distant context.We overcome this limitation with a novel Document Flattening (DOCFLAT) technique that integrates FLAT-BATCH ATTEN-TION (FBA) and NEURAL CONTEXT GATE (NCG) into Transformer model to utilize information beyond the pseudo-document boundaries.FBA allows the model to attend to all the positions in the batch and learns the relationships between positions explicitly and NCG identifies the useful information from the distant context.We conduct comprehensive experiments and analyses on three benchmark datasets for English-German translation, and validate the effectiveness of two variants of DOCFLAT.Empirical results show that our approach outperforms strong baselines with statistical significance on BLEU, COMET and accuracy on the contrastive test set.The analyses highlight that DOCFLAT is highly effective in capturing the long-range information.
Minghao Wu, George F. Foster, Lizhen Qu, Gholamreza Haffari
EACL2
2023 Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
abstract
Kendall's τ is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations.Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be handled, a gray area that has motivated different variants in the literature.We demonstrate that, in settings like modern MT meta-evaluation, existing variants have weaknesses arising from their handling of ties, and in some situations can even be gamed.We propose instead to meta-evaluate metrics with a version of pairwise accuracy that gives metrics credit for correctly predicting ties, in combination with a tie calibration procedure that automatically introduces ties into metric scores, enabling fair comparison between metrics that do and do not predict ties.We argue and provide experimental evidence that these modifications lead to fairer ranking-based assessments of metric performance.1
Daniel Deutsch, George F. Foster, Markus Freitag
EMNLP2
2023 The Unreasonable Effectiveness of Few-shot Learning for Machine Translation
abstract
We demonstrate the potential of few-shot translation systems, trained with unpaired language data, for both high and low-resource language pairs. We show that with only 5 examples of high-quality translation data shown at inference, a transformer decoder-only model trained solely with self-supervised learning, is able to match specialized supervised state-of-the-art models as well as more general commercial translation systems. In particular, we outperform the best performing system on the WMT'21 English-Chinese news translation task by only using five examples of English-Chinese parallel data at inference. Furthermore, the resulting models are two orders of magnitude smaller than state-of-the-art language models. We then analyze the factors which impact the performance of few-shot translation systems, and highlight that the quality of the few-shot demonstrations heavily determines the quality of the translations generated by our models. Finally, we show that the few-shot paradigm also provides a way to control certain attributes of the translation --- we show that we are able to control for regional varieties and formality using only a five examples at inference, paving the way towards controllable machine translation systems.
Xavier Garcia, Yamini Bansal, Colin Cherry, George F. Foster, Maxim Krikun, Melvin Johnson, Orhan Firat
ICML4
2021 Assessing Reference-Free Peer Evaluation for Machine Translation
abstract
Sweta Agrawal, George Foster, Markus Freitag, Colin Cherry. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Sweta Agrawal, George F. Foster, Markus Freitag, Colin Cherry
NAACL-HLT2
2021 Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation
abstract
Abstract Human evaluation of modern high-quality machine translation systems is a difficult problem, and there is increasing evidence that inadequate evaluation procedures can lead to erroneous conclusions. While there has been considerable research on human evaluation, the field still lacks a commonly accepted standard procedure. As a step toward this goal, we propose an evaluation methodology grounded in explicit error analysis, based on the Multidimensional Quality Metrics (MQM) framework. We carry out the largest MQM research study to date, scoring the outputs of top systems from the WMT 2020 shared task in two language pairs using annotations provided by professional translators with access to full document context. We analyze the resulting data extensively, finding among other results a substantially different ranking of evaluated systems from the one established by the WMT crowd workers, exhibiting a clear preference for human over machine output. Surprisingly, we also find that automatic metrics based on pre-trained embeddings can outperform human crowd workers. We make our corpus publicly available for further research.
Markus Freitag, George F. Foster, David Grangier, Viresh Ratnakar, Qijun Tan, Wolfgang Macherey
Trans. Assoc. Comput. Linguistics2
2020 Inference Strategies for Machine Translation with Conditional Masking
abstract
Conditional masked language model (CMLM) training has proven successful for nonautoregressive and semi-autoregressive sequence generation tasks, such as machine translation.Given a trained CMLM, however, it is not clear what the best inference strategy is.We formulate masked inference as a factorization of conditional probabilities of partial sequences, show that this does not harm performance, and investigate a number of simple heuristics motivated by this perspective.We identify a thresholding strategy that has advantages over the standard "mask-predict" algorithm, and provide analyses of its behavior on machine translation tasks.
Julia Kreutzer, George F. Foster, Colin Cherry
EMNLP (1)2
2020 Re-Translation Strategies for Long Form, Simultaneous, Spoken Language Translation
abstract
We investigate the problem of simultaneous machine translation of long-form speech content. We target a continuous speech-to-text scenario, generating translated captions for a live audio feed, such as a lecture or play-by-play commentary. As this scenario allows for revisions to our incremental translations, we adopt a re-translation approach to simultaneous translation, where the source is repeatedly translated from scratch as it grows. This approach naturally exhibits very low latency and high final quality, but at the cost of incremental instability as the output is continuously refined. We experiment with a pipeline of industry-grade speech recognition and translation tools, augmented with simple inference heuristics to improve stability. We use TED Talks as a source of multilingual test data, developing our techniques on English-to-German spoken language translation. Our minimalist approach to simultaneous translation allows us to scale our final evaluation to several other target languages, dramatically improving incremental stability for all of them.
Naveen Arivazhagan, Colin Cherry, Te I, Wolfgang Macherey, Pallavi Baljekar, George F. Foster
ICASSP6
2020 Shaping the Narrative Arc: Information-Theoretic Collaborative DialoguePaper type: Technical Paper
Kory W. Mathewson, Pablo Samuel Castro, Colin Cherry, George F. Foster, Marc G. Bellemare
ICCC4
2018 The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
abstract
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes
ACL (1)6
2018 Revisiting Character-Based Neural Machine Translation with Capacity and Compression
abstract
Translating characters instead of words or word-fragments has the potential to simplify the processing pipeline for neural machine translation (NMT), and improve results by eliminating hyper-parameters and manual feature engineering.However, it results in longer sequences in which each symbol contains less information, creating both modeling and computational challenges.In this paper, we show that the modeling problem can be solved by standard sequence-to-sequence architectures of sufficient depth, and that deep models operating at the character level outperform identical models operating over word fragments.This result implies that alternative architectures for handling character input are better viewed as methods for reducing computation time than as improved ways of modeling longer sequences.From this perspective, we evaluate several techniques for characterlevel NMT, verify that they do not match the performance of our deep character baseline model, and evaluate the performance versus computation time tradeoffs they offer.Within this framework, we also perform the first evaluation for NMT of conditional computation over time, in which the model learns which timesteps can be skipped, rather than having them be dictated by a fixed schedule specified before training begins.
Colin Cherry, George F. Foster, Ankur Bapna, Orhan Firat, Wolfgang Macherey
EMNLP2
2017 A Challenge Set Approach to Evaluating Machine Translation
abstract
Neural machine translation represents an exciting leap forward in translation quality.But what longstanding weaknesses does it resolve, and which remain?We address these questions with a challenge set approach to translation evaluation and error analysis.A challenge set consists of a small set of sentences, each hand-designed to probe a system's capacity to bridge a particular structural divergence between languages.To exemplify this approach, we present an English-French challenge set, and use it to analyze phrase-based and neural systems.The resulting analysis provides not only a more fine-grained picture of the strengths of neural systems, but also insight into which linguistic phenomena remain out of reach.
Pierre Isabelle, Colin Cherry, George F. Foster
EMNLP3
2013 Vector Space Model for Adaptation in Statistical Machine Translation
Boxing Chen, Roland Kuhn 0001, George F. Foster
ACL (1)3
2013 Simulating Discriminative Training for Linear Mixture Adaptation in Statistical Machine Translation
George F. Foster, Boxing Chen, Roland Kuhn 0001
MTSummit1
2013 PEPr: Post-Edit Propagation Using Phrase-based Statistical Machine Translation
Michel Simard, George F. Foster
MTSummit2
2013 Adaptation of Reordering Models for Statistical Machine Translation
Boxing Chen, George F. Foster, Roland Kuhn 0001
HLT-NAACL2
2012 Mixing Multiple Translation Models in Statistical Machine Translation
Majid Razmara, George F. Foster, Baskaran Sankaran, Anoop Sarkar
ACL (1)2
2012 Batch Tuning Strategies for Statistical Machine Translation
Colin Cherry, George F. Foster
HLT-NAACL2
2011 Unpacking and Transforming Feature Functions: New Ways to Smooth Phrase Tables
Boxing Chen, Roland Kuhn 0001, George F. Foster, Howard Johnson
MTSummit3
2010 Bilingual Sense Similarity for Statistical Machine Translation
Boxing Chen, George F. Foster, Roland Kuhn 0001
ACL2
2010 Phrase Clustering for Smoothing TM Probabilities - or, How to Extract Paraphrases from Phrase Tables
Roland Kuhn 0001, Boxing Chen, George F. Foster, Evan Stratford
COLING3
2010 Discriminative Instance Weighting for Domain Adaptation in Statistical Machine Translation
George F. Foster, Cyril Goutte, Roland Kuhn 0001
EMNLP1
2009 Phrase Translation Model Enhanced with Association based Features
Boxing Chen, George F. Foster, Roland Kuhn 0001
MTSummit2
2008 Tighter Integration of Rule-Based and Statistical MT in Serial System Combination
Nicola Ueffing, Jens Stephan, Evgeny Matusov, Loïc Dugast, George F. Foster, Roland Kuhn 0001, Jean Senellart
COLING5
2007 Improving Translation Quality by Discarding Most of the Phrasetable
Howard Johnson, Joel D. Martin, George F. Foster, Roland Kuhn 0001
EMNLP-CoNLL3
2006 Phrasetable Smoothing for Statistical Machine Translation
George F. Foster, Roland Kuhn 0001, Howard Johnson
EMNLP1
2006 Segment Choice Models: Feature-Rich Models for Global Distortion in Statistical Machine Translation
Roland Kuhn 0001, Denis Yuen, Michel Simard, Patrick Paul, George F. Foster, Eric Joanis, Howard Johnson
HLT-NAACL5
2004 Confidence Estimation for Machine Translation
John Blatz, Erin Fitzgerald, George F. Foster, Simona Gandrabur, Cyril Goutte, Alex Kulesza, Alberto Sanchís, Nicola Ueffing
COLING3
2004 Adaptive Language and Translation Models for Interactive Machine Translation
Laurent Nepveu, Guy Lapalme, Philippe Langlais, George F. Foster
EMNLP4
2003 Confidence estimation for translation prediction
Simona Gandrabur, George F. Foster
CoNLL2
2003 Statistical machine translation: rapid development with limited resources
abstract
We describe an experiment in rapid development of a statistical machine translation (SMT) system from scratch, using limited resources: under this heading we include not only training data, but also computing power, linguistic knowledge, programming effort, and absolute time.
George F. Foster, Simona Gandrabur, Philippe Langlais, Pierre Plamondon, Graham Russell, Michel Simard
MTSummit1
2002 User-Friendly Text Prediction For Translators
abstract
Text prediction is a form of interactive machine translation that is well suited to skilled translators. In principle it can assist in the production of a target text with minimal disruption to a translator's normal routine. However, recent evaluations of a prototype prediction system showed that it significantly decreased the productivity of most translators who used it. In this paper, we analyze the reasons for this and propose a solution which consists in seeking predictions that maximize the expected benefit to the translator, rather than just trying to anticipate some amount of upcoming text. Using a model of a "typical translator" constructed from data collected in the evaluations of the prediction prototype, we show that this approach has the potential to turn text prediction into a help rather than a hindrance to a translator.
George F. Foster, Philippe Langlais, Guy Lapalme
EMNLP1
2001 Integrating bilingual lexicons in a probabilistic translation assistant
abstract
In this paper, we present a way to integrate bilingual lexicons into an operational probabilistic translation assistant (TransType). These lexicons could be any resource available to the translator (e.g. terminological lexicons) or any resource statistically derived from training material. We describe a bilingual lexicon acquisition process that we developped and we evaluate from a theoretical point of view its benefits to a translation completion task.
Philippe Langlais, George F. Foster, Guy Lapalme
MTSummit2
2000 A Maximum Entropy/Minimum Divergence Translation Model
abstract
I present empirical comparisons between a linear combination of standard statistical language and translation models and an equivalent Maximum Entropy/Minimum Divergence (MEMD) model, using several different methods for automatic feature selection. The MEMD model significantly outperforms the standard model in test corpus perplexity, even though it has far fewer parameters.
George F. Foster
ACL1
2000 Evaluation of TRANSTYPE, a Computer-aided Translation Typing System: A Comparison of a Theoretical- and a User-oriented Evaluation Procedures
Philippe Langlais, Sébastien Sauvé, George F. Foster, Elliott Macklovitch, Guy Lapalme
LREC3
2000 Unit Completion for a Computer-aided Translation Typing System
Philippe Langlais, George F. Foster, Guy Lapalme
Mach. Transl.2
1997 Target-Text Mediated Interactive Machine Translation
George F. Foster, Pierre Isabelle, Pierre Plamondon
Mach. Transl.1
1996 Word Completion- A First Step Toward Target-Text Mediated IMT
George F. Foster, Pierre Isabelle, Pierre Plamondon
COLING1
1995 French speech recognition in an automatic dictation system for translators: the transtalk project
abstract
This paper describes a system designed for use by professional translators that enables them to dictate their translation. Because the speech recognizer has access to the source text as well as the spoken translation, a statistical translation model can guide recognition. This can be done in many different ways---which is best? We discuss the experiments that led to integration of the translation model in a way that improves both speed and performance. 1 Introduction The TransTalk project attempts to integrate speech recognition and machine translation in a way that makes maximal use of their complementary strengths. Professional translators often dictate their translations first and have them typed afterwards. If they dictate to a speech recognition system instead, and if that system has access to the source language text, it can use probabilistic translation models to aid recognition. For instance, if the speech recognition system is deciding between the acoustically similar French ...
Julie Brousseau, Caroline Drouin, George F. Foster, Pierre Isabelle, Roland Kuhn 0001, Yves Normandin, Pierre Plamondon
EUROSPEECH3
1994 Towards an automatic dictation system for translators : the transtalk project
abstract
Professional translators often dictate their translations orally and have them typed afterwards. The TransTalk project aims at automating the second part of this process. Its originality as a dictation system lies in the fact that both the acoustic signal produced by the translator and the source text under translation are made available to the system. Probable translations of the source text can be predicted and these predictions used to help the speech recognition system in its lexical choices. We present the results of the first prototype, which show a marked improvement in the performance of the speech recognition task when translation predictions are taken into account. 1 Introduction The integration of machine translation and speech technology is currently the focus of major projects in several countries [5, 6, 9]. Usually, the aim of these efforts is some type of speech-to-speech translation, where speech recognition, machine translation and speech synthesis are performed seque...
Marc Dymetman, Julie Brousseau, George F. Foster, Pierre Isabelle, Yves Normandin, Pierre Plamondon
ICSLP3