Marcin Junczys-Dowmunt

dblp:79/7927 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Machine translation · 67% Language models and text generation · 20% Efficient and distributed learning · 10%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
statistical machine translation
0.522016
Phrase-based Machine Translation is State-of-the-Art for Automatic Grammatical Error Correction · EMNLP 2016
Target-Side Context for Discriminative Models in Statistical Machine Translation · ACL (1) 2016
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.512021
Levenshtein Training for Word-level Quality Estimation · EMNLP (1) 2021
Natural language and speech › Machine translation › machine translation evaluation › translation quality estimation
word-level quality estimation
0.512021
Levenshtein Training for Word-level Quality Estimation · EMNLP (1) 2021
Natural language and speech › Language models and text generation › text generation
grammatical error correction
0.522016
Phrase-based Machine Translation is State-of-the-Art for Automatic Grammatical Error Correction · EMNLP 2016
Human Evaluation of Grammatical Error Correction Systems · EMNLP 2015
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent
0.312018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Machine translation
neural machine translation
0.312018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation
0.212016
Phrase-based Machine Translation is State-of-the-Art for Automatic Grammatical Error Correction · EMNLP 2016
Natural language and speech › Language models and text generation › language modeling › language model architecture
levenshtein transformer
0.112021
Levenshtein Training for Word-level Quality Estimation · EMNLP (1) 2021
Natural language and speech › Machine translation
machine translation evaluation
0.122016
Phrase-based Machine Translation is State-of-the-Art for Automatic Grammatical Error Correction · EMNLP 2016
Human Evaluation of Grammatical Error Correction Systems · EMNLP 2015
Machine learning › Optimization for machine learning › learning rate
learning rate scaling
0.112018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Language models and text generation › text evaluation
human evaluation
0.112015
Human Evaluation of Grammatical Error Correction Systems · EMNLP 2015

Methods — techniques the papers use, named apart from their topics

transfer learning · 0.5levenshtein transformer · 0.5iterative decoding · 0.5momentum · 0.3local optimizers · 0.3asynchronous SGD · 0.3parameter tuning · 0.2moses · 0.2dense and sparse features · 0.2correlation analysis · 0.2
YearPublicationVenuePosition
2021 Levenshtein Training for Word-level Quality Estimation
abstract
We propose a novel scheme to use the Levenshtein Transformer to perform the task of word-level quality estimation.A Levenshtein Transformer is a natural fit for this task: trained to perform decoding in an iterative manner, a Levenshtein Transformer can learn to post-edit without explicit supervision.To further minimize the mismatch between the translation task and the word-level QE task, we propose a two-stage transfer learning procedure on both augmented data and human postediting data.We also propose heuristics to construct reference labels that are compatible with subword-level finetuning and inference.Results on WMT 2020 QE shared task dataset show that our proposed method has superior data efficiency under the data-constrained setting and competitive performance under the unconstrained setting.* Shuoyang Ding had a part-time affiliation with Microsoft at the time of this work.
Shuoyang Ding, Marcin Junczys-Dowmunt, Matt Post, Philipp Koehn
EMNLP (1)2
2021 The Curious Case of Hallucinations in Neural Machine Translation
abstract
Vikas Raunak, Arul Menezes, Marcin Junczys-Dowmunt. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Vikas Raunak, Arul Menezes, Marcin Junczys-Dowmunt
NAACL-HLT3
2018 Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation
abstract
In order to extract the best possible performance from asynchronous stochastic gradient descent one must increase the mini-batch size and scale the learning rate accordingly.In order to achieve further speedup we introduce a technique that delays gradient updates effectively increasing the mini-batch size.Unfortunately with the increase of mini-batch size we worsen the stale gradient problem in asynchronous stochastic gradient descent (SGD) which makes the model convergence poor.We introduce local optimizers which mitigate the stale gradient problem and together with fine tuning our momentum we are able to train a shallow machine translation system 27% faster than an optimized baseline with negligible penalty in BLEU.
Nikolay Bogoychev, Kenneth Heafield, Alham Fikri Aji, Marcin Junczys-Dowmunt
EMNLP4
2018 Approaching Neural Grammatical Error Correction as a Low-Resource Machine Translation Task
abstract
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Shubha Guha, Kenneth Heafield. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Shubha Guha, Kenneth Heafield
NAACL-HLT1
2017 An Exploration of Neural Sequence-to-Sequence Architectures for Automatic Post-Editing
abstract
In this work, we explore multiple neural architectures adapted for the task of automatic post-editing of machine translation output. We focus on neural end-to-end models that combine both inputs mt (raw MT output) and src (source language input) in a single neural architecture, modeling \{mt, src\} \rightarrow pe directly. Apart from that, we investigate the influence of hard-attention models which seem to be well-suited for monolingual tasks, as well as combinations of both ideas. We report results on data sets provided during the WMT-2016 shared task on automatic post-editing and can demonstrate that dual-attention models that incorporate all available data in the APE scenario in a single model improve on the best shared task system and on all other published results after the shared task. Dual-attention models that are combined with hard attention remain competitive despite applying fewer changes to the input.
Marcin Junczys-Dowmunt, Roman Grundkiewicz
IJCNLP(1)1
2017 Pushing the Limits of Translation Quality Estimation
abstract
Translation quality estimation is a task of growing importance in NLP, due to its potential to reduce post-editing human effort in disruptive ways. However, this potential is currently limited by the relatively low accuracy of existing systems. In this paper, we achieve remarkable improvements by exploiting synergies between the related tasks of word-level quality estimation and automatic post-editing. First, we stack a new, carefully engineered, neural model into a rich feature-based word-level quality estimation system. Then, we use the output of an automatic post-editing system as an extra feature, obtaining striking results on WMT16: a word-level FMULT1 score of 57.47% (an absolute gain of +7.95% over the current state of the art), and a Pearson correlation score of 65.56% for sentence-level HTER prediction (an absolute gain of +13.36%).
André F. T. Martins, Marcin Junczys-Dowmunt, Fábio N. Kepler, Ramón Fernandez Astudillo, Chris Hokamp, Roman Grundkiewicz
Trans. Assoc. Comput. Linguistics2
2016 Target-Side Context for Discriminative Models in Statistical Machine Translation
abstract
Discriminative translation models utilizing source context have been shown to help statistical machine translation performance.We propose a novel extension of this work using target context information.Surprisingly, we show that this model can be efficiently integrated directly in the decoding process.Our approach scales to large training data sizes and results in consistent improvements in translation quality on four language pairs.We also provide an analysis comparing the strengths of the baseline source-context model with our extended source-context and targetcontext model and we show that our extension allows us to better capture morphological coherence.Our work is freely available as part of Moses.
Ales Tamchyna, Alexander Fraser 0001, Ondrej Bojar, Marcin Junczys-Dowmunt
ACL (1)4
2016 Phrase-based Machine Translation is State-of-the-Art for Automatic Grammatical Error Correction
abstract
In this work, we study parameter tuning towards the M 2 metric, the standard metric for automatic grammar error correction (GEC) tasks.After implementing M 2 as a scorer in the Moses tuning framework, we investigate interactions of dense and sparse features, different optimizers, and tuning strategies for the CoNLL-2014 shared task.We notice erratic behavior when optimizing sparse feature weights with M 2 and offer partial solutions.We find that a bare-bones phrase-based SMT setup with task-specific parameter-tuning outperforms all previously published results for the CoNLL-2014 test set by a large margin (46.37% M 2 over previously 41.75%, by an SMT system with neural features) while being trained on the same, publicly available data.Our newly introduced dense and sparse features widen that gap, and we improve the state-of-the-art to 49.49% M 2 .
Marcin Junczys-Dowmunt, Roman Grundkiewicz
EMNLP1
2016 The United Nations Parallel Corpus v1.0
Michal Ziemski, Marcin Junczys-Dowmunt, Bruno Pouliquen
LREC2
2015 SMT at the International Maritime Organization: experiences with combining in-house corpora with out-of-domain corpora
Bruno Pouliquen, Marcin Junczys-Dowmunt, Blanca Pinero, Michal Ziemski
EAMT2
2015 Human Evaluation of Grammatical Error Correction Systems
abstract
The paper presents the results of the first large-scale human evaluation of automatic grammatical error correction (GEC) systems.Twelve participating systems and the unchanged input of the CoNLL-2014 shared task have been reassessed in a WMT-inspired human evaluation procedure.Methods introduced for the Workshop of Machine Translation evaluation campaigns have been adapted to GEC and extended where necessary.The produced rankings are used to evaluate standard metrics for grammatical error correction in terms of correlation with human judgment.
Roman Grundkiewicz, Marcin Junczys-Dowmunt, Edward Gillian
EMNLP2
2015 Large scale speech-to-text translation with out-of-domain corpora using better context-based models and domain adaptation
abstract
In this paper, we described the process of building a large-scale speech-to-text pipeline. Two target domains, daily conversations and travel-related conversations between two agents, for the English-German language pair (both directions) are examined. The SMT component is built from out-of-domain but freely-available bilingual and monolingual data. We make use of most of the known available resources to examine the effects of unrestricted data and large scale models. A naive baseline delivers solid results in terms of MT-quality. Extending the baseline with context-based translation model features like operations sequence models, higher-order class-based language models, and additional web-scale word-based language models leads to a system that significantly outperforms the baseline. Domain adaption is performed by separately weighting the influence of the out-of-domain subcorpora. This is explored for translation models and language models yielding significant improvements in both cases. Automatic and manual evaluation results are provided for raw MT-quality and ASR+MT-quality.
Marcin Junczys-Dowmunt, Pawel Przybysz, Arleta Staszuk, Eun-Kyoung Kim
INTERSPEECH1
2014 SMT of German patents at WIPO: decompounding and verb structure pre-reordering
Marcin Junczys-Dowmunt, Bruno Pouliquen
EAMT1
2012 A Phrase Table without Phrases: Rank Encoding for Better Phrase Table Compression
Marcin Junczys-Dowmunt
EAMT1
2010 A Maximum Entropy Approach to Syntactic Translation Rule Filtering
Marcin Junczys-Dowmunt
CICLing1