Jörn Wübker

dblp:36/9013 · also Joern Wuebker · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Machine translation · 67% Transfer learning and domain adaptation · 14% Efficient and distributed learning · 14%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
neural machine translation
0.832018
Compact Personalized Models for Neural Machine Translation · EMNLP 2018
Models and Inference for Prefix-Constrained Machine Translation · ACL (1) 2016
Translation Modeling with Bidirectional Recurrent Neural Networks · EMNLP 2014
Natural language and speech › Machine translation
statistical machine translation
0.742020
Hierarchical Incremental Adaptation for Statistical Machine Translation · EMNLP 2015
A Comparison between Count and Neural Network Models Based on Joint Translation and Reordering Sequences · EMNLP 2015
Improving Statistical Machine Translation with Word Class Models · EMNLP 2013
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.722020
End-to-End Neural Word Alignment Outperforms GIZA++ · ACL 2020
A Comparison between Count and Neural Network Models Based on Joint Translation and Reordering Sequences · EMNLP 2015
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.312018
Compact Personalized Models for Neural Machine Translation · EMNLP 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
Compact Personalized Models for Neural Machine Translation · EMNLP 2018
Machine learning › Transfer learning and domain adaptation › model adaptation
personalized model adaptation
0.312018
Compact Personalized Models for Neural Machine Translation · EMNLP 2018
Machine learning › Efficient and distributed learning › model compression › sparsity
structured sparsity
0.312018
Compact Personalized Models for Neural Machine Translation · EMNLP 2018
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation
0.322016
Models and Inference for Prefix-Constrained Machine Translation · ACL (1) 2016
Translation Modeling with Bidirectional Recurrent Neural Networks · EMNLP 2014
Natural language and speech › Machine translation › computer-assisted translation
interactive machine translation
0.212016
Models and Inference for Prefix-Constrained Machine Translation · ACL (1) 2016
Natural language and speech › Machine translation › statistical machine translation
reordering model
0.212015
A Comparison between Count and Neural Network Models Based on Joint Translation and Reordering Sequences · EMNLP 2015
Machine learning › Deep learning architectures and training
recurrent neural network
0.212014
Translation Modeling with Bidirectional Recurrent Neural Networks · EMNLP 2014

Methods — techniques the papers use, named apart from their topics

transformer · 0.4group lasso regularization · 0.3gradient-based adaptation · 0.3n-best extraction · 0.2joint alignment and translation model · 0.2beam search · 0.2recurrent neural network · 0.2n-gram model · 0.2kneser-ney smoothing · 0.2feedforward neural network · 0.2
YearPublicationVenuePosition
2022 Automatic Correction of Human Translations
abstract
Jessy Lin, Geza Kovacs, Aditya Shastry, Joern Wuebker, John DeNero. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jessy Lin, Geza Kovacs, Jörn Wübker, John DeNero
NAACL-HLT4
2020 End-to-End Neural Word Alignment Outperforms GIZA++
abstract
Word alignment was once a core unsupervised learning task in natural language processing because of its essential role in training statistical machine translation (MT) models.Although unnecessary for training neural MT models, word alignment still plays an important role in interactive applications of neural machine translation, such as annotation transfer and lexicon injection.While statistical MT methods have been replaced by neural approaches with superior performance, the twenty-year-old GIZA++ toolkit remains a key component of state-of-the-art word alignment systems.Prior work on neural word alignment has only been able to outperform GIZA++ by using its output during training.We present the first end-to-end neural word alignment method that consistently outperforms GIZA++ on three data sets.Our approach repurposes a Transformer model trained for supervised translation to also serve as an unsupervised word alignment model in a manner that is tightly integrated and does not affect translation quality.
Thomas Zenkel, Jörn Wübker, John DeNero
ACL2
2018 Compact Personalized Models for Neural Machine Translation
abstract
We propose and compare methods for gradientbased domain adaptation of self-attentive neural machine translation models.We demonstrate that a large proportion of model parameters can be frozen during adaptation with minimal or no reduction in translation quality by encouraging structured sparsity in the set of offset tensors during learning via group lasso regularization.We evaluate this technique for both batch and incremental adaptation across multiple data sets and language pairs.Our system architecture-combining a state-of-the-art self-attentive model with compact domain adaptation-provides high quality personalized machine translation that is both space and time efficient.
Jörn Wübker, Patrick Simianer, John DeNero
EMNLP1
2016 Models and Inference for Prefix-Constrained Machine Translation
abstract
We apply phrase-based and neural models to a core task in interactive machine translation: suggesting how to complete a partial translation.For the phrase-based system, we demonstrate improvements in suggestion quality using novel objective functions, learning techniques, and inference algorithms tailored to this task.Our contributions include new tunable metrics, an improved beam search strategy, an n-best extraction method that increases suggestion diversity, and a tuning procedure for a hierarchical joint model of alignment and translation.The combination of these techniques improves next-word suggestion accuracy dramatically from 28.5% to 41.2% in a large-scale English-German experiment.Our recurrent neural translation system increases accuracy yet further to 53.0%, but inference is two orders of magnitude slower.Manual error analysis shows the strengths and weaknesses of both approaches.
Jörn Wübker, Spence Green, John DeNero, Sasa Hasan, Minh-Thang Luong
ACL (1)1
2015 A Comparison between Count and Neural Network Models Based on Joint Translation and Reordering Sequences
abstract
We propose a conversion of bilingual sentence pairs and the corresponding word alignments into novel linear sequences.These are joint translation and reordering (JTR) uniquely defined sequences, combining interdepending lexical and alignment dependencies on the word level into a single framework.They are constructed in a simple manner while capturing multiple alignments and empty words.JTR sequences can be used to train a variety of models.We investigate the performances of ngram models with modified Kneser-Ney smoothing, feed-forward and recurrent neural network architectures when estimated on JTR sequences, and compare them to the operation sequence model (Durrani et al., 2013b).Evaluations on the IWSLT German→English, WMT German→English and BOLT Chinese→English tasks show that JTR models improve state-of-the-art phrasebased systems by up to 2.2 BLEU.
Andreas Guta, Tamer Alkhouli, Jan-Thorsten Peter, Jörn Wübker, Hermann Ney
EMNLP4
2015 Hierarchical Incremental Adaptation for Statistical Machine Translation
abstract
We present an incremental adaptation approach for statistical machine translation that maintains a flexible hierarchical domain structure within a single consistent model.Both weights and rules are updated incrementally on a stream of post-edits.Our multi-level domain hierarchy allows the system to adapt simultaneously towards local context at different levels of granularity, including genres and individual documents.Our experiments show consistent improvements in translation quality from all components of our approach.
Jörn Wübker, Spence Green, John DeNero
EMNLP1
2015 A Comparison of Update Strategies for Large-Scale Maximum Expected BLEU Training
abstract
Joern Wuebker, Sebastian Muehr, Patrick Lehnen, Stephan Peitz, Hermann Ney. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Jörn Wübker, Sebastian Muehr, Patrick Lehnen, Stephan Peitz, Hermann Ney
HLT-NAACL1
2014 Translation Modeling with Bidirectional Recurrent Neural Networks
abstract
This work presents two different trans-lation models using recurrent neural net-works. The first one is a word-based ap-proach using word alignments. Second, we present phrase-based translation mod-els that are more consistent with phrase-based decoding. Moreover, we introduce bidirectional recurrent neural models to the problem of machine translation, allow-ing us to use the full source sentence in our models, which is also of theoretical inter-est. We demonstrate that our translation models are capable of improving strong baselines already including recurrent neu-ral language models on three tasks: IWSLT 2013 German→English, BOLT Arabic→English and Chinese→English. We obtain gains up to 1.6 % BLEU and 1.7 % TER by rescoring 1000-best lists. 1
Martin Sundermeyer, Tamer Alkhouli, Jörn Wübker, Hermann Ney
EMNLP3
2013 Improving Statistical Machine Translation with Word Class Models
abstract
Automatically clustering words from a monolingual or bilingual training corpus into classes is a widely used technique in statistical natural language processing.We present a very simple and easy to implement method for using these word classes to improve translation quality.It can be applied across different machine translation paradigms and with arbitrary types of models.We show its efficacy on a small German→English and a larger French→German translation task with both standard phrase-based and hierarchical phrase-based translation systems for a common set of models.Our results show that with word class models, the baseline can be improved by up to 1.4% BLEU and 1.0% TER on the French→German task and 0.3% BLEU and 1.1% TER on the German→English task.
Jörn Wübker, Stephan Peitz, Felix Rietig, Hermann Ney
EMNLP1
2013 (Hidden) Conditional Random Fields Using Intermediate Classes for Statistical Machine Translation
Patrick Lehnen, Jan-Thorsten Peter, Jörn Wübker, Stephan Peitz, Hermann Ney
MTSummit3
2010 Training Phrase Translation Models with Leaving-One-Out
Jörn Wübker, Arne Mauser, Hermann Ney
ACL1