Kenneth Heafield

dblp:81/3540 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0002-6344-9927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Machine translation · 47% Efficient and distributed learning · 34% Language models and text generation · 10%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
neural machine translation
1.242020
Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation · EMNLP (1) 2020
Multi-Source Syntactic Neural Machine Translation · EMNLP 2018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Machine translation
parallel corpus mining
0.922020
Parallel Sentence Mining by Constrained Decoding · ACL 2020
ParaCrawl: Web-Scale Acquisition of Parallel Corpora · ACL 2020
Natural language and speech › Machine translation
document-level machine translation
0.812024
Document-Level Machine Translation with Large-Scale Public Parallel Corpora · ACL (1) 2024
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.722019
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training · EMNLP/IJCNLP (1) 2019
Sparse Communication for Distributed Gradient Descent · EMNLP 2017
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
attention head pruning
0.412020
Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation · EMNLP (1) 2020
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.412020
Parallel Sentence Mining by Constrained Decoding · ACL 2020
Natural language and speech › Machine translation
low-resource machine translation
0.412020
In Neural Machine Translation, What Does Transfer Learning Transfer? · ACL 2020
Natural language and speech › Machine translation › parallel corpus mining
parallel sentence extraction
0.412020
Parallel Sentence Mining by Constrained Decoding · ACL 2020
Machine learning › Efficient and distributed learning › model compression › pruning
transformer pruning
0.412020
Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation · EMNLP (1) 2020
Information retrieval
cross-language information retrieval
0.412020
ParaCrawl: Web-Scale Acquisition of Parallel Corpora · ACL 2020
Machine learning › Efficient and distributed learning › distributed training
distributed DNN training
0.412019
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training · EMNLP/IJCNLP (1) 2019
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent
0.312018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Machine learning › Efficient and distributed learning
distributed training
0.312017
Sparse Communication for Distributed Gradient Descent · EMNLP 2017
Natural language and speech › Language models and text generation
language modeling
0.212016
Normalized Log-Linear Interpolation of Backoff Language Models is Efficient · ACL (1) 2016
Natural language and speech › Machine translation
parallel corpora
0.212024
Document-Level Machine Translation with Large-Scale Public Parallel Corpora · ACL (1) 2024
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis
0.112020
Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation · EMNLP (1) 2020
Machine learning › Efficient and distributed learning
model compression
0.112020
Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation · EMNLP (1) 2020
Machine learning › Optimization for machine learning
distributed optimization
0.112019
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training · EMNLP/IJCNLP (1) 2019
Machine learning › Optimization for machine learning › learning rate
learning rate scaling
0.112018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.112018
Multi-Source Syntactic Neural Machine Translation · EMNLP 2018
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.012012
Language Model Rest Costs and Space-Efficient Storage · EMNLP-CoNLL 2012

Methods — techniques the papers use, named apart from their topics

web crawling · 0.9automatic alignment · 0.9context-aware model · 0.8prefix tree · 0.4neural machine translation · 0.4lottery ticket hypothesis · 0.4knowledge distillation · 0.4beam search · 0.4autoencoder alignment · 0.4ablation study · 0.4
YearPublicationVenuePosition
2024 Document-Level Machine Translation with Large-Scale Public Parallel Corpora
abstract
Despite the fact that document-level machine translation has inherent advantages over sentence-level machine translation due to additional information available to a model from document context, most translation systems continue to operate at a sentence level.This is primarily due to the severe lack of publicly available large-scale parallel corpora at the document level.We release a large-scale open parallel corpus with document context extracted from ParaCrawl in five language pairs, along with code to compile document-level datasets for any language pair supported by ParaCrawl.We train context-aware models on these datasets and find improvements in terms of overall translation quality and targeted document-level phenomena.We also analyse how much long-range information is useful to model some of these discourse phenomena and find models are able to utilise context from several preceding sentences.
Proyag Pal, Alexandra Birch, Kenneth Heafield
ACL (1)3
2024 Code-Switched Language Identification is Harder Than You Think
abstract
Laurie Burchell, Alexandra Birch, Robert Thompson, Kenneth Heafield. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Laurie Burchell, Alexandra Birch, Robert P. Thompson, Kenneth Heafield
EACL (1)4
2024 Iterative Translation Refinement with Large Language Models
abstract
We propose iteratively prompting a large language model to self-correct a translation, with inspiration from their strong language capability as well as a human-like translation approach. Interestingly, multi-turn querying reduces the output’s string-based metric scores, but neural metrics suggest comparable or improved quality after two or more iterations. Human evaluations indicate better fluency and naturalness compared to initial translations and even human references, all while maintaining quality. Ablation studies underscore the importance of anchoring the refinement to the source and a reasonable seed translation for quality considerations. We also discuss the challenges in evaluation and relation to human performance and translationese.
Pinzhen Chen, Zhicheng Guo, Barry Haddow, Kenneth Heafield
EAMT (1)4
2023 Efficient Methods for Natural Language Processing: A Survey
abstract
Abstract Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources include data, time, storage, or energy, all of which are naturally limited and unevenly distributed. This motivates research into efficient methods that require fewer resources to achieve similar results. This survey synthesizes and relates current methods and findings in efficient NLP. We aim to provide both guidance for conducting NLP under limited resources, and point towards promising research directions for developing more efficient methods.
Marcos V. Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro Henrique Martins, André F. T. Martins, Jessica Zosa Forde, Peter A. Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz 0001
Trans. Assoc. Comput. Linguistics8
2022 Constrained Regeneration for Cross-Lingual Query-Focused Extractive Summarization
abstract
Query-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. However, in a cross-lingual setting, the query term does not necessarily appear in the translations of relevant documents. In this work, we show that constrained machine translation and constrained post-editing can improve human relevance judgments by including a query term in a summary when its translation appears in the source document. We also present several strategies for selecting only certain documents for regeneration which yield further improvements
Elsbeth Turcan, David Wan, Faisal Ladhak, Petra Galuscáková, Sukanta Sen, Svetlana Tchistiakova, Weijia Xu, Marine Carpuat, Kenneth Heafield, Douglas W. Oard, Kathy McKeown
COLING9
2022 The EuroPat Corpus: A Parallel Corpus of European Patent Data
abstract
We present the EuroPat corpus of patent-specific parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. The filtered parallel corpora range in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpora being up to 2 times larger. Access to clean, high quality, parallel data in technical domains such as science, engineering, and medicine is needed for training neural machine translation systems for tasks like online dispute resolution and eProcurement. Our evaluation found that the addition of EuroPat data to a generic baseline improved the performance of machine translation systems on in-domain test data in German, Spanish, French, and Polish; and in translating patent data from Croatian to English. The corpus has been released under Creative Commons Zero, and is expected to be widely useful for training high-quality machine translation systems, and particularly for those targeting technical documents such as patents and contracts.
Kenneth Heafield, Elaine Farrow, Jelmer van der Linde, Gema Ramírez-Sánchez, Dion Wiggins
LREC1
2022 Cheat Codes to Quantify Missing Source Information in Neural Machine Translation
abstract
This paper describes a method to quantify the amount of information H(t|s) added by the target sentence t that is not present in the source s in a neural machine translation system.We do this by providing the model the target sentence in a highly compressed form (a "cheat code"), and exploring the effect of the size of the cheat code.We find that the model is able to capture extra information from just a single float representation of the target and nearly reproduces the target with two 32-bit floats per target token.
Proyag Pal, Kenneth Heafield
NAACL-HLT2
2022 Approaching Neural Chinese Word Segmentation as a Low-Resource Machine Translation Task
Pinzhen Chen, Kenneth Heafield
PACLIC2
2020 In Neural Machine Translation, What Does Transfer Learning Transfer?
abstract
Transfer learning improves quality for lowresource machine translation, but it is unclear what exactly it transfers.We perform several ablation studies that limit information transfer, then measure the quality impact across three language pairs to gain a black-box understanding of transfer learning.Word embeddings play an important role in transfer learning, particularly if they are properly aligned.Although transfer learning can be performed without embeddings, results are sub-optimal.In contrast, transferring only the embeddings but nothing else yields catastrophic results.We then investigate diagonal alignments with auto-encoders over real languages and randomly generated sequences, finding even randomly generated sequences as parents yield noticeable but smaller gains.Finally, transfer learning can eliminate the need for a warmup phase when training transformer models in high resource language pairs.
Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico Sennrich
ACL3
2020 ParaCrawl: Web-Scale Acquisition of Parallel Corpora
abstract
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strong, Brian Thompson 0001, William Waites, Dion Wiggins, Jaume Zaragoza
ACL4
2020 Parallel Sentence Mining by Constrained Decoding
abstract
We present a novel method to extract parallel sentences from two monolingual corpora, using neural machine translation.Our method relies on translating sentences in one corpus, but constraining the decoding by a prefix tree built on the other corpus.We argue that a neural machine translation system by itself can be a sentence similarity scorer and it efficiently approximates pairwise comparison with a modified beam search.When benchmarked on the BUCC shared task, our method achieves results comparable to other submissions.
Pinzhen Chen, Nikolay Bogoychev, Kenneth Heafield, Faheem Kirefu
ACL3
2020 Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation
abstract
The attention mechanism is the crucial component of the transformer architecture. Recent research shows that most attention heads are not confident in their decisions and can be pruned. However, removing them before training a model results in lower quality. In this paper, we apply the lottery ticket hypothesis to prune heads in the early stages of training. Our experiments on machine translation show that it is possible to remove up to three-quarters of attention heads from transformer-big during early training with an average -0.1 change in BLEU for Turkish→English. The pruned model is 1.5 times as fast at inference, albeit at the cost of longer training. Our method is complementary to other approaches, such as teacher-student, with English→German student model gaining an additional 10% speed-up with 75% encoder attention removed and 0.2 BLEU loss.
Maximiliana Behnke, Kenneth Heafield
EMNLP (1)2
2019 Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training
abstract
Alham Fikri Aji, Kenneth Heafield, Nikolay Bogoychev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alham Fikri Aji, Kenneth Heafield, Nikolay Bogoychev
EMNLP/IJCNLP (1)2
2018 Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation
abstract
In order to extract the best possible performance from asynchronous stochastic gradient descent one must increase the mini-batch size and scale the learning rate accordingly.In order to achieve further speedup we introduce a technique that delays gradient updates effectively increasing the mini-batch size.Unfortunately with the increase of mini-batch size we worsen the stale gradient problem in asynchronous stochastic gradient descent (SGD) which makes the model convergence poor.We introduce local optimizers which mitigate the stale gradient problem and together with fine tuning our momentum we are able to train a shallow machine translation system 27% faster than an optimized baseline with negligible penalty in BLEU.
Nikolay Bogoychev, Kenneth Heafield, Alham Fikri Aji, Marcin Junczys-Dowmunt
EMNLP2
2018 Multi-Source Syntactic Neural Machine Translation
abstract
We introduce a novel multi-source technique for incorporating source syntax into neural machine translation using linearized parses.This is achieved by employing separate encoders for the sequential and parsed versions of the same source sentence; the resulting representations are then combined using a hierarchical attention mechanism.The proposed model improves over both seq2seq and parsed baselines by over 1 BLEU on the WMT17 English→German task.Further analysis shows that our multi-source syntactic model is able to translate successfully without any parsed input, unlike standard parsed methods.In addition, performance does not deteriorate as much on long sentences as for the baselines.
Anna Currey, Kenneth Heafield
EMNLP2
2018 Approaching Neural Grammatical Error Correction as a Low-Resource Machine Translation Task
abstract
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Shubha Guha, Kenneth Heafield. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Shubha Guha, Kenneth Heafield
NAACL-HLT4
2017 Sparse Communication for Distributed Gradient Descent
abstract
We make distributed stochastic gradient descent faster by exchanging sparse updates instead of dense updates.Gradient updates are positively skewed as most updates are near zero, so we map the 99% smallest updates (by absolute value) to zero then exchange sparse matrices.This method can be combined with quantization to further improve the compression.We explore different configurations and apply them to neural machine translation and MNIST image classification tasks.Most configurations work on MNIST, whereas different configurations reduce convergence rate on the more complex translation task.Our experiments show that we can achieve up to 49% speed up on MNIST and 22% on NMT without damaging the final accuracy or BLEU.
Alham Fikri Aji, Kenneth Heafield
EMNLP2
2016 Normalized Log-Linear Interpolation of Backoff Language Models is Efficient
abstract
We prove that log-linearly interpolated backoff language models can be efficiently and exactly collapsed into a single normalized backoff model, contradicting Hsu (2007).While prior work reported that log-linear interpolation yields lower perplexity than linear interpolation, normalizing at query time was impractical.We normalize the model offline in advance, which is efficient due to a recurrence relationship between the normalizing factors.To tune interpolation weights, we apply Newton's method to this convex problem and show that the derivatives can be computed efficiently in a batch process.These findings are combined in new open-source interpolation tool, which is distributed with KenLM.With 21 out-of-domain corpora, log-linear interpolation yields 72.58 perplexity on TED talks, compared to 75.91 for linear interpolation.
Kenneth Heafield, Chase Geigle, Sean Massung, Lane Schwartz
ACL (1)1
2014 N-gram Counts and Language Models from the Common Crawl
Christian Buck, Kenneth Heafield, Bas van Ooyen
LREC2
2013 Grouping Language Model Boundary Words to Speed K-Best Extraction from Hypergraphs
Kenneth Heafield, Philipp Koehn, Alon Lavie
HLT-NAACL1
2012 Language Model Rest Costs and Space-Efficient Storage
Kenneth Heafield, Philipp Koehn, Alon Lavie
EMNLP-CoNLL1