Nikolay Bogoychev

dblp:184/3755 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0003-4373-5261ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Machine translation · 35% Language models and text generation · 27% Efficient and distributed learning · 24%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.412020
Parallel Sentence Mining by Constrained Decoding · ACL 2020
Natural language and speech › Machine translation
low-resource machine translation
0.412020
In Neural Machine Translation, What Does Transfer Learning Transfer? · ACL 2020
Natural language and speech › Machine translation
parallel corpus mining
0.412020
Parallel Sentence Mining by Constrained Decoding · ACL 2020
Natural language and speech › Machine translation › parallel corpus mining
parallel sentence extraction
0.412020
Parallel Sentence Mining by Constrained Decoding · ACL 2020
Machine learning › Efficient and distributed learning › distributed training
distributed DNN training
0.412019
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training · EMNLP/IJCNLP (1) 2019
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.412019
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training · EMNLP/IJCNLP (1) 2019
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent
0.312018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Machine translation
neural machine translation
0.312018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.212016
N-gram language models for massively parallel devices · ACL (1) 2016
GPUs and heterogeneous computing › GPU computing
GPU implementation
0.212016
N-gram language models for massively parallel devices · ACL (1) 2016
Machine learning › Optimization for machine learning
distributed optimization
0.112019
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training · EMNLP/IJCNLP (1) 2019
Machine learning › Optimization for machine learning › learning rate
learning rate scaling
0.112018
Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

unargmaxable token detection · 0.6low-rank softmax analysis · 0.6prefix tree · 0.4neural machine translation · 0.4beam search · 0.4autoencoder alignment · 0.4ablation study · 0.4sparse gradients · 0.4local gradients · 0.4gradient compression · 0.4b-tree · 0.2
YearPublicationVenuePosition
2023 HPLT: High Performance Language Technologies
abstract
We describe the High Performance Language Technologies project (HPLT), a 3-year EU-funded project started in September 2022. HPLT will build a space combining petabytes of natural language data with large-scale model training. It will derive monolingual and bilingual datasets from the Internet Archive and CommonCrawl and build efficient and solid machine translation (MT) as well as large language models (LLMs). HPLT aims at providing free, sustainable and reusable datasets, models and workflows at scale using high-performance computing (HPC).
Mikko Aulamo, Nikolay Bogoychev, Shaoxiong Ji, Graeme Nail, Gema Ramírez-Sánchez, Jörg Tiedemann, Jelmer van der Linde, Jaume Zaragoza
EAMT2
2023 The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR
abstract
English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many varieties of English. Although the great many advances in English automatic speech recognition (ASR) over the past decades, results are usually reported based on test datasets which fail to represent the diversity of English as spoken today around the globe. We present the first release of The Edinburgh International Accents of English Corpus (EdAcc). This dataset attempts to better represent the wide diversity of English, encompassing almost 40 hours of dyadic video call conversations between friends. Unlike other datasets, EdAcc includes a wide range of first and second-language varieties of English and a linguistic background profile of each speaker. Results on latest public, and commercial models show that EdAcc highlights shortcomings of current English ASR models. The best performing model, trained on 680 thousand hours of transcribed data, obtains an average of 19.7% word error rate (WER) – in contrast to the 2.7% WER obtained when evaluated on US English clean read speech. Across all models, we observe a drop in performance on Indian, Jamaican, and Nigerian English speakers. Recordings, linguistic backgrounds, data statement, and evaluation scripts are released on our website under CC-BY-SA1license.2We hope that this work will encourage future research on a wider range of English varieties to create more accessible speech technologies.
Ramon Sanabria, Nikolay Bogoychev, Nina Markl, Andrea Carmantini, Ondrej Klejch, Peter Bell 0001
ICASSP2
2022 Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in Practice
abstract
Classifiers in natural language processing (NLP) often have a large number of output classes.For example, neural language models (LMs) and machine translation (MT) models both predict tokens from a vocabulary of thousands.The Softmax output layer of these models typically receives as input a dense feature representation, which has much lower dimensionality than the output.In theory, the result is some words may be impossible to be predicted via argmax, irrespective of input features, and empirically, there is evidence this happens in small language models (Demeter et al., 2020).In this paper we ask whether it can happen in practical large language models and translation models.To do so, we develop algorithms to detect such unargmaxable tokens in public models.We find that 13 out of 150 models do indeed have such tokens; however, they are very infrequent and unlikely to impact model quality.We release our algorithms and code so that others can test their models.1
Andreas Grivas, Nikolay Bogoychev, Adam Lopez
ACL (1)2
2020 In Neural Machine Translation, What Does Transfer Learning Transfer?
abstract
Transfer learning improves quality for lowresource machine translation, but it is unclear what exactly it transfers.We perform several ablation studies that limit information transfer, then measure the quality impact across three language pairs to gain a black-box understanding of transfer learning.Word embeddings play an important role in transfer learning, particularly if they are properly aligned.Although transfer learning can be performed without embeddings, results are sub-optimal.In contrast, transferring only the embeddings but nothing else yields catastrophic results.We then investigate diagonal alignments with auto-encoders over real languages and randomly generated sequences, finding even randomly generated sequences as parents yield noticeable but smaller gains.Finally, transfer learning can eliminate the need for a warmup phase when training transformer models in high resource language pairs.
Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico Sennrich
ACL2
2020 Parallel Sentence Mining by Constrained Decoding
abstract
We present a novel method to extract parallel sentences from two monolingual corpora, using neural machine translation.Our method relies on translating sentences in one corpus, but constraining the decoding by a prefix tree built on the other corpus.We argue that a neural machine translation system by itself can be a sentence similarity scorer and it efficiently approximates pairwise comparison with a modified beam search.When benchmarked on the BUCC shared task, our method achieves results comparable to other submissions.
Pinzhen Chen, Nikolay Bogoychev, Kenneth Heafield, Faheem Kirefu
ACL2
2019 Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training
abstract
Alham Fikri Aji, Kenneth Heafield, Nikolay Bogoychev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alham Fikri Aji, Kenneth Heafield, Nikolay Bogoychev
EMNLP/IJCNLP (1)3
2018 Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation
abstract
In order to extract the best possible performance from asynchronous stochastic gradient descent one must increase the mini-batch size and scale the learning rate accordingly.In order to achieve further speedup we introduce a technique that delays gradient updates effectively increasing the mini-batch size.Unfortunately with the increase of mini-batch size we worsen the stale gradient problem in asynchronous stochastic gradient descent (SGD) which makes the model convergence poor.We introduce local optimizers which mitigate the stale gradient problem and together with fine tuning our momentum we are able to train a shallow machine translation system 27% faster than an optimized baseline with negligible penalty in BLEU.
Nikolay Bogoychev, Kenneth Heafield, Alham Fikri Aji, Marcin Junczys-Dowmunt
EMNLP1
2016 N-gram language models for massively parallel devices
abstract
For many applications, the query speed of N -gram language models is a computational bottleneck.Although massively parallel hardware like GPUs offer a potential solution to this bottleneck, exploiting this hardware requires a careful rethinking of basic algorithms and data structures.We present the first language model designed for such hardware, using B-trees to maximize data parallelism and minimize memory footprint and latency.Compared with a single-threaded instance of KenLM (Heafield, 2011), a highly optimized CPUbased language model, our GPU implementation produces identical results with a smaller memory footprint and a sixfold increase in throughput on a batch query task.When we saturate both devices, the GPU delivers nearly twice the throughput per hardware dollar even when the CPU implementation uses faster data structures.
Nikolay Bogoychev, Adam Lopez
ACL (1)1