VLDB 2026 Research / reviewers in the wild / expert
Sergey Edunov
dblp:166/8381
· DBLP profile ↗
16ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Machine translation · 34% Language models and text generation · 22% Deep learning architectures and training · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 40% Graph data management · 40% Distributed and cloud data management · 20% |
Topics — the 20 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Law of the Weakest Link: Cross Capabilities of Large Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.7 | 1 | 2023 | LegoNN: Building Modular Encoder-Decoder Models · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Natural language and speech › Machine translation
parallel corpus mining |
0.5 | 1 | 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web · ACL/IJCNLP (1) 2021 |
Natural language and speech › Machine translation › parallel corpus mining
parallel sentence extraction |
0.5 | 1 | 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web · ACL/IJCNLP (1) 2021 |
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis |
0.4 | 1 | 2020 | Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP · ICLR 2020 |
Natural language and speech › Machine translation
machine translation evaluation |
0.4 | 1 | 2020 | On The Evaluation of Machine Translation SystemsTrained With Back-Translation · ACL 2020 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.4 | 1 | 2020 | Dense Passage Retrieval for Open-Domain Question Answering · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
passage retrieval |
0.4 | 1 | 2020 | Dense Passage Retrieval for Open-Domain Question Answering · EMNLP (1) 2020 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.4 | 1 | 2020 | Dense Passage Retrieval for Open-Domain Question Answering · EMNLP (1) 2020 |
Machine learning › Representation and self-supervised learning
pre-training |
0.4 | 1 | 2019 | Cloze-driven Pretraining of Self-attention Networks · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Machine translation › monolingual data augmentation
back-translation |
0.3 | 1 | 2018 | Understanding Back-Translation at Scale · EMNLP 2018 |
Machine learning › Deep learning architectures and training
data augmentation |
0.3 | 1 | 2018 | Understanding Back-Translation at Scale · EMNLP 2018 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2018 | Understanding Back-Translation at Scale · EMNLP 2018 |
Natural language and speech › Machine translation
synthetic parallel data |
0.3 | 1 | 2018 | Understanding Back-Translation at Scale · EMNLP 2018 |
Graph data management › graph processing
graph processing systems |
0.2 | 1 | 2015 | One Trillion Edges: Graph Processing at Facebook-Scale · Proc. VLDB Endow. 2015 |
Graph data management › graph processing
large-scale graph processing |
0.2 | 1 | 2015 | One Trillion Edges: Graph Processing at Facebook-Scale · Proc. VLDB Endow. 2015 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.2 | 1 | 2023 | LegoNN: Building Modular Encoder-Decoder Models · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.2 | 1 | 2023 | LegoNN: Building Modular Encoder-Decoder Models · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Machine learning › Representation and self-supervised learning › text embedding
cross-lingual representation |
0.1 | 1 | 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web · ACL/IJCNLP (1) 2021 |
Machine learning › Deep learning architectures and training › attention mechanism › attention network
self-attention network |
0.1 | 1 | 2019 | Cloze-driven Pretraining of Self-attention Networks · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
human annotation · 0.9benchmark construction · 0.9dual encoder · 0.9dense passage retrieval · 0.9marginal distribution interface · 0.7length control mechanism · 0.7gradient isolation · 0.7pregel model extensions · 0.4language model score · 0.4back-translation · 0.4BLEU · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Law of the Weakest Link: Cross Capabilities of Large Language ModelsabstractThe development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term **cross capabilities**. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce *CrossEval*, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the ``Law of the Weakest Link,'' where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our [project website](https://www.llm-cross-capabilities.org). Ming Zhong 0005, Aston Zhang, Wenhan Xiong, Chenguang Zhu 0001, Zhengxing Chen, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan 0001, Sergey Edunov, Jiawei Han 0001, Laurens van der Maaten |
ICLR | 15 |
| 2024 | Effective Long-Context Scaling of Foundation ModelsabstractWenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wenhan Xiong, Igor Molybog, Prajjwal Bhargava, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma 0001 |
NAACL-HLT | 18 |
| 2023 | LegoNN: Building Modular Encoder-Decoder ModelsabstractState-of-the-art encoder-decoder models (e.g. for machine translation (MT) or automatic speech recognition (ASR)) are constructed and trained end-to-end as an atomic unit. No component of the model can be (re-)used without the others, making it impossible to share parts, e.g. a high resourced decoder, across tasks. We describe LegoNN, a procedure for building encoder-decoder architectures in a way so that its parts can be applied to other tasks without the need for any fine-tuning. To achieve this reusability, the interface between encoder and decoder modules is grounded to a sequence of marginal distributions over a pre-defined discrete vocabulary. We present two approaches for ingesting these marginals; one is differentiable, allowing the flow of gradients across the entire network, and the other is gradient-isolating. To enable the portability of decoder modules between MT tasks for different source languages and across other tasks like ASR, we introduce a modality agnostic encoder which consists of a length control mechanism to dynamically adapt encoders' output lengths in order to match the expected input length range of pre-trained decoders. We present several experiments to demonstrate the effectiveness of LegoNN models: a trained language generation LegoNN decoder module from German-English (De-En) MT task can be reused without any fine-tuning for the Europarl English ASR and the Romanian-English (Ro-En) MT tasks, matching or beating the performance of baseline. After fine-tuning, LegoNN models improve the Ro-En MT task by 1.5 BLEU points and achieve 12.5% relative WER reduction on the Europarl ASR task. To show how the approach generalizes, we compose a LegoNN ASR model from three modules – each has been learned within different end-to-end trained models on three different datasets – achieving an overall WER reduction of 19.5%. Siddharth Dalmia, Dmytro Okhonko, Mike Lewis, Sergey Edunov, Shinji Watanabe 0001, Florian Metze, Luke Zettlemoyer, Abdel-rahman Mohamed |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebabstractHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan |
ACL/IJCNLP (1) | 3 |
| 2021 | Beyond English-Centric Multilingual Machine TranslationabstractExisting work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric, training only on data which was translated from or to English.While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open-source a training data set that covers thousands of language directions with parallel data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems from the Workshop on Machine Translation (WMT). We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model. Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal 0001, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, Armand Joulin |
J. Mach. Learn. Res. | 14 |
| 2020 | On The Evaluation of Machine Translation SystemsTrained With Back-TranslationabstractBack-translation is a widely used data augmentation technique which leverages target monolingual data.However, its effectiveness has been challenged since automatic metrics such as BLEU only show significant improvements for test examples where the source itself is a translation, or translationese.This is believed to be due to translationese inputs better matching the back-translated training data.In this work, we show that this conjecture is not empirically supported and that backtranslation improves translation quality of both naturally occurring text as well as translationese according to professional human translators.We provide empirical evidence to support the view that back-translation is preferred by humans because it produces more fluent outputs.BLEU cannot capture human preferences because references are translationese when source sentences are natural text.We recommend complementing BLEU with a language model score to measure fluency. Sergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael Auli |
ACL | 1 |
| 2020 | Dense Passage Retrieval for Open-Domain Question AnsweringabstractVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen 0001, Scott Yih |
EMNLP (1) | 6 |
| 2020 | Training ASR Models By Generation of Contextual InformationabstractSupervised ASR models have reached unprecedented levels of accuracy, thanks in part to ever-increasing amounts of labelled training data. However, in many applications and locales, only moderate amounts of data are available, which has led to a surge in semi- and weakly-supervised learning research. In this paper, we conduct a large-scale study evaluating the effectiveness of weakly-supervised learning for speech recognition by using loosely related contextual information as a surrogate for ground-truth labels. For weakly supervised training, we use 50k hours of public English social media videos along with their respective titles and post text to train an encoder-decoder transformer model. Our best encoder-decoder models achieve an average of 20.8% WER reduction over a 1000 hours supervised baseline, and an average of 13.4% WER reduction when using only the weakly supervised encoder for CTC fine-tuning. Our results show that our setup for weak supervision improved both the encoder acoustic representations as well as the decoder language generation abilities. Kritika Singh, Dmytro Okhonko, Yongqiang Wang 0005, Frank Zhang 0001, Ross B. Girshick, Sergey Edunov, Fuchun Peng, Yatharth Saraf, Geoffrey Zweig, Abdel-rahman Mohamed |
ICASSP | 7 |
| 2020 | Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP
Haonan Yu, Sergey Edunov, Yuandong Tian, Ari S. Morcos |
ICLR | 2 |
| 2020 | Large Scale Weakly and Semi-Supervised Learning for Low-Resource Video ASRabstractMany semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of transcribing social media videos in low-resource conditions, we conduct a large scale systematic comparison between two self-labeling methods on one hand, and weakly-supervised pretraining using contextual metadata on the other. We investigate distillation methods at the frame level and the sequence level for hybrid, encoder-only CTC-based, and encoder-decoder speech recognition systems on Dutch and Romanian languages using 27,000 and 58,000 hours of unlabeled audio respectively. Although all approaches improved upon their respective baseline WERs by more than 8%, sequence-level distillation for encoder-decoder models provided the largest relative WER reduction of 20% compared to the strongest data-augmented supervised baseline. Kritika Singh, Vimal Manohar, Alex Xiao, Sergey Edunov, Ross B. Girshick, Vitaliy Liptchinsky, Christian Fügen, Yatharth Saraf, Geoffrey Zweig, Abdel-rahman Mohamed |
INTERSPEECH | 4 |
| 2020 | Multilingual Denoising Pre-training for Neural Machine TranslationabstractThis paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using the BART objective (Lewis et al., 2019 ). mBART is the first method for pre-training a complete sequence-to-sequence model by denoising full texts in multiple languages, whereas previous approaches have focused only on the encoder, decoder, or reconstructing parts of the text. Pre-training a complete model allows it to be directly fine-tuned for supervised (both sentence-level and document-level) and unsupervised machine translation, with no task- specific modifications. We demonstrate that adding mBART initialization produces performance gains in all but the highest-resource settings, including up to 12 BLEU points for low resource MT and over 5 BLEU points for many document-level and unsupervised models. We also show that it enables transfer to language pairs with no bi-text or that were not in the pre-training corpus, and present extensive analysis of which factors contribute the most to effective pre-training. 1 Yinhan Liu, Jiatao Gu, Naman Goyal 0001, Xian Li 0003, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, Luke Zettlemoyer |
Trans. Assoc. Comput. Linguistics | 5 |
| 2019 | Cloze-driven Pretraining of Self-attention NetworksabstractAlexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, Michael Auli. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, Michael Auli |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Understanding Back-Translation at ScaleabstractAn effective method to improve neural machine translation with monolingual data is to augment the parallel training corpus with back-translations of target language sentences.This work broadens the understanding of back-translation and investigates a number of methods to generate synthetic source sentences.We find that in all but resource poor settings back-translations obtained via sampling or noised beam outputs are most effective.Our analysis shows that sampling or noisy synthetic data gives a much stronger training signal than data generated by beam or greedy search.We also compare how synthetic data compares to genuine bitext and study various domain effects.Finally, we scale to hundreds of millions of monolingual sentences and achieve a new state of the art of 35 BLEU on the WMT'14 English-German test set. Sergey Edunov, Myle Ott, Michael Auli, David Grangier |
EMNLP | 1 |
| 2018 | Generating Synthetic Social Graphs with DarwiniabstractSynthetic graph generators facilitate research in graph algorithms and graph processing systems by providing access to graphs that resemble real social networks while addressing privacy and security concerns. Nevertheless, their practical value lies in their ability to capture important metrics of real graphs, such as degree distribution and clustering properties. Graph generators must also be able to produce such graphs at the scale of real-world industry graphs, that is, hundreds of billions or trillions of edges. In this paper, we propose Darwini, a graph generator that captures a number of core characteristics of real graphs. Importantly, given a source graph, it can reproduce the degree distribution and, unlike existing approaches, the local clustering coefficient distribution. Furthermore, Darwini maintains a number of metrics, such as graph assortativity, eigenvalues, and others. Comparing Darwini with state-of-the-art generative models, we show that it can reproduce these characteristics more accurately. Finally, we provide an open source implementation of Darwini on the vertex-centric Apache Giraph model that can generate synthetic graphs with up to 3 trillion edges. Sergey Edunov, Dionysios Logothetis, Cheng Wang 0001, Avery Ching, Maja Kabiljo |
ICDCS | 1 |
| 2018 | Classical Structured Prediction Losses for Sequence to Sequence LearningabstractSergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc’Aurelio Ranzato. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato |
NAACL-HLT | 1 |
| 2015 | One Trillion Edges: Graph Processing at Facebook-ScaleabstractAnalyzing large graphs provides valuable insights for social networking and web companies in content ranking and recommendations. While numerous graph processing systems have been developed and evaluated on available benchmark graphs of up to 6.6B edges, they often face significant difficulties in scaling to much larger graphs. Industry graphs can be two orders of magnitude larger - hundreds of billions or up to one trillion edges. In addition to scalability challenges, real world applications often require much more complex graph processing workflows than previously evaluated. In this paper, we describe the usability, performance, and scalability improvements we made to Apache Giraph, an open-source graph processing system, in order to use it on Facebook-scale graphs of up to one trillion edges. We also describe several key extensions to the original Pregel model that make it possible to develop a broader range of production graph applications and workflows as well as improve code reuse. Finally, we report on real-world operations as well as performance characteristics of several large-scale production applications. Avery Ching, Sergey Edunov, Maja Kabiljo, Dionysios Logothetis, Sambavi Muthukrishnan |
Proc. VLDB Endow. | 2 |