VLDB 2026 Research / reviewers in the wild / expert
Guillaume Wenzek
dblp:169/3295
· DBLP profile ↗
7ranked-venue papers
1as first author
3since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 33% Machine translation · 32% Language models and text generation · 18% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
parallel corpus mining |
0.5 | 1 | 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web · ACL/IJCNLP (1) 2021 |
Natural language and speech › Machine translation › parallel corpus mining
parallel sentence extraction |
0.5 | 1 | 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web · ACL/IJCNLP (1) 2021 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning |
0.4 | 1 | 2020 | Unsupervised Cross-lingual Representation Learning at Scale · ACL 2020 |
Natural language and speech › Information extraction and text analysis
fact-checking |
0.4 | 1 | 2020 | Generating Fact Checking Briefs · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
masked language modeling |
0.4 | 1 | 2020 | Unsupervised Cross-lingual Representation Learning at Scale · ACL 2020 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding |
0.2 | 1 | 2015 | Trans-gram, Fast Cross-lingual Word-embeddings · EMNLP 2015 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 1 | 2015 | Trans-gram, Fast Cross-lingual Word-embeddings · EMNLP 2015 |
Machine learning › Representation and self-supervised learning › text embedding
cross-lingual representation |
0.1 | 1 | 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation
text generation |
0.1 | 1 | 2020 | Generating Fact Checking Briefs · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › text classification › transfer learning for text classification
cross-lingual text classification |
0.1 | 1 | 2015 | Trans-gram, Fast Cross-lingual Word-embeddings · EMNLP 2015 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine TranslationabstractAbstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the Flores-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are fully aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond. Naman Goyal 0001, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzmán, Angela Fan |
Trans. Assoc. Comput. Linguistics | 5 |
| 2021 | CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebabstractHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave, Armand Joulin, Angela Fan |
ACL/IJCNLP (1) | 2 |
| 2021 | Beyond English-Centric Multilingual Machine TranslationabstractExisting work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric, training only on data which was translated from or to English.While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open-source a training data set that covers thousands of language directions with parallel data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems from the Workshop on Machine Translation (WMT). We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model. Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal 0001, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, Armand Joulin |
J. Mach. Learn. Res. | 9 |
| 2020 | Unsupervised Cross-lingual Representation Learning at ScaleabstractAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Alexis Conneau, Kartikay Khandelwal, Naman Goyal 0001, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov |
ACL | 5 |
| 2020 | Generating Fact Checking BriefsabstractAngela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, Sebastian Riedel. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos 0001, Antoine Bordes, Sebastian Riedel 0001 |
EMNLP (1) | 4 |
| 2020 | CCNet: Extracting High Quality Monolingual Datasets from Web Crawl DataabstractPre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved. In this paper, we describe an automatic pipeline to extract massive high-quality monolingual datasets from Common Crawl for a variety of languages. Our pipeline follows the data processing introduced in fastText (Mikolov et al., 2017; Grave et al., 2018), that deduplicates documents and identifies their language. We augment this pipeline with a filtering step to select documents that are close to high quality corpora like Wikipedia. Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, Edouard Grave |
LREC | 1 |
| 2015 | Trans-gram, Fast Cross-lingual Word-embeddingsabstractWe introduce Trans-gram, a simple and computationally-efficient method to simultaneously learn and align wordembeddings for a variety of languages, using only monolingual data and a smaller set of sentence-aligned data.We use our new method to compute aligned wordembeddings for twenty-one languages using English as a pivot language.We show that some linguistic features are aligned across languages for which we do not have aligned data, even though those properties do not exist in the pivot language.We also achieve state of the art results on standard cross-lingual text classification and word translation tasks. Jocelyn Coulmance, Jean-Marc Marty, Guillaume Wenzek, Amine Benhalloum |
EMNLP | 3 |