Michael Auli

dblp:11/9768 · DBLP profile ↗
← Back
66ranked-venue papers
5as first author
31since 2021 · last 2025
0000-0001-5974-4459ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 5 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 17 since 2021
YearPublicationVenuePosition
2025 Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
abstract
Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical settings. Automatic Spoken Language Identification (SLID) models are not perfect and misclassifications have a substantial impact on the final ASR accuracy. In this paper, we present a simple and effective N-best re-ranking approach to improve multilingual ASR accuracy for several prominent acoustic models by employing external features such as language models and text-based language identification models. Our results on FLEURS using the MMS and Whisper models show spoken language identification accuracy improvements of 8.7% and 6.1%, respectively and word error rates which are 3.3% and 2.0% lower on these benchmarks. The code is available at: https://github.com/facebookresearch/fairseq/tree/main/examples/mms/lid_rerank.
Brian Yan, Vineel Pratap, Shinji Watanabe 0001, Michael Auli
ICASSP4
2025 Scaling A Simple Approach to Zero-Shot Speech Recognition
abstract
Despite rapid progress in increasing the language coverage of automatic speech recognition, the field is still far from covering all languages with a known writing script. Recent work showed promising results with a zero-shot approach requiring only a small amount of text data, however, accuracy heavily depends on the quality of the used phonemizer which is often weak for unseen languages. In this paper, we present MMS Zero-shot, a conceptually simpler approach based on romanization and an acoustic model trained on data in 1,078 different languages or three orders of magnitude more than prior art. MMS Zero-shot reduces the average character error rate by a relative 46% over 100 unseen languages compared to the best previous work. Moreover, the error rate of our approach is only 2.5x higher than in-domain supervised baselines, while MMS Zero-shot uses no labeled data for the evaluation languages at all.
Jinming Zhao, Vineel Pratap, Michael Auli
ICASSP3
2025 Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
Vineel Pratap, Michael Auli, Jean Maillard
INTERSPEECH3
2024 Scaling Speech Technology to 1, 000+ Languages
abstract
Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the over 7,000 languages spoken around the world. The Massively Multilingual Speech (MMS) project increases the number of supported languages by 10-40x, depending on the task while providing improved accuracy compared to prior work. The main ingredients are a new dataset based on readings of publicly available religious texts and effectively leveraging self-supervised learning. We built pre-trained wav2vec 2.0 models covering 1,406 languages, a single multilingual automatic speech recognition model for 1,107 languages, speech synthesis models for the same number of languages, as well as a language identification model for 4,017 languages. Experiments show that our multilingual speech recognition model more than halves the word error rate of Whisper on 54 languages of the FLEURS benchmark while being trained on a small fraction of the labeled data.
Vineel Pratap, Andros Tjandra, Bowen Shi 0002, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiaohui Zhang 0007, Wei-Ning Hsu, Alexis Conneau, Michael Auli
J. Mach. Learn. Res.16
2023 Simple and Effective Unsupervised Speech Translation
abstract
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang 0002, Wei-Ning Hsu, Michael Auli, Juan Pino 0001
ACL (1)7
2023 Av-Data2Vec: Self-Supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
abstract
Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint representations of both modalities. In this paper, we introduce AV-data2vec which addresses these challenges and builds audio-visual representations based on predicting contextualized representations which has been successful in the uni-modal case. The model uses a shared transformer encoder for both audio and video and can combine both modalities to improve speech recognition. Results on LRS3 show that AV-data2vec consistently outperforms existing methods under all settings with the same amount of data and model size.
Jiachen Lian, Alexei Baevski, Wei-Ning Hsu, Michael Auli
ASRU4
2023 Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
abstract
Current self-supervised learning algorithms are often modality-specific and require large amounts of computational resources. To address these issues, we increase the training efficiency of data2vec, a learning objective that generalizes across several modalities. We do not encode masked tokens, use a fast convolutional decoder and amortize the effort to build teacher representations. data2vec 2.0 benefits from the rich contextualized target representations introduced in data2vec which enable a fast self-supervised learner. Experiments on ImageNet-1K image classification show that data2vec 2.0 matches the accuracy of Masked Autoencoders in 16.4x lower pre-training time, on Librispeech speech recognition it performs as well as wav2vec 2.0 in 10.6x less time, and on GLUE natural language understanding it matches a retrained RoBERTa model in half the time. Trading some speed for accuracy results in ImageNet-1K top-1 accuracy of 86.8% with a ViT-L model trained for 150 epochs.
Alexei Baevski, Arun Babu, Wei-Ning Hsu, Michael Auli
ICML4
2023 DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
abstract
In this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units.
Alexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu, James R. Glass
NeurIPS3
2022 Unified Speech-Text Pre-training for Speech Translation and Recognition
abstract
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, Juan Pino. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yun Tang 0002, Hongyu Gong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li 0003, Abdel-rahman Mohamed, Michael Auli, Juan Pino 0001
ACL (1)10
2022 Improved Language Identification Through Cross-Lingual Self-Supervised Learning
abstract
Language identification greatly impacts the success of downstream tasks such as automatic speech recognition. Recently, self-supervised speech representations learned by wav2vec 2.0 have been shown to be very effective for a range of speech tasks. We extend previous self-supervised work on language identification by experimenting with pre-trained models which were learned on real-world unconstrained speech in multiple languages and not just on English. We show that models pre-trained on many languages perform better and enable language identification systems that require very little labeled data to perform well. Results on a 26 languages setup show that with only 10 minutes of labeled data per language, a cross-lingually pre-trained model can achieve over 89.2% accuracy.
Andros Tjandra, Diptanu Gon Choudhury, Frank Zhang 0001, Kritika Singh, Alexis Conneau, Alexei Baevski, Assaf Sela, Yatharth Saraf, Michael Auli
ICASSP9
2022 data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
abstract
While the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind. To get us closer to general self-supervised learning, we present data2vec, a framework that uses the same learning method for either speech, NLP or computer vision. The core idea is to predict latent representations of the full input data based on a masked view of the input in a self-distillation setup using a standard Transformer architecture. Instead of predicting modality-specific targets such as words, visual tokens or units of human speech which are local in nature, data2vec predicts contextualized latent representations that contain information from the entire input. Experiments on the major benchmarks of speech recognition, image classification, and natural language understanding demonstrate a new state of the art or competitive performance to predominant approaches.
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, Michael Auli
ICML6
2022 XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
abstract
This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0.We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in 128 languages, an order of magnitude more public data than the largest known prior work.Our evaluation covers a wide range of tasks, domains, data regimes and languages, both high and low-resource.On the CoVoST-2 speech translation benchmark, we improve the previous state of the art by an average of 7.4 BLEU over 21 translation directions into English.For speech recognition, XLS-R improves over the best known prior work on BABEL, MLS, CommonVoice as well as VoxPopuli, lowering error rates by 14-34% relative on average.XLS-R also sets a new state of the art on VoxLin-gua107 language identification.Moreover, we show that with sufficient model size, cross-lingual pretraining can perform as well as English-only pretraining when translating English speech into other languages, a setting which favors monolingual pretraining.We hope XLS-R can help to improve speech processing tasks for many more languages of the world.Models and code are available at www.github.
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal 0001, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino 0001, Alexei Baevski, Alexis Conneau, Michael Auli
INTERSPEECH13
2022 XTREME-S: Evaluating Cross-lingual Speech Representations
abstract
We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages.XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval.Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning.This paper describes the new benchmark and establishes the first speech-only and speechtext baselines using XLS-R and mSLAM on all downstream tasks.We motivate the design choices and detail how to use the benchmark.Datasets and fine-tuning scripts are made easily accessible through the HuggingFace platform. 1
Alexis Conneau, Ankur Bapna, Yu Zhang 0033, Patrick von Platen, Anton Lozhkov, Colin Cherry, Ye Jia, Clara Rivera, Mihir Kale, Daan van Esch, Vera Axelrod, Simran Khanuja, Jonathan H. Clark, Orhan Firat, Michael Auli, Sebastian Ruder, Jason Riesa, Melvin Johnson
INTERSPEECH16
2022 Simple and Effective Unsupervised Speech Synthesis
abstract
We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe.The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesis.Using only unlabeled speech audio and unlabeled text as well as a lexicon, our method enables speech synthesis without the need for a human-labeled corpus.Experiments demonstrate the unsupervised system can synthesize speech similar to a supervised counterpart in terms of naturalness and intelligibility measured by human evaluation.
Alexander H. Liu, Cheng-I Lai, Wei-Ning Hsu, Michael Auli, Alexei Baevski, James R. Glass
INTERSPEECH4
2022 Wav2Vec-Aug: Improved self-supervised training with limited data
Anuroop Sriram, Michael Auli, Alexei Baevski
INTERSPEECH2
2022 On-demand compute reduction with stochastic wav2vec 2.0
abstract
Squeeze and Efficient Wav2vec (SEW) is a recently proposed architecture [1] that squeezes the input to the transformer encoder for compute efficient pre-training and inference with wav2vec 2.0 (W2V2) models.In this work, we propose stochastic compression for on-demand compute reduction for W2V2 models.As opposed to using a fixed squeeze factor, we sample it uniformly during training.We further introduce query and key-value pooling mechanisms that can be applied to each transformer layer for further compression.Our results for models pre-trained on 960h Librispeech dataset and fine-tuned on 10h of transcribed data show that using the same stochastic model, we get a smooth trade-off between word error rate (WER) and inference time with only marginal WER degradation compared to the W2V2 and SEW models trained for a specific setting.We further show that we can fine-tune the same stochastically pretrained model to a specific configuration to recover the WER difference resulting in significant computational savings on pretraining models from scratch.
Apoorv Vyas, Wei-Ning Hsu, Michael Auli, Alexei Baevski
INTERSPEECH3
2022 Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
abstract
Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data.However, in many cases there is labeled data available for related languages which is not utilized by these methods.This paper extends previous work on zero-shot cross-lingual transfer learning by fine-tuning a multilingually pretrained wav2vec 2.0 model to transcribe unseen languages.This is done by mapping phonemes of the training languages to the target language using articulatory features.Experiments show that this simple method significantly outperforms prior work which introduced task-specific architectures and used only part of a monolingually pretrained model.
Qiantong Xu, Alexei Baevski, Michael Auli
INTERSPEECH3
2022 Masked Autoencoders that Listen
abstract
This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only the non-masked tokens through encoder layers. The decoder then re-orders and decodes the encoded context padded with mask tokens, in order to reconstruct the input spectrogram. We find it beneficial to incorporate local window attention in the decoder, as audio spectrograms are highly correlated in local time and frequency bands. We then fine-tune the encoder with a lower masking ratio on target datasets. Empirically, Audio-MAE sets new state-of-the-art performance on six audio and speech classification tasks, outperforming other recent models that use external supervised pre-training. Our code and models is available at https://github.com/facebookresearch/AudioMAE.
Po-Yao Huang 0001, Hu Xu 0001, Juncheng Li 0001, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, Christoph Feichtenhofer
NeurIPS5
2022 Towards End-to-End Unsupervised Speech Recognition
abstract
Unsupervised speech recognition has shown great potential to make Automatic Speech Recognition (ASR) systems accessible to every language. However, existing methods still heavily rely on hand-crafted pre-processing. Similar to the trend of making supervised speech recognition end-to-end, we introduce wav2vec-U 2.0 which does away with all audio-side pre-processing and improves accuracy through better architecture. In addition, we introduce an auxiliary self-supervised objective that ties model predictions back to the input. Experiments show that wav2vec-U 2.0 improves unsupervised recognition results across different languages while being conceptually simpler.
Alexander H. Liu, Wei-Ning Hsu, Michael Auli, Alexei Baevski
SLT3
2021 Discriminative Reranking for Neural Machine Translation
abstract
Ann Lee, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ann Lee 0001, Michael Auli, Marc'Aurelio Ranzato
ACL/IJCNLP (1)2
2021 Multilingual Speech Translation from Efficient Finetuning of Pretrained Models
abstract
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, Michael Auli. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xian Li 0003, Changhan Wang, Yun Tang 0002, Chau Tran, Juan Pino 0001, Alexei Baevski, Alexis Conneau, Michael Auli
ACL/IJCNLP (1)9
2021 Reservoir Transformers
abstract
Sheng Shen, Alexei Baevski, Ari Morcos, Kurt Keutzer, Michael Auli, Douwe Kiela. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Sheng Shen 0001, Alexei Baevski, Ari S. Morcos, Kurt Keutzer, Michael Auli, Douwe Kiela
ACL/IJCNLP (1)5
2021 The Source-Target Domain Mismatch Problem in Machine Translation
abstract
Jiajun Shen, Peng-Jen Chen, Matthew Le, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Peng-Jen Chen, Matt Le 0001, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc'Aurelio Ranzato
EACL7
2021 Self-Training and Pre-Training are Complementary for Speech Recognition
abstract
Self-training and unsupervised pre-training have emerged as effective approaches to improve speech recognition systems using unlabeled data. However, it is not clear whether they learn similar patterns or if they can be effectively combined. In this paper, we show that pseudo-labeling and pre-training with wav2vec 2.0 are complementary in a variety of labeled data setups. Using just 10 minutes of labeled data from Libri-light as well as 53k hours of unlabeled data from LibriVox achieves word error rates (WER) of 2.8%/4.8% on the clean and other test sets of Librispeech – rivaling the best published systems trained on 960 hours of labeled data only a year ago. Training on all labeled data of Librispeech achieves WERs of 1.5%/3.1%.
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, Michael Auli
ICASSP8
2021 A Comparison of Discrete Latent Variable Models for Speech Representation Learning
abstract
Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly based on predicting future time-steps or auto-encoding the input signal. Our study compares the representations learned by vq-vae and vq-wav2vec in terms of sub-word unit discovery and phoneme recognition performance. Results show that future time-step prediction with vq-wav2vec achieves better performance. The best system achieves an error rate of 13.22 on the ZeroSpeech 2019 ABX phoneme discrimination challenge.
Henry Zhou, Alexei Baevski, Michael Auli
ICASSP3
2021 Unsupervised Cross-Lingual Representation Learning for Speech Recognition
abstract
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The resulting model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. On the CommonVoice benchmark, XLSR shows a relative phoneme error rate reduction of 72% compared to the best known results. On BABEL, our approach improves word error rate by 16% relative compared to a comparable system. Our approach enables a single multilingual speech recognition model which is competitive to strong individual models. Analysis shows that the latent discrete speech representations are shared across languages with increased sharing for related languages. We hope to catalyze research in low-resource speech understanding by releasing XLSR-53, a large model pretrained in 53 languages.
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdel-rahman Mohamed, Michael Auli
Interspeech5
2021 Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
abstract
Self-supervised learning of speech representations has been a very active research area but most work is focused on a single domain such as read audio books for which there exist large quantities of labeled and unlabeled data. In this paper, we explore more general setups where the domain of the unlabeled data for pre-training data differs from the domain of the labeled data for fine-tuning, which in turn may differ from the test data domain. Our experiments show that using target domain data during pre-training leads to large performance improvements across a variety of setups. On a large-scale competitive setup, we show that pre-training on unlabeled in-domain data reduces the gap between models trained on in-domain and out-of-domain labeled data by 66%-73%. This has obvious practical implications since it is much easier to obtain unlabeled target domain data than labeled data. Moreover, we find that pre-training on multiple domains improves generalization performance on domains not seen during training. Code and models will be made available at https://github.com/pytorch/fairseq.
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee 0001, Ronan Collobert, Gabriel Synnaeve, Michael Auli
Interspeech11
2021 Large-Scale Self- and Semi-Supervised Learning for Speech Translation
abstract
In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore both pretraining and self-training by using the large Libri-Light speech audio corpus and language modeling with CommonCrawl. Our experiments improve over the previous state of the art by 2.6 BLEU on average on all four considered CoVoST 2 language pairs via a simple recipe of combining wav2vec 2.0 pretraining, a single iteration of self-training and decoding with a language model. Different to existing work, our approach does not leverage any other supervision than ST data. Code and models will be publicly released.
Changhan Wang, Anne Wu, Juan Pino 0001, Alexei Baevski, Michael Auli, Alexis Conneau
Interspeech5
2021 Self-training Improves Pre-training for Natural Language Understanding
abstract
Jingfei Du, Edouard Grave, Beliz Gunel, Vishrav Chaudhary, Onur Celebi, Michael Auli, Veselin Stoyanov, Alexis Conneau. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jingfei Du, Edouard Grave, Beliz Gunel, Vishrav Chaudhary, Onur Celebi, Michael Auli, Veselin Stoyanov, Alexis Conneau
NAACL-HLT6
2021 Unsupervised Speech Recognition
abstract
Despite rapid progress in the recent past, current speech recognition systems still require labeled training data which limits this technology to a small fraction of the languages spoken around the globe. This paper describes wav2vec-U, short for wav2vec Unsupervised, a method to train speech recognition models without any labeled data. We leverage self-supervised speech representations to segment unlabeled audio and learn a mapping from these representations to phonemes via adversarial training. The right representations are key to the success of our method. Compared to the best previous unsupervised work, wav2vec-U reduces the phone error rate on the TIMIT benchmark from 26.1 to 11.3. On the larger English Librispeech benchmark, wav2vec-U achieves a word error rate of 5.9 on test-other, rivaling some of the best published systems trained on 960 hours of labeled data from only two years ago. We also experiment on nine other languages, including low-resource languages such as Kyrgyz, Swahili and Tatar.
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael Auli
NeurIPS4
2021 Beyond English-Centric Multilingual Machine Translation
abstract
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric, training only on data which was translated from or to English.While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open-source a training data set that covers thousands of language directions with parallel data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems from the Workshop on Machine Translation (WMT). We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model.
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal 0001, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, Armand Joulin
J. Mach. Learn. Res.15
2020 On The Evaluation of Machine Translation SystemsTrained With Back-Translation
abstract
Back-translation is a widely used data augmentation technique which leverages target monolingual data.However, its effectiveness has been challenged since automatic metrics such as BLEU only show significant improvements for test examples where the source itself is a translation, or translationese.This is believed to be due to translationese inputs better matching the back-translated training data.In this work, we show that this conjecture is not empirically supported and that backtranslation improves translation quality of both naturally occurring text as well as translationese according to professional human translators.We provide empirical evidence to support the view that back-translation is preferred by humans because it produces more fluent outputs.BLEU cannot capture human preferences because references are translationese when source sentences are natural text.We recommend complementing BLEU with a language model score to measure fluency.
Sergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael Auli
ACL4
2020 Robust and On-the-Fly Dataset Denoising for Image Classification
Jiaming Song, Yann N. Dauphin, Michael Auli, Tengyu Ma 0001
ECCV (29)3
2020 vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Alexei Baevski, Steffen Schneider 0001, Michael Auli
ICLR3
2020 Depth-Adaptive Transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, Michael Auli
ICLR4
2020 wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
abstract
We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler. wav2vec 2.0 masks the speech input in the latent space and solves a contrastive task defined over a quantization of the latent representations which are jointly learned. Experiments using all labeled data of Librispeech achieve 1.8/3.3 WER on the clean/other test sets. When lowering the amount of labeled data to one hour, wav2vec 2.0 outperforms the previous state of the art on the 100 hour subset while using 100 times less labeled data. Using just ten minutes of labeled data and pre-training on 53k hours of unlabeled data still achieves 4.8/8.2 WER. This demonstrates the feasibility of speech recognition with limited amounts of labeled data.
Alexei Baevski, Abdel-rahman Mohamed, Michael Auli
NeurIPS4
2020 Modeling Human Motion with Quaternion-Based Neural Networks
Dario Pavllo, Christoph Feichtenhofer, Michael Auli, David Grangier
Int. J. Comput. Vis.3
2019 ELI5: Long Form Question Answering
abstract
We introduce the first large-scale corpus for long-form question answering, a task requiring elaborate and in-depth answers to openended questions.The dataset comprises 270K threads from the Reddit forum "Explain Like I'm Five" (ELI5) where an online community provides answers to questions which are comprehensible by five year olds.Compared to existing datasets, ELI5 comprises diverse questions requiring multi-sentence answers.We provide a large set of web documents to help answer the question.Automatic and human evaluations show that an abstractive model trained with a multi-task objective outperforms conventional Seq2Seq, language modeling, as well as a strong extractive baseline.However, our best model is still far from human performance since raters prefer gold responses in over 86% of cases, leaving ample opportunity for future improvement.1
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, Michael Auli
ACL (1)6
2019 3D Human Pose Estimation in Video With Temporal Convolutions and Semi-Supervised Training
abstract
In this work, we demonstrate that 3D poses in video can be effectively estimated with a fully convolutional model based on dilated temporal convolutions over 2D keypoints. We also introduce back-projection, a simple and effective semi-supervised training method that leverages unlabeled video data. We start with predicted 2D keypoints for unlabeled video, then estimate 3D poses and finally back-project to the input 2D keypoints. In the supervised setting, our fully-convolutional model outperforms the previous best result from the literature by 6 mm mean per-joint position error on Human3.6M, corresponding to an error reduction of 11%, and the model also shows significant improvements on HumanEva-I. Moreover, experiments with back-projection show that it comfortably outperforms previous state-of-the-art results in semi-supervised settings where labeled data is scarce. Code and models are available at https://github.com/facebookresearch/VideoPose3D.
Dario Pavllo, Christoph Feichtenhofer, David Grangier, Michael Auli
CVPR4
2019 Cloze-driven Pretraining of Self-attention Networks
abstract
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, Michael Auli. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, Michael Auli
EMNLP/IJCNLP (1)5
2019 Simple and Effective Noisy Channel Modeling for Neural Machine Translation
abstract
Kyra Yee, Yann Dauphin, Michael Auli. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Kyra Yee, Yann N. Dauphin, Michael Auli
EMNLP/IJCNLP (1)3
2019 Adaptive Input Representations for Neural Language Modeling
Alexei Baevski, Michael Auli
ICLR (Poster)2
2019 Wizard of Wikipedia: Knowledge-Powered Conversational Agents
Emily Dinan, Stephen Roller, Kurt Shuster 0001, Angela Fan, Michael Auli, Jason Weston
ICLR (Poster)5
2019 Pay Less Attention with Lightweight and Dynamic Convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, Michael Auli
ICLR5
2019 Mixture Models for Diverse Machine Translation: Tricks of the Trade
abstract
Mixture models trained via EM are among the simplest, most widely used and well understood latent variable models in the machine learning literature. Surprisingly, these models have been hardly explored in text generation applications such as machine translation. In principle, they provide a latent variable to control generation and produce a diverse set of hypotheses. In practice, however, mixture models are prone to degeneracies—often only one component gets trained or the latent variable is simply ignored. We find that disabling dropout noise in responsibility computation is critical to successful training. In addition, the design choices of parameterization, prior distribution, hard versus soft EM and online versus offline assignment can dramatically affect model performance. We develop an evaluation protocol to assess both quality and diversity of generations against multiple references, and provide an extensive empirical study of several mixture model variants. Our analysis shows that certain types of mixture models are more robust and offer the best trade-off between translation quality and diversity compared to variational models and diverse decoding approaches.\footnote{Code to reproduce the results in this paper is available at \url{https://github.com/pytorch/fairseq}}
Tianxiao Shen, Myle Ott, Michael Auli, Marc'Aurelio Ranzato
ICML3
2019 wav2vec: Unsupervised Pre-Training for Speech Recognition
abstract
We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model training. We pre-train a simple multi-layer convolutional neural network optimized via a noise contrastive binary classification task. Our experiments on WSJ reduce WER of a strong character-based log-mel filterbank baseline by up to 36% when only a few hours of transcribed data is available. Our approach achieves 2.43% WER on the nov92 test set. This outperforms Deep Speech 2, the best reported character-based system in the literature while using two orders of magnitude less labeled training data.
Steffen Schneider 0001, Alexei Baevski, Ronan Collobert, Michael Auli
INTERSPEECH4
2018 QuaterNet: A Quaternion-based Recurrent Model for Human Motion
Dario Pavllo, David Grangier, Michael Auli
BMVC3
2018 Understanding Back-Translation at Scale
abstract
An effective method to improve neural machine translation with monolingual data is to augment the parallel training corpus with back-translations of target language sentences.This work broadens the understanding of back-translation and investigates a number of methods to generate synthetic source sentences.We find that in all but resource poor settings back-translations obtained via sampling or noised beam outputs are most effective.Our analysis shows that sampling or noisy synthetic data gives a much stronger training signal than data generated by beam or greedy search.We also compare how synthetic data compares to genuine bitext and study various domain effects.Finally, we scale to hundreds of millions of monolingual sentences and achieve a new state of the art of 35 BLEU on the WMT'14 English-German test set.
Sergey Edunov, Myle Ott, Michael Auli, David Grangier
EMNLP3
2018 Analyzing Uncertainty in Neural Machine Translation
abstract
Machine translation is a popular test bed for research in neural sequence-to-sequence models but despite much recent research, there is still a lack of understanding of these models. Practitioners report performance degradation with large beams, the under-estimation of rare words and a lack of diversity in the final translations. Our study relates some of these issues to the inherent uncertainty of the task, due to the existence of multiple valid translations for a single source sentence, and to the extrinsic uncertainty caused by noisy training data. We propose tools and metrics to assess how uncertainty in the data is captured by the model distribution and how it affects search strategies that generate translations. Our results show that search works remarkably well but that the models tend to spread too much probability mass over the hypothesis space. Next, we propose tools to assess model calibration and show how to easily fix some shortcomings of current models. We release both code and multiple human reference translations for two popular benchmarks.
Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato
ICML2
2018 Classical Structured Prediction Losses for Sequence to Sequence Learning
abstract
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc’Aurelio Ranzato. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato
NAACL-HLT3
2018 QuickEdit: Editing Text & Translations by Crossing Words Out
abstract
David Grangier, Michael Auli. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
David Grangier, Michael Auli
NAACL-HLT2
2017 A Convolutional Encoder Model for Neural Machine Translation
abstract
The prevalent approach to neural machine translation relies on bi-directional LSTMs to encode the source sentence.We present a faster and simpler architecture based on a succession of convolutional layers.This allows to encode the source sentence simultaneously compared to recurrent networks for which computation is constrained by temporal dependencies.On WMT'16 English-Romanian translation we achieve competitive accuracy to the state-of-the-art and on WMT'15 English-German we outperform several recently published results.Our models obtain almost the same accuracy as a very deep LSTM setup on WMT'14 English-French translation.We speed up CPU decoding by more than two times at the same or higher accuracy as a strong bidirectional LSTM. 1
Jonas Gehring, Michael Auli, David Grangier, Yann N. Dauphin
ACL (1)2
2017 Language Modeling with Gated Convolutional Networks
abstract
The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a finite context approach through stacked convolutions, which can be more efficient since they allow parallelization over sequential tokens. We propose a novel simplified gating mechanism that outperforms Oord et al. (2016) and investigate the impact of key architectural decisions. The proposed approach achieves state-of-the-art on the WikiText-103 benchmark, even though it features long-term dependencies, as well as competitive results on the Google Billion Words benchmark. Our model reduces the latency to score a sentence by an order of magnitude compared to a recurrent baseline. To our knowledge, this is the first time a non-recurrent approach is competitive with strong recurrent models on these large scale language tasks.
Yann N. Dauphin, Angela Fan, Michael Auli, David Grangier
ICML3
2017 Convolutional Sequence to Sequence Learning
abstract
The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully parallelized during training to better exploit the GPU hardware and optimization is easier since the number of non-linearities is fixed and independent of the input length. Our use of gated linear units eases gradient propagation and we equip each decoder layer with a separate attention module. We outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT’14 English-German and WMT’14 English-French translation at an order of magnitude faster speed, both on GPU and CPU.
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin
ICML2
2016 Strategies for Training Large Vocabulary Neural Language Models
abstract
Training neural network language models over large vocabularies is computationally costly compared to count-based models such as Kneser-Ney.We present a systematic comparison of neural strategies to represent and train large vocabularies, including softmax, hierarchical softmax, target sampling, noise contrastive estimation and self normalization.We extend self normalization to be a proper estimator of likelihood and introduce an efficient variant of softmax.We evaluate each method on three popular benchmarks, examining performance on rare words, the speed/accuracy trade-off and complementarity to Kneser-Ney.
David Grangier, Michael Auli
ACL (1)3
2016 Neural Text Generation from Structured Data with Application to the Biography Domain
abstract
This paper introduces a neural model for concept-to-text generation that scales to large, rich domains.It generates biographical sentences from fact tables on a new dataset of biographies from Wikipedia.This set is an order of magnitude larger than existing resources with over 700k samples and a 400k vocabulary.Our model builds on conditional neural language models for text generation.To deal with the large vocabulary, we extend these models to mix a fixed vocabulary with copy actions that transfer sample-specific words from the input database to the generated output sentence.To deal with structured data, we allow the model to embed words differently depending on the data fields in which they occur.Our neural model significantly outperforms a Templated Kneser-Ney language model by nearly 15 BLEU.
Rémi Lebret, David Grangier, Michael Auli
EMNLP3
2016 Abstractive Sentence Summarization with Attentive Recurrent Neural Networks
abstract
Abstractive Sentence Summarization generates a shorter version of a given sentence while attempting to preserve its meaning.We introduce a conditional recurrent neural network (RNN) which generates a summary of an input sentence.The conditioning is provided by a novel convolutional attention-based encoder which ensures that the decoder focuses on the appropriate input words at each step of generation.Our model relies only on learned features and is easy to train in an end-to-end fashion on large data sets.Our experiments show that the model significantly outperforms the recently proposed state-of-the-art method on the Gigaword corpus while performing competitively on the DUC-2004 shared task.
Sumit Chopra, Michael Auli, Alexander M. Rush
HLT-NAACL2
2016 Expected F-Measure Training for Shift-Reduce Parsing with Recurrent Neural Networks
abstract
We present expected F-measure training for shift-reduce parsing with RNNs, which enables the learning of a global parsing model optimized for sentence-level F1.We apply the model to CCG parsing, where it improves over a strong greedy RNN baseline, by 1.47% F1, yielding state-of-the-art results for shiftreduce CCG parsing.
Wenduan Xu, Michael Auli, Stephen Clark
HLT-NAACL2
2015 A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
abstract
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao 0001, William B. Dolan
HLT-NAACL3
2015 Learning Translation Models from Monolingual Continuous Representations
abstract
Translation models often fail to generate good translations for infrequent words or phrases. Previous work attacked this problem by inducing new translation rules from monolingual data with a semi-supervised algorithm. However, this approach does not scale very well since it is very computationally expensive to generate new translation rules for only a few thousand sentences. We propose a much faster and simpler method that directly hallucinates translation rules for infrequent phrases based on phrases with similar continuous representations for which a translation is known. To speed up the retrieval of similar phrases, we investigate approximated nearest neighbor search with redundant bit vectors which we find to be three times faster and significantly more accurate than locality sensitive hashing. Our approach of learning new translation rules improves a phrase-based baseline by up to 1.6 BLEU on Arabic-English translation, it is three-orders of magnitudes faster than existing semi-supervised methods and 0.5 BLEU more accurate.
Hany Hassan, Michael Auli
HLT-NAACL3
2014 Minimum Translation Modeling with Recurrent Neural Networks
abstract
We introduce recurrent neural networkbased Minimum Translation Unit (MTU) models which make predictions based on an unbounded history of previous bilingual contexts.Traditional back-off n-gram models suffer under the sparse nature of MTUs which makes estimation of highorder sequence models challenging.We tackle the sparsity problem by modeling MTUs both as bags-of-words and as a sequence of individual source and target words.Our best results improve the output of a phrase-based statistical machine translation system trained on WMT 2012 French-English data by up to 1.5 BLEU, and we outperform the traditional n-gram based MTU approach by up to 0.8 BLEU.
Yuening Hu, Michael Auli, Qin Gao, Jianfeng Gao 0001
EACL2
2014 Large-scale Expected BLEU Training of Phrase-based Reordering Models
abstract
Recent work by Cherry (2013) has shown that directly optimizing phrase-based re-ordering models towards BLEU can lead to significant gains. Their approach is lim-ited to small training sets of a few thou-sand sentences and a similar number of sparse features. We show how the ex-pected BLEU objective allows us to train a simple linear discriminative reordering model with millions of sparse features on hundreds of thousands of sentences re-sulting in significant improvements. A comparison to likelihood training demon-strates that expected BLEU is vastly more effective. Our best results improve a hi-erarchical lexicalized reordering baseline by up to 2.0 BLEU in a single-reference setting on a French-English WMT 2012 setup. 1
Michael Auli, Michel Galley, Jianfeng Gao 0001
EMNLP1
2013 Joint Language and Translation Modeling with Recurrent Neural Networks
abstract
We present a joint language and translation model based on a recurrent neural network which predicts target words based on an unbounded history of both source and target words.The weaker independence assumptions of this model result in a vastly larger search space compared to related feedforward-based language or translation models.We tackle this issue with a new lattice rescoring algorithm and demonstrate its effectiveness empirically.Our joint model builds on a well known recurrent neural network language model (Mikolov, 2012) augmented by a layer of additional inputs from the source language.We show competitive accuracy compared to the traditional channel model features.Our best results improve the output of a system trained on WMT 2012 French-English data by up to 1.5 BLEU, and by 1.1 BLEU on average across several test sets.
Michael Auli, Michel Galley, Chris Quirk, Geoffrey Zweig
EMNLP1
2011 A Comparison of Loopy Belief Propagation and Dual Decomposition for Integrated CCG Supertagging and Parsing
Michael Auli, Adam Lopez
ACL1
2011 Efficient CCG Parsing: A* versus Adaptive Supertagging
Michael Auli, Adam Lopez
ACL1
2011 Training a Log-Linear Parser with Loss Functions via Softmax-Margin
Michael Auli, Adam Lopez
EMNLP1