Noam Shazeer

dblp:80/4668 · also Noam M. Shazeer · DBLP profile ↗
← Back
28ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Efficient and distributed learning · 29% Language models and text generation · 24% Deep learning architectures and training · 18%
Computer graphics and multimedia
2 papers
Audio and music processing · 54% Image and video processing · 46%

Topics — the 30 heaviest of 60, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.432023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · J. Mach. Learn. Res. 2022
Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation
1.132021
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding · ICLR 2021
HydraNets: Specialized Dynamic Architectures for Efficient Inference · CVPR 2018
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR (Poster) 2017
Machine learning › Efficient and distributed learning
distributed training
1.032023
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding · ICLR 2021
Mesh-TensorFlow: Deep Learning for Supercomputers · NeurIPS 2018
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
mixture of experts
0.922022
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · J. Mach. Learn. Res. 2022
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR (Poster) 2017
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.822023
Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding · ICLR 2021
Natural language and speech › Machine translation
neural machine translation
0.732018
The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation · ACL (1) 2018
Attention is All you Need · NIPS 2017
Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018
Machine learning › Generative modeling
autoregressive model
0.722019
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Image Transformer · ICML 2018
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning › distributed training
distributed training systems
0.712023
Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
instruction following
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding
0.722018
Blockwise Parallel Decoding for Deep Autoregressive Models · NeurIPS 2018
Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018
Machine learning › Deep learning architectures and training
scaling laws
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
transformer
0.632021
Image Transformer · ICML 2018
Searching for Efficient Transformers for Language Modeling · NeurIPS 2021
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity
0.612022
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning › large-scale learning
model scaling
0.612022
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · J. Mach. Learn. Res. 2022
Machine learning › Representation and self-supervised learning
pre-training
0.612022
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · J. Mach. Learn. Res. 2022
Machine learning › Optimization for machine learning
sparse model
0.612022
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.522023
Mesh-TensorFlow: Deep Learning for Supercomputers · NeurIPS 2018
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.512021
Searching for Efficient Transformers for Language Modeling · NeurIPS 2021
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.512021
Searching for Efficient Transformers for Language Modeling · NeurIPS 2021
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
transformer architecture search
0.512021
Searching for Efficient Transformers for Language Modeling · NeurIPS 2021
Machine learning › Learning paradigms › multi-task learning
multi-task transfer learning
0.412020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer · J. Mach. Learn. Res. 2020
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.412020
How Much Knowledge Can You Pack Into the Parameters of a Language Model? · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pre-training and fine-tuning
0.412020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer · J. Mach. Learn. Res. 2020
Machine learning › Transfer learning and domain adaptation
transfer learning for NLP
0.412020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer · J. Mach. Learn. Res. 2020
Audio and music processing
music generation
0.412019
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.312018
Generating Wikipedia by Summarizing Long Sequences · ICLR (Poster) 2018
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate
0.312018
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost · ICML 2018
Natural language and speech › Language models and text generation › decoding
autoregressive decoding
0.312018
Blockwise Parallel Decoding for Deep Autoregressive Models · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

self-attention · 1.9transformer · 1.0autoregressive modeling · 1.0fine-tuning · 0.9pathways · 0.7sparsity · 0.6mixture-of-experts routing · 0.6bfloat16 training · 0.6reproducibility study · 0.5ablation study · 0.5relative positional encoding · 0.4collective communication · 0.3allreduce · 0.3SPMD programming · 0.3
YearPublicationVenuePosition
2025 Hot ChipsKeynote
Noam Shazeer
HCS1
2023 PaLM: Scaling Language Modeling with Pathways
abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel
J. Mach. Learn. Res.18
2023 Scaling Up Models and Data with t5x and seqio
abstract
Scaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio.
Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo
J. Mach. Learn. Res.34
2022 Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
abstract
In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) models defy this and instead select different parameters for each incoming example. The result is a sparsely-activated model---with an outrageous number of parameters---but a constant computational cost. However, despite several notable successes of MoE, widespread adoption has been hindered by complexity, communication costs, and training instability. We address these with the introduction of the Switch Transformer. We simplify the MoE routing algorithm and design intuitive improved models with reduced communication and computational costs. Our proposed training techniques mitigate the instabilities, and we show large sparse models may be trained, for the first time, with lower precision (bfloat16) formats. We design models based off T5-Base and T5-Large to obtain up to 7x increases in pre-training speed with the same computational resources. These improvements extend into multilingual settings where we measure gains over the mT5-Base version across all 101 languages. Finally, we advance the current scale of language models by pre-training up to trillion parameter models on the "Colossal Clean Crawled Corpus", and achieve a 4x speedup over the T5-XXL model.
William Fedus, Barret Zoph, Noam Shazeer
J. Mach. Learn. Res.3
2021 Do Transformer Modifications Transfer Across Implementations and Applications?
abstract
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, Colin Raffel. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhen-Zhong Lan, Yanqi Zhou, Wei Li 0133, Nan Ding 0002, Jake Marcus, Adam Roberts, Colin Raffel
EMNLP (1)9
2021 GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer
ICLR8
2021 Searching for Efficient Transformers for Language Modeling
abstract
Large Transformer models have been central to recent advances in natural language processing. The training and inference costs of these models, however, have grown rapidly and become prohibitively expensive. Here we aim to reduce the costs of Transformers by searching for a more efficient variant. Compared to previous approaches, our search is performed at a lower level, over the primitives that define a Transformer TensorFlow program. We identify an architecture, named Primer, that has a smaller training cost than the original Transformer and other variants for auto-regressive language modeling. Primer’s improvements can be mostly attributed to two simple modifications: squaring ReLU activations and adding a depthwise convolution layer after each Q, K, and V projection in self-attention.Experiments show Primer’s gains over Transformer increase as compute scale grows and follow a power law with respect to quality at optimal model sizes. We also verify empirically that Primer can be dropped into different codebases to significantly speed up training without additional tuning. For example, at a 500M parameter size, Primer improves the original T5 architecture on C4 auto-regressive language modeling, reducing the training cost by 4X. Furthermore, the reduced training cost means Primer needs much less compute to reach a target one-shot performance. For instance, in a 1.9B parameter configuration similar to GPT-3 XL, Primer uses 1/3 of the training compute to achieve the same one-shot performance as Transformer. We open source our models and several comparisons in T5 to help with reproducibility.
David R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, Quoc V. Le
NeurIPS5
2020 How Much Knowledge Can You Pack Into the Parameters of a Language Model?
abstract
It has recently been observed that neural language models trained on unstructured text can implicitly store and retrieve knowledge using natural language queries.In this short paper, we measure the practical utility of this approach by fine-tuning pre-trained models to answer questions without access to any external context or knowledge.We show that this approach scales with model size and performs competitively with open-domain systems that explicitly retrieve answers from an external knowledge source when answering questions.To facilitate reproducibility and future work, we release our code and trained models. 1 * Equal contribution.Noam suggested trying T5 on open-domain QA and coded and ran initial experiments on TriviaQA showing improved performance with model size.Adam wrote the code and ran most experiments.Colin set the research scope, wrote the paper, and ran a few experiments.1 https://goo.gle/t5-cbqa
Adam Roberts, Colin Raffel, Noam Shazeer
EMNLP (1)3
2020 Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
abstract
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning has given rise to a diversity of approaches, methodology, and practice. In this paper, we explore the landscape of transfer learning techniques for NLP by introducing a unified framework that converts all text-based language problems into a text-to-text format. Our systematic study compares pre-training objectives, architectures, unlabeled data sets, transfer approaches, and other factors on dozens of language understanding tasks. By combining the insights from our exploration with scale and our new “Colossal Clean Crawled Corpus”, we achieve state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more. To facilitate future work on transfer learning for NLP, we release our data set, pre-trained models, and code.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li 0133, Peter J. Liu
J. Mach. Learn. Res.2
2019 Music Transformer: Generating Music with Long-Term Structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew Hoffman 0001, Monica Dinculescu, Douglas Eck
ICLR (Poster)6
2018 The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
abstract
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes
ACL (1)9
2018 HydraNets: Specialized Dynamic Architectures for Efficient Inference
abstract
There is growing interest in improving the design of deep network architectures to be both accurate and low cost. This paper explores semantic specialization as a mechanism for improving the computational efficiency (accuracy-per-unit-cost) of inference in the context of image classification. Specifically, we propose a network architecture template called HydraNet, which enables state-of-the-art architectures for image classification to be transformed into dynamic architectures which exploit conditional execution for efficient inference. HydraNets are wide networks containing distinct components specialized to compute features for visually similar classes, but they retain efficiency by dynamically selecting only a small number of components to evaluate for any one input image. This design is made possible by a soft gating mechanism that encourages component specialization during training and accurately performs component selection during inference. We evaluate the HydraNet approach on both the CIFAR-100 and ImageNet classification tasks. On CIFAR, applying the HydraNet template to the ResNet and DenseNet family of models reduces inference cost by 2-4× while retaining the accuracy of the baseline architectures. On ImageNet, applying the HydraNet template improves accuracy up to 2.5% when compared to an efficient baseline architecture with similar inference cost.
Ravi Teja Mullapudi, William R. Mark, Noam Shazeer, Kayvon Fatahalian
CVPR3
2018 Generating Wikipedia by Summarizing Long Sequences
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, Noam Shazeer
ICLR (Poster)7
2018 Fast Decoding in Sequence Models Using Discrete Latent Variables
abstract
Autoregressive sequence models based on deep neural networks, such as RNNs, Wavenet and Transformer are the state-of-the-art on many tasks. However, they lack parallelism and are thus slow for long sequences. RNNs lack parallelism both during training and decoding, while architectures like WaveNet and Transformer are much more parallel during training, but still lack parallelism during decoding. We present a method to extend sequence models using discrete latent variables that makes decoding much more parallel. The main idea behind this approach is to first autoencode the target sequence into a shorter discrete latent sequence, which is generated autoregressively, and finally decode the full sequence from this shorter latent sequence in a parallel manner. To this end, we introduce a new method for constructing discrete latent variables and compare it with previously introduced methods. Finally, we verify that our model works on the task of neural machine translation, where our models are an order of magnitude faster than comparable autoregressive models and, while lower in BLEU than purely autoregressive models, better than previously proposed non-autogregressive translation.
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, Noam Shazeer
ICML7
2018 Image Transformer
abstract
Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran
ICML5
2018 Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
abstract
In several recently proposed stochastic optimization methods (e.g. RMSProp, Adam, Adadelta), parameter updates are scaled by the inverse square roots of exponential moving averages of squared past gradients. Maintaining these per-parameter second-moment estimators requires memory equal to the number of parameters. For the case of neural network weight matrices, we propose maintaining only the per-row and per-column sums of these moving averages, and estimating the per-parameter second moments based on these sums. We demonstrate empirically that this method produces similar results to the baseline. Secondly, we show that adaptive methods can produce larger-than-desired updates when the decay rate of the second moment accumulator is too slow. We propose update clipping and a gradually increasing decay rate scheme as remedies. Combining these methods and dropping momentum, we achieve comparable results to the published Adam regime in training the Transformer model on the WMT 2014 English-German machine translation task, while using very little auxiliary storage in the optimizer. Finally, we propose scaling the parameter updates based on the scale of the parameters themselves.
Noam Shazeer, Mitchell Stern
ICML1
2018 Mesh-TensorFlow: Deep Learning for Supercomputers
abstract
Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to memory constraints), high latency, and inefficiency at small batch sizes. All of these can be solved by more general distribution strategies (model-parallelism). Unfortunately, efficient model-parallel algorithms tend to be complicated to discover, describe, and to implement, particularly on large clusters. We introduce Mesh-TensorFlow, a language for specifying a general class of distributed tensor computations. Where data-parallelism can be viewed as splitting tensors and operations along the "batch" dimension, in Mesh-TensorFlow, the user can specify any tensor-dimensions to be split across any dimensions of a multi-dimensional mesh of processors. A Mesh-TensorFlow graph compiles into a SPMD program consisting of parallel operations coupled with collective communication primitives such as Allreduce. We use Mesh-TensorFlow to implement an efficient data-parallel, model-parallel version of the Transformer sequence-to-sequence model. Using TPU meshes of up to 512 cores, we train Transformer models with up to 5 billion parameters, surpassing SOTA results on WMT'14 English-to-French translation task and the one-billion-word Language modeling benchmark. Mesh-Tensorflow is available at https://github.com/tensorflow/mesh
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, Blake A. Hechtman
NeurIPS1
2018 Blockwise Parallel Decoding for Deep Autoregressive Models
abstract
Deep autoregressive sequence-to-sequence models have demonstrated impressive performance across a wide variety of tasks in recent years. While common architecture classes such as recurrent, convolutional, and self-attention networks make different trade-offs between the amount of computation needed per layer and the length of the critical path at training time, generation still remains an inherently sequential process. To overcome this limitation, we propose a novel blockwise parallel decoding scheme in which we make predictions for multiple time steps in parallel then back off to the longest prefix validated by a scoring model. This allows for substantial theoretical improvements in generation speed when applied to architectures that can process output sequences in parallel. We verify our approach empirically through a series of experiments using state-of-the-art self-attention models for machine translation and image super-resolution, achieving iteration reductions of up to 2x over a baseline greedy decoder with no loss in quality, or up to 7x in exchange for a slight decrease in performance. In terms of wall-clock time, our fastest models exhibit real-time speedups of up to 4x over standard greedy decoding.
Mitchell Stern, Noam Shazeer, Jakob Uszkoreit
NeurIPS2
2017 Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, Jeffrey Dean
ICLR (Poster)1
2017 Attention is All you Need
abstract
The dominant sequence transduction models are based on complex recurrent orconvolutional neural networks in an encoder and decoder configuration. The best performing such models also connect the encoder and decoder through an attentionm echanisms. We propose a novel, simple network architecture based solely onan attention mechanism, dispensing with recurrence and convolutions entirely.Experiments on two machine translation tasks show these models to be superiorin quality while being more parallelizable and requiring significantly less timeto train. Our single model with 165 million parameters, achieves 27.5 BLEU onEnglish-to-German translation, improving over the existing best ensemble result by over 1 BLEU. On English-to-French translation, we outperform the previoussingle state-of-the-art with model by 0.7 BLEU, achieving a BLEU score of 41.1.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin
NIPS2
2016 End-to-end text-dependent speaker verification
abstract
In this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the system's components using the same evaluation protocol and metric as at test time. Such an approach will result in simple and efficient systems, requiring little domain-specific knowledge and making few model assumptions. We implement the idea by formulating the problem as a single neural network architecture, including the estimation of a speaker model on only a few utterances, and evaluate it on our internal "Ok Google" benchmark for text-dependent speaker verification. The proposed approach appears to be very effective for big data applications Like ours that require highly accurate, easy-to-maintain systems with a small footprint.
Georg Heigold, Ignacio Moreno, Samy Bengio, Noam Shazeer
ICASSP4
2016 NN-Grams: Unifying Neural Network and n-Gram Language Models for Speech Recognition
abstract
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural networks. We report experiments where the model is trained on 26B words. NN-grams are efficient at run-time since they do not include an output soft-max layer. The model is trained using noise contrastive estimation (NCE), an approach that transforms the estimation problem of neural networks into one of binary classification between data samples and noise samples. We present results with noise samples derived from either an n-gram distribution or from speech recognition lattices. NN-grams outperforms an n-gram model on an Italian speech recognition dictation task.
Babak Damavandi, Shankar Kumar, Noam Shazeer, Antoine Bruguier
INTERSPEECH3
2016 Sparse Non-negative Matrix Language Modeling
abstract
We present Sparse Non-negative Matrix (SNM) estimation, a novel probability estimation technique for language modeling that can efficiently incorporate arbitrary features. We evaluate SNM language models on two corpora: the One Billion Word Benchmark and a subset of the LDC English Gigaword corpus. Results show that SNM language models trained with n-gram features are a close match for the well-established Kneser-Ney models. The addition of skip-gram features yields a model that is in the same league as the state-of-the-art recurrent neural network language models, as well as complementary: combining the two modeling techniques yields the best known result on the One Billion Word Benchmark. On the Gigaword corpus further improvements are observed using features that cross sentence boundaries. The computational advantages of SNM estimation over both maximum entropy and neural network estimation are probably its main strength, promising an approach that has large flexibility in combining arbitrary features and yet scales gracefully to large amounts of data.
Joris Pelemans, Noam Shazeer, Ciprian Chelba
Trans. Assoc. Comput. Linguistics2
2015 Sparse non-negative matrix language modeling for geo-annotated query session data
abstract
The paper investigates the impact on query language modeling when using skip-grams within query as well as across queries in a given search session, in conjunction with the geo-annotation available for the query stream data. As modeling tool we use the recently proposed sparse non-negative matrix estimation technique, since it offers the same expressive power as the well-established maximum entropy approach in combining arbitrary context features. Experiments on the google.com query stream show that using session-level and geo-location context we can expect reductions in perplexity of 34% relative over the Kneser-Ney N-gram baseline; when evaluating on the '"local" subset of the query stream, the relative reduction in PPL is 51% - more than a bit. Both sources of context information (geo-location, and previous queries in session) are about equally valuable in building a language model for the query stream.
Ciprian Chelba, Noam Shazeer
ASRU2
2015 Pruning sparse non-negative matrix n-gram language models
abstract
Copyright © 2015 ISCA. In this paper we present a pruning algorithm and experimental results for our recently proposed Sparse Non-negative Matrix (SNM) family of language models (LMs). We show that when trained with only n-gram features SNMLM pruning based on a mutual information criterion yields the best known pruned model on the One Billion Word Language Model Benchmark, reducing perplexity with 18% and 57% over Katz and Kneser- Ney LMs, respectively. We also present a method for converting an SNMLM to ARPA back-off format which can be readily used in a single-pass decoder for Automatic Speech Recognition.
Joris Pelemans, Noam Shazeer, Ciprian Chelba
INTERSPEECH2
2015 Sparse non-negative matrix language modeling for skip-grams
abstract
Copyright © 2015 ISCA. We present a novel family of language model (LM) estimation techniques named Sparse Non-negative Matrix (SNM) estimation. A first set of experiments empirically evaluating these techniques on the One BillionWord Benchmark [3] shows that with skip-gram features SNMLMs are able to match the state-of-theart recurrent neural network (RNN) LMs; combining the two modeling techniques yields the best known result on the benchmark. The computational advantages of SNM over both maximum entropy and RNNLM estimation are probably its main strength, promising an approach that has the same flexibility in combining arbitrary features effectively and yet should scale to very large amounts of data as gracefully as n-gram LMs do.
Noam Shazeer, Joris Pelemans, Ciprian Chelba
INTERSPEECH1
2015 Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
abstract
Recurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning. The current approach to training them consists of maximizing the likelihood of each token in the sequence given the current (recurrent) state and the previous token. At inference, the unknown previous token is then replaced by a token generated by the model itself. This discrepancy between training and inference can yield errors that can accumulate quickly along the generated sequence. We propose a curriculum learning strategy to gently change the training process from a fully guided scheme using the true previous token, towards a less guided scheme which mostly uses the generated token instead. Experiments on several sequence prediction tasks show that this approach yields significant improvements. Moreover, it was used successfully in our winning bid to the MSCOCO image captioning challenge, 2015.
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer
NIPS4
2002 A probabilistic approach to solving crossword puzzles
Michael L. Littman, Greg A. Keim, Noam Shazeer
Artif. Intell.3