Jeffrey Dean

dblp:d/JeffreyDean · also Jeff Dean · DBLP profile ↗
← Back
51ranked-venue papers
16as first author
8since 2021 · last 2025
0000-0002-2769-2049ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 13 · 5 first-authorSystems, architecture and hardware · 10 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
20 papers
Efficient and distributed learning · 32% Deep learning architectures and training · 27% Language models and text generation · 21%
Computer architecture, parallel and distributed computing, and storage systems
17 papers
Distributed systems · 51% Hardware accelerators and domain-specific architectures · 20% Performance modeling and evaluation · 14%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%
Databases, data mining, and information retrieval
5 papers
Machine learning and data management · 46% Indexing and storage engines · 34% Information retrieval · 8%

Topics — the 30 heaviest of 88, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.982023
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
A Hierarchical Model for Device Placement · ICLR (Poster) 2018
Dynamic control flow in large-scale machine learning · EuroSys 2018
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.422024
Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
1.422024
Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
mixture of experts
0.922023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR (Poster) 2017
Natural language and speech › Language models and text generation
large language model
0.932023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Large Language Models in Machine Translation · EMNLP-CoNLL 2007
Machine learning › Efficient and distributed learning
model compression
0.912025
Matryoshka Quantization · ICML 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
Matryoshka Quantization · ICML 2025
Algorithmic game theory and mechanism design › auction theory
budget-constrained auction
0.912025
Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025
Algorithmic game theory and mechanism design
market design
0.912025
Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025
Algorithmic game theory and mechanism design
price of anarchy
0.912025
Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025
Algorithmic game theory and mechanism design › auction theory › auction mechanism
procurement auction
0.912025
Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.822023
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
instruction tuning
0.812024
Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Natural language and speech › Language models and text generation
instruction following
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Deep learning architectures and training
scaling laws
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Optimization for machine learning
sparse model
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
Brainformers: Trading Simplicity for Efficiency · ICML 2023
Distributed systems › distributed machine learning
device placement
0.622018
A Hierarchical Model for Device Placement · ICLR (Poster) 2018
Device Placement Optimization with Reinforcement Learning · ICML 2017
Machine learning › Deep learning architectures and training
backpropagation
0.612022
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
Machine learning › Deep learning architectures and training › neural network training
local learning
0.612022
Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022
Distributed systems
distributed machine learning
0.522018
Dynamic control flow in large-scale machine learning · EuroSys 2018
Large Scale Distributed Deep Networks · NIPS 2012
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement
0.312018
A Hierarchical Model for Device Placement · ICLR (Poster) 2018
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.312018
Efficient Neural Architecture Search via Parameter Sharing · ICML 2018
Machine learning and data management
learned database components
0.312018
The Case for Learned Index Structures · SIGMOD Conference 2018
Indexing and storage engines
learned index
0.312018
The Case for Learned Index Structures · SIGMOD Conference 2018
Performance modeling and evaluation › system modeling
hierarchical modeling
0.312018
A Hierarchical Model for Device Placement · ICLR (Poster) 2018
Distributed systems
distributed database
0.322013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Spanner: Google's Globally-Distributed Database · OSDI 2012
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation
0.312017
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR (Poster) 2017
Cloud and datacenter computing
cluster resource management and scheduling
0.312017
Device Placement Optimization with Reinforcement Learning · ICML 2017

Methods — techniques the papers use, named apart from their topics

no-regret learning · 1.7knapsack · 1.7bang-per-buck algorithm · 1.7quantization-aware training · 0.9knowledge distillation · 0.9transformer · 0.7pathways · 0.7neural architecture search · 0.7layer normalization · 0.7deep learning · 0.6data flow graphs · 0.6data flow graph · 0.6neural network placement · 0.3hierarchical reinforcement learning · 0.3sequence-to-sequence model · 0.3reinforcement learning · 0.3distributed systems infrastructure · 0.2truetime · 0.2
YearPublicationVenuePosition
2025 Matryoshka Quantization
abstract
Quantizing model weights is critical for reducing the communication and inference costs of large models. However, quantizing models – especially to low precisions like int4 or int2 – requires a trade-off in model quality; int2, in particular, is known to severely degrade model quality. Consequently, practitioners are often forced to maintain multiple models with different quantization levels or serve a single model that best satisfies the quality-latency trade-off. On the other hand, integer data types, such as int8, inherently possess a nested (Matryoshka) structure where smaller bit-width integers, like int4 or int2, are nested within the most significant bits. Leveraging this insight, in this paper, we propose Matryoshka Quantization (MatQuant), a novel multi-scale quantization technique that alleviates the aforementioned challenge. This technique allows us to train and maintain a single quantized model but serve it with the precision demanded by the deployment. Furthermore, leveraging MatQuant’s co-training and co-distillation, int2 precision models extracted by MatQuant outperform standard int2 quantization by up to 4% and 7% with OmniQuant and QAT as base algorithms respectively. Finally, we demonstrate that by using an extra bit to represent outliers, a model with an effective precision of 2.05-bit improves further by 6% with OmniQuant as the base algorithm.
Pranav Ajit Nair, Puranjay Datta, Jeffrey Dean, Prateek Jain 0002, Aditya Kusupati
ICML3
2025 DataRater: Meta-Learned Dataset Curation
abstract
The quality of foundation models depends heavily on their training data. Consequently, great efforts have been put into dataset curation. Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics. An approach that is ultimately more scalable (let alone more satisfying) is to \emph{learn} which data is actually valuable for training. This type of meta-learning could allow more sophisticated, fine-grained, and effective curation. Our proposed \emph{DataRater} is an instance of this idea. It estimates the value of training on any particular data point. This is done by meta-learning using `meta-gradients', with the objective of improving training efficiency on held out data. In extensive experiments across a range of model scales and datasets, we find that using our DataRater to filter data is highly effective, resulting in significantly improved compute efficiency.
Dan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf, Matteo Hessel, Jeremy Shar, Junhyuk Oh, András György 0001, Tom Schaul, Jeffrey Dean, Hado van Hasselt, David Silver 0001
NeurIPS10
2025 Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces
abstract
We envision a marketplace where diverse entities offer specialized "modules" through APIs, allowing users to compose the outputs of these modules for complex tasks within a given budget. This paper studies the market design problem in such an ecosystem, where module owners strategically set prices for their APIs (to maximize their profit) and a central platform orchestrates the aggregation of module outputs at query-time. One can also think about this as a first-price procurement auction with budgets. The first observation is that if the platform's algorithm is to find the optimal set of modules then this could result in a poor outcome, in the sense that there are price equilibria which provide arbitrarily low value for the user. We show that under a suitable version of the "bang-per-buck" algorithm for the knapsack problem, an ε-approximate equilibrium always exists, for any arbitrary ε > 0. Further, our first main result shows that with this algorithm any such equilibrium provides a constant approximation to the optimal value that the buyer could get under various constraints including (i) a budget constraint and (ii) a budget and a matroid constraint. This can also be interpreted as a price of anarchy result for budget-constrained first-price procurement auctions. Finally, we demonstrate that these approximately efficient equilibria can be learned through decentralized price adjustments by module owners using no-regret learning algorithms.
Kshipra Bhawalkar, Jeffrey Dean, Christopher Liaw, Aranyak Mehta
EC2
2024 Scaling Instruction-Finetuned Language Models
abstract
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation, RealToxicityPrompts). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PaLM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks (at time of release), such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints,1 which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models.
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang 0002, Mostafa Dehghani 0001, Siddhartha Brahma, Albert Webson, Shixiang Gu, Zhuyun Dai, Mirac Suzgun, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu 0001, Slav Petrov, Ed H. Chi, Jeffrey Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei
J. Mach. Learn. Res.30
2023 Exciting Directions for ML Models and the Implications for Computing Hardware
abstract
In recent years, ML has completely changed our expectations of what is possible with computers
Jeffrey Dean, Amin Vahdat
HCS1
2023 Brainformers: Trading Simplicity for Efficiency
abstract
Transformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations.
Yanqi Zhou, Nan Du 0002, Yanping Huang, Daiyi Peng, Chang Lan, Siamak Shakeri, David R. So, Andrew M. Dai, Yifeng Lu, Quoc V. Le, Claire Cui, James Laudon, Jeffrey Dean
ICML15
2023 PaLM: Scaling Language Modeling with Pathways
abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel
J. Mach. Learn. Res.65
2022 Interlocking Backpropagation: Improving depthwise model-parallelism
abstract
The number of parameters in state of the art neural networks has drastically increased in recent years. This surge of interest in large scale neural networks has motivated the development of new distributed training strategies enabling such models. One such strategy is model-parallel distributed training. Unfortunately, model-parallelism can suffer from poor resource utilisation, which leads to wasted resources. In this work, we improve upon recent developments in an idealised model-parallel optimisation setting: local learning. Motivated by poor resource utilisation in the global setting and poor task performance in the local setting, we introduce a class of intermediary strategies between local and global learning referred to as interlocking backpropagation. These strategies preserve many of the compute-efficiency advantages of local optimisation, while recovering much of the task performance achieved by global optimisation. We assess our strategies on both image classification ResNets and Transformer language models, finding that our strategy consistently out-performs local learning in terms of task performance, and out-performs global learning in training efficiency.
Aidan N. Gomez, Oscar Key, Kuba Perlin, Stephen Gou, Nicholas Frosst, Jeffrey Dean, Yarin Gal
J. Mach. Learn. Res.6
2019 Deep Learning for Solving Important Problems
abstract
In this keynote we describe progress in work that our research teams have been doing over the past years, including advances in difficult problems in artificial intelligence, on building large-scale computer systems for machine learning research, and, in collaboration with many teams at Google, on applying our research and systems to dozens of Google products. Our group has open-sourced the TensorFlow system [2], a widely popular system designed to easily express machine learning ideas, and to quickly train, evaluate and deploy machine learning systems. We then highlight some of our research accomplishments, and relate them to the National Academy of Engineering's Grand Engineering Challenges for the 21st Century.
Jeffrey Dean
WWW1
2018 Dynamic control flow in large-scale machine learning
abstract
Many recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments.
Martín Abadi, Paul Barham 0001, Eugene Brevdo, Michael Burrows, Andy Davis, Jeffrey Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Gordon Murray, Xiaoqiang Zheng
EuroSys7
2018 A Hierarchical Model for Device Placement
Azalia Mirhoseini, Anna Goldie, Hieu Pham 0001, Benoit Steiner, Quoc V. Le, Jeffrey Dean
ICLR (Poster)6
2018 Efficient Neural Architecture Search via Parameter Sharing
Hieu Pham 0001, Melody Y. Guan, Barret Zoph, Quoc V. Le, Jeffrey Dean
ICML5
2018 The Case for Learned Index Structures
abstract
Indexes are models: a \btree-Index can be seen as a model to map a key to the position of a record within a sorted array, a Hash-Index as a model to map a key to a position of a record within an unsorted array, and a BitMap-Index as a model to indicate if a data record exists or not. In this exploratory research paper, we start from this premise and posit that all existing index structures can be replaced with other types of models, including deep-learning models, which we term \em learned indexes. We theoretically analyze under which conditions learned indexes outperform traditional index structures and describe the main challenges in designing learned index structures. Our initial results show that our learned indexes can have significant advantages over traditional indexes. More importantly, we believe that the idea of replacing core components of a data management system through learned models has far reaching implications for future systems designs and that this work provides just a glimpse of what might be possible.
Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, Neoklis Polyzotis
SIGMOD Conference4
2017 Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, Jeffrey Dean
ICLR (Poster)7
2017 Device Placement Optimization with Reinforcement Learning
abstract
The past few years have witnessed a growth in size and computational requirements for training and inference with neural networks. Currently, a common approach to address these requirements is to use a heterogeneous distributed environment with a mixture of hardware devices such as CPUs and GPUs. Importantly, the decision of placing parts of the neural models on devices is often made by human experts based on simple heuristics and intuitions. In this paper, we propose a method which learns to optimize device placement for TensorFlow computational graphs. Key to our method is the use of a sequence-to-sequence model to predict which subsets of operations in a TensorFlow graph should run on which of the available devices. The execution time of the predicted placements is then used as the reward signal to optimize the parameters of the sequence-to-sequence model. Our main result is that on Inception-V3 for ImageNet classification, and on RNN LSTM, for language modeling and neural machine translation, our model finds non-trivial device placements that outperform hand-crafted heuristics and traditional algo-rithmic methods.
Azalia Mirhoseini, Hieu Pham 0001, Quoc V. Le, Benoit Steiner, Rasmus Larsen 0002, Yuefeng Zhou, Mohammad Norouzi 0002, Samy Bengio, Jeffrey Dean
ICML10
2017 In-Datacenter Performance Analysis of a Tensor Processing Unit
abstract
Many architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU) --- deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU's deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters' NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X -- 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X -- 80X higher. Moreover, using the CPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU.
Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson 0001, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, Richard Ho 0001, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, Doe Hyun Yoon
ISCA18
2017 Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
abstract
We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no changes to the model architecture from a standard NMT system but instead introduces an artificial token at the beginning of the input sentence to specify the required target language. Using a shared wordpiece vocabulary, our approach enables Multilingual NMT systems using a single model. On the WMT’14 benchmarks, a single multilingual model achieves comparable performance for English→French and surpasses state-of-theart results for English→German. Similarly, a single multilingual model surpasses state-of-the-art results for French→English and German→English on WMT’14 and WMT’15 benchmarks, respectively. On production corpora, multilingual models of up to twelve language pairs allow for better translation of many individual pairs. Our models can also learn to perform implicit bridging between language pairs never seen explicitly during training, showing that transfer learning and zero-shot translation is possible for neural translation. Finally, we show analyses that hints at a universal interlingua representation in our models and also show some interesting examples when mixing languages.
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Nikhil Thorat, Fernanda B. Viégas, Martin Wattenberg, Gregory S. Corrado, Macduff Hughes, Jeffrey Dean
Trans. Assoc. Comput. Linguistics12
2016 TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng
OSDI6
2016 Building Machine Learning Systems that Understand
abstract
Over the past five years, deep learning and large-scale neural networks have made significant advances in speech recognition, computer vision, language understanding and translation, robotics, and many other fields. Deep learning allows the use of very raw forms of data in order to build higher-level understanding of data automatically, and can also be used to learn to accomplish complex tasks. In the next decade, it is likely that a fruitful direction for research in data management will be in how to seamlessly integrate these kinds of machine learning models into systems that store and manage data. In this talk, I will highlight some of the advances that have been made in deep learning and suggest some interesting directions for future research.
Jeffrey Dean
SIGMOD Conference1
2016 Large-Scale Deep Learning For Building Intelligent Computer Systems
abstract
For the past five years, the Google Brain team has focused on conducting research in difficult problems in artificial intelligence, on building large-scale computer systems for machine learning research, and, in collaboration with many teams at Google, on applying our research and systems to dozens of Google products. Our group has recently open-sourced the TensorFlow system (tensorflow.org), a system designed to easily express machine ideas, and to quickly train, evaluate and deploy machine learning systems. In this talk, I'll highlight some of the design decisions we made in building TensorFlow, discuss research results produced within our group, and describe ways in which these ideas have been applied to a variety of problems in Google's products, usually in close collaboration with other teams.
Jeffrey Dean
WSDM1
2013 Multilingual acoustic models using distributed deep neural networks
abstract
Today's speech recognition technology is mature enough to be useful for many practical applications. In this context, it is of paramount importance to train accurate acoustic models for many languages within given resource constraints such as data, processing power, and time. Multilingual training has the potential to solve the data issue and close the performance gap between resource-rich and resource-scarce languages. Neural networks lend themselves naturally to parameter sharing across languages, and distributed implementations have made it feasible to train large networks. In this paper, we present experimental results for cross- and multi-lingual network training of eleven Romance languages on 10k hours of data in total. The average relative gains over the monolingual baselines are 4%/2% (data-scarce/data-rich languages) for cross- and 7%/2% for multi-lingual training. However, the additional gain from jointly training the languages on all data comes at an increased training time of roughly four weeks, compared to two weeks (monolingual) and one week (crosslingual).
Georg Heigold, Vincent Vanhoucke, Andrew W. Senior, Patrick Nguyen, Marc'Aurelio Ranzato, Matthieu Devin, Jeffrey Dean
ICASSP7
2013 On rectified linear units for speech processing
abstract
Deep neural networks have recently become the gold standard for acoustic modeling in speech recognition systems. The key computational unit of a deep network is a linear projection followed by a point-wise non-linearity, which is typically a logistic function. In this work, we show that we can improve generalization and make training of deep networks faster and simpler by substituting the logistic units with rectified linear units. These units are linear when their input is positive and zero otherwise. In a supervised setting, we can successfully train very deep nets from random initialization on a large vocabulary speech recognition task achieving lower word error rates than using a logistic network with the same topology. Similarly in an unsupervised setting, we show how we can learn sparse features that can be useful for discriminative tasks. All our experiments are executed in a distributed environment using several hundred machines and several hundred hours of speech data.
Matthew D. Zeiler, Marc'Aurelio Ranzato, Rajat Monga, Mark Z. Mao, Quoc V. Le, Patrick Nguyen, Andrew W. Senior, Vincent Vanhoucke, Jeffrey Dean, Geoffrey E. Hinton
ICASSP10
2013 DeViSE: A Deep Visual-Semantic Embedding Model
abstract
Modern visual recognition systems are often limited in their ability to scale to large numbers of object categories. This limitation is in part due to the increasing difficulty of acquiring sufficient training data in the form of labeled images as the number of object categories grows. One remedy is to leverage data from other sources -- such as text data -- both to train visual models and to constrain their predictions. In this paper we present a new deep visual-semantic embedding model trained to identify visual objects using both labeled image data as well as semantic information gleaned from unannotated text. We demonstrate that this model matches state-of-the-art performance on the 1000-class ImageNet object recognition challenge while making more semantically reasonable errors, and also show that the semantic information can be exploited to make predictions about tens of thousands of image labels not observed during training. Semantic knowledge improves such zero-shot predictions by up to 65%, achieving hit rates of up to 10% across thousands of novel labels never seen by the visual model.
Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc'Aurelio Ranzato, Tomás Mikolov
NIPS5
2013 Distributed Representations of Words and Phrases and their Compositionality
abstract
The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several improvements that make the Skip-gram model more expressive and enable it to learn higher quality vectors more rapidly. We show that by subsampling frequent words we obtain significant speedup, and also learn higher quality representations as measured by our tasks. We also introduce Negative Sampling, a simplified variant of Noise Contrastive Estimation (NCE) that learns more accurate vectors for frequent words compared to the hierarchical softmax. An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases. For example, the meanings of Canada'' and "Air'' cannot be easily combined to obtain "Air Canada''. Motivated by this example, we present a simple and efficient method for finding phrases, and show that their vector representations can be accurately learned by the Skip-gram model. "
Tomás Mikolov, Ilya Sutskever, Kai Chen 0010, Gregory S. Corrado, Jeffrey Dean
NIPS5
2013 Spanner: Google's Globally Distributed Database
abstract
Spanner is Google’s scalable, multiversion, globally distributed, and synchronously replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. This article describes how Spanner is structured, its feature set, the rationale underlying various design decisions, and a novel time API that exposes clock uncertainty. This API and its implementation are critical to supporting external consistency and a variety of powerful features: nonblocking reads in the past, lock-free snapshot transactions, and atomic schema changes, across all of Spanner.
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford
ACM Trans. Comput. Syst.2
2012 Building high-level features using large scale unsupervised learning
Quoc V. Le, Marc'Aurelio Ranzato, Rajat Monga, Matthieu Devin, Gregory S. Corrado, Kai Chen 0010, Jeffrey Dean, Andrew Y. Ng
ICML7
2012 Large Scale Distributed Deep Networks
abstract
Recent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm.
Jeffrey Dean, Gregory S. Corrado, Rajat Monga, Kai Chen 0010, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Andrew Y. Ng
NIPS1
2012 Spanner: Google's Globally-Distributed Database
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford
OSDI2
2010 Evolution and future directions of large-scale storage and computation systems at Google
abstract
No abstract available.
Jeffrey Dean
SoCC1
2009 Back-off language model compression
Boulos Harb, Ciprian Chelba, Jeffrey Dean, Sanjay Ghemawat
INTERSPEECH3
2009 Challenges in building large-scale information retrieval systems: invited talk
abstract
Building and operating large-scale information retrieval systems used by hundreds of millions of people around the world provides a number of interesting challenges. Designing such systems requires making complex design tradeoffs in a number of dimensions, including (a) the number of user queries that must be handled per second and the response latency to these requests, (b) the number and size of various corpora that are searched, (c) the latency and frequency with which documents are updated or added to the corpora, and (d) the quality and cost of the ranking algorithms that are used for retrieval. In this talk I will discuss the evolution of Google's hardware infrastructure and information retrieval systems and some of the design challenges that arise from ever-increasing demands in all of these dimensions. I will also describe how we use various pieces of distributed systems infrastructure when building these retrieval systems. Finally, I will describe some future challenges and open research problems in this area.
Jeffrey Dean
WSDM1
2008 Bigtable: A Distributed Storage System for Structured Data
abstract
Bigtable is a distributed storage system for managing structured data that is designed to scale to a very large size: petabytes of data across thousands of commodity servers. Many projects at Google store data in Bigtable, including web indexing, Google Earth, and Google Finance. These applications place very different demands on Bigtable, both in terms of data size (from URLs to web pages to satellite imagery) and latency requirements (from backend bulk processing to real-time data serving). Despite these varied demands, Bigtable has successfully provided a flexible, high-performance solution for all of these Google products. In this article, we describe the simple data model provided by Bigtable, which gives clients dynamic control over data layout and format, and we describe the design and implementation of Bigtable.
Fay Chang, Jeffrey Dean, Sanjay Ghemawat, Wilson C. Hsieh, Deborah A. Wallach, Michael Burrows, Tushar Chandra, Andrew Fikes, Robert Gruber
ACM Trans. Comput. Syst.2
2007 Large Language Models in Machine Translation
Thorsten Brants, Ashok C. Popat, Franz Josef Och, Jeffrey Dean
EMNLP-CoNLL5
2007 MapReduce and Other Building Blocks for Large-Scale Distributed Systems at Google
Jeffrey Dean
USENIX ATC1
2006 Experiences with MapReduce, an abstraction for large-scale computation
abstract
MapReduce is a programming model and an associated implementation for processing and generating large data sets. Users specify a Map function that processes a key/value pair to generate a set of intermediate key/value pairs, and a Reduce function that merges all intermediate values associated with the same intermediate key. Many real world tasks are expressible in this model. Programs written in this functional style are automatically parallelized and executed on a large cluster of commodity machines.The MapReduce run-time system takes care of the details of partitioning the input data, scheduling the program's execution across a set of machines, handling machine failures, and managing the required intermachine communication. This allows programmers without any experience with parallel and distributed systems to easily utilize the resources of a large distributed system.Our implementation of MapReduce runs on a large cluster of commodity machines and is highly scalable: a typical MapReduce computation processes many terabytes of data on thousands of machines. Programmers find the system easy to use: thousands of MapReduce programs have been implemented and several thousand thousand MapReduce jobs are executed on Google's clusters every day.In this talk I'll describe the basic programming model, discuss our experience using it in a variety of domains, and talk about the implications of programming models like MapReduce as one paradigm to simplify development of parallel software for multi-core microprocessors.
Jeffrey Dean
PACT1
2006 Bigtable: A Distributed Storage System for Structured Data (Awarded Best Paper!)
Fay Chang, Jeffrey Dean, Sanjay Ghemawat, Wilson C. Hsieh, Deborah A. Wallach, Michael Burrows, Tushar Chandra, Andrew Fikes, Robert Gruber
OSDI2
2004 MapReduce: Simplified Data Processing on Large Clusters
Jeffrey Dean, Sanjay Ghemawat
OSDI1
2000 A comparison of techniques to find mirrored hosts on the WWW
abstract
We compare several algorithms for identifying mirrored hosts on the World Wide Web. The algorithms operate on the basis of URL strings and linkage data: the type of information about Web pages easily available from Web proxies and crawlers. Identification of mirrored hosts can improve Web-based information retrieval in several ways: first, by identifying mirrored hosts, search engines can avoid storing and returning duplicate documents. Second, several new information retrieval techniques for the Web make inferences based on the explicit links among hypertext documents—mirroring perturbs their graph model and degrades performance. Third, mirroring information can be used to redirect users to alternate mirror sites to compensate for various failures, and can thus improve the performance of Web browsers and proxies. We evaluated four classes of “top-down” algorithms for detecting mirrored host pairs (that is, algorithms that are based on page attributes such as URL, IP address, and hyperlinks between pages, and not on the page content) on a collection of 140 million URLs (on 230,000 hosts) and their associated connectivity information. Our best approach is one which combines five algorithms and achieved a precision of 0.57 for a recall of 0.86 considering 100,000 ranked host pairs.
Krishna Bharat, Andrei Z. Broder, Jeffrey Dean, Monika Henzinger
J. Am. Soc. Inf. Sci.3
1999 Finding Related Pages in the World Wide Web
Jeffrey Dean, Monika Henzinger
Comput. Networks1
1998 Walknet--a biologically inspired network to control six-legged walking
Holk Cruse, Thomas Kindermann, Michael Schumm, Jeffrey Dean, Josef Schmitz
Neural Networks4
1997 ProfileMe: Hardware Support for Instruction-Level Profiling on Out-of-Order Processors
abstract
Profile data is valuable for identifying performance bottlenecks and guiding optimizations. Periodic sampling of a processor's performance monitoring hardware is an effective, unobtrusive way to obtain detailed profiles. Unfortunately, existing hardware simply counts events, such as cache misses and branch mispredictions, and cannot accurately attribute these events to instructions, especially on out-of-order machines. We propose an alternative approach, called ProfileMe, that samples instructions. As a sampled instruction moves through the processor pipeline, a detailed record of all interesting events and pipeline stage latencies is collected. ProfileMe also supports paired sampling, which captures information about the interactions between concurrent instructions, revealing information about useful concurrency and the utilization of various pipeline stages while an instruction is in flight. We describe an inexpensive hardware implementation of ProfileMe, outline a variety of software techniques to extract useful profile information from the hardware, and explain several ways in which this information can provide valuable feedback for programmers and optimizers.
Jeffrey Dean, James E. Hicks, Carl A. Waldspurger, William E. Weihl, George Z. Chrysos
MICRO1
1997 Call Graph Construction in Object-Oriented Languages
abstract
Interprocedural analyses enable optimizing compilers to more precisely model the effects of non-inlined procedure calls, potentially resulting in substantial increases in application performance. Applying interprocedural analysis to programs written in object-oriented or functional languages is complicated by the difficulty of constructing an accurate program call graph. This paper presents a parameterized algorithmic framework for call graph construction in the presence of message sends and/or first class functions. We use this framework to describe and to implement a number of well-known and new algorithms. We then empirically assess these algorithms by applying them to a suite of medium-sized programs written in Cecil and Java, reporting on the relative cost of the analyses, the relative precision of the constructed call graphs, and the impact of this precision on the effectiveness of a number of interprocedural optimizations.
David Grove, Greg DeFouw, Jeffrey Dean, Craig Chambers
OOPSLA3
1997 Continuous Profiling: Where Have All the Cycles Gone?
abstract
Article Continuous profiling: where have all the cycles gone? Share on Authors: Jennifer M. Anderson View Profile , Lance M. Berc View Profile , Jeffrey Dean View Profile , Sanjay Ghemawat View Profile , Monika R. Henzinger View Profile , Shun-Tak A. Leung View Profile , Richard L. Sites View Profile , Mark T. Vandevoorde View Profile , Carl A. Waldspurger View Profile , William E. Weihl View Profile Authors Info & Claims SOSP '97: Proceedings of the sixteenth ACM symposium on Operating systems principlesOctober 1997 Pages 1–14https://doi.org/10.1145/268998.266637Published:01 October 1997 175citation1,209DownloadsMetricsTotal Citations175Total Downloads1,209Last 12 Months17Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jennifer-Ann M. Anderson, Lance M. Berc, Jeffrey Dean, Sanjay Ghemawat, Monika Henzinger, Shun-Tak Leung, Richard L. Sites, Mark T. Vandevoorde, Carl A. Waldspurger, William E. Weihl
SOSP3
1997 Continuous Profiling: Where Have All the Cycles Gone?
abstract
This article describes the Digital Continuous Profiling Infrastructure, a sampling-based profiling system designed to run continuously on production systems. The system supports multiprocessors, works on unmodified executables, and collects profiles for entire systems, including user programs, shared libraries, and the operating system kernel. Samples are collected at a high rate (over 5200 samples/sec. per 333MHz processor), yet with low overhead (1–3% slowdown for most workloads). Analysis tools supplied with the profiling system use the sample data to produce a precise and accurate accounting, down to the level of pipeline stalls incurred by individual instructions, of where time is bring spent. When instructions incur stalls, the tools identify possible reasons, such as cache misses, branch mispredictions, and functional unit contention. The fine-grained instruction-level analysis guides users and automated optimizers to the causes of performance problems and provides important insights for fixing them.
Jennifer-Ann M. Anderson, Lance M. Berc, Jeffrey Dean, Sanjay Ghemawat, Monika Henzinger, Shun-Tak Leung, Richard L. Sites, Mark T. Vandevoorde, Carl A. Waldspurger, William E. Weihl
ACM Trans. Comput. Syst.3
1996 Simplifying Neural Networks for Controlling Walking by Exploiting Physical Properties
Holk Cruse, Christian Bartling, Jeffrey Dean, Thomas Kindermann, Josef Schmitz, Michael Schumm, Hendrik Wagner
ICANN3
1996 Vortex: An Optimizing Compiler for Object-Oriented Languages
abstract
Previously, techniques such as class hierarchy analysis and profile-guided receiver class prediction have been demonstrated to greatly improve the performance of applications written in pure object-oriented languages, but the degree to which these results are transferable to applications written in hybrid languages has been unclear. In part to answer this question, we have developed the Vortex compiler infrastructure, a language-independent optimizing compiler for object-oriented languages, with front-ends for Cecil, C++, Java, and Modula-3. In this paper, we describe the Vortex compiler's intermediate language, internal structure, and optimization suite, and then we report the results of experiments assessing the effectiveness of different combinations of optimizations on sizable applications across these four languages. We characterize the benchmark programs in terms of a collection of static and dynamic metrics, intended to quantify aspects of the "object-orientedness" of a program.
Jeffrey Dean, Greg DeFouw, David Grove, Vassily Litvinov, Craig Chambers
OOPSLA1
1995 Optimization of Object-Oriented Programs Using Static Class Hierarchy Analysis
Jeffrey Dean, David Grove, Craig Chambers
ECOOP1
1995 A Framework for Selective Recompilation in the Presence of Complex Intermodule Dependencies
abstract
Compilersand other programming environment tools derive information from the source code of programs; derived information includes compiled code, interprocedurrd summary information, and call graph views.If the source program changes, the derived information needs to be updated.We present a simple framework for maintaining interrnodule dependencies, embodying different tradeoffs in terms of space usage, speed of processing, and selectivity of invalidation, that eases the implementation of incremental update of derived information.Our framework augments a directed acyclic graph representation of dependencies with factoring nodes (to save space) and$ltering nodes (to increase selectivity), and it includes an algorithm for efficient invalidation processing.We show how several schemes for selective recompilation, such as smart recompilation, filter sets for interprocedural summary information, and dependencies for whole-program optimization of object-oriented languages, map naturally onto our framework.For this latter application, by exploiting the facilities of our framework, we are able to reduce the number of lines of source code recompiled by a factor of seven over a header file-based scheme, and by a factor of two over the previous state-of-the-art selective dependency mechanism without consuming additional space.
Craig Chambers, Jeffrey Dean, David Grove
ICSE2
1995 Profile-Guided Receiver Class Prediction
abstract
The use of dynamically-dispatched procedure calls is a key mechanism for writing extensible and flexible code in object-oriented languages. Unfortunately, dynamic dispatching imposes a runtime performance penalty. Some recent implementations of pure object-oriented languages have utilized profile-guided receiver class prediction to reduce this performance penalty, and some researchers have argued for applying receiver class prediction in hybrid languages like C++. We performed a detailed examination of the dynamic profiles of eight large object-oriented applications written in C++ and Cecil, determining that the receiver class distributions are strongly peaked and stable across both inputs and program versions through time. We describe techniques for gathering and manipulating profile information at varying degrees of precision, particularly in the presence of optimizations such as inlining. Our implementation of profile-guided receiver class prediction improves the performance of large Cecil applications by more than a factor of two over solely static optimizations.
David Grove, Jeffrey Dean, Charles Garrett, Craig Chambers
OOPSLA2
1995 Selective Specialization for Object-Oriented Languages
abstract
Dynamic dispatching is a major source of run-time overhead in object-oriented languages, due both to the direct cost of method lookup and to the indirect effect of preventing other optimizations. To reduce this overhead, optimizing compilers for object-oriented languages analyze the classes of objects stored in program variables, with the goal of bounding the possible classes of message receivers enough so that the compiler can uniquely determine the target of a message send at compile time and replace the message send with a direct procedure call. Specialization is one important technique for improving the precision of this static class information: by compiling multiple versions of a method, each applicable to a subset of the possible argument classes of the method, more precise static information about the classes of the method's arguments is obtained. Previous specialization strategies have not been selective about where this technique is applied, and therefore tended to significantly increase compile time and code space usage, particularly for large applications. In this paper, we present a more general framework for specialization in object-oriented languages and describe a goal directed specialization algorithm that makes selective decisions to apply specialization to those cases where it provides the highest benefit. Our results show that our algorithm improves the performance of a group of sizeable programs by 65% to 275% while increasing compiled code space requirements by only 4% to 10%. Moreover, when compared to the previous state-of-the-art specialization scheme, our algorithm improves performance by 11% to 67% while simultaneously reducing code space requirements by 65% to 73%.
Jeffrey Dean, Craig Chambers, David Grove
PLDI1
1994 Identifying Profitable Specialization in Object-Oriented Languages
Jeffrey Dean, Craig Chambers, David Grove
PEPM1