EDBT 2026 Demo / reviewers in the wild / expert
Jeffrey Dean
dblp:d/JeffreyDean · also Jeff Dean
· DBLP profile ↗
51ranked-venue papers
16as first author
8since 2021 · last 2025
0000-0002-2769-2049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 13 · 5 first-authorSystems, architecture and hardware · 10 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
20 papers |
Efficient and distributed learning · 32% Deep learning architectures and training · 27% Language models and text generation · 21% | |
| Computer architecture, parallel and distributed computing, and storage systems
17 papers |
Distributed systems · 51% Hardware accelerators and domain-specific architectures · 20% Performance modeling and evaluation · 14% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% | |
| Databases, data mining, and information retrieval
5 papers |
Machine learning and data management · 46% Indexing and storage engines · 34% Information retrieval · 8% |
Topics — the 30 heaviest of 88, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.9 | 8 | 2023 | Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022 A Hierarchical Model for Device Placement · ICLR (Poster) 2018 Dynamic control flow in large-scale machine learning · EuroSys 2018 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.4 | 2 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
1.4 | 2 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 2 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR (Poster) 2017 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 3 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 Brainformers: Trading Simplicity for Efficiency · ICML 2023 Large Language Models in Machine Translation · EMNLP-CoNLL 2007 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Matryoshka Quantization · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Matryoshka Quantization · ICML 2025 |
Algorithmic game theory and mechanism design › auction theory
budget-constrained auction |
0.9 | 1 | 2025 | Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025 |
Algorithmic game theory and mechanism design
market design |
0.9 | 1 | 2025 | Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025 |
Algorithmic game theory and mechanism design
price of anarchy |
0.9 | 1 | 2025 | Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025 |
Algorithmic game theory and mechanism design › auction theory › auction mechanism
procurement auction |
0.9 | 1 | 2025 | Equilibrium Efficiency and Learning in Budgeted Procurement Marketplaces · EC 2025 |
Machine learning › Efficient and distributed learning › distributed training
model parallelism |
0.8 | 2 | 2023 | Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022 PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Natural language and speech › Language models and text generation
instruction tuning |
0.8 | 1 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.7 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Natural language and speech › Language models and text generation
instruction following |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Deep learning architectures and training
scaling laws |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Optimization for machine learning
sparse model |
0.7 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Machine learning › Deep learning architectures and training
transformer |
0.7 | 1 | 2023 | Brainformers: Trading Simplicity for Efficiency · ICML 2023 |
Distributed systems › distributed machine learning
device placement |
0.6 | 2 | 2018 | A Hierarchical Model for Device Placement · ICLR (Poster) 2018 Device Placement Optimization with Reinforcement Learning · ICML 2017 |
Machine learning › Deep learning architectures and training
backpropagation |
0.6 | 1 | 2022 | Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022 |
Machine learning › Deep learning architectures and training › neural network training
local learning |
0.6 | 1 | 2022 | Interlocking Backpropagation: Improving depthwise model-parallelism · J. Mach. Learn. Res. 2022 |
Distributed systems
distributed machine learning |
0.5 | 2 | 2018 | Dynamic control flow in large-scale machine learning · EuroSys 2018 Large Scale Distributed Deep Networks · NIPS 2012 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement |
0.3 | 1 | 2018 | A Hierarchical Model for Device Placement · ICLR (Poster) 2018 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.3 | 1 | 2018 | Efficient Neural Architecture Search via Parameter Sharing · ICML 2018 |
Machine learning and data management
learned database components |
0.3 | 1 | 2018 | The Case for Learned Index Structures · SIGMOD Conference 2018 |
Indexing and storage engines
learned index |
0.3 | 1 | 2018 | The Case for Learned Index Structures · SIGMOD Conference 2018 |
Performance modeling and evaluation › system modeling
hierarchical modeling |
0.3 | 1 | 2018 | A Hierarchical Model for Device Placement · ICLR (Poster) 2018 |
Distributed systems
distributed database |
0.3 | 2 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 Spanner: Google's Globally-Distributed Database · OSDI 2012 |
Machine learning › Efficient and distributed learning › adaptive computation
conditional computation |
0.3 | 1 | 2017 | Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer · ICLR (Poster) 2017 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.3 | 1 | 2017 | Device Placement Optimization with Reinforcement Learning · ICML 2017 |
Methods — techniques the papers use, named apart from their topics
no-regret learning · 1.7knapsack · 1.7bang-per-buck algorithm · 1.7quantization-aware training · 0.9knowledge distillation · 0.9transformer · 0.7pathways · 0.7neural architecture search · 0.7layer normalization · 0.7deep learning · 0.6data flow graphs · 0.6data flow graph · 0.6neural network placement · 0.3hierarchical reinforcement learning · 0.3sequence-to-sequence model · 0.3reinforcement learning · 0.3distributed systems infrastructure · 0.2truetime · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Matryoshka QuantizationabstractQuantizing model weights is critical for reducing
the communication and inference costs of large
models. However, quantizing models – especially
to low precisions like int4 or int2 – requires a
trade-off in model quality; int2, in particular, is
known to severely degrade model quality. Consequently, practitioners are often forced to maintain
multiple models with different quantization levels or serve a single model that best satisfies the
quality-latency trade-off. On the other hand, integer data types, such as int8, inherently possess
a nested (Matryoshka) structure where smaller
bit-width integers, like int4 or int2, are nested
within the most significant bits. Leveraging this
insight, in this paper, we propose Matryoshka
Quantization (MatQuant), a novel multi-scale
quantization technique that alleviates the aforementioned challenge. This technique allows us to
train and maintain a single quantized model but
serve it with the precision demanded by the deployment. Furthermore, leveraging MatQuant’s
co-training and co-distillation, int2 precision models extracted by MatQuant outperform standard
int2 quantization by up to 4% and 7% with OmniQuant and QAT as base algorithms respectively.
Finally, we demonstrate that by using an extra bit
to represent outliers, a model with an effective
precision of 2.05-bit improves further by 6% with
OmniQuant as the base algorithm. Pranav Ajit Nair, Puranjay Datta, Jeffrey Dean, Prateek Jain 0002, Aditya Kusupati |
ICML | 3 |
| 2025 | DataRater: Meta-Learned Dataset CurationabstractThe quality of foundation models depends heavily on their training data.
Consequently, great efforts have been put into dataset curation.
Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics.
An approach that is ultimately more scalable (let alone more satisfying) is to \emph{learn} which data is actually valuable for training.
This type of meta-learning could allow more sophisticated, fine-grained, and effective curation.
Our proposed \emph{DataRater} is an instance of this idea. It estimates the value of training on any particular data point. This is done by meta-learning using `meta-gradients', with the objective of improving training efficiency on held out data.
In extensive experiments across a range of model scales and datasets, we find that using our DataRater to filter data is highly effective, resulting in significantly improved compute efficiency. Dan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf, Matteo Hessel, Jeremy Shar, Junhyuk Oh, András György 0001, Tom Schaul, Jeffrey Dean, Hado van Hasselt, David Silver 0001 |
NeurIPS | 10 |
| 2025 | Equilibrium Efficiency and Learning in Budgeted Procurement MarketplacesabstractWe envision a marketplace where diverse entities offer specialized "modules" through APIs, allowing users to compose the outputs of these modules for complex tasks within a given budget. This paper studies the market design problem in such an ecosystem, where module owners strategically set prices for their APIs (to maximize their profit) and a central platform orchestrates the aggregation of module outputs at query-time. One can also think about this as a first-price procurement auction with budgets. The first observation is that if the platform's algorithm is to find the optimal set of modules then this could result in a poor outcome, in the sense that there are price equilibria which provide arbitrarily low value for the user. We show that under a suitable version of the "bang-per-buck" algorithm for the knapsack problem, an ε-approximate equilibrium always exists, for any arbitrary ε > 0. Further, our first main result shows that with this algorithm any such equilibrium provides a constant approximation to the optimal value that the buyer could get under various constraints including (i) a budget constraint and (ii) a budget and a matroid constraint. This can also be interpreted as a price of anarchy result for budget-constrained first-price procurement auctions. Finally, we demonstrate that these approximately efficient equilibria can be learned through decentralized price adjustments by module owners using no-regret learning algorithms. Kshipra Bhawalkar, Jeffrey Dean, Christopher Liaw, Aranyak Mehta |
EC | 2 |
| 2024 | Scaling Instruction-Finetuned Language ModelsabstractFinetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation, RealToxicityPrompts). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PaLM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks (at time of release), such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints,1 which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models. Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang 0002, Mostafa Dehghani 0001, Siddhartha Brahma, Albert Webson, Shixiang Gu, Zhuyun Dai, Mirac Suzgun, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu 0001, Slav Petrov, Ed H. Chi, Jeffrey Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei |
J. Mach. Learn. Res. | 30 |
| 2023 | Exciting Directions for ML Models and the Implications for Computing HardwareabstractIn recent years, ML has completely changed our expectations of what is possible with computers Jeffrey Dean, Amin Vahdat |
HCS | 1 |
| 2023 | Brainformers: Trading Simplicity for EfficiencyabstractTransformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations. Yanqi Zhou, Nan Du 0002, Yanping Huang, Daiyi Peng, Chang Lan, Siamak Shakeri, David R. So, Andrew M. Dai, Yifeng Lu, Quoc V. Le, Claire Cui, James Laudon, Jeffrey Dean |
ICML | 15 |
| 2023 | PaLM: Scaling Language Modeling with PathwaysabstractLarge language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel |
J. Mach. Learn. Res. | 65 |
| 2022 | Interlocking Backpropagation: Improving depthwise model-parallelismabstractThe number of parameters in state of the art neural networks has drastically increased in recent years. This surge of interest in large scale neural networks has motivated the development of new distributed training strategies enabling such models. One such strategy is model-parallel distributed training. Unfortunately, model-parallelism can suffer from poor resource utilisation, which leads to wasted resources. In this work, we improve upon recent developments in an idealised model-parallel optimisation setting: local learning. Motivated by poor resource utilisation in the global setting and poor task performance in the local setting, we introduce a class of intermediary strategies between local and global learning referred to as interlocking backpropagation. These strategies preserve many of the compute-efficiency advantages of local optimisation, while recovering much of the task performance achieved by global optimisation. We assess our strategies on both image classification ResNets and Transformer language models, finding that our strategy consistently out-performs local learning in terms of task performance, and out-performs global learning in training efficiency. Aidan N. Gomez, Oscar Key, Kuba Perlin, Stephen Gou, Nicholas Frosst, Jeffrey Dean, Yarin Gal |
J. Mach. Learn. Res. | 6 |
| 2019 | Deep Learning for Solving Important ProblemsabstractIn this keynote we describe progress in work that our research teams have been doing over the past years, including advances in difficult problems in artificial intelligence, on building large-scale computer systems for machine learning research, and, in collaboration with many teams at Google, on applying our research and systems to dozens of Google products. Our group has open-sourced the TensorFlow system [2], a widely popular system designed to easily express machine learning ideas, and to quickly train, evaluate and deploy machine learning systems. We then highlight some of our research accomplishments, and relate them to the National Academy of Engineering's Grand Engineering Challenges for the 21st Century. Jeffrey Dean |
WWW | 1 |
| 2018 | Dynamic control flow in large-scale machine learningabstractMany recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments. Martín Abadi, Paul Barham 0001, Eugene Brevdo, Michael Burrows, Andy Davis, Jeffrey Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Gordon Murray, Xiaoqiang Zheng |
EuroSys | 7 |
| 2018 | A Hierarchical Model for Device Placement
Azalia Mirhoseini, Anna Goldie, Hieu Pham 0001, Benoit Steiner, Quoc V. Le, Jeffrey Dean |
ICLR (Poster) | 6 |
| 2018 | Efficient Neural Architecture Search via Parameter Sharing
Hieu Pham 0001, Melody Y. Guan, Barret Zoph, Quoc V. Le, Jeffrey Dean |
ICML | 5 |
| 2018 | The Case for Learned Index StructuresabstractIndexes are models: a \btree-Index can be seen as a model to map a key to the position of a record within a sorted array, a Hash-Index as a model to map a key to a position of a record within an unsorted array, and a BitMap-Index as a model to indicate if a data record exists or not. In this exploratory research paper, we start from this premise and posit that all existing index structures can be replaced with other types of models, including deep-learning models, which we term \em learned indexes. We theoretically analyze under which conditions learned indexes outperform traditional index structures and describe the main challenges in designing learned index structures. Our initial results show that our learned indexes can have significant advantages over traditional indexes. More importantly, we believe that the idea of replacing core components of a data management system through learned models has far reaching implications for future systems designs and that this work provides just a glimpse of what might be possible. Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, Neoklis Polyzotis |
SIGMOD Conference | 4 |
| 2017 | Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, Jeffrey Dean |
ICLR (Poster) | 7 |
| 2017 | Device Placement Optimization with Reinforcement LearningabstractThe past few years have witnessed a growth in size and computational requirements for training and inference with neural networks. Currently, a common approach to address these requirements is to use a heterogeneous distributed environment with a mixture of hardware devices such as CPUs and GPUs. Importantly, the decision of placing parts of the neural models on devices is often made by human experts based on simple heuristics and intuitions. In this paper, we propose a method which learns to optimize device placement for TensorFlow computational graphs. Key to our method is the use of a sequence-to-sequence model to predict which subsets of operations in a TensorFlow graph should run on which of the available devices. The execution time of the predicted placements is then used as the reward signal to optimize the parameters of the sequence-to-sequence model. Our main result is that on Inception-V3 for ImageNet classification, and on RNN LSTM, for language modeling and neural machine translation, our model finds non-trivial device placements that outperform hand-crafted heuristics and traditional algo-rithmic methods. Azalia Mirhoseini, Hieu Pham 0001, Quoc V. Le, Benoit Steiner, Rasmus Larsen 0002, Yuefeng Zhou, Mohammad Norouzi 0002, Samy Bengio, Jeffrey Dean |
ICML | 10 |
| 2017 | In-Datacenter Performance Analysis of a Tensor Processing UnitabstractMany architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU) --- deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU's deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters' NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X -- 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X -- 80X higher. Moreover, using the CPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU. Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson 0001, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, Richard Ho 0001, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, Doe Hyun Yoon |
ISCA | 18 |
| 2017 | Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot TranslationabstractWe propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no changes to the model architecture from a standard NMT system but instead introduces an artificial token at the beginning of the input sentence to specify the required target language. Using a shared wordpiece vocabulary, our approach enables Multilingual NMT systems using a single model. On the WMT’14 benchmarks, a single multilingual model achieves comparable performance for English→French and surpasses state-of-theart results for English→German. Similarly, a single multilingual model surpasses state-of-the-art results for French→English and German→English on WMT’14 and WMT’15 benchmarks, respectively. On production corpora, multilingual models of up to twelve language pairs allow for better translation of many individual pairs. Our models can also learn to perform implicit bridging between language pairs never seen explicitly during training, showing that transfer learning and zero-shot translation is possible for neural translation. Finally, we show analyses that hints at a universal interlingua representation in our models and also show some interesting examples when mixing languages. Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Nikhil Thorat, Fernanda B. Viégas, Martin Wattenberg, Gregory S. Corrado, Macduff Hughes, Jeffrey Dean |
Trans. Assoc. Comput. Linguistics | 12 |
| 2016 | TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng |
OSDI | 6 |
| 2016 | Building Machine Learning Systems that UnderstandabstractOver the past five years, deep learning and large-scale neural networks have made significant advances in speech recognition, computer vision, language understanding and translation, robotics, and many other fields. Deep learning allows the use of very raw forms of data in order to build higher-level understanding of data automatically, and can also be used to learn to accomplish complex tasks. In the next decade, it is likely that a fruitful direction for research in data management will be in how to seamlessly integrate these kinds of machine learning models into systems that store and manage data. In this talk, I will highlight some of the advances that have been made in deep learning and suggest some interesting directions for future research. Jeffrey Dean |
SIGMOD Conference | 1 |
| 2016 | Large-Scale Deep Learning For Building Intelligent Computer SystemsabstractFor the past five years, the Google Brain team has focused on conducting research in difficult problems in artificial intelligence, on building large-scale computer systems for machine learning research, and, in collaboration with many teams at Google, on applying our research and systems to dozens of Google products. Our group has recently open-sourced the TensorFlow system (tensorflow.org), a system designed to easily express machine ideas, and to quickly train, evaluate and deploy machine learning systems. In this talk, I'll highlight some of the design decisions we made in building TensorFlow, discuss research results produced within our group, and describe ways in which these ideas have been applied to a variety of problems in Google's products, usually in close collaboration with other teams. Jeffrey Dean |
WSDM | 1 |
| 2013 | Multilingual acoustic models using distributed deep neural networksabstractToday's speech recognition technology is mature enough to be useful for many practical applications. In this context, it is of paramount importance to train accurate acoustic models for many languages within given resource constraints such as data, processing power, and time. Multilingual training has the potential to solve the data issue and close the performance gap between resource-rich and resource-scarce languages. Neural networks lend themselves naturally to parameter sharing across languages, and distributed implementations have made it feasible to train large networks. In this paper, we present experimental results for cross- and multi-lingual network training of eleven Romance languages on 10k hours of data in total. The average relative gains over the monolingual baselines are 4%/2% (data-scarce/data-rich languages) for cross- and 7%/2% for multi-lingual training. However, the additional gain from jointly training the languages on all data comes at an increased training time of roughly four weeks, compared to two weeks (monolingual) and one week (crosslingual). Georg Heigold, Vincent Vanhoucke, Andrew W. Senior, Patrick Nguyen, Marc'Aurelio Ranzato, Matthieu Devin, Jeffrey Dean |
ICASSP | 7 |
| 2013 | On rectified linear units for speech processingabstractDeep neural networks have recently become the gold standard for acoustic modeling in speech recognition systems. The key computational unit of a deep network is a linear projection followed by a point-wise non-linearity, which is typically a logistic function. In this work, we show that we can improve generalization and make training of deep networks faster and simpler by substituting the logistic units with rectified linear units. These units are linear when their input is positive and zero otherwise. In a supervised setting, we can successfully train very deep nets from random initialization on a large vocabulary speech recognition task achieving lower word error rates than using a logistic network with the same topology. Similarly in an unsupervised setting, we show how we can learn sparse features that can be useful for discriminative tasks. All our experiments are executed in a distributed environment using several hundred machines and several hundred hours of speech data. Matthew D. Zeiler, Marc'Aurelio Ranzato, Rajat Monga, Mark Z. Mao, Quoc V. Le, Patrick Nguyen, Andrew W. Senior, Vincent Vanhoucke, Jeffrey Dean, Geoffrey E. Hinton |
ICASSP | 10 |
| 2013 | DeViSE: A Deep Visual-Semantic Embedding ModelabstractModern visual recognition systems are often limited in their ability to scale to large numbers of object categories. This limitation is in part due to the increasing difficulty of acquiring sufficient training data in the form of labeled images as the number of object categories grows. One remedy is to leverage data from other sources -- such as text data -- both to train visual models and to constrain their predictions. In this paper we present a new deep visual-semantic embedding model trained to identify visual objects using both labeled image data as well as semantic information gleaned from unannotated text. We demonstrate that this model matches state-of-the-art performance on the 1000-class ImageNet object recognition challenge while making more semantically reasonable errors, and also show that the semantic information can be exploited to make predictions about tens of thousands of image labels not observed during training. Semantic knowledge improves such zero-shot predictions by up to 65%, achieving hit rates of up to 10% across thousands of novel labels never seen by the visual model. Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc'Aurelio Ranzato, Tomás Mikolov |
NIPS | 5 |
| 2013 | Distributed Representations of Words and Phrases and their CompositionalityabstractThe recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several improvements that make the Skip-gram model more expressive and enable it to learn higher quality vectors more rapidly. We show that by subsampling frequent words we obtain significant speedup, and also learn higher quality representations as measured by our tasks. We also introduce Negative Sampling, a simplified variant of Noise Contrastive Estimation (NCE) that learns more accurate vectors for frequent words compared to the hierarchical softmax. An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases. For example, the meanings of Canada'' and "Air'' cannot be easily combined to obtain "Air Canada''. Motivated by this example, we present a simple and efficient method for finding phrases, and show that their vector representations can be accurately learned by the Skip-gram model. " Tomás Mikolov, Ilya Sutskever, Kai Chen 0010, Gregory S. Corrado, Jeffrey Dean |
NIPS | 5 |
| 2013 | Spanner: Google's Globally Distributed DatabaseabstractSpanner is Google’s scalable, multiversion, globally distributed, and synchronously replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. This article describes how Spanner is structured, its feature set, the rationale underlying various design decisions, and a novel time API that exposes clock uncertainty. This API and its implementation are critical to supporting external consistency and a variety of powerful features: nonblocking reads in the past, lock-free snapshot transactions, and atomic schema changes, across all of Spanner. James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford |
ACM Trans. Comput. Syst. | 2 |
| 2012 | Building high-level features using large scale unsupervised learning
Quoc V. Le, Marc'Aurelio Ranzato, Rajat Monga, Matthieu Devin, Gregory S. Corrado, Kai Chen 0010, Jeffrey Dean, Andrew Y. Ng |
ICML | 7 |
| 2012 | Large Scale Distributed Deep NetworksabstractRecent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm. Jeffrey Dean, Gregory S. Corrado, Rajat Monga, Kai Chen 0010, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Andrew Y. Ng |
NIPS | 1 |
| 2012 | Spanner: Google's Globally-Distributed Database
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford |
OSDI | 2 |
| 2010 | Evolution and future directions of large-scale storage and computation systems at GoogleabstractNo abstract available. Jeffrey Dean |
SoCC | 1 |
| 2009 | Back-off language model compression
Boulos Harb, Ciprian Chelba, Jeffrey Dean, Sanjay Ghemawat |
INTERSPEECH | 3 |
| 2009 | Challenges in building large-scale information retrieval systems: invited talkabstractBuilding and operating large-scale information retrieval systems used by hundreds of millions of people around the world provides a number of interesting challenges. Designing such systems requires making complex design tradeoffs in a number of dimensions, including (a) the number of user queries that must be handled per second and the response latency to these requests, (b) the number and size of various corpora that are searched, (c) the latency and frequency with which documents are updated or added to the corpora, and (d) the quality and cost of the ranking algorithms that are used for retrieval. In this talk I will discuss the evolution of Google's hardware infrastructure and information retrieval systems and some of the design challenges that arise from ever-increasing demands in all of these dimensions. I will also describe how we use various pieces of distributed systems infrastructure when building these retrieval systems. Finally, I will describe some future challenges and open research problems in this area. Jeffrey Dean |
WSDM | 1 |
| 2008 | Bigtable: A Distributed Storage System for Structured DataabstractBigtable is a distributed storage system for managing structured data that is designed to scale to a very large size: petabytes of data across thousands of commodity servers. Many projects at Google store data in Bigtable, including web indexing, Google Earth, and Google Finance. These applications place very different demands on Bigtable, both in terms of data size (from URLs to web pages to satellite imagery) and latency requirements (from backend bulk processing to real-time data serving). Despite these varied demands, Bigtable has successfully provided a flexible, high-performance solution for all of these Google products. In this article, we describe the simple data model provided by Bigtable, which gives clients dynamic control over data layout and format, and we describe the design and implementation of Bigtable. Fay Chang, Jeffrey Dean, Sanjay Ghemawat, Wilson C. Hsieh, Deborah A. Wallach, Michael Burrows, Tushar Chandra, Andrew Fikes, Robert Gruber |
ACM Trans. Comput. Syst. | 2 |
| 2007 | Large Language Models in Machine Translation
Thorsten Brants, Ashok C. Popat, Franz Josef Och, Jeffrey Dean |
EMNLP-CoNLL | 5 |
| 2007 | MapReduce and Other Building Blocks for Large-Scale Distributed Systems at Google
Jeffrey Dean |
USENIX ATC | 1 |
| 2006 | Experiences with MapReduce, an abstraction for large-scale computationabstractMapReduce is a programming model and an associated implementation for processing and generating large data sets. Users specify a Map function that processes a key/value pair to generate a set of intermediate key/value pairs, and a Reduce function that merges all intermediate values associated with the same intermediate key. Many real world tasks are expressible in this model. Programs written in this functional style are automatically parallelized and executed on a large cluster of commodity machines.The MapReduce run-time system takes care of the details of partitioning the input data, scheduling the program's execution across a set of machines, handling machine failures, and managing the required intermachine communication. This allows programmers without any experience with parallel and distributed systems to easily utilize the resources of a large distributed system.Our implementation of MapReduce runs on a large cluster of commodity machines and is highly scalable: a typical MapReduce computation processes many terabytes of data on thousands of machines. Programmers find the system easy to use: thousands of MapReduce programs have been implemented and several thousand thousand MapReduce jobs are executed on Google's clusters every day.In this talk I'll describe the basic programming model, discuss our experience using it in a variety of domains, and talk about the implications of programming models like MapReduce as one paradigm to simplify development of parallel software for multi-core microprocessors. Jeffrey Dean |
PACT | 1 |
| 2006 | Bigtable: A Distributed Storage System for Structured Data (Awarded Best Paper!)
Fay Chang, Jeffrey Dean, Sanjay Ghemawat, Wilson C. Hsieh, Deborah A. Wallach, Michael Burrows, Tushar Chandra, Andrew Fikes, Robert Gruber |
OSDI | 2 |
| 2004 | MapReduce: Simplified Data Processing on Large Clusters
Jeffrey Dean, Sanjay Ghemawat |
OSDI | 1 |
| 2000 | A comparison of techniques to find mirrored hosts on the WWWabstractWe compare several algorithms for identifying mirrored hosts on the World Wide Web. The algorithms operate on the basis of URL strings and linkage data: the type of information about Web pages easily available from Web proxies and crawlers. Identification of mirrored hosts can improve Web-based information retrieval in several ways: first, by identifying mirrored hosts, search engines can avoid storing and returning duplicate documents. Second, several new information retrieval techniques for the Web make inferences based on the explicit links among hypertext documents—mirroring perturbs their graph model and degrades performance. Third, mirroring information can be used to redirect users to alternate mirror sites to compensate for various failures, and can thus improve the performance of Web browsers and proxies. We evaluated four classes of “top-down” algorithms for detecting mirrored host pairs (that is, algorithms that are based on page attributes such as URL, IP address, and hyperlinks between pages, and not on the page content) on a collection of 140 million URLs (on 230,000 hosts) and their associated connectivity information. Our best approach is one which combines five algorithms and achieved a precision of 0.57 for a recall of 0.86 considering 100,000 ranked host pairs. Krishna Bharat, Andrei Z. Broder, Jeffrey Dean, Monika Henzinger |
J. Am. Soc. Inf. Sci. | 3 |
| 1999 | Finding Related Pages in the World Wide Web
Jeffrey Dean, Monika Henzinger |
Comput. Networks | 1 |
| 1998 | Walknet--a biologically inspired network to control six-legged walking
Holk Cruse, Thomas Kindermann, Michael Schumm, Jeffrey Dean, Josef Schmitz |
Neural Networks | 4 |
| 1997 | ProfileMe: Hardware Support for Instruction-Level Profiling on Out-of-Order ProcessorsabstractProfile data is valuable for identifying performance bottlenecks and guiding optimizations. Periodic sampling of a processor's performance monitoring hardware is an effective, unobtrusive way to obtain detailed profiles. Unfortunately, existing hardware simply counts events, such as cache misses and branch mispredictions, and cannot accurately attribute these events to instructions, especially on out-of-order machines. We propose an alternative approach, called ProfileMe, that samples instructions. As a sampled instruction moves through the processor pipeline, a detailed record of all interesting events and pipeline stage latencies is collected. ProfileMe also supports paired sampling, which captures information about the interactions between concurrent instructions, revealing information about useful concurrency and the utilization of various pipeline stages while an instruction is in flight. We describe an inexpensive hardware implementation of ProfileMe, outline a variety of software techniques to extract useful profile information from the hardware, and explain several ways in which this information can provide valuable feedback for programmers and optimizers. Jeffrey Dean, James E. Hicks, Carl A. Waldspurger, William E. Weihl, George Z. Chrysos |
MICRO | 1 |
| 1997 | Call Graph Construction in Object-Oriented LanguagesabstractInterprocedural analyses enable optimizing compilers to more precisely model the effects of non-inlined procedure calls, potentially resulting in substantial increases in application performance. Applying interprocedural analysis to programs written in object-oriented or functional languages is complicated by the difficulty of constructing an accurate program call graph. This paper presents a parameterized algorithmic framework for call graph construction in the presence of message sends and/or first class functions. We use this framework to describe and to implement a number of well-known and new algorithms. We then empirically assess these algorithms by applying them to a suite of medium-sized programs written in Cecil and Java, reporting on the relative cost of the analyses, the relative precision of the constructed call graphs, and the impact of this precision on the effectiveness of a number of interprocedural optimizations. David Grove, Greg DeFouw, Jeffrey Dean, Craig Chambers |
OOPSLA | 3 |
| 1997 | Continuous Profiling: Where Have All the Cycles Gone?abstractArticle Continuous profiling: where have all the cycles gone? Share on Authors: Jennifer M. Anderson View Profile , Lance M. Berc View Profile , Jeffrey Dean View Profile , Sanjay Ghemawat View Profile , Monika R. Henzinger View Profile , Shun-Tak A. Leung View Profile , Richard L. Sites View Profile , Mark T. Vandevoorde View Profile , Carl A. Waldspurger View Profile , William E. Weihl View Profile Authors Info & Claims SOSP '97: Proceedings of the sixteenth ACM symposium on Operating systems principlesOctober 1997 Pages 1–14https://doi.org/10.1145/268998.266637Published:01 October 1997 175citation1,209DownloadsMetricsTotal Citations175Total Downloads1,209Last 12 Months17Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Jennifer-Ann M. Anderson, Lance M. Berc, Jeffrey Dean, Sanjay Ghemawat, Monika Henzinger, Shun-Tak Leung, Richard L. Sites, Mark T. Vandevoorde, Carl A. Waldspurger, William E. Weihl |
SOSP | 3 |
| 1997 | Continuous Profiling: Where Have All the Cycles Gone?abstractThis article describes the Digital Continuous Profiling Infrastructure, a sampling-based profiling system designed to run continuously on production systems. The system supports multiprocessors, works on unmodified executables, and collects profiles for entire systems, including user programs, shared libraries, and the operating system kernel. Samples are collected at a high rate (over 5200 samples/sec. per 333MHz processor), yet with low overhead (1–3% slowdown for most workloads). Analysis tools supplied with the profiling system use the sample data to produce a precise and accurate accounting, down to the level of pipeline stalls incurred by individual instructions, of where time is bring spent. When instructions incur stalls, the tools identify possible reasons, such as cache misses, branch mispredictions, and functional unit contention. The fine-grained instruction-level analysis guides users and automated optimizers to the causes of performance problems and provides important insights for fixing them. Jennifer-Ann M. Anderson, Lance M. Berc, Jeffrey Dean, Sanjay Ghemawat, Monika Henzinger, Shun-Tak Leung, Richard L. Sites, Mark T. Vandevoorde, Carl A. Waldspurger, William E. Weihl |
ACM Trans. Comput. Syst. | 3 |
| 1996 | Simplifying Neural Networks for Controlling Walking by Exploiting Physical Properties
Holk Cruse, Christian Bartling, Jeffrey Dean, Thomas Kindermann, Josef Schmitz, Michael Schumm, Hendrik Wagner |
ICANN | 3 |
| 1996 | Vortex: An Optimizing Compiler for Object-Oriented LanguagesabstractPreviously, techniques such as class hierarchy analysis and profile-guided receiver class prediction have been demonstrated to greatly improve the performance of applications written in pure object-oriented languages, but the degree to which these results are transferable to applications written in hybrid languages has been unclear. In part to answer this question, we have developed the Vortex compiler infrastructure, a language-independent optimizing compiler for object-oriented languages, with front-ends for Cecil, C++, Java, and Modula-3. In this paper, we describe the Vortex compiler's intermediate language, internal structure, and optimization suite, and then we report the results of experiments assessing the effectiveness of different combinations of optimizations on sizable applications across these four languages. We characterize the benchmark programs in terms of a collection of static and dynamic metrics, intended to quantify aspects of the "object-orientedness" of a program. Jeffrey Dean, Greg DeFouw, David Grove, Vassily Litvinov, Craig Chambers |
OOPSLA | 1 |
| 1995 | Optimization of Object-Oriented Programs Using Static Class Hierarchy Analysis
Jeffrey Dean, David Grove, Craig Chambers |
ECOOP | 1 |
| 1995 | A Framework for Selective Recompilation in the Presence of Complex Intermodule DependenciesabstractCompilersand other programming environment tools derive information from the source code of programs; derived information includes compiled code, interprocedurrd summary information, and call graph views.If the source program changes, the derived information needs to be updated.We present a simple framework for maintaining interrnodule dependencies, embodying different tradeoffs in terms of space usage, speed of processing, and selectivity of invalidation, that eases the implementation of incremental update of derived information.Our framework augments a directed acyclic graph representation of dependencies with factoring nodes (to save space) and$ltering nodes (to increase selectivity), and it includes an algorithm for efficient invalidation processing.We show how several schemes for selective recompilation, such as smart recompilation, filter sets for interprocedural summary information, and dependencies for whole-program optimization of object-oriented languages, map naturally onto our framework.For this latter application, by exploiting the facilities of our framework, we are able to reduce the number of lines of source code recompiled by a factor of seven over a header file-based scheme, and by a factor of two over the previous state-of-the-art selective dependency mechanism without consuming additional space. Craig Chambers, Jeffrey Dean, David Grove |
ICSE | 2 |
| 1995 | Profile-Guided Receiver Class PredictionabstractThe use of dynamically-dispatched procedure calls is a key mechanism for writing extensible and flexible code in object-oriented languages. Unfortunately, dynamic dispatching imposes a runtime performance penalty. Some recent implementations of pure object-oriented languages have utilized profile-guided receiver class prediction to reduce this performance penalty, and some researchers have argued for applying receiver class prediction in hybrid languages like C++. We performed a detailed examination of the dynamic profiles of eight large object-oriented applications written in C++ and Cecil, determining that the receiver class distributions are strongly peaked and stable across both inputs and program versions through time. We describe techniques for gathering and manipulating profile information at varying degrees of precision, particularly in the presence of optimizations such as inlining. Our implementation of profile-guided receiver class prediction improves the performance of large Cecil applications by more than a factor of two over solely static optimizations. David Grove, Jeffrey Dean, Charles Garrett, Craig Chambers |
OOPSLA | 2 |
| 1995 | Selective Specialization for Object-Oriented LanguagesabstractDynamic dispatching is a major source of run-time overhead in object-oriented languages, due both to the direct cost of method lookup and to the indirect effect of preventing other optimizations. To reduce this overhead, optimizing compilers for object-oriented languages analyze the classes of objects stored in program variables, with the goal of bounding the possible classes of message receivers enough so that the compiler can uniquely determine the target of a message send at compile time and replace the message send with a direct procedure call. Specialization is one important technique for improving the precision of this static class information: by compiling multiple versions of a method, each applicable to a subset of the possible argument classes of the method, more precise static information about the classes of the method's arguments is obtained. Previous specialization strategies have not been selective about where this technique is applied, and therefore tended to significantly increase compile time and code space usage, particularly for large applications. In this paper, we present a more general framework for specialization in object-oriented languages and describe a goal directed specialization algorithm that makes selective decisions to apply specialization to those cases where it provides the highest benefit. Our results show that our algorithm improves the performance of a group of sizeable programs by 65% to 275% while increasing compiled code space requirements by only 4% to 10%. Moreover, when compared to the previous state-of-the-art specialization scheme, our algorithm improves performance by 11% to 67% while simultaneously reducing code space requirements by 65% to 73%. Jeffrey Dean, Craig Chambers, David Grove |
PLDI | 1 |
| 1994 | Identifying Profitable Specialization in Object-Oriented Languages
Jeffrey Dean, Craig Chambers, David Grove |
PEPM | 1 |