VLDB 2026 Research / reviewers in the wild / expert
Joel Hestness
dblp:60/3063
· DBLP profile ↗
11ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0001-6920-0906ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 since 2021Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Deep learning architectures and training · 46% Efficient and distributed learning · 21% Language models and text generation · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 30% Performance modeling and evaluation · 30% Interconnection networks and networks-on-chip · 25% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
hyperparameter transfer |
1.6 | 2 | 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformers · NeurIPS 2025 Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
1.6 | 2 | 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformers · NeurIPS 2025 Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › efficient training
compute-efficient training |
0.9 | 1 | 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformers · NeurIPS 2025 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.9 | 1 | 2025 | Power Lines: Scaling laws for weight decay and batch size in LLM pre-training · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Power Lines: Scaling laws for weight decay and batch size in LLM pre-training · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs · ICLR 2025 |
Machine learning › Optimization for machine learning
learning rate schedule |
0.9 | 1 | 2025 | Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs · ICLR 2025 |
Machine learning › Deep learning architectures and training › training optimization
large-batch training |
0.8 | 1 | 2024 | Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.8 | 1 | 2024 | Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Natural language and speech › Language models and text generation
compositional generalization |
0.4 | 1 | 2019 | Compositional Generalization for Primitive Substitutions · EMNLP/IJCNLP (1) 2019 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.4 | 1 | 2019 | Compositional Generalization for Primitive Substitutions · EMNLP/IJCNLP (1) 2019 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2019 | Beyond human-level accuracy: computational challenges in deep learning · PPoPP 2019 |
Performance modeling and evaluation
workload characterization |
0.4 | 1 | 2019 | Beyond human-level accuracy: computational challenges in deep learning · PPoPP 2019 |
Interconnection networks and networks-on-chip
flow control |
0.1 | 1 | 2011 | Kilo-NOC: a heterogeneous network-on-chip architecture for scalability and service guarantees · ISCA 2011 |
Cloud and datacenter computing
quality of service |
0.1 | 1 | 2011 | Kilo-NOC: a heterogeneous network-on-chip architecture for scalability and service guarantees · ISCA 2011 |
Machine learning › Deep learning architectures and training
scaling laws |
0.1 | 1 | 2019 | Beyond human-level accuracy: computational challenges in deep learning · PPoPP 2019 |
Interconnection networks and networks-on-chip
network topology |
0.1 | 1 | 2009 | Express Cube Topologies for on-Chip Interconnects · HPCA 2009 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2011 | Kilo-NOC: a heterogeneous network-on-chip architecture for scalability and service guarantees · ISCA 2011 Express Cube Topologies for on-Chip Interconnects · HPCA 2009 |
Methods — techniques the papers use, named apart from their topics
maximal update parameterization · 1.6scaling laws · 0.9power law fitting · 0.9linear decay-to-zero · 0.9lazy learning theory · 0.9adamw · 0.9sparsity-aware reparameterization · 0.8per-example gradient norms · 0.8layernorm backward pass · 0.8custom kernel · 0.8scaling projection · 0.4simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMsabstractLLMs are commonly trained with a learning rate (LR) warmup, followed by cosine decay to 10% of the maximum (10x decay). In a large-scale empirical study, we show that under an optimal peak LR, a simple linear decay-to-zero (D2Z) schedule consistently outperforms other schedules when training at compute-optimal dataset sizes. D2Z is superior across a range of model sizes, batch sizes, datasets, and vocabularies. Benefits increase as dataset size increases. Leveraging a novel interpretation of AdamW as an exponential moving average of weight updates, we show how linear D2Z optimally balances the demands of early training (moving away from initial conditions) and late training (averaging over more updates in order to mitigate gradient noise). In experiments, a 610M-parameter model trained for 80 tokens-per-parameter (TPP) using D2Z achieves lower loss than when trained for 200 TPP using 10x decay, corresponding to an astonishing 60% compute savings. Models such as Llama2-7B, trained for 286 TPP with 10x decay, could likely have saved a majority of compute by training with D2Z. Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness |
ICLR | 6 |
| 2025 | Power Lines: Scaling laws for weight decay and batch size in LLM pre-trainingabstractEfficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N, dataset size D, and batch size B. Recent work suggests the AdamW timescale, τ = B/(ηλD), should remain constant across training settings, and we verify the implication that optimal λ scales linearly with B, for a fixed N and D. However, as N and D scale, we show optimal τ obeys a precise power law in the tokens-per-parameter ratio, D/N. This law thus provides a method to accurately predict λopt in advance of large-scale training. We also study scaling laws for optimal batch size Bopt (the B enabling lowest loss at a given N,D) and critical batch size Bcrit (the B beyond which further data parallelism becomes ineffective). In contrast to prior work, we find both Bopt and Bcrit scale as power laws in D, independent of model size, N. Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives. Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness |
NeurIPS | 6 |
| 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformersabstractWe study compute efficiency of LLM training when using different parameterizations, i.e., rules for adjusting model and optimizer hyperparameters (HPs) as model size changes. Some parameterizations fail to transfer optimal base HPs (such as learning rate) across changes in model depth, requiring practitioners to either re-tune these HPs as they scale up (expensive), or accept sub-optimal training when re-tuning is prohibitive. Even when they achieve HP transfer, we develop theory to show parameterizations may still exist in the lazy learning regime where layers learn only features close to their linearization, preventing effective use of depth and nonlinearity. Finally, we identify and adopt the parameterization we call CompleteP that achieves both depth-wise HP transfer and non-lazy learning in all layers. CompleteP enables a wider range of model width/depth ratios to remain compute-efficient, unlocking shapes better suited for different hardware settings and operational contexts. Moreover, CompleteP enables 12-34% compute efficiency improvements over the prior state-of-the-art. All experiments were run on Cerebras CS-3 systems. A minimal implementation is available at https://github.com/EleutherAI/nanoGPT-mup/tree/completep. Nolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li, Blake Bordelon, Shane Bergsma, Cengiz Pehlevan, Boris Hanin, Joel Hestness |
NeurIPS | 9 |
| 2024 | Sparse maximal update parameterization: A holistic approach to sparse training dynamicsabstractSeveral challenges make it difficult for sparse neural networks to compete with dense models. First, setting a large fraction of weights to zero impairs forward and gradient signal propagation. Second, sparse studies often need to test multiple sparsity levels, while also introducing new hyperparameters (HPs), leading to prohibitive tuning costs. Indeed, the standard practice is to re-use the learning HPs originally crafted for dense models. Unfortunately, we show sparse and
dense networks do not share the same optimal HPs. Without stable dynamics and effective training recipes, it is costly to test sparsity at scale, which is key to surpassing dense networks and making the business case for sparsity acceleration in hardware.
A holistic approach is needed to tackle these challenges and we propose S$\textmu$Par as one such approach. For random unstructured static sparsity, S$\textmu$Par ensures activations, gradients, and weight updates all scale independently of sparsity level. Further, by reparameterizing the HPs, S$\textmu$Par enables the same HP values to be optimal as we vary both sparsity level and model width. HPs can be tuned on small dense networks and transferred to large sparse models, greatly reducing tuning costs. On large-scale language modeling, S$\textmu$Par shows increasing improvements over standard parameterization as sparsity increases, leading up to 11.9\% relative loss improvement at 99.2\% sparsity. A minimal implementation of S$\textmu$Par is available at https://github.com/EleutherAI/nanoGPT-mup/tree/supar. Nolan Dey, Shane Bergsma, Joel Hestness |
NeurIPS | 3 |
| 2024 | Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in TransformersabstractPer-example gradient norms are a vital ingredient for estimating gradient noise scale (GNS) with minimal variance. Observing the tensor contractions required to compute them, we propose a method with minimal FLOPs in 3D or greater tensor regimes by simultaneously computing the norms while computing the parameter gradients. Using this method we are able to observe the GNS of different layers at higher accuracy than previously possible. We find that the total GNS of contemporary transformer models is predicted well by the GNS of only the normalization layers. As a result, focusing only on the normalization layer, we develop a custom kernel to compute the per-example gradient norms while performing the LayerNorm backward pass with zero throughput overhead. Tracking GNS on only those layers, we are able to guide a practical batch size schedule that reduces training time by 18% on a Chinchilla-optimal language model. Gavia Gray, Aman Tiwari, Shane Bergsma, Joel Hestness |
NeurIPS | 4 |
| 2019 | Compositional Generalization for Primitive SubstitutionsabstractYuanpeng Li, Liang Zhao, Jianyu Wang, Joel Hestness. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yuanpeng Li 0001, Joel Hestness |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Beyond human-level accuracy: computational challenges in deep learningabstractDeep learning (DL) research yields accuracy and product improvements from both model architecture changes and scale: larger data sets and models, and more computation. For hardware design, it is difficult to predict DL model changes. However, recent prior work shows that as dataset sizes grow, DL model accuracy and model size grow predictably. This paper leverages the prior work to project the dataset and model size growth required to advance DL accuracy beyond human-level, to frontier targets defined by machine learning experts. Datasets will need to grow 33--971×, while models will need to grow 6.6--456× to achieve target accuracies. Joel Hestness, Newsha Ardalani, Gregory Frederick Diamos |
PPoPP | 1 |
| 2019 | A survey of 25 years of evaluationabstractAbstract Evaluation was not a thing when the first author was a graduate student in the late 1970s. There was an Artificial Intelligence (AI) boom then, but that boom was quickly followed by a bust and a long AI Winter. Charles Wayne restarted funding in the mid-1980s by emphasizing evaluation. No other sort of program could have been funded at the time, at least in America. His program was so successful that these days, shared tasks and leaderboards have become common place in speech and language (and Vision and Machine Learning). It is hard to remember that evaluation was a tough sell 25 years ago. That said, we may be a bit too satisfied with current state of the art. This paper will survey considerations from other fields such as reliability and validity from psychology and generalization from systems. There has been a trend for publications to report better and better numbers, but what do these numbers mean? Sometimes the numbers are too good to be true, and sometimes the truth is better than the numbers. It is one thing for an evaluation to fail to find a difference between man and machine, and quite another thing to pass the Turing Test. As Feynman said, “the first principle is that you must not fool yourself–and you are the easiest person to fool.” Kenneth Church 0001, Joel Hestness |
Nat. Lang. Eng. | 2 |
| 2017 | Convolutional Recurrent Neural Networks for Small-Footprint Keyword SpottingabstractKeyword spotting (KWS) constitutes a major component of human-technology interfaces.Maximizing the detection accuracy at a low false alarm (FA) rate, while minimizing the footprint size, latency and complexity are the goals for KWS.Towards achieving them, we study Convolutional Recurrent Neural Networks (CRNNs).Inspired by large-scale state-ofthe-art speech recognition systems, we combine the strengths of convolutional layers and recurrent layers to exploit local structure and long-range context.We analyze the effect of architecture parameters, and propose training strategies to improve performance.With only ~230k parameters, our CRNN model yields acceptably low latency, and achieves 97.71% accuracy at 0.5 FA/hour for 5 dB signal-to-noise ratio. Sercan Ö. Arik, Markus Kliegl, Rewon Child, Joel Hestness, Andrew Gibiansky, Christopher Fougner, Ryan Prenger, Adam Coates 0002 |
INTERSPEECH | 4 |
| 2011 | Kilo-NOC: a heterogeneous network-on-chip architecture for scalability and service guaranteesabstractToday's chip-level multiprocessors (CMPs) feature up to a hundred discrete cores, and with increasing levels of integration, CMPs with hundreds of cores, cache tiles, and specialized accelerators are anticipated in the near future. In this paper, we propose and evaluate technologies to enable networks-on-chip (NOCs) to support a thousand connected components (Kilo-NOC) with high area and energy efficiency, good performance, and strong quality-of-service (QOS) guarantees. Our analysis shows that QOS support burdens the network with high area and energy costs. In response, we propose a new lightweight topology-aware QOS architecture that provides service guarantees for applications such as consolidated servers on CMPs and real-time SOCs. Unlike prior NOC quality-of-service proposals which require QOS support at every network node, our scheme restricts the extent of hardware support to portions of the die, reducing router complexity in the rest of the chip. We further improve network area- and energy-efficiency through a novel flow control mechanism that enables a single-network, low-cost elastic buffer implementation. Together, these techniques yield a heterogeneous Kilo-NOC architecture that consumes 45% less area and 29% less power than a state-of-the-art QOS-enabled NOC without these features. Boris Grot, Joel Hestness, Stephen W. Keckler, Onur Mutlu |
ISCA | 2 |
| 2009 | Express Cube Topologies for on-Chip InterconnectsabstractDriven by continuing scaling of Moore's law, chip multi-processors and systems-on-a-chip are expected to grow the core count from dozens today to hundreds in the near future. Scalability of on-chip interconnect topologies is critical to meeting these demands. In this work, we seek to develop a better understanding of how network topologies scale with regard to cost, performance, and energy considering the advantages and limitations afforded on a die. Our contributions are three-fold. First, we propose a new topology, called Multidrop Express Channels (MECS), that uses a one-to-many communication model enabling a high degree of connectivity in a bandwidth-efficient manner. In a 64-terminal network, MECS enjoys a 9% latency advantage over other topologies at low network loads, which extends to over 20% in a 256-terminal network. Second, we demonstrate that partitioning the available wires among multiple networks and channels enables new opportunities for trading-off performance, area, and energy-efficiency that depend on the partitioning scheme. Third, we introduce Generalized Express Cubes - a framework for expressing the space of on-chip interconnects - and demonstrate how existing and proposed topologies can be mapped to it. Boris Grot, Joel Hestness, Stephen W. Keckler, Onur Mutlu |
HPCA | 2 |