VLDB 2026 Research / reviewers in the wild / expert
Nolan Dey
dblp:263/9353 · also Nolan S. Dey, Nolan Simran Dey
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 36% Efficient and distributed learning · 26% Optimization for machine learning · 19% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
hyperparameter transfer |
1.6 | 2 | 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformers · NeurIPS 2025 Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
training dynamics |
1.6 | 2 | 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformers · NeurIPS 2025 Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › efficient training
compute-efficient training |
0.9 | 1 | 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformers · NeurIPS 2025 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.9 | 1 | 2025 | Power Lines: Scaling laws for weight decay and batch size in LLM pre-training · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Power Lines: Scaling laws for weight decay and batch size in LLM pre-training · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs · ICLR 2025 |
Machine learning › Optimization for machine learning
learning rate schedule |
0.9 | 1 | 2025 | Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.8 | 1 | 2024 | Sparse maximal update parameterization: A holistic approach to sparse training dynamics · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
maximal update parameterization · 1.6scaling laws · 0.9power law fitting · 0.9linear decay-to-zero · 0.9lazy learning theory · 0.9adamw · 0.9sparsity-aware reparameterization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMsabstractLLMs are commonly trained with a learning rate (LR) warmup, followed by cosine decay to 10% of the maximum (10x decay). In a large-scale empirical study, we show that under an optimal peak LR, a simple linear decay-to-zero (D2Z) schedule consistently outperforms other schedules when training at compute-optimal dataset sizes. D2Z is superior across a range of model sizes, batch sizes, datasets, and vocabularies. Benefits increase as dataset size increases. Leveraging a novel interpretation of AdamW as an exponential moving average of weight updates, we show how linear D2Z optimally balances the demands of early training (moving away from initial conditions) and late training (averaging over more updates in order to mitigate gradient noise). In experiments, a 610M-parameter model trained for 80 tokens-per-parameter (TPP) using D2Z achieves lower loss than when trained for 200 TPP using 10x decay, corresponding to an astonishing 60% compute savings. Models such as Llama2-7B, trained for 286 TPP with 10x decay, could likely have saved a majority of compute by training with D2Z. Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness |
ICLR | 2 |
| 2025 | Power Lines: Scaling laws for weight decay and batch size in LLM pre-trainingabstractEfficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate η and weight decay λ. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N, dataset size D, and batch size B. Recent work suggests the AdamW timescale, τ = B/(ηλD), should remain constant across training settings, and we verify the implication that optimal λ scales linearly with B, for a fixed N and D. However, as N and D scale, we show optimal τ obeys a precise power law in the tokens-per-parameter ratio, D/N. This law thus provides a method to accurately predict λopt in advance of large-scale training. We also study scaling laws for optimal batch size Bopt (the B enabling lowest loss at a given N,D) and critical batch size Bcrit (the B beyond which further data parallelism becomes ineffective). In contrast to prior work, we find both Bopt and Bcrit scale as power laws in D, independent of model size, N. Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives. Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness |
NeurIPS | 2 |
| 2025 | Don't be lazy: CompleteP enables compute-efficient deep transformersabstractWe study compute efficiency of LLM training when using different parameterizations, i.e., rules for adjusting model and optimizer hyperparameters (HPs) as model size changes. Some parameterizations fail to transfer optimal base HPs (such as learning rate) across changes in model depth, requiring practitioners to either re-tune these HPs as they scale up (expensive), or accept sub-optimal training when re-tuning is prohibitive. Even when they achieve HP transfer, we develop theory to show parameterizations may still exist in the lazy learning regime where layers learn only features close to their linearization, preventing effective use of depth and nonlinearity. Finally, we identify and adopt the parameterization we call CompleteP that achieves both depth-wise HP transfer and non-lazy learning in all layers. CompleteP enables a wider range of model width/depth ratios to remain compute-efficient, unlocking shapes better suited for different hardware settings and operational contexts. Moreover, CompleteP enables 12-34% compute efficiency improvements over the prior state-of-the-art. All experiments were run on Cerebras CS-3 systems. A minimal implementation is available at https://github.com/EleutherAI/nanoGPT-mup/tree/completep. Nolan Dey, Bin Claire Zhang, Lorenzo Noci, Mufan Bill Li, Blake Bordelon, Shane Bergsma, Cengiz Pehlevan, Boris Hanin, Joel Hestness |
NeurIPS | 1 |
| 2024 | Sparse maximal update parameterization: A holistic approach to sparse training dynamicsabstractSeveral challenges make it difficult for sparse neural networks to compete with dense models. First, setting a large fraction of weights to zero impairs forward and gradient signal propagation. Second, sparse studies often need to test multiple sparsity levels, while also introducing new hyperparameters (HPs), leading to prohibitive tuning costs. Indeed, the standard practice is to re-use the learning HPs originally crafted for dense models. Unfortunately, we show sparse and
dense networks do not share the same optimal HPs. Without stable dynamics and effective training recipes, it is costly to test sparsity at scale, which is key to surpassing dense networks and making the business case for sparsity acceleration in hardware.
A holistic approach is needed to tackle these challenges and we propose S$\textmu$Par as one such approach. For random unstructured static sparsity, S$\textmu$Par ensures activations, gradients, and weight updates all scale independently of sparsity level. Further, by reparameterizing the HPs, S$\textmu$Par enables the same HP values to be optimal as we vary both sparsity level and model width. HPs can be tuned on small dense networks and transferred to large sparse models, greatly reducing tuning costs. On large-scale language modeling, S$\textmu$Par shows increasing improvements over standard parameterization as sparsity increases, leading up to 11.9\% relative loss improvement at 99.2\% sparsity. A minimal implementation of S$\textmu$Par is available at https://github.com/EleutherAI/nanoGPT-mup/tree/supar. Nolan Dey, Shane Bergsma, Joel Hestness |
NeurIPS | 1 |