Jaehoon Lee 0001

dblp:95/386-1 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Deep learning architectures and training · 28% Learning theory · 20% Probabilistic and Bayesian machine learning · 12%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 30 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
1.842022
Fast Neural Kernel Embeddings for General Activations · NeurIPS 2022
Finite Versus Infinite Neural Networks: an Empirical Study · NeurIPS 2020
Neural Tangents: Fast and Easy Infinite Neural Networks in Python · ICLR 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural network gaussian process
1.532022
Fast Neural Kernel Embeddings for General Activations · NeurIPS 2022
Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width Limit · ICLR 2021
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes · ICLR (Poster) 2019
Machine learning › Kernel, tree and ensemble methods
kernel methods
1.332021
Dataset Distillation with Infinitely Wide Convolutional Networks · NeurIPS 2021
Finite Versus Infinite Neural Networks: an Empirical Study · NeurIPS 2020
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent · NeurIPS 2019
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning · ICLR 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning · ICLR 2025
Machine learning › Optimization for machine learning › adaptive optimization
adam
0.812024
Scaling Exponents Across Parameterizations and Optimizers · ICML 2024
Machine learning › Deep learning architectures and training
hyperparameter transfer
0.812024
Scaling Exponents Across Parameterizations and Optimizers · ICML 2024
Machine learning › Learning theory › neural network theory › neural network parameterization
maximal update parameterization
0.812024
Scaling Exponents Across Parameterizations and Optimizers · ICML 2024
Machine learning › Deep learning architectures and training
scaling laws
0.812024
Scaling Exponents Across Parameterizations and Optimizers · ICML 2024
Machine learning › Deep learning architectures and training › training dynamics
training instability
0.812024
Small-scale proxies for large-scale Transformer training instabilities · ICLR 2024
Machine learning › Deep learning architectures and training › transformer
transformer training
0.812024
Small-scale proxies for large-scale Transformer training instabilities · ICLR 2024
Machine learning › Learning theory › neural network theory
neural network kernels
0.612022
Fast Neural Kernel Embeddings for General Activations · NeurIPS 2022
Algorithms and data structures
sketching
0.612022
Fast Neural Kernel Embeddings for General Activations · NeurIPS 2022
Algorithms and data structures › numerical linear algebra › dimensionality reduction
subspace embedding
0.612022
Fast Neural Kernel Embeddings for General Activations · NeurIPS 2022
Machine learning › Efficient and distributed learning
dataset distillation
0.512021
Dataset Distillation with Infinitely Wide Convolutional Networks · NeurIPS 2021
Machine learning › Learning theory › neural network theory
infinite-width limit
0.512021
Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width Limit · ICLR 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.512021
Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width Limit · ICLR 2021
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks
0.412020
Finite Versus Infinite Neural Networks: an Empirical Study · NeurIPS 2020
Programming languages and type systems
library design
0.412020
Neural Tangents: Fast and Easy Infinite Neural Networks in Python · ICLR 2020
Machine learning › Deep learning architectures and training › training dynamics
batch size scaling
0.412019
Measuring the Effects of Data Parallelism on Neural Network Training · J. Mach. Learn. Res. 2019
Machine learning › Deep learning architectures and training › convolutional neural network
bayesian convolutional neural networks
0.412019
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes · ICLR (Poster) 2019
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412019
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes · ICLR (Poster) 2019
Machine learning › Efficient and distributed learning › distributed training
data parallel training
0.412019
Measuring the Effects of Data Parallelism on Neural Network Training · J. Mach. Learn. Res. 2019
Machine learning › Efficient and distributed learning
distributed training
0.412019
Measuring the Effects of Data Parallelism on Neural Network Training · J. Mach. Learn. Res. 2019
Machine learning › Deep learning architectures and training › training optimization
large-batch training
0.412019
Measuring the Effects of Data Parallelism on Neural Network Training · J. Mach. Learn. Res. 2019
Machine learning › Deep learning architectures and training
training dynamics
0.412019
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.312018
Deep Neural Networks as Gaussian Processes · ICLR (Poster) 2018
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.312025
Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning · ICLR 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.312025
Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning · ICLR 2025
Computer vision › Image recognition and object detection
image classification
0.112021
Dataset Distillation with Infinitely Wide Convolutional Networks · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

sketching · 1.1hermite expansion · 1.1verifier · 0.9process reward model · 0.9best-of-n sampling · 0.9weight decay · 0.8warm-up · 0.8muparam · 0.8gaussian process · 0.8infinite-width analysis · 0.5neural tangent kernel · 0.4
YearPublicationVenuePosition
2025 Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning
abstract
Enabling LLMs to improve their outputs by using more test-time compute is a critical step towards building self-improving agents that can operate on open-ended natural language. In this paper, we scale up inference-time computation in LLMs, with a focus on answering: if an LLM is allowed to use a fixed but non-trivial amount of inference-time compute, how much can it improve its performance on a challenging prompt? Answering this question has implications not only on performance, but also on the future of LLM pretraining and how to tradeoff inference-time and pre-training compute. Little research has attempted to understand the scaling behaviors of test-time inference methods, with current work largely providing negative results for a number of these strategies. In this work, we analyze two primary mechanisms to scale test-time computation: (1) searching against dense, process-based verifier reward models (PRMs); and (2) updating the model's distribution over a response adaptively, given the prompt at test time. We find that in both cases, the effectiveness of different approaches to scaling test-time compute critically varies depending on the difficulty of the prompt. This observation motivates applying a "compute-optimal" scaling strategy, which acts to, as effectively as possible, allocate test-time compute per prompt in an adaptive manner. Using this compute-optimal strategy, we can improve the efficiency of test-time compute scaling for math reasoning problems by more than 4x compared to a best-of-N baseline. Additionally, in a FLOPs-matched evaluation, we find that on problems where a smaller base model attains somewhat non-trivial success rates, test-time compute can be used to outperform a 14x larger model.
Charlie Snell, Jaehoon Lee 0001, Kelvin Xu, Aviral Kumar
ICLR2
2024 Small-scale proxies for large-scale Transformer training instabilities
abstract
Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smaller scales. Although the causes of such instabilities are of scientific interest, the amount of resources required to reproduce them has made investigation difficult. In this work, we seek ways to reproduce and study training instability at smaller scales. First, we focus on two sources of training instability described in previous work: the growth of logits in attention layers (Dehghani et al., 2023) and divergence of the output logits from the log probabilities (Chowdhery et al., 2022). By measuring the relationship between learning rate and loss across scales, we show that these instabilities also appear in small models when training at high learning rates, and that mitigations previously employed at large scales are equally effective in this regime. This prompts us to investigate the extent to which other known optimizer and model interventions influence the sensitivity of the final loss to changes in the learning rate. To this end, we study methods such as warm-up, weight decay, and the MuParam (Yang et al., 2022), and combine techniques to train small models that achieve similar losses across orders of magnitude of learning rate variation. Finally, to conclude our exploration we study two cases where instabilities can be predicted before they emerge by examining the scaling behavior of model characteristics such as activation and gradient norms.
Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie Everett, Alexander A. Alemi, Ben Adlam, John D. Co-Reyes, Izzeddin Gur, Roman Novak, Jeffrey Pennington, Jascha Sohl-Dickstein, Kelvin Xu, Jaehoon Lee 0001, Justin Gilmer, Simon Kornblith
ICLR14
2024 Scaling Exponents Across Parameterizations and Optimizers
abstract
Robust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization and optimizer choices. In this work, we propose a new perspective on parameterization by investigating a key assumption in prior work about the alignment between parameters and data and derive new theoretical results under weaker assumptions and a broader set of optimizers. Our extensive empirical investigation includes *tens of thousands* of models trained with *all combinations of* three optimizers, four parameterizations, several alignment assumptions, more than a dozen learning rates, and fourteen model sizes up to 27B parameters. We find that the best learning rate scaling prescription would often have been excluded by the assumptions in prior work. Our results show that all parameterizations, not just maximal update parameterization (muP), can achieve hyperparameter transfer; moreover, our novel per-layer learning rate prescription for standard parameterization outperforms muP. Finally, we demonstrate that an overlooked aspect of parameterization, the epsilon parameter in Adam, must be scaled correctly to avoid gradient underflow and propose *Adam-atan2*, a new numerically stable, scale-invariant version of Adam that eliminates the epsilon hyperparameter entirely.
Katie Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi, Roman Novak, Peter J. Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee 0001, Jeffrey Pennington
ICML10
2022 Fast Neural Kernel Embeddings for General Activations
abstract
Infinite width limit has shed light on generalization and optimization aspects of deep learning by establishing connections between neural networks and kernel methods. Despite their importance, the utility of these kernel methods was limited in large-scale learning settings due to their (super-)quadratic runtime and memory complexities. Moreover, most prior works on neural kernels have focused on the ReLU activation, mainly due to its popularity but also due to the difficulty of computing such kernels for general activations. In this work, we overcome such difficulties by providing methods to work with general activations. First, we compile and expand the list of activation functions admitting exact dual activation expressions to compute neural kernels. When the exact computation is unknown, we present methods to effectively approximate them. We propose a fast sketching method that approximates any multi-layered Neural Network Gaussian Process (NNGP) kernel and Neural Tangent Kernel (NTK) matrices for a wide range of activation functions, going beyond the commonly analyzed ReLU activation. This is done by showing how to approximate the neural kernels using the truncated Hermite expansion of any desired activation functions. While most prior works require data points on the unit sphere, our methods do not suffer from such limitations and are applicable to any dataset of points in $\mathbb{R}^d$. Furthermore, we provide a subspace embedding for NNGP and NTK matrices with near input-sparsity runtime and near-optimal target dimension which applies to any \emph{homogeneous} dual activation functions with rapidly convergent Taylor expansion. Empirically, with respect to exact convolutional NTK (CNTK) computation, our method achieves $106\times$ speedup for approximate CNTK of a 5-layer Myrtle network on CIFAR-10 dataset.
Insu Han, Amir Zandieh, Jaehoon Lee 0001, Roman Novak, Lechao Xiao, Amin Karbasi
NeurIPS3
2021 Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width Limit
Ben Adlam, Jaehoon Lee 0001, Lechao Xiao, Jeffrey Pennington, Jasper Snoek
ICLR2
2021 Dataset Distillation with Infinitely Wide Convolutional Networks
abstract
The effectiveness of machine learning algorithms arises from being able to extract useful features from large amounts of data. As model and dataset sizes increase, dataset distillation methods that compress large datasets into significantly smaller yet highly performant ones will become valuable in terms of training efficiency and useful feature extraction. To that end, we apply a novel distributed kernel-based meta-learning framework to achieve state-of-the-art results for dataset distillation using infinitely wide convolutional neural networks. For instance, using only 10 datapoints (0.02% of original dataset), we obtain over 65% test accuracy on CIFAR-10 image classification task, a dramatic improvement over the previous best test accuracy of 40%. Our state-of-the-art results extend across many other settings for MNIST, Fashion-MNIST, CIFAR-10, CIFAR-100, and SVHN. Furthermore, we perform some preliminary analyses of our distilled datasets to shed light on how they differ from naturally occurring data.
Timothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon Lee 0001
NeurIPS4
2020 Neural Tangents: Fast and Easy Infinite Neural Networks in Python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 0001, Alexander A. Alemi, Jascha Sohl-Dickstein, Samuel S. Schoenholz
ICLR4
2020 Finite Versus Infinite Neural Networks: an Empirical Study
abstract
We perform a careful, thorough, and large scale empirical study of the correspondence between wide neural networks and kernel methods. By doing so, we resolve a variety of open questions related to the study of infinitely wide neural networks. Our experimental results include: kernel methods outperform fully-connected finite-width networks, but underperform convolutional finite width networks; neural network Gaussian process (NNGP) kernels frequently outperform neural tangent (NT) kernels; centered and ensembled finite networks have reduced posterior variance and behave more similarly to infinite networks; weight decay and the use of a large learning rate break the correspondence between finite and infinite networks; the NTK parameterization outperforms the standard parameterization for finite width networks; diagonal regularization of kernels acts similarly to early stopping; floating point precision limits kernel performance beyond a critical dataset size; regularized ZCA whitening improves accuracy; finite network performance depends non-monotonically on width in ways not captured by double descent phenomena; equivariance of CNNs is only beneficial for narrow networks far from the kernel regime. Our experiments additionally motivate an improved layer-wise scaling for weight decay which improves generalization in finite-width networks. Finally, we develop improved best practices for using NNGP and NT kernels for prediction, including a novel ensembling technique. Using these best practices we achieve state-of-the-art results on CIFAR-10 classification for kernels corresponding to each architecture class we consider.
Jaehoon Lee 0001, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, Jascha Sohl-Dickstein
NeurIPS1
2019 Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee 0001, Greg Yang, Jiri Hron, Daniel A. Abolafia, Jeffrey Pennington, Jascha Sohl-Dickstein
ICLR (Poster)4
2019 Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
abstract
A longstanding goal in deep learning research has been to precisely characterize training and generalization. However, the often complex loss landscapes of neural networks have made a theory of learning dynamics elusive. In this work, we show that for wide neural networks the learning dynamics simplify considerably and that, in the infinite width limit, they are governed by a linear model obtained from the first-order Taylor expansion of the network around its initial parameters. Furthermore, mirroring the correspondence between wide Bayesian neural networks and Gaussian processes, gradient-based training of wide neural networks with a squared loss produces test set predictions drawn from a Gaussian process with a particular compositional kernel. While these theoretical results are only exact in the infinite width limit, we nevertheless find excellent empirical agreement between the predictions of the original network and those of the linearized version even for finite practically-sized networks. This agreement is robust across different architectures, optimization methods, and loss functions.
Jaehoon Lee 0001, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, Jeffrey Pennington
NeurIPS1
2019 Measuring the Effects of Data Parallelism on Neural Network Training
abstract
Recent hardware developments have dramatically increased the scale of data parallelism available for neural network training. Among the simplest ways to harness next-generation hardware is to increase the batch size in standard mini-batch neural network training algorithms. In this work, we aim to experimentally characterize the effects of increasing the batch size on training time, as measured by the number of steps necessary to reach a goal out-of-sample error. We study how this relationship varies with the training algorithm, model, and data set, and find extremely large variation between workloads. Along the way, we show that disagreements in the literature on how batch size affects model quality can largely be explained by differences in metaparameter tuning and compute budgets at different batch sizes. We find no evidence that larger batch sizes degrade out-of-sample performance. Finally, we discuss the implications of our results on efforts to train neural networks much faster in the future. Our experimental data is publicly available as a database of 71,638,836 loss measurements taken over the course of training for 168,160 individual models across 35 workloads.
Christopher J. Shallue, Jaehoon Lee 0001, Joseph M. Antognini, Jascha Sohl-Dickstein, Roy Frostig, George E. Dahl
J. Mach. Learn. Res.2
2018 Deep Neural Networks as Gaussian Processes
Jaehoon Lee 0001, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, Jascha Sohl-Dickstein
ICLR (Poster)1