VLDB 2026 Research / reviewers in the wild / expert
Katie Everett
dblp:270/9991 · also Katie E. Everett
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Deep learning architectures and training · 35% Optimization for machine learning · 29% Probabilistic and Bayesian machine learning · 18% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
scaling laws |
1.6 | 2 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 Scaling Exponents Across Parameterizations and Optimizers · ICML 2024 |
Machine learning › Generative modeling
generative flow networks |
1.3 | 2 | 2023 | GFlowNet-EM for Learning Compositional Latent Variable Models · ICML 2023 GFlowNets and variational inference · ICLR 2023 |
Machine learning › Optimization for machine learning › stochastic gradient descent
stochastic gradient descent with momentum |
0.9 | 1 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.9 | 1 | 2025 | Dimension-adapted Momentum Outscales SGD · NeurIPS 2025 |
Machine learning › Optimization for machine learning › adaptive optimization
adam |
0.8 | 1 | 2024 | Scaling Exponents Across Parameterizations and Optimizers · ICML 2024 |
Machine learning › Deep learning architectures and training
hyperparameter transfer |
0.8 | 1 | 2024 | Scaling Exponents Across Parameterizations and Optimizers · ICML 2024 |
Machine learning › Learning theory › neural network theory › neural network parameterization
maximal update parameterization |
0.8 | 1 | 2024 | Scaling Exponents Across Parameterizations and Optimizers · ICML 2024 |
Machine learning › Deep learning architectures and training › training dynamics
training instability |
0.8 | 1 | 2024 | Small-scale proxies for large-scale Transformer training instabilities · ICLR 2024 |
Machine learning › Deep learning architectures and training › transformer
transformer training |
0.8 | 1 | 2024 | Small-scale proxies for large-scale Transformer training instabilities · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
amortized variational inference |
0.7 | 1 | 2023 | GFlowNet-EM for Learning Compositional Latent Variable Models · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.7 | 1 | 2023 | GFlowNet-EM for Learning Compositional Latent Variable Models · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.7 | 1 | 2023 | GFlowNets and variational inference · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
stochastic gradient descent · 0.9nesterov acceleration · 0.9weight decay · 0.8warm-up · 0.8muparam · 0.8variational inference · 0.7expectation-maximization · 0.7GFlowNets · 0.7GFlowNet sampling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dimension-adapted Momentum Outscales SGDabstractWe investigate scaling laws for stochastic momentum algorithms on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying data-target complexities. While traditional stochastic gradient descent with momentum (SGD-M) yields identical scaling law exponents to SGD, dimension-adapted Nesterov acceleration (DANA) improves these exponents by scaling momentum hyperparameters based on model size and data complexity. This outscaling phenomenon, which also improves compute-optimal scaling behavior, is achieved by DANA across a broad range of data and target complexities, while traditional methods fall short. Extensive experiments on high-dimensional synthetic quadratics validate our theoretical predictions and large-scale text experiments with LSTMs show DANA's improved loss exponents over SGD hold in a practical setting. Damien Ferbach, Katie Everett, Gauthier Gidel, Elliot Paquette, Courtney Paquette |
NeurIPS | 2 |
| 2024 | Small-scale proxies for large-scale Transformer training instabilitiesabstractTeams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smaller scales. Although the causes of such instabilities are of scientific interest, the amount of resources required to reproduce them has made investigation difficult. In this work, we seek ways to reproduce and study training instability at smaller scales. First, we focus on two sources of training instability described in previous work: the growth of logits in attention layers (Dehghani et al., 2023) and divergence of the output logits from the log probabilities (Chowdhery et al., 2022). By measuring the relationship between learning rate and loss across scales, we show that these instabilities also appear in small models when training at high learning rates, and that mitigations previously employed at large scales are equally effective in this regime. This prompts us to investigate the extent to which other known optimizer and model interventions influence the sensitivity of the final loss to changes in the learning rate. To this end, we study methods such as warm-up, weight decay, and the MuParam (Yang et al., 2022), and combine techniques to train small models that achieve similar losses across orders of magnitude of learning rate variation. Finally, to conclude our exploration we study two cases where instabilities can be predicted before they emerge by examining the scaling behavior of model characteristics such as activation and gradient norms. Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie Everett, Alexander A. Alemi, Ben Adlam, John D. Co-Reyes, Izzeddin Gur, Roman Novak, Jeffrey Pennington, Jascha Sohl-Dickstein, Kelvin Xu, Jaehoon Lee 0001, Justin Gilmer, Simon Kornblith |
ICLR | 4 |
| 2024 | Scaling Exponents Across Parameterizations and OptimizersabstractRobust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization and optimizer choices. In this work, we propose a new perspective on parameterization by investigating a key assumption in prior work about the alignment between parameters and data and derive new theoretical results under weaker assumptions and a broader set of optimizers. Our extensive empirical investigation includes *tens of thousands* of models trained with *all combinations of* three optimizers, four parameterizations, several alignment assumptions, more than a dozen learning rates, and fourteen model sizes up to 27B parameters. We find that the best learning rate scaling prescription would often have been excluded by the assumptions in prior work. Our results show that all parameterizations, not just maximal update parameterization (muP), can achieve hyperparameter transfer; moreover, our novel per-layer learning rate prescription for standard parameterization outperforms muP. Finally, we demonstrate that an overlooked aspect of parameterization, the epsilon parameter in Adam, must be scaled correctly to avoid gradient underflow and propose *Adam-atan2*, a new numerically stable, scale-invariant version of Adam that eliminates the epsilon hyperparameter entirely. Katie Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi, Roman Novak, Peter J. Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee 0001, Jeffrey Pennington |
ICML | 1 |
| 2023 | GFlowNets and variational inference
Nikolay Malkin, Salem Lahlou, Tristan Deleu, Edward J. Hu, Katie Everett, Dinghuai Zhang, Yoshua Bengio |
ICLR | 6 |
| 2023 | GFlowNet-EM for Learning Compositional Latent Variable ModelsabstractLatent variable models (LVMs) with discrete compositional latents are an important but challenging setting due to a combinatorially large number of possible configurations of the latents. A key tradeoff in modeling the posteriors over latents is between expressivity and tractable optimization. For algorithms based on expectation-maximization (EM), the E-step is often intractable without restrictive approximations to the posterior. We propose the use of GFlowNets, algorithms for sampling from an unnormalized density by learning a stochastic policy for sequential construction of samples, for this intractable E-step. By training GFlowNets to sample from the posterior over latents, we take advantage of their strengths as amortized variational inference algorithms for complex distributions over discrete structures. Our approach, GFlowNet-EM, enables the training of expressive LVMs with discrete compositional latents, as shown by experiments on non-context-free grammar induction and on images using discrete variational autoencoders (VAEs) without conditional independence enforced in the encoder. Edward J. Hu, Nikolay Malkin, Moksh Jain, Katie Everett, Alexandros Graikos, Yoshua Bengio |
ICML | 4 |