VLDB 2026 Research / reviewers in the wild / expert
Jessica Grogan
dblp:317/7017
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 77% Deep learning architectures and training · 19% Language models and text generation · 4% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 50% Mathematical optimization · 50% |
Topics — the 6 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression
structured matrices |
1.2 | 2 | 2023 | Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture · NeurIPS 2023 Monarch: Expressive Structured Matrices for Efficient and Accurate Training · ICML 2022 |
Mathematical optimization
least squares |
0.9 | 1 | 2025 | Towards Learning High-Precision Least Squares Algorithms with Sequence Models · ICLR 2025 |
Algorithms and data structures
numerical algorithms |
0.9 | 1 | 2025 | Towards Learning High-Precision Least Squares Algorithms with Sequence Models · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
efficient architecture design |
0.7 | 1 | 2023 | Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | Monarch: Expressive Structured Matrices for Efficient and Accurate Training · ICML 2022 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.6 | 1 | 2022 | Monarch: Expressive Structured Matrices for Efficient and Accurate Training · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.7polynomial architectures · 1.7linear attention · 1.7high-precision training · 1.7gated convolution · 1.7monarch matrices · 0.7GEMM · 0.7structured matrix approximation · 0.6block-diagonal matrix factorization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Learning High-Precision Least Squares Algorithms with Sequence ModelsabstractThis paper investigates whether sequence models can learn to perform numerical algorithms, e.g. gradient descent, on the fundamental problem of least squares. Our goal is to inherit two properties of standard algorithms from numerical analysis: (1) machine precision, i.e. we want to obtain solutions that are accurate to near floating point error, and (2) numerical generality, i.e. we want them to apply broadly across problem instances. We find that prior approaches using Transformers fail to meet these criteria, and identify limitations present in existing architectures and training procedures. First, we show that softmax Transformers struggle to perform high-precision multiplications, which prevents them from precisely learning numerical algorithms. Second, we identify an alternate class of architectures, comprised entirely of polynomials, that can efficiently represent high-precision gradient descent iterates. Finally, we investigate precision bottlenecks during training and address them via a high-precision training recipe that reduces stochastic gradient noise. Our recipe enables us to train two polynomial architectures, gated convolutions and linear attention, to perform gradient descent iterates on least squares problems. For the first time, we demonstrate the ability to train to near machine precision. Applied iteratively, our models obtain $100,000\times$ lower MSE than standard Transformers trained end-to-end and they incur a $10,000\times$ smaller generalization gap on out-of-distribution problems. We make progress towards end-to-end learning of numerical algorithms for least squares. Jerry W. Liu, Jessica Grogan, Owen Dugan, Ashish Rao, Simran Arora, Atri Rudra, Christopher Ré |
ICLR | 2 |
| 2023 | Monarch Mixer: A Simple Sub-Quadratic GEMM-Based ArchitectureabstractMachine learning models are increasingly being scaled in both sequence length and model dimension to reach longer contexts and better performance. However, existing architectures such as Transformers scale quadratically along both these axes. We ask: are there performant architectures that can scale sub-quadratically along sequence length and model dimension? We introduce Monarch Mixer (M2), a new architecture that uses the same sub-quadratic primitive along both sequence length and model dimension: Monarch matrices, a simple class of expressive structured matrices that captures many linear transforms, achieves high hardware efficiency on GPUs, and scales sub-quadratically. As a proof of concept, we explore the performance of M2 in three domains: non-causal BERT-style language modeling, ViT-style image classification, and causal GPT-style language modeling. For non-causal BERT-style modeling, M2 matches BERT-base and BERT-large in downstream GLUE quality with up to 27% fewer parameters, and achieves up to 9.1$\times$ higher throughput at sequence length 4K. On ImageNet, M2 outperforms ViT-b by 1% in accuracy, with only half the parameters. Causal GPT-style models introduce a technical challenge: enforcing causality via masking introduces a quadratic bottleneck. To alleviate this bottleneck, we develop a novel theoretical view of Monarch matrices based on multivariate polynomial evaluation and interpolation, which lets us parameterize M2 to be causal while remaining sub-quadratic. Using this parameterization, M2 matches GPT-style Transformers at 360M parameters in pretraining perplexity on The PILE—showing for the first time that it may be possible to match Transformer quality without attention or MLPs. Daniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson, Sabri Eyuboglu, Armin W. Thomas, Benjamin Spector, Michael Poli, Atri Rudra, Christopher Ré |
NeurIPS | 3 |
| 2022 | Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingabstractLarge neural networks excel in many domains, but they are expensive to train and fine-tune. A popular approach to reduce their compute or memory requirements is to replace dense weight matrices with structured ones (e.g., sparse, low-rank, Fourier transform). These methods have not seen widespread adoption (1) in end-to-end training due to unfavorable efficiency–quality tradeoffs, and (2) in dense-to-sparse fine-tuning due to lack of tractable algorithms to approximate a given dense weight matrix. To address these issues, we propose a class of matrices (Monarch) that is hardware-efficient (they are parameterized as products of two block-diagonal matrices for better hardware utilization) and expressive (they can represent many commonly used transforms). Surprisingly, the problem of approximating a dense weight matrix with a Monarch matrix, though nonconvex, has an analytical optimal solution. These properties of Monarch matrices unlock new ways to train and fine-tune sparse and dense models. We empirically validate that Monarch can achieve favorable accuracy-efficiency tradeoffs in several end-to-end sparse training applications: speeding up ViT and GPT-2 training on ImageNet classification and Wikitext-103 language modeling by 2x with comparable model quality, and reducing the error on PDE solving and MRI reconstruction tasks by 40%. In sparse-to-dense training, with a simple technique called "reverse sparsification," Monarch matrices serve as a useful intermediate representation to speed up GPT-2 pretraining on OpenWebText by 2x without quality drop. The same technique brings 23% faster BERT pretraining than even the very optimized implementation from Nvidia that set the MLPerf 1.1 record. In dense-to-sparse fine-tuning, as a proof-of-concept, our Monarch approximation algorithm speeds up BERT fine-tuning on GLUE by 1.7x with comparable accuracy. Tri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai, Michael Poli, Jessica Grogan, Aniruddh Rao, Atri Rudra, Christopher Ré |
ICML | 6 |