EDBT 2026 Demo / reviewers in the wild / expert
Mansheej Paul
dblp:277/6622
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Efficient and distributed learning · 55% Deep learning architectures and training · 21% Language models and text generation · 9% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis |
1.2 | 2 | 2023 | Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023 Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › model compression
pruning |
1.2 | 2 | 2023 | Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023 Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022 |
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry |
1.0 | 2 | 2022 | Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022 Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
training dynamics |
1.0 | 2 | 2022 | Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022 Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › data selection
data pruning |
0.9 | 1 | 2025 | Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models · ICLR 2025 |
Machine learning › Efficient and distributed learning › low-precision training
FP8 training |
0.9 | 1 | 2025 | µnit Scaling: Simple and Scalable FP8 LLM Training · ICML 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | µnit Scaling: Simple and Scalable FP8 LLM Training · ICML 2025 |
Machine learning › Efficient and distributed learning
low-precision training |
0.9 | 1 | 2025 | µnit Scaling: Simple and Scalable FP8 LLM Training · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.9 | 1 | 2025 | Scaling Laws for Precision · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Scaling Laws for Precision · ICLR 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | Scaling Laws for Precision · ICLR 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.7 | 1 | 2023 | Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023 |
Machine learning › Efficient and distributed learning › model compression › pruning
iterative magnitude pruning |
0.6 | 1 | 2022 | Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022 |
Machine learning › Efficient and distributed learning
data selection |
0.5 | 1 | 2021 | Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
noisy label detection |
0.5 | 1 | 2021 | Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › data-centric deep learning
training data pruning |
0.5 | 1 | 2021 | Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.4 | 1 | 2020 | Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining |
0.3 | 1 | 2025 | Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.2 | 1 | 2023 | Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.2 | 1 | 2023 | Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.1 | 1 | 2020 | Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
unit scaling · 0.9small reference models · 0.9scaling laws · 0.9perplexity-based pruning · 0.9FP8 quantization · 0.9linear mode connectivity · 0.6gradient norm scoring · 0.5error l2-norm scoring · 0.5loss landscape measures · 0.4NTK · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference ModelsabstractIn this work, we investigate whether small language models can determine high-quality subsets of large-scale text datasets that improve the performance of larger language models. While existing work has shown that pruning based on the perplexity of a larger model can yield high-quality data, we investigate whether smaller models can be used for perplexity-based pruning and how pruning is affected by the domain composition of the data being pruned. We demonstrate that for multiple dataset compositions, perplexity-based pruning of pretraining data can significantly improve downstream task performance: pruning based on perplexities computed with a 125 million parameter model improves the average performance on downstream tasks of a 3 billion parameter model by up to 2.04 and achieves up to a 1.45× reduction in pretraining steps to reach commensurate baseline performance. Furthermore, we demonstrate that such perplexity-based data pruning also yields downstream performance gains in the over-trained and data-constrained regimes. Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L. Leavitt, Mansheej Paul |
ICLR | 6 |
| 2025 | Scaling Laws for PrecisionabstractLow precision training and inference affect both the quality and cost of language models, but current scaling laws do not account for this. In this work, we devise "precision-aware" scaling laws for both training and inference. We propose that training in lower precision reduces the model's "effective parameter count," allowing us to predict the additional loss incurred from training in low precision and post-train quantization. For inference, we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful. For training, our scaling laws allow us to predict the loss of a model with different parts in different precisions, and suggest that training larger models in lower precision can be compute optimal. We unify the scaling laws for post and pretraining quantization to arrive at a single functional form that predicts degradation from training and inference in varied precisions. We fit on over 465 pretraining runs and validate our predictions on model sizes up to 1.7B parameters trained on up to 26B tokens. Tanishq Kumar, Zachary Ankner, Benjamin Spector, Blake Bordelon, Niklas Muennighoff, Mansheej Paul, Cengiz Pehlevan, Christopher Ré, Aditi Raghunathan |
ICLR | 6 |
| 2025 | µnit Scaling: Simple and Scalable FP8 LLM TrainingabstractLarge language model training with 8-bit floating point (FP8) formats promises significant efficiency improvements, but reduced numerical precision makes training challenging. It is currently possible to train in FP8 only if one is willing to tune various hyperparameters, reduce model scale, or accept the overhead of computing dynamic scale factors. We demonstrate simple, scalable FP8 training that requires no dynamic scaling factors or special hyperparameters, even at large model sizes. Our method, \textit{µnit Scaling (µS)}, also enables simple hyperparameter transfer across model widths, matched numerics across training and inference, and other desirable properties. µnit Scaling is straightforward to implement, consisting of a set of minimal interventions based on a first-principles analysis of transformer operations. We validate our method by training models with parameters ranging from 1B to 13B, performing all hidden linear layer computations in FP8. We achieve quality equal to higher-precision baselines while also training up to 33% faster. Saaketh Narayan, Abhay Gupta, Mansheej Paul, Davis W. Blalock |
ICML | 3 |
| 2023 | Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask?
Mansheej Paul, Feng Chen 0046, Brett W. Larsen, Jonathan Frankle, Surya Ganguli, Gintare Karolina Dziugaite |
ICLR | 1 |
| 2023 | Pretraining task diversity and the emergence of non-Bayesian in-context learning for regressionabstractPretrained transformers exhibit the remarkable ability of in-context learning (ICL): they can learn tasks from just a few examples provided in the prompt without updating any weights. This raises a foundational question: can ICL solve fundamentally _new_ tasks that are very different from those seen during pretraining? To probe this question, we examine ICL’s performance on linear regression while varying the diversity of tasks in the pretraining dataset. We empirically demonstrate a _task diversity threshold_ for the emergence of ICL. Below this threshold, the pretrained transformer cannot solve unseen regression tasks, instead behaving like a Bayesian estimator with the _non-diverse pretraining task distribution_ as the prior. Beyond this threshold, the transformer significantly outperforms this estimator; its behavior aligns with that of ridge regression, corresponding to a Gaussian prior over _all tasks_, including those not seen during pretraining. Thus, when pretrained on data with task diversity greater than the threshold, transformers _can_ optimally solve fundamentally new tasks in-context. Importantly, this capability hinges on it deviating from the Bayes optimal estimator with the pretraining distribution as the prior. This study also explores the effect of regularization, model capacity and task structure and underscores, in a concrete example, the critical role of task diversity, alongside data and model scale, in the emergence of ICL. Allan Raventós, Mansheej Paul, Feng Chen 0046, Surya Ganguli |
NeurIPS | 2 |
| 2022 | Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable NetworksabstractA striking observation about iterative magnitude pruning (IMP; Frankle et al. 2020) is that—after just a few hundred steps of dense training—the method can find a sparse sub-network that can be trained to the same accuracy as the dense network. However, the same does not hold at step 0, i.e. random initialization. In this work, we seek to understand how this early phase of pre-training leads to a good initialization for IMP both through the lens of the data distribution and the loss landscape geometry. Empirically we observe that, holding the number of pre-training iterations constant, training on a small fraction of (randomly chosen) data suffices to obtain an equally good initialization for IMP. We additionally observe that by pre-training only on "easy" training data, we can decrease the number of steps necessary to find a good initialization for IMP compared to training on the full dataset or a randomly chosen subset. Finally, we identify novel properties of the loss landscape of dense networks that are predictive of IMP performance, showing in particular that more examples being linearly mode connected in the dense network correlates well with good initializations for IMP. Combined, these results provide new insight into the role played by the early phase training in IMP. Mansheej Paul, Brett W. Larsen, Surya Ganguli, Jonathan Frankle, Gintare Karolina Dziugaite |
NeurIPS | 1 |
| 2021 | Deep Learning on a Data Diet: Finding Important Examples Early in TrainingabstractRecent success in deep learning has partially been driven by training increasingly overparametrized networks on ever larger datasets. It is therefore natural to ask: how much of the data is superfluous, which examples are important for generalization, and how do we find them? In this work, we make the striking observation that, in standard vision datasets, simple scores averaged over several weight initializations can be used to identify important examples very early in training. We propose two such scores—the Gradient Normed (GraNd) and the Error L2-Norm (EL2N) scores—and demonstrate their efficacy on a range of architectures and datasets by pruning significant fractions of training data without sacrificing test accuracy. In fact, using EL2N scores calculated a few epochs into training, we can prune half of the CIFAR10 training set while slightly improving test accuracy. Furthermore, for a given dataset, EL2N scores from one architecture or hyperparameter configuration generalize to other configurations. Compared to recent work that prunes data by discarding examples that are rarely forgotten over the course of training, our scores use only local information early in training. We also use our scores to detect noisy examples and study training dynamics through the lens of important examples—we investigate how the data distribution shapes the loss surface and identify subspaces of the model’s data representation that are relatively stable over training. Mansheej Paul, Surya Ganguli, Gintare Karolina Dziugaite |
NeurIPS | 1 |
| 2020 | Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelabstractIn suitably initialized wide networks, small learning rates transform deep neural networks (DNNs) into neural tangent kernel (NTK) machines, whose training dynamics is well-approximated by a linear weight expansion of the network at initialization. Standard training, however, diverges from its linearization in ways that are poorly understood. We study the relationship between the training dynamics of nonlinear deep networks, the geometry of the loss landscape, and the time evolution of a data-dependent NTK. We do so through a large-scale phenomenological analysis of training, synthesizing diverse measures characterizing loss landscape geometry and NTK dynamics. In multiple neural architectures and datasets, we find these diverse measures evolve in a highly correlated manner, revealing a universal picture of the deep learning process. In this picture, deep network training exhibits a highly chaotic rapid initial transient that within 2 to 3 epochs determines the final linearly connected basin of low loss containing the end point of training. During this chaotic transient, the NTK changes rapidly, learning useful features from the training data that enables it to outperform the standard initial NTK by a factor of 3 in less than 3 to 4 epochs. After this rapid chaotic transient, the NTK changes at constant velocity, and its performance matches that of full network training in 15\% to 45\% of training time. Overall, our analysis reveals a striking correlation between a diverse set of metrics over training time, governed by a rapid chaotic to stable transition in the first few epochs, that together poses challenges and opportunities for the development of more accurate theories of deep learning. Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M. Roy 0001, Surya Ganguli |
NeurIPS | 3 |