Mansheej Paul

dblp:277/6622 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 55% Deep learning architectures and training · 21% Language models and text generation · 9%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis
1.222023
Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023
Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression
pruning
1.222023
Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023
Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry
1.022022
Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020
Machine learning › Deep learning architectures and training
training dynamics
1.022022
Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020
Machine learning › Efficient and distributed learning › data selection
data pruning
0.912025
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models · ICLR 2025
Machine learning › Efficient and distributed learning › low-precision training
FP8 training
0.912025
µnit Scaling: Simple and Scalable FP8 LLM Training · ICML 2025
Natural language and speech › Language models and text generation
large language model training
0.912025
µnit Scaling: Simple and Scalable FP8 LLM Training · ICML 2025
Machine learning › Efficient and distributed learning
low-precision training
0.912025
µnit Scaling: Simple and Scalable FP8 LLM Training · ICML 2025
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.912025
Scaling Laws for Precision · ICLR 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
Scaling Laws for Precision · ICLR 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Scaling Laws for Precision · ICLR 2025
Natural language and speech › Language models and text generation
in-context learning
0.712023
Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression · NeurIPS 2023
Machine learning › Efficient and distributed learning
model compression
0.712023
Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023
Machine learning › Efficient and distributed learning › model compression › pruning
iterative magnitude pruning
0.612022
Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks · NeurIPS 2022
Machine learning › Efficient and distributed learning
data selection
0.512021
Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
noisy label detection
0.512021
Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021
Machine learning › Deep learning architectures and training › data-centric deep learning
training data pruning
0.512021
Deep Learning on a Data Diet: Finding Important Examples Early in Training · NeurIPS 2021
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.412020
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining
0.312025
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.212023
Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.212023
Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? · ICLR 2023
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.112020
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

unit scaling · 0.9small reference models · 0.9scaling laws · 0.9perplexity-based pruning · 0.9FP8 quantization · 0.9linear mode connectivity · 0.6gradient norm scoring · 0.5error l2-norm scoring · 0.5loss landscape measures · 0.4NTK · 0.4
YearPublicationVenuePosition
2025 Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
abstract
In this work, we investigate whether small language models can determine high-quality subsets of large-scale text datasets that improve the performance of larger language models. While existing work has shown that pruning based on the perplexity of a larger model can yield high-quality data, we investigate whether smaller models can be used for perplexity-based pruning and how pruning is affected by the domain composition of the data being pruned. We demonstrate that for multiple dataset compositions, perplexity-based pruning of pretraining data can significantly improve downstream task performance: pruning based on perplexities computed with a 125 million parameter model improves the average performance on downstream tasks of a 3 billion parameter model by up to 2.04 and achieves up to a 1.45× reduction in pretraining steps to reach commensurate baseline performance. Furthermore, we demonstrate that such perplexity-based data pruning also yields downstream performance gains in the over-trained and data-constrained regimes.
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L. Leavitt, Mansheej Paul
ICLR6
2025 Scaling Laws for Precision
abstract
Low precision training and inference affect both the quality and cost of language models, but current scaling laws do not account for this. In this work, we devise "precision-aware" scaling laws for both training and inference. We propose that training in lower precision reduces the model's "effective parameter count," allowing us to predict the additional loss incurred from training in low precision and post-train quantization. For inference, we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful. For training, our scaling laws allow us to predict the loss of a model with different parts in different precisions, and suggest that training larger models in lower precision can be compute optimal. We unify the scaling laws for post and pretraining quantization to arrive at a single functional form that predicts degradation from training and inference in varied precisions. We fit on over 465 pretraining runs and validate our predictions on model sizes up to 1.7B parameters trained on up to 26B tokens.
Tanishq Kumar, Zachary Ankner, Benjamin Spector, Blake Bordelon, Niklas Muennighoff, Mansheej Paul, Cengiz Pehlevan, Christopher Ré, Aditi Raghunathan
ICLR6
2025 µnit Scaling: Simple and Scalable FP8 LLM Training
abstract
Large language model training with 8-bit floating point (FP8) formats promises significant efficiency improvements, but reduced numerical precision makes training challenging. It is currently possible to train in FP8 only if one is willing to tune various hyperparameters, reduce model scale, or accept the overhead of computing dynamic scale factors. We demonstrate simple, scalable FP8 training that requires no dynamic scaling factors or special hyperparameters, even at large model sizes. Our method, \textit{µnit Scaling (µS)}, also enables simple hyperparameter transfer across model widths, matched numerics across training and inference, and other desirable properties. µnit Scaling is straightforward to implement, consisting of a set of minimal interventions based on a first-principles analysis of transformer operations. We validate our method by training models with parameters ranging from 1B to 13B, performing all hidden linear layer computations in FP8. We achieve quality equal to higher-precision baselines while also training up to 33% faster.
Saaketh Narayan, Abhay Gupta, Mansheej Paul, Davis W. Blalock
ICML3
2023 Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask?
Mansheej Paul, Feng Chen 0046, Brett W. Larsen, Jonathan Frankle, Surya Ganguli, Gintare Karolina Dziugaite
ICLR1
2023 Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression
abstract
Pretrained transformers exhibit the remarkable ability of in-context learning (ICL): they can learn tasks from just a few examples provided in the prompt without updating any weights. This raises a foundational question: can ICL solve fundamentally _new_ tasks that are very different from those seen during pretraining? To probe this question, we examine ICL’s performance on linear regression while varying the diversity of tasks in the pretraining dataset. We empirically demonstrate a _task diversity threshold_ for the emergence of ICL. Below this threshold, the pretrained transformer cannot solve unseen regression tasks, instead behaving like a Bayesian estimator with the _non-diverse pretraining task distribution_ as the prior. Beyond this threshold, the transformer significantly outperforms this estimator; its behavior aligns with that of ridge regression, corresponding to a Gaussian prior over _all tasks_, including those not seen during pretraining. Thus, when pretrained on data with task diversity greater than the threshold, transformers _can_ optimally solve fundamentally new tasks in-context. Importantly, this capability hinges on it deviating from the Bayes optimal estimator with the pretraining distribution as the prior. This study also explores the effect of regularization, model capacity and task structure and underscores, in a concrete example, the critical role of task diversity, alongside data and model scale, in the emergence of ICL.
Allan Raventós, Mansheej Paul, Feng Chen 0046, Surya Ganguli
NeurIPS2
2022 Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks
abstract
A striking observation about iterative magnitude pruning (IMP; Frankle et al. 2020) is that—after just a few hundred steps of dense training—the method can find a sparse sub-network that can be trained to the same accuracy as the dense network. However, the same does not hold at step 0, i.e. random initialization. In this work, we seek to understand how this early phase of pre-training leads to a good initialization for IMP both through the lens of the data distribution and the loss landscape geometry. Empirically we observe that, holding the number of pre-training iterations constant, training on a small fraction of (randomly chosen) data suffices to obtain an equally good initialization for IMP. We additionally observe that by pre-training only on "easy" training data, we can decrease the number of steps necessary to find a good initialization for IMP compared to training on the full dataset or a randomly chosen subset. Finally, we identify novel properties of the loss landscape of dense networks that are predictive of IMP performance, showing in particular that more examples being linearly mode connected in the dense network correlates well with good initializations for IMP. Combined, these results provide new insight into the role played by the early phase training in IMP.
Mansheej Paul, Brett W. Larsen, Surya Ganguli, Jonathan Frankle, Gintare Karolina Dziugaite
NeurIPS1
2021 Deep Learning on a Data Diet: Finding Important Examples Early in Training
abstract
Recent success in deep learning has partially been driven by training increasingly overparametrized networks on ever larger datasets. It is therefore natural to ask: how much of the data is superfluous, which examples are important for generalization, and how do we find them? In this work, we make the striking observation that, in standard vision datasets, simple scores averaged over several weight initializations can be used to identify important examples very early in training. We propose two such scores—the Gradient Normed (GraNd) and the Error L2-Norm (EL2N) scores—and demonstrate their efficacy on a range of architectures and datasets by pruning significant fractions of training data without sacrificing test accuracy. In fact, using EL2N scores calculated a few epochs into training, we can prune half of the CIFAR10 training set while slightly improving test accuracy. Furthermore, for a given dataset, EL2N scores from one architecture or hyperparameter configuration generalize to other configurations. Compared to recent work that prunes data by discarding examples that are rarely forgotten over the course of training, our scores use only local information early in training. We also use our scores to detect noisy examples and study training dynamics through the lens of important examples—we investigate how the data distribution shapes the loss surface and identify subspaces of the model’s data representation that are relatively stable over training.
Mansheej Paul, Surya Ganguli, Gintare Karolina Dziugaite
NeurIPS1
2020 Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
abstract
In suitably initialized wide networks, small learning rates transform deep neural networks (DNNs) into neural tangent kernel (NTK) machines, whose training dynamics is well-approximated by a linear weight expansion of the network at initialization. Standard training, however, diverges from its linearization in ways that are poorly understood. We study the relationship between the training dynamics of nonlinear deep networks, the geometry of the loss landscape, and the time evolution of a data-dependent NTK. We do so through a large-scale phenomenological analysis of training, synthesizing diverse measures characterizing loss landscape geometry and NTK dynamics. In multiple neural architectures and datasets, we find these diverse measures evolve in a highly correlated manner, revealing a universal picture of the deep learning process. In this picture, deep network training exhibits a highly chaotic rapid initial transient that within 2 to 3 epochs determines the final linearly connected basin of low loss containing the end point of training. During this chaotic transient, the NTK changes rapidly, learning useful features from the training data that enables it to outperform the standard initial NTK by a factor of 3 in less than 3 to 4 epochs. After this rapid chaotic transient, the NTK changes at constant velocity, and its performance matches that of full network training in 15\% to 45\% of training time. Overall, our analysis reveals a striking correlation between a diverse set of metrics over training time, governed by a rapid chaotic to stable transition in the first few epochs, that together poses challenges and opportunities for the development of more accurate theories of deep learning.
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M. Roy 0001, Surya Ganguli
NeurIPS3