Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rudrajit Das

dblp:227/2712 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-9818-0518ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Efficient and distributed learning · 18% Optimization for machine learning · 16% Transfer learning and domain adaptation · 16%
Network and information security
2 papers
Privacy and data protection · 100%

Topics — the 26 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › model adaptation
model retraining
1.722025
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing · NeurIPS 2025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
1.522025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025
Understanding Self-Distillation in the Presence of Label Noise · ICML 2023
Privacy and data protection
differential privacy
1.522025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025
Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023
Machine learning › Learning theory › classification
binary classification
1.122025
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing · NeurIPS 2025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025
Machine learning › Graph learning › graph neural network › message passing
approximate message passing
0.912025
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing · NeurIPS 2025
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.912025
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting · ICML 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.912025
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting · ICML 2025
Machine learning › Deep learning architectures and training › data-centric deep learning
sample weighting
0.912025
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting · ICML 2025
Privacy and data protection › differential privacy › relaxed differential privacy
label differential privacy
0.912025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling
0.812024
Understanding the Training Speedup from Sampling with Approximate Losses · ICML 2024
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
sample selection
0.812024
Understanding the Training Speedup from Sampling with Approximate Losses · ICML 2024
Machine learning › Optimization for machine learning
stochastic optimization
0.812024
Understanding the Training Speedup from Sampling with Approximate Losses · ICML 2024
Machine learning › Optimization for machine learning
gradient clipping
0.712023
Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation
0.712023
Understanding Self-Distillation in the Presence of Label Noise · ICML 2023
Machine learning › Optimization for machine learning
stochastic gradient descent
0.712023
Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023
Machine learning › Learning paradigms
supervised learning
0.712023
Understanding Self-Distillation in the Presence of Label Noise · ICML 2023
Privacy and data protection › differential privacy › differentially private deep learning
DP-SGD
0.712023
Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.612022
On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Efficient and distributed learning › federated learning
decentralized federated learning
0.612022
On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Efficient and distributed learning
federated learning
0.612022
On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.612022
On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Learning theory
linear separability
0.312025
Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025
Machine learning › Optimization for machine learning
convergence analysis
0.212023
Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023
Machine learning › Learning theory › statistical estimation › regularized estimation
regularized linear regression
0.212023
Understanding Self-Distillation in the Presence of Label Noise · ICML 2023
Machine learning › Optimization for machine learning
non-convex optimization
0.212022
On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Optimization for machine learning › non-convex optimization
polyak-łojasiewicz condition
0.212022
On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

predicted hard labels · 1.7consensus-based retraining · 1.7theoretical analysis · 0.9loss-based sample weighting · 0.9generalized linear model · 0.9gaussian mixture model · 0.9approximate message passing · 0.9greedy sample selection · 0.8early exiting · 0.8convex optimization · 0.8per-sample lipschitz constants · 0.7gradient clipping · 0.7DP-SGD · 0.7
YearPublicationVenuePosition
2025 Retraining with Predicted Hard Labels Provably Increases Model Accuracy
abstract
The performance of a model trained with noisy labels is often improved by simply *retraining* the model with its *own predicted hard labels* (i.e., $1$/$0$ labels). Yet, a detailed theoretical characterization of this phenomenon is lacking. In this paper, we theoretically analyze retraining in a linearly separable binary classification setting with randomly corrupted labels given to us and prove that retraining can improve the population accuracy obtained by initially training with the given (noisy) labels. To the best of our knowledge, this is the first such theoretical result. Retraining finds application in improving training with local label differential privacy (DP), which involves training with noisy labels. We empirically show that retraining selectively on the samples for which the predicted label matches the given label significantly improves label DP training at no extra privacy cost; we call this consensus-based retraining. For example, when training ResNet-18 on CIFAR-100 with $\epsilon=3$ label DP, we obtain more than $6$% improvement in accuracy with consensus-based retraining.
Rudrajit Das, Inderjit S. Dhillon, Alessandro Epasto, Adel Javanmard, Jieming Mao, Vahab S. Mirrokni, Sujay Sanghavi, Peilin Zhong
ICML1
2025 Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting
abstract
Fine-tuning a pre-trained model on a downstream task often degrades its original capabilities, a phenomenon known as "catastrophic forgetting". This is especially an issue when one does not have access to the data and recipe used to develop the pre-trained model. Under this constraint, most existing methods for mitigating forgetting are inapplicable. To address this challenge, we propose a sample weighting scheme for the fine-tuning data solely based on the pre-trained model’s losses. Specifically, we upweight the easy samples on which the pre-trained model’s loss is low and vice versa to limit the drift from the pre-trained model. Our approach is orthogonal and yet complementary to existing methods; while such methods mostly operate on parameter or gradient space, we concentrate on the sample space. We theoretically analyze the impact of fine-tuning with our method in a linear setting, showing that it stalls learning in a certain subspace, which inhibits overfitting to the target task. We empirically demonstrate the efficacy of our method on both language and vision tasks. As an example, when fine-tuning Gemma 2 2B on MetaMathQA, our method results in only a $0.8$% drop in accuracy on GSM8K (another math dataset) compared to standard fine-tuning, while preserving $5.4$% more accuracy on the pre-training datasets.
Sunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis, Sujay Sanghavi
ICML3
2025 Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
abstract
Retraining a model using its own predictions together with the original, potentially noisy labels is a well-known strategy for improving the model’s performance. While prior works have demonstrated the benefits of specific heuristic retraining schemes, the question of how to optimally combine the model's predictions and the provided labels remains largely open. This paper addresses this fundamental question for binary classification tasks. We develop a principled framework based on approximate message passing (AMP) to analyze iterative retraining procedures for two ground truth settings: Gaussian mixture model (GMM) and generalized linear model (GLM). Our main contribution is the derivation of the Bayes optimal aggregator function to combine the current model's predictions and the given labels, which when used to retrain the same model, minimizes its prediction error. We also quantify the performance of this optimal retraining strategy over multiple rounds. We complement our theoretical results by proposing a practically usable version of the theoretically-optimal aggregator function and demonstrate its superiority over baseline methods under different label noise models.
Adel Javanmard, Rudrajit Das, Alessandro Epasto, Vahab S. Mirrokni
NeurIPS2
2024 Understanding the Training Speedup from Sampling with Approximate Losses
abstract
It is well known that selecting samples with large losses/gradients can significantly reduce the number of training steps. However, the selection overhead is often too high to yield any meaningful gains in terms of overall training time. In this work, we focus on the greedy approach of selecting samples with large approximate losses instead of exact losses in order to reduce the selection overhead. For smooth convex losses, we show that such a greedy strategy can converge to a constant factor of the minimum value of the average loss in fewer iterations than the standard approach of random selection. We also theoretically quantify the effect of the approximation level. We then develop SIFT which uses early exiting to obtain approximate losses with an intermediate layer’s representations for sample selection. We evaluate SIFT on the task of training a 110M parameter 12 layer BERT base model, and show significant gains (in terms of training hours and number of backpropagation steps) without any optimized implementation over vanilla training. For e.g., to reach 64% validation accuracy, SIFT with exit at the first layer takes $\sim$ 43 hours compared to $\sim$ 57 hours of vanilla training.
Rudrajit Das, Bertram Ieong, Parikshit Bansal, Sujay Sanghavi
ICML1
2023 Beyond Uniform Lipschitz Condition in Differentially Private Optimization
abstract
Most prior results on differentially private stochastic gradient descent (DP-SGD) are derived under the simplistic assumption of uniform Lipschitzness, i.e., the per-sample gradients are uniformly bounded. We generalize uniform Lipschitzness by assuming that the per-sample gradients have sample-dependent upper bounds, i.e., per-sample Lipschitz constants, which themselves may be unbounded. We provide principled guidance on choosing the clip norm in DP-SGD for convex over-parameterized settings satisfying our general version of Lipschitzness when the per-sample Lipschitz constants are bounded; specifically, we recommend tuning the clip norm only till values up to the minimum per-sample Lipschitz constant. This finds application in the private training of a softmax layer on top of a deep network pre-trained on public data. We verify the efficacy of our recommendation via experiments on 8 datasets. Furthermore, we provide new convergence results for DP-SGD on convex and nonconvex functions when the Lipschitz constants are unbounded but have bounded moments, i.e., they are heavy-tailed.
Rudrajit Das, Satyen Kale, Zheng Xu 0002, Tong Zhang 0001, Sujay Sanghavi
ICML1
2023 Understanding Self-Distillation in the Presence of Label Noise
abstract
Self-distillation (SD) is the process of first training a "teacher" model and then using its predictions to train a "student" model that has the *same* architecture. Specifically, the student's loss is $\big(\xi*\ell(\text{teacher's predictions}, \text{ student's predictions}) + (1-\xi)*\ell(\text{given labels}, \text{ student's predictions})\big)$, where $\ell$ is the loss function and $\xi$ is some parameter $\in [0,1]$. SD has been empirically observed to provide performance gains in several settings. In this paper, we theoretically characterize the effect of SD in two supervised learning problems with *noisy labels*. We first analyze SD for regularized linear regression and show that in the high label noise regime, the optimal value of $\xi$ that minimizes the expected error in estimating the ground truth parameter is surprisingly greater than 1. Empirically, we show that $\xi > 1$ works better than $\xi \leq 1$ even with the cross-entropy loss for several classification datasets when 50% or 30% of the labels are corrupted. Further, we quantify when optimal SD is better than optimal regularization. Next, we analyze SD in the case of logistic regression for binary classification with random label corruption and quantify the range of label corruption in which the student outperforms the teacher (w.r.t. accuracy). To our knowledge, this is the first result of its kind for the cross-entropy loss.
Rudrajit Das, Sujay Sanghavi
ICML1
2022 Faster non-convex federated learning via global and local momentum
abstract
We propose \texttt{FedGLOMO}, a novel federated learning (FL) algorithm with an iteration complexity of $\mathcal{O}(\epsilon^{-1.5})$ to converge to an $\epsilon$-stationary point (i.e., $\mathbb{E}[\|\nabla f(x)\|^2] \leq \epsilon$) for smooth non-convex functions – under arbitrary client heterogeneity and compressed communication – compared to the $\mathcal{O}(\epsilon^{-2})$ complexity of most prior works. Our key algorithmic idea that enables achieving this improved complexity is based on the observation that the convergence in FL is hampered by two sources of high variance: (i) the global server aggregation step with multiple local updates, exacerbated by client heterogeneity, and (ii) the noise of the local client-level stochastic gradients. The first issue is particularly detrimental to FL algorithms that perform plain averaging at the server. By modeling the server aggregation step as a generalized gradient-type update, we propose a variance-reducing momentum-based global update at the server, which when applied in conjunction with variance-reduced local updates at the clients, enables \texttt{FedGLOMO} to enjoy an improved convergence rate. Our experiments illustrate the intrinsic variance reduction effect of \texttt{FedGLOMO}, which implicitly suppresses client-drift in heterogeneous data distribution settings and promotes communication efficiency.
Rudrajit Das, Anish Acharya, Abolfazl Hashemi, Sujay Sanghavi, Inderjit S. Dhillon, Ufuk Topcu
UAI1
2022 On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning
abstract
Federated learning (FL) is an emerging collaborative machine learning (ML) framework that enables training of predictive models in a distributed fashion where the communication among the participating nodes are facilitated by a central server. To deal with the communication bottleneck at the server, decentralized FL (DFL) methods advocate rely on local communication of nodes with their neighbors according to a specific communication network. In DFL, it is common algorithmic practice to have nodes interleave (local) gradient descent iterations with gossip (i.e., averaging over the network) steps. As the size of the ML models grows, the limited communication bandwidth among the nodes does not permit communication of full-precision messages; hence, it is becoming increasingly common to require that messages belossy, compressedversions of the local parameters. The requirement of communicating compressed messages gives rise to the important question:given a fixed communication budget, what should be our communication strategy to minimize the (training) loss as much as possible?In this article, we explore this direction, and show that in such compressed DFL settings, there are benefits to havingmultiplegossip steps between subsequent gradient iterations, even when the cost of doing so is appropriately accounted for, e.g., by means of reducing the precision of compressed information. In particular, we show that having${\mathcal O}(\log \frac{1}{\epsilon })$gradient iterations with constant step size - and${\mathcal O}(\log \frac{1}{\epsilon })$gossip steps between every pair of these iterations - enables convergence to within$\epsilon$of the optimal value for a class of non-convex problems that arise in the training of deep learning models, namely, smooth non-convex objectives satisfying Polyak-Łojasiewicz condition. Empirically, we show that our proposed scheme bridges the gap between centralized gradient descent and DFL on various machine learning tasks across different network topologies and compression operators.
Abolfazl Hashemi, Anish Acharya, Rudrajit Das, Haris Vikalo, Sujay Sanghavi, Inderjit S. Dhillon
IEEE Trans. Parallel Distributed Syst.3
2019 Nonlinear Blind Compressed Sensing Under Signal-Dependent Noise
abstract
In this paper, we consider the problem of nonlinear blind compressed sensing, i.e. jointly estimating the sparse codes and sparsity-promoting basis, under signal-dependent noise. We focus our efforts on the Poisson noise model, though other signal-dependent noise models can be considered. By employing a well-known variance stabilizing transform such as the Anscombe transform, we formulate our task as a nonlinear least squares problem with the ℓ1penalty imposed for promoting sparsity. We solve this objective function under non-negativity constraints imposed on both the sparse codes and the basis. To this end, we propose a multiplicative update rule, similar to that used in non-negative matrix factorization (NMF), for our alternating minimization algorithm. To the best of our knowledge, this is the first attempt at a formulation for nonlinear blind compressed sensing, with and without the Poisson noise model. Further, we also provide some theoretical bounds on the performance of our algorithm.
Rudrajit Das, Ajit Rajwade 0001
ICIP1
2018 Sparse Kernel PCA for Outlier Detection
abstract
In this paper, we propose a new method to perform Sparse Kernel Principal Component Analysis (SKPCA) and also mathematically analyze the validity of SKPCA. We formulate SKPCA as a constrained optimization problem with elastic net regularization in kernel feature space and solve it. We consider outlier detection (where KPCA is employed) as an application for SKPCA, using the RBF kernel. We test it on 5 real world datasets and show that by using just 4% (or even less) of the principal components (PCs), where each PC has on average less than 12% non-zero elements in the worst case among all 5 datasets, we are able to nearly match and in 3 datasets even outperform KPCA. We also compare the performance of our method with a recently proposed method for SKPCA and show that our method performs better in terms of both accuracy and sparsity. We also provide a novel probabilistic proof to justify the existence of sparse solutions for KPCA using the RBF kernel. To the best of our knowledge, this is the first attempt at theoretically analyzing the validity of SKPCA.
Rudrajit Das, Aditya Golatkar, Suyash P. Awate
ICMLA1