EDBT 2026 Demo / reviewers in the wild / expert
Rudrajit Das
dblp:227/2712
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-9818-0518ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Efficient and distributed learning · 18% Optimization for machine learning · 16% Transfer learning and domain adaptation · 16% | |
| Network and information security
2 papers |
Privacy and data protection · 100% |
Topics — the 26 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation › model adaptation
model retraining |
1.7 | 2 | 2025 | Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing · NeurIPS 2025 Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
1.5 | 2 | 2025 | Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025 Understanding Self-Distillation in the Presence of Label Noise · ICML 2023 |
Privacy and data protection
differential privacy |
1.5 | 2 | 2025 | Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025 Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023 |
Machine learning › Learning theory › classification
binary classification |
1.1 | 2 | 2025 | Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing · NeurIPS 2025 Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025 |
Machine learning › Graph learning › graph neural network › message passing
approximate message passing |
0.9 | 1 | 2025 | Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing · NeurIPS 2025 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.9 | 1 | 2025 | Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting · ICML 2025 |
Machine learning › Deep learning architectures and training › data-centric deep learning
sample weighting |
0.9 | 1 | 2025 | Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting · ICML 2025 |
Privacy and data protection › differential privacy › relaxed differential privacy
label differential privacy |
0.9 | 1 | 2025 | Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling |
0.8 | 1 | 2024 | Understanding the Training Speedup from Sampling with Approximate Losses · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
sample selection |
0.8 | 1 | 2024 | Understanding the Training Speedup from Sampling with Approximate Losses · ICML 2024 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.8 | 1 | 2024 | Understanding the Training Speedup from Sampling with Approximate Losses · ICML 2024 |
Machine learning › Optimization for machine learning
gradient clipping |
0.7 | 1 | 2023 | Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation |
0.7 | 1 | 2023 | Understanding Self-Distillation in the Presence of Label Noise · ICML 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.7 | 1 | 2023 | Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023 |
Machine learning › Learning paradigms
supervised learning |
0.7 | 1 | 2023 | Understanding Self-Distillation in the Presence of Label Noise · ICML 2023 |
Privacy and data protection › differential privacy › differentially private deep learning
DP-SGD |
0.7 | 1 | 2023 | Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.6 | 1 | 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Efficient and distributed learning › federated learning
decentralized federated learning |
0.6 | 1 | 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Efficient and distributed learning
federated learning |
0.6 | 1 | 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Efficient and distributed learning › distributed training
gradient compression |
0.6 | 1 | 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Learning theory
linear separability |
0.3 | 1 | 2025 | Retraining with Predicted Hard Labels Provably Increases Model Accuracy · ICML 2025 |
Machine learning › Optimization for machine learning
convergence analysis |
0.2 | 1 | 2023 | Beyond Uniform Lipschitz Condition in Differentially Private Optimization · ICML 2023 |
Machine learning › Learning theory › statistical estimation › regularized estimation
regularized linear regression |
0.2 | 1 | 2023 | Understanding Self-Distillation in the Presence of Label Noise · ICML 2023 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.2 | 1 | 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Optimization for machine learning › non-convex optimization
polyak-łojasiewicz condition |
0.2 | 1 | 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Methods — techniques the papers use, named apart from their topics
predicted hard labels · 1.7consensus-based retraining · 1.7theoretical analysis · 0.9loss-based sample weighting · 0.9generalized linear model · 0.9gaussian mixture model · 0.9approximate message passing · 0.9greedy sample selection · 0.8early exiting · 0.8convex optimization · 0.8per-sample lipschitz constants · 0.7gradient clipping · 0.7DP-SGD · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Retraining with Predicted Hard Labels Provably Increases Model AccuracyabstractThe performance of a model trained with noisy labels is often improved by simply *retraining* the model with its *own predicted hard labels* (i.e., $1$/$0$ labels). Yet, a detailed theoretical characterization of this phenomenon is lacking. In this paper, we theoretically analyze retraining in a linearly separable binary classification setting with randomly corrupted labels given to us and prove that retraining can improve the population accuracy obtained by initially training with the given (noisy) labels. To the best of our knowledge, this is the first such theoretical result. Retraining finds application in improving training with local label differential privacy (DP), which involves training with noisy labels. We empirically show that retraining selectively on the samples for which the predicted label matches the given label significantly improves label DP training at no extra privacy cost; we call this consensus-based retraining. For example, when training ResNet-18 on CIFAR-100 with $\epsilon=3$ label DP, we obtain more than $6$% improvement in accuracy with consensus-based retraining. Rudrajit Das, Inderjit S. Dhillon, Alessandro Epasto, Adel Javanmard, Jieming Mao, Vahab S. Mirrokni, Sujay Sanghavi, Peilin Zhong |
ICML | 1 |
| 2025 | Upweighting Easy Samples in Fine-Tuning Mitigates ForgettingabstractFine-tuning a pre-trained model on a downstream task often degrades its original capabilities, a phenomenon known as "catastrophic forgetting". This is especially an issue when one does not have access to the data and recipe used to develop the pre-trained model. Under this constraint, most existing methods for mitigating forgetting are inapplicable. To address this challenge, we propose a sample weighting scheme for the fine-tuning data solely based on the pre-trained model’s losses. Specifically, we upweight the easy samples on which the pre-trained model’s loss is low and vice versa to limit the drift from the pre-trained model. Our approach is orthogonal and yet complementary to existing methods; while such methods mostly operate on parameter or gradient space, we concentrate on the sample space. We theoretically analyze the impact of fine-tuning with our method in a linear setting, showing that it stalls learning in a certain subspace, which inhibits overfitting to the target task. We empirically demonstrate the efficacy of our method on both language and vision tasks. As an example, when fine-tuning Gemma 2 2B on MetaMathQA, our method results in only a $0.8$% drop in accuracy on GSM8K (another math dataset) compared to standard fine-tuning, while preserving $5.4$% more accuracy on the pre-training datasets. Sunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis, Sujay Sanghavi |
ICML | 3 |
| 2025 | Self-Boost via Optimal Retraining: An Analysis via Approximate Message PassingabstractRetraining a model using its own predictions together with the original, potentially noisy labels is a well-known strategy for improving the model’s performance. While prior works have demonstrated the benefits of specific heuristic retraining schemes, the question of how to optimally combine the model's predictions and the provided labels remains largely open. This paper addresses this fundamental question for binary classification tasks. We develop a principled framework based on approximate message passing (AMP) to analyze iterative retraining procedures for two ground truth settings: Gaussian mixture model (GMM) and generalized linear model (GLM). Our main contribution is the derivation of the Bayes optimal aggregator function to combine the current model's predictions and the given labels, which when used to retrain the same model, minimizes its prediction error. We also quantify the performance of this optimal retraining strategy over multiple rounds. We complement our theoretical results by proposing a practically usable version of the theoretically-optimal aggregator function and demonstrate its superiority over baseline methods under different label noise models. Adel Javanmard, Rudrajit Das, Alessandro Epasto, Vahab S. Mirrokni |
NeurIPS | 2 |
| 2024 | Understanding the Training Speedup from Sampling with Approximate LossesabstractIt is well known that selecting samples with large losses/gradients can significantly reduce the number of training steps. However, the selection overhead is often too high to yield any meaningful gains in terms of overall training time. In this work, we focus on the greedy approach of selecting samples with large approximate losses instead of exact losses in order to reduce the selection overhead. For smooth convex losses, we show that such a greedy strategy can converge to a constant factor of the minimum value of the average loss in fewer iterations than the standard approach of random selection. We also theoretically quantify the effect of the approximation level. We then develop SIFT which uses early exiting to obtain approximate losses with an intermediate layer’s representations for sample selection. We evaluate SIFT on the task of training a 110M parameter 12 layer BERT base model, and show significant gains (in terms of training hours and number of backpropagation steps) without any optimized implementation over vanilla training. For e.g., to reach 64% validation accuracy, SIFT with exit at the first layer takes $\sim$ 43 hours compared to $\sim$ 57 hours of vanilla training. Rudrajit Das, Bertram Ieong, Parikshit Bansal, Sujay Sanghavi |
ICML | 1 |
| 2023 | Beyond Uniform Lipschitz Condition in Differentially Private OptimizationabstractMost prior results on differentially private stochastic gradient descent (DP-SGD) are derived under the simplistic assumption of uniform Lipschitzness, i.e., the per-sample gradients are uniformly bounded. We generalize uniform Lipschitzness by assuming that the per-sample gradients have sample-dependent upper bounds, i.e., per-sample Lipschitz constants, which themselves may be unbounded. We provide principled guidance on choosing the clip norm in DP-SGD for convex over-parameterized settings satisfying our general version of Lipschitzness when the per-sample Lipschitz constants are bounded; specifically, we recommend tuning the clip norm only till values up to the minimum per-sample Lipschitz constant. This finds application in the private training of a softmax layer on top of a deep network pre-trained on public data. We verify the efficacy of our recommendation via experiments on 8 datasets. Furthermore, we provide new convergence results for DP-SGD on convex and nonconvex functions when the Lipschitz constants are unbounded but have bounded moments, i.e., they are heavy-tailed. Rudrajit Das, Satyen Kale, Zheng Xu 0002, Tong Zhang 0001, Sujay Sanghavi |
ICML | 1 |
| 2023 | Understanding Self-Distillation in the Presence of Label NoiseabstractSelf-distillation (SD) is the process of first training a "teacher" model and then using its predictions to train a "student" model that has the *same* architecture. Specifically, the student's loss is $\big(\xi*\ell(\text{teacher's predictions}, \text{ student's predictions}) + (1-\xi)*\ell(\text{given labels}, \text{ student's predictions})\big)$, where $\ell$ is the loss function and $\xi$ is some parameter $\in [0,1]$. SD has been empirically observed to provide performance gains in several settings. In this paper, we theoretically characterize the effect of SD in two supervised learning problems with *noisy labels*. We first analyze SD for regularized linear regression and show that in the high label noise regime, the optimal value of $\xi$ that minimizes the expected error in estimating the ground truth parameter is surprisingly greater than 1. Empirically, we show that $\xi > 1$ works better than $\xi \leq 1$ even with the cross-entropy loss for several classification datasets when 50% or 30% of the labels are corrupted. Further, we quantify when optimal SD is better than optimal regularization. Next, we analyze SD in the case of logistic regression for binary classification with random label corruption and quantify the range of label corruption in which the student outperforms the teacher (w.r.t. accuracy). To our knowledge, this is the first result of its kind for the cross-entropy loss. Rudrajit Das, Sujay Sanghavi |
ICML | 1 |
| 2022 | Faster non-convex federated learning via global and local momentumabstractWe propose \texttt{FedGLOMO}, a novel federated learning (FL) algorithm with an iteration complexity of $\mathcal{O}(\epsilon^{-1.5})$ to converge to an $\epsilon$-stationary point (i.e., $\mathbb{E}[\|\nabla f(x)\|^2] \leq \epsilon$) for smooth non-convex functions – under arbitrary client heterogeneity and compressed communication – compared to the $\mathcal{O}(\epsilon^{-2})$ complexity of most prior works. Our key algorithmic idea that enables achieving this improved complexity is based on the observation that the convergence in FL is hampered by two sources of high variance: (i) the global server aggregation step with multiple local updates, exacerbated by client heterogeneity, and (ii) the noise of the local client-level stochastic gradients. The first issue is particularly detrimental to FL algorithms that perform plain averaging at the server. By modeling the server aggregation step as a generalized gradient-type update, we propose a variance-reducing momentum-based global update at the server, which when applied in conjunction with variance-reduced local updates at the clients, enables \texttt{FedGLOMO} to enjoy an improved convergence rate. Our experiments illustrate the intrinsic variance reduction effect of \texttt{FedGLOMO}, which implicitly suppresses client-drift in heterogeneous data distribution settings and promotes communication efficiency. Rudrajit Das, Anish Acharya, Abolfazl Hashemi, Sujay Sanghavi, Inderjit S. Dhillon, Ufuk Topcu |
UAI | 1 |
| 2022 | On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Federated LearningabstractFederated learning (FL) is an emerging collaborative machine learning (ML) framework that enables training of predictive models in a distributed fashion where the communication among the participating nodes are facilitated by a central server. To deal with the communication bottleneck at the server, decentralized FL (DFL) methods advocate rely on local communication of nodes with their neighbors according to a specific communication network. In DFL, it is common algorithmic practice to have nodes interleave (local) gradient descent iterations with gossip (i.e., averaging over the network) steps. As the size of the ML models grows, the limited communication bandwidth among the nodes does not permit communication of full-precision messages; hence, it is becoming increasingly common to require that messages belossy, compressedversions of the local parameters. The requirement of communicating compressed messages gives rise to the important question:given a fixed communication budget, what should be our communication strategy to minimize the (training) loss as much as possible?In this article, we explore this direction, and show that in such compressed DFL settings, there are benefits to havingmultiplegossip steps between subsequent gradient iterations, even when the cost of doing so is appropriately accounted for, e.g., by means of reducing the precision of compressed information. In particular, we show that having${\mathcal O}(\log \frac{1}{\epsilon })$gradient iterations with constant step size - and${\mathcal O}(\log \frac{1}{\epsilon })$gossip steps between every pair of these iterations - enables convergence to within$\epsilon$of the optimal value for a class of non-convex problems that arise in the training of deep learning models, namely, smooth non-convex objectives satisfying Polyak-Łojasiewicz condition. Empirically, we show that our proposed scheme bridges the gap between centralized gradient descent and DFL on various machine learning tasks across different network topologies and compression operators. Abolfazl Hashemi, Anish Acharya, Rudrajit Das, Haris Vikalo, Sujay Sanghavi, Inderjit S. Dhillon |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Nonlinear Blind Compressed Sensing Under Signal-Dependent NoiseabstractIn this paper, we consider the problem of nonlinear blind compressed sensing, i.e. jointly estimating the sparse codes and sparsity-promoting basis, under signal-dependent noise. We focus our efforts on the Poisson noise model, though other signal-dependent noise models can be considered. By employing a well-known variance stabilizing transform such as the Anscombe transform, we formulate our task as a nonlinear least squares problem with the ℓ1penalty imposed for promoting sparsity. We solve this objective function under non-negativity constraints imposed on both the sparse codes and the basis. To this end, we propose a multiplicative update rule, similar to that used in non-negative matrix factorization (NMF), for our alternating minimization algorithm. To the best of our knowledge, this is the first attempt at a formulation for nonlinear blind compressed sensing, with and without the Poisson noise model. Further, we also provide some theoretical bounds on the performance of our algorithm. Rudrajit Das, Ajit Rajwade 0001 |
ICIP | 1 |
| 2018 | Sparse Kernel PCA for Outlier DetectionabstractIn this paper, we propose a new method to perform Sparse Kernel Principal Component Analysis (SKPCA) and also mathematically analyze the validity of SKPCA. We formulate SKPCA as a constrained optimization problem with elastic net regularization in kernel feature space and solve it. We consider outlier detection (where KPCA is employed) as an application for SKPCA, using the RBF kernel. We test it on 5 real world datasets and show that by using just 4% (or even less) of the principal components (PCs), where each PC has on average less than 12% non-zero elements in the worst case among all 5 datasets, we are able to nearly match and in 3 datasets even outperform KPCA. We also compare the performance of our method with a recently proposed method for SKPCA and show that our method performs better in terms of both accuracy and sparsity. We also provide a novel probabilistic proof to justify the existence of sparse solutions for KPCA using the RBF kernel. To the best of our knowledge, this is the first attempt at theoretically analyzing the validity of SKPCA. Rudrajit Das, Aditya Golatkar, Suyash P. Awate |
ICMLA | 1 |