VLDB 2026 Research / reviewers in the wild / expert
Supratim Shit
dblp:185/5317
· DBLP profile ↗
5ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-6602-6436ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Theoretical computer science
2 papers |
Algorithms and data structures · 62% Mathematical optimization · 38% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › subset selection
coreset |
0.9 | 1 | 2025 | Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025 |
Machine learning › Efficient and distributed learning › data selection
coreset selection |
0.9 | 1 | 2025 | Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025 |
Machine learning › Efficient and distributed learning
federated learning |
0.9 | 1 | 2025 | Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025 |
Machine learning › Efficient and distributed learning › federated learning › federated learning architecture
vertical federated learning |
0.9 | 1 | 2025 | Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025 |
Algorithms and data structures › data summarization
coresets |
0.9 | 2 | 2020 | Streaming Coresets for Symmetric Tensor Factorization · ICML 2020 On Coresets for Regularized Regression · ICML 2020 |
Mathematical optimization › continuous optimization
convex optimization |
0.4 | 1 | 2020 | On Coresets for Regularized Regression · ICML 2020 |
Algorithms and data structures › numerical linear algebra › matrix and tensor decomposition › tensor decomposition
CP decomposition |
0.4 | 1 | 2020 | Streaming Coresets for Symmetric Tensor Factorization · ICML 2020 |
Mathematical optimization › statistical estimation › regression › sparse regression
lasso |
0.4 | 1 | 2020 | On Coresets for Regularized Regression · ICML 2020 |
Mathematical optimization › statistical estimation › regression
regularized regression |
0.4 | 1 | 2020 | On Coresets for Regularized Regression · ICML 2020 |
Algorithms and data structures › data streams
streaming algorithms |
0.4 | 1 | 2020 | Streaming Coresets for Symmetric Tensor Factorization · ICML 2020 |
Algorithms and data structures › numerical linear algebra › matrix and tensor decomposition
tensor decomposition |
0.4 | 1 | 2020 | Streaming Coresets for Symmetric Tensor Factorization · ICML 2020 |
Methods — techniques the papers use, named apart from their topics
regularized logistic regression · 0.9regularized linear regression · 0.9online row sampling · 0.4online filtering · 0.4kernelization · 0.4coreset construction · 0.4convex relaxation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic RegressionsabstractCoreset, as a summary of training data, offers an efficient approach for reducing data processing and storage complexity during training. In the emerging vertical federated learning (VFL) setting, where scattered clients store different data features, it directly reduces communication complexity. In this work, we introduce coresets construction for regularized logistic regression both in centralized and VFL settings. Additionally, we improve the coreset size for regularized linear regression in the VFL setting. We also eliminate the dependency of the coreset size on a property of the data due to the VFL setting. The improvement in the coreset sizes is due to our novel coreset construction algorithms that capture the reduced model complexity due to the added regularization and its subsequent analysis. In experiments, we provide extensive empirical evaluation that backs our theoretical claims. We also report the performance of our coresets by comparing the models trained on the complete data and on the coreset. Supratim Shit, Gurmehak Kaur Chadha, Bapi Chatterjee |
ICML | 1 |
| 2022 | On Coresets for Fair Regression and Individually Fair ClusteringabstractIn this paper we present coresets for Fair Regression with Statistical Parity (SP) constraints and for Individually Fair Clustering. Due to the fairness constraints, the classical coreset definition is not enough for these problems. We first define coresets for both the problems. We show that to obtain such coresets, it is sufficient to sample points based on the probabilities dependent on combination of sensitivity score and a carefully chosen term according to the fairness constraints. We give provable guarantees with relative error in preserving the cost and a small additive error in preserving fairness constraints for both problems. Since our coresets are much smaller in size as compared to $n$, the number of points, they can give huge benefits in computational costs (from polynomial to polylogarithmic in $n$), especially when $n \gg d$, where $d$ is the input dimension. We support our theoretical claims with experimental evaluations. Rachit Chhaya, Anirban Dasgupta 0001, Jayesh Choudhari, Supratim Shit |
AISTATS | 4 |
| 2020 | On Coresets for Regularized RegressionabstractWe study the effect of norm based regularization on the size of coresets for regression problems. Specifically, given a matrix $ \mathbf{A} \in {\mathbb{R}}^{n \times d}$ with $n\gg d$ and a vector $\mathbf{b} \in \mathbb{R} ^ n $ and $\lambda > 0$, we analyze the size of coresets for regularized versions of regression of the form $\|\mathbf{Ax}-\mathbf{b}\|_p^r + \lambda\|{\mathbf{x}}\|_q^s$. Prior work has shown that for ridge regression (where $p,q,r,s=2$) we can obtain a coreset that is smaller than the coreset for the unregularized counterpart i.e. least squares regression \cite{avron2017sharper}. We show that when $r \neq s$, no coreset for regularized regression can have size smaller than the optimal coreset of the unregularized version. The well known lasso problem falls under this category and hence does not allow a coreset smaller than the one for least squares regression. We propose a modified version of the lasso problem and obtain for it a coreset of size smaller than the least square regression. We empirically show that the modified version of lasso also induces sparsity in solution, similar to the original lasso. We also obtain smaller coresets for $\ell_p$ regression with $\ell_p$ regularization. We extend our methods to multi response regularized regression. Finally, we empirically demonstrate the coreset performance for the modified lasso and the $\ell_1$ regression with $\ell_1$ regularization. Rachit Chhaya, Anirban Dasgupta 0001, Supratim Shit |
ICML | 3 |
| 2020 | Streaming Coresets for Symmetric Tensor FactorizationabstractFactorizing tensors has recently become an important optimization module in a number of machine learning pipelines, especially in latent variable models. We show how to do this efficiently in the streaming setting. Given a set of $n$ vectors, each in $\mathbb{R}^d$, we present algorithms to select a sublinear number of these vectors as coreset, while guaranteeing that the CP decomposition of the $p$-moment tensor of the coreset approximates the corresponding decomposition of the $p$-moment tensor computed from the full data. We introduce two novel algorithmic techniques: online filtering and kernelization. Using these two, we present four algorithms that achieve different tradeoffs of coreset size, update time and working space, beating or matching various state of the art algorithms. In the case of matrices (2-ordered tensor), our online row sampling algorithm guarantees $(1 \pm \epsilon)$ relative error spectral approximation. We show applications of our algorithms in learning single topic modeling. Rachit Chhaya, Jayesh Choudhari, Anirban Dasgupta 0001, Supratim Shit |
ICML | 4 |
| 2016 | Consensus-Aware Sociopsychological Trust Model for Wireless Sensor NetworksabstractSecurity plays a vital role in Wireless Sensor Networks (WSN) for providing reliability to the network. In WSN, where nodes, in addition to having their inbuilt capability of sensing, processing, and communicating data, also possess certain risks. These risks expose them to attacks and bring in many security challenges. Many researchers are engaged in developing innovative design paradigms to address security issues by developing trust management systems. In WSN, trust is important for the establishment of cooperation among the sensor nodes. The article presents a sociopsychological model for detecting fraudulent nodes in WSN. The three factors, viz. ability, benevolence, and integrity, are used for the computation of trust. Furthermore, the article provides a novel consensus-aware sociopsychological approach to deal even in the presence of higher number of fraudulent nodes than benevolent nodes. The proposed work has been implemented in the LabVIEW platform and extensive simulations were carried out to study its performance. Additionally, it is experimentally evaluated on a testbed of size 16 nodes to obtain results that demonstrate the accuracy and robustness of the proposed model. Heena Rathore, Venkataramana Badarla, Supratim Shit |
ACM Trans. Sens. Networks | 3 |