Supratim Shit

dblp:185/5317 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-6602-6436ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Theoretical computer science
2 papers
Algorithms and data structures · 62% Mathematical optimization · 38%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › subset selection
coreset
0.912025
Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025
Machine learning › Efficient and distributed learning › data selection
coreset selection
0.912025
Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025
Machine learning › Efficient and distributed learning
federated learning
0.912025
Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025
Machine learning › Efficient and distributed learning › federated learning › federated learning architecture
vertical federated learning
0.912025
Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions · ICML 2025
Algorithms and data structures › data summarization
coresets
0.922020
Streaming Coresets for Symmetric Tensor Factorization · ICML 2020
On Coresets for Regularized Regression · ICML 2020
Mathematical optimization › continuous optimization
convex optimization
0.412020
On Coresets for Regularized Regression · ICML 2020
Algorithms and data structures › numerical linear algebra › matrix and tensor decomposition › tensor decomposition
CP decomposition
0.412020
Streaming Coresets for Symmetric Tensor Factorization · ICML 2020
Mathematical optimization › statistical estimation › regression › sparse regression
lasso
0.412020
On Coresets for Regularized Regression · ICML 2020
Mathematical optimization › statistical estimation › regression
regularized regression
0.412020
On Coresets for Regularized Regression · ICML 2020
Algorithms and data structures › data streams
streaming algorithms
0.412020
Streaming Coresets for Symmetric Tensor Factorization · ICML 2020
Algorithms and data structures › numerical linear algebra › matrix and tensor decomposition
tensor decomposition
0.412020
Streaming Coresets for Symmetric Tensor Factorization · ICML 2020

Methods — techniques the papers use, named apart from their topics

regularized logistic regression · 0.9regularized linear regression · 0.9online row sampling · 0.4online filtering · 0.4kernelization · 0.4coreset construction · 0.4convex relaxation · 0.4
YearPublicationVenuePosition
2025 Improved Coresets for Vertical Federated Learning: Regularized Linear and Logistic Regressions
abstract
Coreset, as a summary of training data, offers an efficient approach for reducing data processing and storage complexity during training. In the emerging vertical federated learning (VFL) setting, where scattered clients store different data features, it directly reduces communication complexity. In this work, we introduce coresets construction for regularized logistic regression both in centralized and VFL settings. Additionally, we improve the coreset size for regularized linear regression in the VFL setting. We also eliminate the dependency of the coreset size on a property of the data due to the VFL setting. The improvement in the coreset sizes is due to our novel coreset construction algorithms that capture the reduced model complexity due to the added regularization and its subsequent analysis. In experiments, we provide extensive empirical evaluation that backs our theoretical claims. We also report the performance of our coresets by comparing the models trained on the complete data and on the coreset.
Supratim Shit, Gurmehak Kaur Chadha, Bapi Chatterjee
ICML1
2022 On Coresets for Fair Regression and Individually Fair Clustering
abstract
In this paper we present coresets for Fair Regression with Statistical Parity (SP) constraints and for Individually Fair Clustering. Due to the fairness constraints, the classical coreset definition is not enough for these problems. We first define coresets for both the problems. We show that to obtain such coresets, it is sufficient to sample points based on the probabilities dependent on combination of sensitivity score and a carefully chosen term according to the fairness constraints. We give provable guarantees with relative error in preserving the cost and a small additive error in preserving fairness constraints for both problems. Since our coresets are much smaller in size as compared to $n$, the number of points, they can give huge benefits in computational costs (from polynomial to polylogarithmic in $n$), especially when $n \gg d$, where $d$ is the input dimension. We support our theoretical claims with experimental evaluations.
Rachit Chhaya, Anirban Dasgupta 0001, Jayesh Choudhari, Supratim Shit
AISTATS4
2020 On Coresets for Regularized Regression
abstract
We study the effect of norm based regularization on the size of coresets for regression problems. Specifically, given a matrix $ \mathbf{A} \in {\mathbb{R}}^{n \times d}$ with $n\gg d$ and a vector $\mathbf{b} \in \mathbb{R} ^ n $ and $\lambda > 0$, we analyze the size of coresets for regularized versions of regression of the form $\|\mathbf{Ax}-\mathbf{b}\|_p^r + \lambda\|{\mathbf{x}}\|_q^s$. Prior work has shown that for ridge regression (where $p,q,r,s=2$) we can obtain a coreset that is smaller than the coreset for the unregularized counterpart i.e. least squares regression \cite{avron2017sharper}. We show that when $r \neq s$, no coreset for regularized regression can have size smaller than the optimal coreset of the unregularized version. The well known lasso problem falls under this category and hence does not allow a coreset smaller than the one for least squares regression. We propose a modified version of the lasso problem and obtain for it a coreset of size smaller than the least square regression. We empirically show that the modified version of lasso also induces sparsity in solution, similar to the original lasso. We also obtain smaller coresets for $\ell_p$ regression with $\ell_p$ regularization. We extend our methods to multi response regularized regression. Finally, we empirically demonstrate the coreset performance for the modified lasso and the $\ell_1$ regression with $\ell_1$ regularization.
Rachit Chhaya, Anirban Dasgupta 0001, Supratim Shit
ICML3
2020 Streaming Coresets for Symmetric Tensor Factorization
abstract
Factorizing tensors has recently become an important optimization module in a number of machine learning pipelines, especially in latent variable models. We show how to do this efficiently in the streaming setting. Given a set of $n$ vectors, each in $\mathbb{R}^d$, we present algorithms to select a sublinear number of these vectors as coreset, while guaranteeing that the CP decomposition of the $p$-moment tensor of the coreset approximates the corresponding decomposition of the $p$-moment tensor computed from the full data. We introduce two novel algorithmic techniques: online filtering and kernelization. Using these two, we present four algorithms that achieve different tradeoffs of coreset size, update time and working space, beating or matching various state of the art algorithms. In the case of matrices (2-ordered tensor), our online row sampling algorithm guarantees $(1 \pm \epsilon)$ relative error spectral approximation. We show applications of our algorithms in learning single topic modeling.
Rachit Chhaya, Jayesh Choudhari, Anirban Dasgupta 0001, Supratim Shit
ICML4
2016 Consensus-Aware Sociopsychological Trust Model for Wireless Sensor Networks
abstract
Security plays a vital role in Wireless Sensor Networks (WSN) for providing reliability to the network. In WSN, where nodes, in addition to having their inbuilt capability of sensing, processing, and communicating data, also possess certain risks. These risks expose them to attacks and bring in many security challenges. Many researchers are engaged in developing innovative design paradigms to address security issues by developing trust management systems. In WSN, trust is important for the establishment of cooperation among the sensor nodes. The article presents a sociopsychological model for detecting fraudulent nodes in WSN. The three factors, viz. ability, benevolence, and integrity, are used for the computation of trust. Furthermore, the article provides a novel consensus-aware sociopsychological approach to deal even in the presence of higher number of fraudulent nodes than benevolent nodes. The proposed work has been implemented in the LabVIEW platform and extensive simulations were carried out to study its performance. Additionally, it is experimentally evaluated on a testbed of size 16 nodes to obtain results that demonstrate the accuracy and robustness of the proposed model.
Heena Rathore, Venkataramana Badarla, Supratim Shit
ACM Trans. Sens. Networks3