Subha Maity

dblp:278/2922 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021
YearPublicationVenuePosition
2025 Microfoundation inference for strategic prediction
abstract
Often in prediction tasks, the predictive model itself can influence the distribution of the target variable, a phenomenon termed \emph{performative prediction}. Generally, this influence stems from strategic actions taken by stakeholders with a vested interest in predictive models. A key challenge that hinders the widespread adaptation of performative prediction in machine learning is that practitioners are generally unaware of the social impacts of their predictions. To address this gap, we propose a methodology for learning the distribution map that encapsulates the long-term impacts of predictive models on the population. Specifically, we model agents’ responses as a cost-adjusted utility maximization problem and propose estimates for said cost. Our approach leverages optimal transport to align pre-model exposure (\emph{ex ante}) and post-model exposure (\emph{ex post}) distributions. We provide a rate of convergence for this proposed estimate and assess its quality through empirical demonstrations on a credit scoring dataset.
Daniele Bracale, Subha Maity, Felipe Maia Polo, Seamus Somerstep, Moulinath Banerjee, Yuekai Sun
AISTATS2
2025 Learning the Distribution Map in Reverse Causal Performative Prediction
abstract
In numerous predictive scenarios, the predictive model affects the sampling distribution; for example, job applicants often meticulously craft their resumes to navigate through a screening system. Such shifts in distribution are particularly prevalent in social computing, yet, the strategies to learn these shifts from data remain remarkably limited. Inspired by a microeconomic model that adeptly characterizes agents’ behavior within labor markets, we introduce a novel approach to learning the distribution shift. Our method is predicated on a \emph{reverse causal model}, wherein the predictive model instigates a distribution shift exclusively through a finite set of agents’ actions. Within this framework, we employ a microfoundation model for the agents’ actions and develop a statistically justified methodology to learn the distribution shift map, which we demonstrate to effectively minimize the performative prediction risk.
Daniele Bracale, Subha Maity, Yuekai Sun, Moulinath Banerjee
AISTATS2
2024 An Investigation of Representation and Allocation Harms in Contrastive Learning
abstract
The effect of underrepresentation on the performance of minority groups is known to be a serious problem in supervised learning settings; however, it has been underexplored so far in the context of self-supervised learning (SSL). In this paper, we demonstrate that contrastive learning (CL), a popular variant of SSL, tends to collapse representations of minority groups with certain majority groups. We refer to this phenomenon as representation harm and demonstrate it on image and text datasets using the corresponding popular CL methods. Furthermore, our causal mediation analysis of allocation harm on a downstream classification task reveals that representation harm is partly responsible for it, thus emphasizing the importance of studying and mitigating representation harm. Finally, we provide a theoretical explanation for representation harm using a stochastic block model that leads to a representational neural collapse in a contrastive learning setting.
Subha Maity, Mayank Agarwal, Mikhail Yurochkin, Yuekai Sun
ICLR1
2024 Weak Supervision Performance Evaluation via Partial Identification
abstract
Programmatic Weak Supervision (PWS) enables supervised model training without direct access to ground truth labels, utilizing weak labels from heuristics, crowdsourcing, or pre-trained models. However, the absence of ground truth complicates model evaluation, as traditional metrics such as accuracy, precision, and recall cannot be directly calculated. In this work, we present a novel method to address this challenge by framing model evaluation as a partial identification problem and estimating performance bounds using Fréchet bounds. Our approach derives reliable bounds on key metrics without requiring labeled data, overcoming core limitations in current weak supervision evaluation techniques. Through scalable convex optimization, we obtain accurate and computationally efficient bounds for metrics including accuracy, precision, recall, and F1-score, even in high-dimensional settings. This framework offers a robust approach to assessing model quality without ground truth labels, enhancing the practicality of weakly supervised learning for real-world applications.
Felipe Maia Polo, Subha Maity, Mikhail Yurochkin, Moulinath Banerjee, Yuekai Sun
NeurIPS2
2023 Predictor-corrector algorithms for stochastic optimization under gradual distribution shift
Subha Maity, Debarghya Mukherjee, Moulinath Banerjee, Yuekai Sun
ICLR1
2023 Understanding new tasks through the lens of training data via exponential tilting
Subha Maity, Mikhail Yurochkin, Moulinath Banerjee, Yuekai Sun
ICLR1
2023 Simple Disentanglement of Style and Content in Visual Representations
abstract
Learning visual representations with interpretable features, i.e., disentangled representations, remains a challenging problem. Existing methods demonstrate some success but are hard to apply to large-scale vision datasets like ImageNet. In this work, we propose a simple post-processing framework to disentangle content and style in learned representations from pre-trained vision models. We model the pre-trained features probabilistically as linearly entangled combinations of the latent content and style factors and develop a simple disentanglement algorithm based on the probabilistic model. We show that the method provably disentangles content and style features and verify its efficacy empirically. Our post-processed features yield significant domain generalization performance improvements when the distribution shift occurs due to style changes or style-related spurious correlations.
Lilian Ngweta, Subha Maity, Alex Gittens, Yuekai Sun, Mikhail Yurochkin
ICML2
2022 Meta-analysis of heterogeneous data: integrative sparse regression in high-dimensions
abstract
We consider the task of meta-analysis in high-dimensional settings in which the data sources are similar but non-identical. To borrow strength across such heterogeneous datasets, we introduce a global parameter that emphasizes interpretability and statistical efficiency in the presence of heterogeneity. We also propose a one-shot estimator of the global parameter that preserves the anonymity of the data sources and converges at a rate that depends on the size of the combined dataset. For high-dimensional linear model settings, we demonstrate the superiority of our identification restrictions in adapting to a previously seen data distribution as well as predicting for a new/unseen data distribution. Finally, we demonstrate the benefits of our approach on a large-scale drug treatment dataset involving several different cancer cell-lines.
Subha Maity, Yuekai Sun, Moulinath Banerjee
J. Mach. Learn. Res.1
2022 Minimax optimal approaches to the label shift problem in non-parametric settings
abstract
We study the minimax rates of the label shift problem in non-parametric classification. In addition to the unsupervised setting in which the learner only has access to unlabeled examples from the target domain, we also consider the setting in which a small number of labeled examples from the target domain is available to the learner. Our study reveals a difference in the difficulty of the label shift problem in the two settings, and we attribute this difference to the availability of data from the target domain to estimate the class conditional distributions in the latter setting. We also show that a class proportion estimation approach is minimax rate-optimal in the unsupervised setting.
Subha Maity, Yuekai Sun, Moulinath Banerjee
J. Mach. Learn. Res.1
2021 Statistical inference for individual fairness
Subha Maity, Songkai Xue, Mikhail Yurochkin, Yuekai Sun
ICLR1
2021 Does enforcing fairness mitigate biases caused by subpopulation shift?
abstract
Many instances of algorithmic bias are caused by subpopulation shifts. For example, ML models often perform worse on demographic groups that are underrepresented in the training data. In this paper, we study whether enforcing algorithmic fairness during training improves the performance of the trained model in the \emph{target domain}. On one hand, we conceive scenarios in which enforcing fairness does not improve performance in the target domain. In fact, it may even harm performance. On the other hand, we derive necessary and sufficient conditions under which enforcing algorithmic fairness leads to the Bayes model in the target domain. We also illustrate the practical implications of our theoretical results in simulations and on real data.
Subha Maity, Debarghya Mukherjee, Mikhail Yurochkin, Yuekai Sun
NeurIPS1