Huy Mai

dblp:165/3174 · DBLP profile ↗
← Back
2ranked-venue papers in the field
2as first author
2since 2021 · last 2024
0000-0002-4945-4316ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (2 first)
YearPublicationVenuePosition
2024 Federated Learning under Sample Selection Heterogeneity
abstract
Despite having benefits of privacy preservation and secure computation, federated learning (FL) faces the issue of data heterogeneity. Specifically, the performance of FL systems can be degraded due to sample selection heterogeneity. We define the sample selection heterogeneity scenario with two points. First, each local training set is subject to missing-not-at-random (MNAR) sample selection bias where some labels are non-randomly missing. This requires incorporating a sample selection model that utilizes two equations to account for the prediction and selection of samples. Choosing selection features is a challenging task, especially when the number of observed features is large. Second, the sample selection mechanism is not the same across all clients. This implies that clients do not share the same set of selection features. In this work, we propose FL-MNAR to address FL under sample selection heterogeneity. The framework integrates an existing sample selection model that robustly handles sample selection bias for each client. FL-MNAR also trains an assignment function that gives a set of selection features to each client based on how well the features fit selection. Experimental results show that FL-MNAR achieves state-of-the-art performance under sample selection heterogeneity.
Huy Mai, Xintao Wu
IEEE Big Data1
2023 A Robust Classifier under Missing-Not-at-Random Sample Selection Bias
abstract
The shift between the training and testing distributions is commonly due to sample selection bias, a type of bias caused by non-random sampling of examples to be included in the training set. Although there are many approaches proposed to learn a classifier under sample selection bias, few address the case where a subset of labels in the training set are missing-not-at-random (MNAR) as a result of the selection process. In statistics, Greene’s method formulates this type of sample selection with logistic regression as the prediction model. However, we find that simply integrating this method into a robust classification framework is not effective for this bias setting. In this paper, we propose BiasCorr, an algorithm that improves on Greene’s method by modifying the original training set in order for a classifier to learn under MNAR sample selection bias. We provide theoretical guarantee for the improvement of BiasCorr over Greene’s method by analyzing its bias. Experimental results on real-world datasets demonstrate that BiasCorr produces robust classifiers and can be extended to outperform state-of-the-art classifiers that have been proposed to train under sample selection bias.
Huy Mai, Wen Huang 0003, Wei Du 0009, Xintao Wu
IEEE Big Data1