Lovedeep Gondara

dblp:177/1788 · also Lovedeep Singh Gondara · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
4since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 PubSub-ML: A Model Streaming Alternative to Federated Learning
abstract
Federated learning is a decentralized learning framework where participating sites are engaged in a tight collaboration, forcing them into symmetric sharing and the agreement in terms of data samples, feature spaces, model types and architectures, privacy settings, and training processes. We propose PubSub-ML, Publish-Subscribe for Machine Learning, as a solution in a loose collaboration setting where each site maintains local autonomy on these decisions. In PubSub-ML, each site is either a publisher or a subscriber or both. The publishers publish differentially private machine learning models and the subscribers subscribe to published models in order to construct customized models for local use, essentially benefiting from other sites' data by distilling knowledge from publishers' models while respecting data privacy. The term “model streaming” comes from the extension of PubSub-ML to decentralized data streams with concept drift. Our extensive empirical evaluation shows that PubSub-ML outperforms federated learning methods by a significant margin.
Lovedeep Gondara, Ke Wang 0001
Proc. Priv. Enhancing Technol.1
2022 Incorporating Item Frequency for Differentially Private Set Union
abstract
We study the problem of releasing the set union of users' items subject to differential privacy. Previous approaches consider only the set of items for each user as the input. We propose incorporating the item frequency, which is typically available in set union problems, to boost the utility of private mechanisms. However, using the global item frequency over all users would largely increase privacy loss. We propose to use the local item frequency of each user to approximate the global item frequency without incurring additional privacy loss. Local item frequency allows us to design greedy set union mechanisms that are differentially private, which is impossible for previous greedy proposals. Moreover, while all previous works have to use uniform sampling to limit the number of items each user would contribute to, our construction eliminates the sampling step completely and allows our mechanisms to consider all of the users' items. Finally, we propose to transfer the knowledge of the global item frequency from a public dataset into our mechanism, which further boosts utility even when the public and private datasets are from different domains. We evaluate the proposed methods on multiple real-life datasets.
Ricardo Silva Carvalho, Ke Wang 0001, Lovedeep Gondara
AAAI3
2022 Differentially Private Ensemble Classifiers for Data Streams
abstract
Learning from continuous data streams via classification/regression is prevalent in many domains. Adapting to evolving data characteristics (concept drift) while protecting data owners' private information is an open challenge. We present a differentially private ensemble solution to this problem with two distinguishing features: it allows anunbounded number of ensemble updates to deal with the potentially never-ending data streams under a fixed privacy budget, and it ismodel agnostic, in that it treats any pre-trained differentially private classification/regression model as a black-box. Our method outperforms competitors on real-world and simulated datasets for varying settings of privacy, concept drift, and data distribution.
Lovedeep Gondara, Ke Wang 0001, Ricardo Silva Carvalho
WSDM1
2021 Training Differentially Private Neural Networks with Lottery Tickets
Lovedeep Gondara, Ricardo Silva Carvalho, Ke Wang 0001
ESORICS (2)1
2020 Differentially Private Top-k Selection via Stability on Unknown Domain
abstract
We propose a new method that satisfies approximate differential privacy for top-$k$ selection with unordered output in the unknown data domain setting, not relying on the full knowledge of the domain universe. Our algorithm only requires looking at the top-$\bar{k}$ elements for any given $\bar{k} \geq k$, thus, enforcing the principle of minimal privilege. Unlike previous methods, our privacy parameter $\varepsilon$ does not scale with $k$, giving improved applicability for scenarios of very large $k$. Moreover, our novel construction, which combines the sparse vector technique and stability efficiently, can be applied as a general framework to any type of query, thus being of independent interest. We extensively compare our algorithm to previous work of top-$k$ selection on the unknown domain, and show, both analytically and on experiments, settings where we outperform the current state-of-the-art.
Ricardo Silva Carvalho, Ke Wang 0001, Lovedeep Gondara, Chunyan Miao
UAI3
2020 Differentially Private Small Dataset Release Using Random Projections
abstract
Small datasets form a significant portion of releasable data in high sensitivity domains such as healthcare. But, providing differential privacy for small dataset release is a hard task, where current state-of-the-art methods suffer from severe utility loss. As a solution, we propose DPRP (Differentially Private Data Release via Random Projections), a reconstruction based approach for releasing differentially private small datasets. DPRP has several key advantages over the state-of-the-art. Using seven diverse real-life datasets, we show that DPRP outperforms the current state-of-the-art on a variety of tasks, under varying conditions, and for all privacy budgets.
Lovedeep Gondara, Ke Wang 0001
UAI1
2018 MIDA: Multiple Imputation Using Denoising Autoencoders
Lovedeep Gondara, Ke Wang 0001
PAKDD (3)1
2017 Recovering loss to followup information using denoising autoencoders
abstract
Loss to followup is a significant issue in healthcare and has serious consequences for a study's validity and cost. Methods available at present for recovering loss to followup information are restricted by their expressive capabilities and struggle to model highly non-linear relations and complex interactions. In this paper we propose a model based on overcomplete denoising autoencoders to recover loss to followup information. Designed to work with high volume data, results on various simulated and real life datasets show our model is appropriate under varying dataset and loss to followup conditions and outperforms the state-of-the-art methods by a wide margin (≥ 20% in some scenarios) while preserving the dataset utility for final analysis.
Lovedeep Gondara, Ke Wang 0001
IEEE BigData1
2015 RPC: An Efficient Classifier Ensemble Using Random Projections
abstract
We propose a classifier ensemble called RPC based on principles of rotation forest using random projections. Random projections project the original high dimensional data into lower dimensions while preserving the dataset's geometrical structure reducing classifier's complexity. Random projections are also an efficient dimensionality reduction tool, removing noisy features from dataset and representing the information using only small number of features. Training set for RPC is created by applying random projection on random subsets of the feature set. The randomness of random projection coupled with random sampling adds diversity to RPC. Initial evaluation using datasets from UCI machine learning repository shows that RPC performs equally well or better than Random Forest, Bagging and AdaBoost. We demonstrate that using dimensionality reduction with RPC we can dramatically reduce datasets dimensions without any loss in classification accuracy and significantly enhance computational performance. Finally, we experiment building RPC with different base learners.
Lovedeep Gondara
ICMLA1
2015 Random Forest with Random Projection to Impute Missing Gene Expression Data
abstract
Measurement error or lack of proper experimental setup often results in invalid or missing data in gene expression studies. Small sample size and cost of re-running the experiment presents a need for an efficient missing data imputation technique. In this paper, we propose a method based on Random forest using Random projection as a data pre-processing filter. Initial results using varying missing data proportions on variety of real datasets show that the imputation process based on Random forest performs equally well or better than K-Nearest Neighbor & Support Vector Regression based methods. Using Random projection we show that dimensionality of a dataset can be reduced by 50 percent without affecting the imputation process.
Lovedeep Gondara
ICMLA1