EDBT 2026 Demo / reviewers in the wild / expert
Sagar Sharma
dblp:167/2117
· DBLP profile ↗
16ranked-venue papers
10as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 10 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Auditing Differentially Private Interactive Database SystemsabstractThis paper introduces an empirical framework for auditing privacy leakage in interactive database systems (DBS) that implement differential privacy (DP). Without making any assumptions about the formal DP mechanisms or parameters in use, we simulate an auditor with black-box access to query outputs and evaluate privacy leakage using membership inference attacks (MIAs). Our framework provides empirical lower bounds on the privacy loss parameter e based on attack success, providing a signal of privacy risk even when a theoretical analysis is not available or verifiable. We implement this framework in a system modeled after a major social media company's production environment and show how factors like data distribution, target selection, and query specificity affect the observed privacy. Our work offers a valuable and practical tool for red-teaming and auditing privacy in large-scale opaque DBS. Sagar Sharma, Wanrong Zhang 0004, Florian Tramèr |
AsiaCCS | 1 |
| 2025 | Meeting Utility Constraints in Differential Privacy: A Privacy-Boosting ApproachabstractData engineering often requires accuracy (utility) constraints on results, posing significant challenges in designing differentially private (DP) mechanisms, particularly under stringent privacy parameter$\epsilon$. In this paper, we propose a privacy-boosting framework that is compatible with most noise-adding DP mechanisms. Our framework enhances the likelihood of outputs falling within a preferred subset of the support to meet utility requirements while enlarging the overall variance to reduce privacy leakage. We characterize the privacy loss distribution of our framework and present the privacy profile formulation for$(\epsilon,\ \delta)-\mathbf{DP}$and Rényi DP (RDP) guarantees. We study special cases involving data-dependent and data-independent utility formulations. Through extensive experiments, we demonstrate that our framework achieves lower privacy loss than standard DP mechanisms under utility constraints. Notably, our approach is particularly effective in reducing privacy loss with large query sensitivity relative to the true answer, offering a more practical and flexible approach to designing differentially private mechanisms that meet specific utility constraints. Wanrong Zhang 0004, Donghang Lu, Sagar Sharma |
SP | 5 |
| 2024 | Budget Recycling Differential PrivacyabstractDifferential Privacy (DP) mechanisms usually force reduction in data utility by producing "out-of-bound" noisy results for a tight privacy budget. We introduce the Budget Recycling Differential Privacy (BR-DP) framework, designed to provide soft-bounded noisy outputs for a broad range of existing DP mechanisms. By "soft-bounded," we refer to the mechanism’s ability to release most outputs within a predefined error boundary, thereby improving utility and maintaining privacy simultaneously. The core of BR-DP consists of two components: a DP kernel responsible for generating a noisy answer per iteration, and a recycler that probabilistically recycles/regenerates or releases the noisy answer. We delve into the privacy accounting of BR-DP, culminating in the development of a budgeting principle that optimally sub-allocates the available budget between the DP kernel and the recycler. Furthermore, we introduce algorithms for tight BR-DP accounting in composition scenarios, and our findings indicate that BR-DP achieves reduced privacy leakage post-composition compared to DP. Additionally, we explore the concept of privacy amplification via subsampling within the BR-DP framework and propose optimal sampling rates for BR-DP across various queries. We experiment with real data, and the results demonstrate BR-DP’s effectiveness in lifting the utility-privacy tradeoff provided by DP mechanisms. Sagar Sharma |
SP | 3 |
| 2023 | Demo: Image Disguising for Scalable GPU-accelerated Confidential Deep LearningabstractDeep learning training involves large training data and expensive model tweaking, for which cloud GPU resources can be a popular option. However, outsourcing data often raises privacy concerns. The challenge is to preserve data and model confidentiality without sacrificing GPU-based scalable training and low-cost client-side preprocessing, which is difficult for conventional cryptographic solutions to achieve. This demonstration shows a new approach, image disguising, represented by recent work: DisguisedNets, NeuraCrypt, and InstaHide, which aim to securely transform training images while still enabling the desired scalability and efficiency. We present an interactive system for visually and comparatively exploring these methods. Users can view disguised images, note low client-side processing costs, and observe the maintained efficiency and model quality during server-side GPU-accelerated training. This demo aids researchers and practitioners in swiftly grasping the advantages and limitations of image-disguising methods. Yuechun Gu, Sagar Sharma, Keke Chen |
CCS | 2 |
| 2023 | DisguisedNets: Secure Image Outsourcing for Confidential Model Training in CloudsabstractLarge training data and expensive model tweaking are standard features of deep learning with images. As a result, data owners often utilize cloud resources to develop large-scale complex models, which also raises privacy concerns. Existing cryptographic solutions for training deep neural networks (DNNs) are too expensive, cannot effectively utilize cloud GPU resources, and also put a significant burden on client-side pre-processing. This article presents an image disguising approach: DisguisedNets, which allows users to securely outsource images to the cloud and enables confidential, efficient GPU-based model training. DisguisedNets uses a novel combination of image blocktization, block-level random permutation, and block-level secure transformations: random multidimensional projection (RMT) or AES pixel-level encryption (AES) to transform training data. Users can use existing DNN training methods and GPU resources without any modification to training models with disguised images. We have analyzed and evaluated the methods under a multi-level threat model and compared them with another similar method—InstaHide. We also show that the image disguising approach, including both DisguisedNets and InstaHide, can effectively protect models from model-targeted attacks. Keke Chen, Yuechun Gu, Sagar Sharma |
ACM Trans. Internet Techn. | 3 |
| 2021 | Image Disguising for Protecting Data and Model Confidentiality in Outsourced Deep LearningabstractLarge training data and expensive model tweaking are common features of deep learning development for images. As a result, data owners often utilize cloud resources or machine learning service providers for developing large-scale complex models. This practice, however, raises serious privacy concerns. Existing solutions are either too expensive to be practical, or do not sufficiently protect the confidentiality of data and model. In this paper, we aim to achieve a better trade-off among the level of protection for outsourced DNN model training, the expenses, and the utility of data, using novel image disguising mechanisms. We design a suite of image disguising methods that are efficient to implement and then analyze them to understand multiple levels of tradeoffs between data utility and protection of confidentiality. The experimental evaluation shows the surprising ability of DNN modeling methods in discovering patterns in disguised images and the flexibility of these image disguising mechanisms in achieving different levels of resilience to attacks. Sagar Sharma, A. K. M. Mubashwir Alam, Keke Chen |
CLOUD | 1 |
| 2021 | Confidential machine learning on untrusted platforms: a surveyabstractWith the ever-growing data and the need for developing powerful machine learning models, data owners increasingly depend on various untrusted platforms (e.g., public clouds, edges, and machine learning service providers) for scalable processing or collaborative learning. Thus, sensitive data and models are in danger of unauthorized access, misuse, and privacy compromises. A relatively new body of research confidentially trains machine learning models on protected data to address these concerns. In this survey, we summarize notable studies in this emerging area of research. With a unified framework, we highlight the critical challenges and innovations in outsourcing machine learning confidentially. We focus on the cryptographic approaches for confidential machine learning (CML), primarily on model training, while also covering other directions such as perturbation-based approaches and CML in the hardware-assisted computing environment. The discussion will take a holistic way to consider a rich context of the related threat models, security assumptions, design principles, and associated trade-offs amongst data utility, cost, and confidentiality. Sagar Sharma, Keke Chen |
Cybersecur. | 1 |
| 2021 | SGX-MR: Regulating Dataflows for Protecting Access Patterns of Data-Intensive SGX ApplicationsabstractAbstract Intel SGX has been a popular trusted execution environment (TEE) for protecting the integrity and confidentiality of applications running on untrusted platforms such as cloud. However, the access patterns of SGX-based programs can still be observed by adversaries, which may leak important information for successful attacks. Researchers have been experimenting with Oblivious RAM (ORAM) to address the privacy of access patterns. ORAM is a powerful low-level primitive that provides application-agnostic protection for any I/O operations, however, at a high cost. We find that some application-specific access patterns, such as sequential block I/O, do not provide additional information to adversaries. Others, such as sorting, can be replaced with specific oblivious algorithms that are more efficient than ORAM. The challenge is that developers may need to look into all the details of application-specific access patterns to design suitable solutions, which is time-consuming and error-prone. In this paper, we present the lightweight SGX based MapReduce (SGX-MR) approach that regulates the dataflow of data-intensive SGX applications for easier application-level access-pattern analysis and protection. It uses the MapReduce framework to cover a large class of data-intensive applications, and the entire framework can be implemented with a small memory footprint. With this framework, we have examined the stages of data processing, identified the access patterns that need protection, and designed corresponding efficient protection methods. Our experiments show that SGX-MR based applications are much more efficient than the ORAM-based implementations. A. K. M. Mubashwir Alam, Sagar Sharma, Keke Chen |
Proc. Priv. Enhancing Technol. | 2 |
| 2019 | Confidential Boosting with Random Linear Classifiers for Outsourced User-Generated Data
Sagar Sharma, Keke Chen |
ESORICS (1) | 1 |
| 2019 | PrivateGraph: Privacy-Preserving Spectral Analysis of Encrypted Graphs in the CloudabstractBig graphs, such as user interactions in social networks and customer rating matrices in collaborative filters, possess great values for both businesses and research. They are not only big but often keep evolving, which requires a large amount of computing resources to maintain. With the wide deployment of public cloud resources, owners of big graphs may want to use cloud resources to obtain storage and computation scalability. However, privacy and ownership of the graphs in the cloud has become a major concern. In this paper, we study privacy-preserving algorithms for one of the important graph analysis techniques-graph spectral analysis for outsourced graph in the cloud. The core operation: eigendecomposition of large matrix is also important to many data mining algorithms. We consider a cloud-centric framework with three collaborative parties: data contributors, data owner, and cloud provider. Graphs are represented as matrices such as adjacency matrix and Laplacian matrix, the elements of which are encrypted and submitted by distributed contributors. The data owner then interacts with the cloud-side programs to conduct spectral analysis, while protecting data privacy from the honest-but-curious cloud provider. For a N × N graph matrix, we aim to design algorithms with the cloud handling expensive storage and computation in O(N2) complexity, while data owner and data contributors' algorithms take only O(N). To achieve this goal, we develop the privacy-preserving versions of the two approximate eigendecomposition algorithms: the Lanczos algorithm and the Nyström algorithm, considering two encryption methods: additive homomorphic encryption (AHE) methods and somewhat homomorphic encryption (SHE) methods. Both dense and sparse matrices are studied, while sparse matrices also involve a differentially private data submission protocol to allow the trade-off between data sparsity and privacy. Experimental results show that the Nyströ algorithm with sparse encoding can dramatically reduce data owners' costs. SHE-based methods have lower computational time while AHE-based methods have lower communication/storage costs. Sagar Sharma, James Powers, Keke Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Image Disguising for Privacy-preserving Deep LearningabstractDue to the high training costs of deep learning, model developers often rent cloud GPU servers to achieve better efficiency. However, this practice raises privacy concerns. An adversarial party may be interested in 1) personal identifiable information encoded in the training data and the learned models, 2) misusing the sensitive models for its own benefits, or 3) launching model inversion (MIA) and generative adversarial network (GAN) attacks to reconstruct replicas of training data (e.g., sensitive images). Learning from encrypted data seems impractical due to the large training data and expensive learning algorithms, while differential-privacy based approaches have to make significant trade-offs between privacy and model quality. We investigate the use of image disguising techniques to protect both data and model privacy. Our preliminary results show that with block-wise permutation and transformations, surprisingly, disguised images still give reasonably well performing deep neural networks (DNN). The disguised images are also resilient to the deep-learning enhanced visual discrimination attack and provide an extra layer of protection from MIA and GAN attacks. Sagar Sharma, Keke Chen |
CCS | 1 |
| 2018 | Privacy-Preserving Boosting with Random Linear ClassifiersabstractWe propose SecureBoost, a privacy-preserving predictive modeling framework, that allows service providers (SPs) to build powerful boosting models over encrypted or randomly masked user submitted data. SecureBoost uses random linear classifiers (RLCs) as the base classifiers. A Cryptographic Service Provider (CSP) manages keys and assists the SP's processing to reduce the complexity of the protocol constructions. The SP learns only the base models (i.e., RLCs) and the CSP learns only the weights of the base models and a limited leakage function. This separated parameter holding avoids any party from abusing the final model or conducting model-based attacks. We evaluate two constructions of SecureBoost: HE+GC and SecSh+GC using combinations of primitives - homomorphic encryption, garbled circuits, and random masking. We show that SecureBoost efficiently learns high-quality boosting models from protected user-generated data with practical costs. Sagar Sharma, Keke Chen |
CCS | 1 |
| 2018 | Device to monitor quiet breath of CRD patientsabstractRising pollution worldwide is a major reason for increasing number of patients of chronic respiratory diseases (CRD) like Asthma. CRD are irreversible, can only be managed through proper care and medication. Spirometry is the only known test for finding the functional health of lungs but 90% of CRD patients fail to perform the test. Moreover, Spirometry is non reproducible and difference of about 20% were reported over two tests in reasonably close time instance. This is a painful test as a patient has to make full capacity forced exhalation without coughing, a rare possibility with CRD patients. Device prototype tested here is S-MASK, envisioned to assess the level of breathing discomfort experienced by a CRD patient. S-MASK is a compact and sensitive device that can capture the normal or quiet breath. Design objective of S-MASK is to develop a device that can be used on regular basis to assess the wellness and relative/absolute relief from CRD symptoms.The design of S-MASK and its validation is presented in this paper. Breath rate computed from the device has R-squared = 0.9 and p-Value-4with respect to manually observed values. Breath samples collected using S-MASK found to have μ = 0.308 Hz and σ = 0.073 Hz which corresponds to μ = 18.66 bpm (breath per minute) and σ = 4.38 bpm. These values are significantly close to manually observed breath rate i.e. μ = 17.8 bpm and σ = 4.22 bpm.These results are statistically significant to validate the capabilities of S-MASK to capture the quite breath and identify the rate. S-MASK is device of its kind that computes the breath rate using frequency spectrum. Sagar Sharma, Rakesh Kumar Mishra |
HealthCom | 1 |
| 2017 | PrivateGraph: A Cloud-Centric System for Spectral Analysis of Large Encrypted GraphsabstractGraph datasets have invaluable use in business applications and scientific research. Because of the growing size and dynamically changing nature of graphs, graph data owners may want to use public cloud infrastructures to store, process, and perform graph analytics. However, when outsourcing data and computation, data owners are at burden to develop methods to preserve data privacy and data ownership from curious cloud providers. This demonstration exhibits a prototype system for privacy-preserving spectral analysis framework for large graphs in public clouds (PrivateGraph) that allows data owners to collect graph data from data contributors, and store and conduct secure graph spectral analysis in the cloud with preserved privacy and ownership. This demo system lets its audience interactively learn the major cloud-client interaction protocols: the privacy-preserving data submission, the secure Lanczos and Nyström approximate eigen-decomposition algorithms that work over encrypted data, and the outcome of an important application of spectral analysis - spectral clustering. In the process of demonstration the audience will understand the intrinsic relationship amongst costs, result quality, privacy, and scalability of the framework. Sagar Sharma, Keke Chen |
ICDCS | 1 |
| 2016 | Privacy-Preserving Spectral Analysis of Large Graphs in Public CloudsabstractLarge graph datasets have become invaluable assets for studying problems in business applications and scientific research. These datasets, collected and owned by data owners, may also contain privacy-sensitive information. When using public clouds for elastic processing, data owners have to protect both data ownership and privacy from curious cloud providers. We propose a cloud-centric framework that allows data owners to efficiently collect graph data from the distributed data contributors, and privately store and analyze graph data in the cloud. Data owners can conduct expensive operations in untrusted public clouds with privacy and scalability preserved. The major contributions of this work include two privacy-preserving approximate eigen decomposition algorithms (the secure Lanczos and Nystrom methods) for spectral analysis of large graph matrices, and a personalized privacy-preserving data submission method based on differential privacy that allows for the trade-off between data sparsity and privacy. For a N-node graph, the proposed approach allows a data owner to finish the core operations with only O(N) client-side costs in computation, storage, and communication. The expensive O(N2) operations are performed in the cloud with the proposed privacy-preserving algorithms. We prove that our approach can satisfactorily preserve data privacy against the untrusted cloud providers. We have conducted an extensive experimental study to investigate these algorithms in terms of the intrinsic relationships among costs, privacy, scalability, and result quality. Sagar Sharma, James Powers, Keke Chen |
AsiaCCS | 1 |
| 2015 | Scalable Euclidean Embedding for Big DataabstractEuclidean embedding algorithms transform data defined in an arbitrary metric space to the Euclidean space, which is critical to many visualization techniques. At big-data scale, these algorithms need to be scalable to massive data-parallel infrastructures. Designing such scalable algorithms and understanding the factors affecting the algorithms are important research problems for visually analyzing big data. We propose a framework that extends the existing Euclidean embedding algorithms to scalable ones. Specifically, it decomposes an existing algorithm into naturally parallel components and non-parallelizable components. Then, data parallel implementations such as MapReduce and data reduction techniques are applied to the two categories of components, respectively. We show that this can be possibly done for a collection of embedding algorithms. Extensive experiments are conducted to understand the important factors in these scalable algorithms: scalability, time cost, and the effect of data reduction to result quality. The result on sample algorithms: Fast Map-MR and LMDS-MR shows that with the proposed approach the derived algorithms can preserve result quality well, while achieving desirable scalability. Zohreh Alavi, Sagar Sharma, Keke Chen |
CLOUD | 2 |