EDBT 2026 Demo / reviewers in the wild / expert
Mehmet Emre Gursoy
dblp:165/5531 · also Mehmet Emre Gürsoy
· DBLP profile ↗
34ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-7676-0167ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 18 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Computer networks · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Artificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Post-Processing for Utility Improvement under Personalized Local Differential PrivacyabstractLocal Differential Privacy (LDP) has become a widely adopted paradigm for collecting and analyzing sensitive data from user devices. Personalized LDP (PLDP) extends LDP by enabling users to operate under different privacy budgets, better reflecting users' diverse privacy preferences that arise in practice. While PLDP provides greater flexibility, it introduces new challenges for the data collector, particularly in how estimates obtained from users with different privacy budgets should be combined and post-processed to maximize utility. In particular, while post-processing methods have been explored in LDP, similar studies remain lacking for PLDP. In this paper, we present a systematic study of combination and post-processing methods under PLDP. We consider two combination strategies: Simple Averaging (SA) and Inverse Variance Weighting (IVW), as well as three end-to-end post-processing architectures (No-PP, Combine-First, PP-First) that differ in whether post-processing is applied before combination, after combination, or not at all. Through extensive experiments, we show that IVW consistently outperforms SA. We further demonstrate that applying post-processing at the group level before aggregation (PP-First) generally yields higher utility than alternative architectures, although the gap narrows when IVW is used. Our results also reveal that no single post-processing method is universally optimal under PLDP; however, normalization-based methods such as Norm-Sub and Norm-Mul provide strongest performance. Finally, we analyze the impact of population-level privacy preferences and show how the distribution of privacy budgets affects overall utility and user incentives. Together, our results and findings provide practical guidance for designing effective pipelines that improve utility under PLDP. Cagdas Parlak, Dicle Ceylan, Berkay Kemal Balioglu, Alireza Khodaie, Mehmet Emre Gursoy |
CODASPY | 5 |
| 2026 | Understanding Stylistic and Syntactic Backdoor Attacks in Text Classification
Omer Faruk Atasoy, Emirhan Elibol, A. Dilara Yavuz, Mehmet Emre Gursoy |
COMPSAC | 4 |
| 2026 | Budget Inference Attacks and Countermeasures in Locally Differentially Private Data CollectionabstractLocal differential privacy (LDP) has recently become a popular notion for privacy-preserving data collection from user devices. It has been applied in numerous contexts related to the Internet of Things (IoT) and cyber-physical systems to enable privacy-preserving edge data analytics. The strength of privacy protection in LDP deployments depends on the privacy budget ɛ, and there are several scenarios in which it is desirable for the value of ɛ to remain hidden from untrusted third parties, or the inference of ɛ by an untrusted third party may constitute a privacy leakage. In this article, we propose a new class of attacks called budget inference attacks (BIAs), which enable an adversary to infer the ɛ budget value from the outputs of an LDP protocol. We develop BIAs for two types of adversaries: informed adversaries who have knowledge of the statistical data distribution, and uninformed adversaries who do not. We apply our BIAs on five popular LDP protocols and experimentally evaluate them using multiple datasets, varying ɛ budgets, population sizes, and attack settings and parameters. Results show that our BIAs are highly effective, as they enable the adversary to infer the ɛ value with low errors. We also propose three potential countermeasures against our BIAs. Analyses show that while our countermeasures can be effective in reducing BIA accuracy, they also increase utility loss; therefore, the tradeoff between BIA accuracy and utility loss needs to be carefully considered. Berkay Kemal Balioglu, Mehmet Emre Gursoy |
ACM Trans. Internet Techn. | 2 |
| 2025 | Don't Hash Me Like That: Exposing and Mitigating Hash-Induced Unfairness in Local Differential Privacy
Berkay Kemal Balioglu, Alireza Khodaie, Mehmet Emre Gursoy |
ESORICS (4) | 3 |
| 2025 | Post-processing in Local Differential Privacy: An Extensive Evaluation and Benchmark Platform
Alireza Khodaie, Berkay Kemal Balioglu, Mehmet Emre Gursoy |
SEC (1) | 3 |
| 2025 | Learning Bayesian Networks Under Local Differential PrivacyabstractBayesian networks are widely used for causal discovery and probabilistic modeling across diverse domains including healthcare, multi-dimensional data analysis, environmental modeling, and industrial processes. Although previous work has studied the learning of Bayesian networks under centralized differential privacy, to the best of our knowledge, the problem of learning Bayesian networks under local differential privacy (LDP) remains open. In this paper, we address this problem by proposing two solution methods for learning Bayesian networks under LDP: LDP-BN and LDP-BN+. Our first solution called LDP-BN utilizes a novel algorithm for computing mutual information values necessary for building a Bayesian network under LDP, but it suffers from high utility loss since the privacy budget needs to be divided into many pairs of attributes and candidate parent sets. To reduce the amount of noise, we propose LDP-BN+ which utilizes a novel density-aware covering design algorithm that ensures all necessary mutual information values will be computed while the privacy budget is used more effectively. We experimentally evaluate LDP-BN and LDP-BN+ using multiple utility metrics and datasets. Results show that LDP-BN+ outperforms LDP-BN and enables the generation of high-utility Bayesian networks that can be used in practice. Alireza Khodaie, Mehmet Emre Gursoy |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Capacity Planning Under Local Differential Privacy With Optimized Budget SelectionabstractWith the growing popularity of local differential privacy (LDP), there is increasing interest in its deployment in industrial applications, smart homes, and smart cities. However, the main premise of LDP is that data are perturbed to protect privacy, and therefore consumption statistics estimated via LDP are inherently noisy. When noisy estimates are used for capacity planning, they can lead to false positives (false claims of capacity exceedance) or false negatives (actual exceedances are neglected). To address these concerns, this article proposes a system called CAPRI for capacity planning and optimized budget selection in smart city applications under LDP. Based on a specified set of conditions (e.g., number of clients, possible consumption values, LDP protocol) and constraints (e.g., false positive probability should be below 0.01), CAPRI is able to determine the$\varepsilon$privacy budget, which simultaneously satisfies the desired constraints and maximizes clients' privacy. To do so, CAPRI proposes an optimization-based problem formulation and a search-based solution, which relies on LDP simulations. We experimentally validate and demonstrate the effectiveness of CAPRI using real-world and synthetic datasets, three popular LDP protocols, and various constraints and conditions. Seyedpayam Seyedkazemi, Mehmet Emre Gursoy, Yücel Saygin |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Hierarchical Spatial Decompositions under Local Differential PrivacyabstractThe popularity of smartphones, GPS-enabled devices, social networks, and connected vehicles all contribute to the increasing volume of spatial data. Spatial decompositions assist in handling big spatial data, and they have been commonly used in the Differential Privacy (DP) literature for range query answering, spatial indexing, count-of-counts histograms, data summarization, and visualization. However, their applications under the emerging Local DP (LDP) notion are scarce. In this article, we study the problem of building hierarchical spatial decompositions under LDP, focusing on two methods: quadtrees and kd-trees. We develop two solutions for quadtrees: a baseline solution that is inspired by the centralized DP literature, and a proposed solution that utilizes a single data collection step from users, propagates density estimates to remaining nodes, and performs structural corrections to the quadtree. Since kd-trees rely on node medians which are data-dependent, we observe that it is not feasible to build kd-trees using a single data collection step. We therefore propose an iterative solution that constructs kd-trees in top-down fashion by utilizing a novel algorithm for estimating node medians at each tree depth. We experimentally evaluate our quadtree and kd-tree algorithms using four real-world spatial datasets, multiple utility metrics, varying privacy budgets, and tree parameters. Results demonstrate that our algorithms enable the building of accurate spatial decompositions that provide high utility in practice. Notably, our quadtrees and kd-trees achieve substantially lower errors in answering spatial density queries (up to 10-fold improvement) when compared with a state-of-the-art method. Ece Alptekin, Berkay Kemal Balioglu, Mehmet Emre Gursoy |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Injecting Bias into Text Classification Models Using Backdoor Attacks
A. Dilara Yavuz, Mehmet Emre Gursoy |
CRiSIS | 2 |
| 2024 | LuxTrack: Activity Inference Attacks via Smartphone Ambient Light Sensors and CountermeasuresabstractAmbient Light Sensors (ALS) are integrated into mobile devices to enable various functionalities such as automatic adjustment of screen brightness and background color. ALSs can be used to record the light intensity in the surrounding environment without requiring permission from the user; however, this ability raises novel privacy risks. In this paper, we propose LuxTrack, a side-channel privacy attack that uses the ALS of a smartphone to infer the user’s activity on a nearby laptop using the light emitted from the laptop screen. To demonstrate LuxTrack, we developed an Android app that records the light intensity data from the ALS of a mobile device, and used this app to create an ALS light intensity dataset in a controlled environment with real human subjects. From this dataset, LuxTrack extracts a total of 187 features under 6 categories and trains 6 different machine learning models for activity inference. Experiments show that LuxTrack can achieve up to 80% accuracy in inferring the sites/apps the user is viewing on their laptop. We then propose three countermeasures against LuxTrack: binning, smoothing, and noise addition. We demonstrate that while these countermeasures are effective in reducing attack accuracy, they also yield a reduction in the accuracy of legitimate tasks (e.g., adjusting screen background color). By conducting a trade-off analysis between attack accuracy and legitimate task accuracy, we show that the choice of the right countermeasure and parameters can enable the reduction of attack accuracy to below 30% while only incurring 3% loss in legitimate task accuracy. Seyedpayam Seyedkazemi, Mehmet Emre Gursoy, Yücel Saygin |
IEEE Internet Things J. | 2 |
| 2024 | Answering Spatial Density Queries Under Local Differential PrivacyabstractSpatial density queries are fundamental in many geospatial data analysis and crowdsourcing tasks. However, answering spatial density queries may violate users’ privacy by exposing their locations to an untrusted data collector. In this paper, we propose a solution for answering spatial density queries under Local Differential Privacy (LDP), a state-of-the-art privacy protection standard. Our solution consists of four main steps: partitioning, finding sensitivity, user-side noisy response computation, and server-side estimation. For the first step, we propose and analyze three basic partitioning strategies, and based on our analysis, we design an improved strategy called Advanced Partitioning. For the second step, we adapt graph-based modeling of query sets from the centralized DP literature. Advanced Partitioning also leverages and extends this technique by formulating the partitioning problem using vertex coloring. For the third and fourth steps, in addition to adapting two popular LDP protocols (GRR, RAPPOR), we propose a novel extension for the OUE protocol. Our new protocol (OBE) is not only applicable to our problem but can also be used in other LDP problems with bitvector encodings. Finally, we perform an extensive experimental evaluation of different partitioning strategies and protocols using multiple real-world datasets. Results show that Advanced Partitioning and OBE yield the lowest error, demonstrating the superiority of our proposed methods. Ekin Tire, Mehmet Emre Gursoy |
IEEE Internet Things J. | 2 |
| 2023 | Building Quadtrees for Spatial Data Under Local Differential Privacy
Ece Alptekin, Mehmet Emre Gursoy |
DBSec | 2 |
| 2023 | Learning Markov Chain Models from Sequential Data Under Local Differential Privacy
Efehan Guner, Mehmet Emre Gursoy |
ESORICS (2) | 2 |
| 2023 | On the Effectiveness of Re-Identification Attacks and Local Differential Privacy-Based Solutions for Smart Meter Data
Zeynep Sila Kaya, Mehmet Emre Gursoy |
SECRYPT | 2 |
| 2023 | Utility-Aware and Privacy-Preserving Mobile Query ServicesabstractLocation-based queries enable fundamental services for mobile users. While the benefits of location-based services (LBS) are numerous, exposure of mobile users’ locations to untrusted LBS providers may lead to privacy concerns. This article proposesStarCloak, a utility-aware and attack-resilient location anonymization service for privacy-preserving LBS usage.StarCloakcombines several desirable properties. First, unlike conventional approaches which are indifferent to underlying road network structure,StarCloakuses the concept ofstars and proposescloaking graphsfor effective location cloaking on road networks. Second,StarCloaksupports user-specified$k$-user anonymity and$l$-segment indistinguishability, for enabling personalized privacy protection and for serving users with varying privacy preferences. Third,StarCloakachieves strong attack-resilience against replay and query injection attacks through randomized star selection and pruning. Finally, to enable efficient query processing with high throughput and low bandwidth overhead,StarCloakmakes cost-aware star selection decisions by considering query evaluation and network communication costs. We evaluateStarCloakon two datasets using real-world road networks, under various privacy and utility constraints. Results show thatStarCloakachieves improved query success rate and throughput, reduced anonymization time and network usage, and higher attack-resilience in comparison toXStar, its most relevant competitor. Emre Yigitoglu, Mehmet Emre Gursoy, Ling Liu 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | An Adversarial Approach to Protocol Analysis and Selection in Local Differential PrivacyabstractLocal Differential Privacy (LDP) is a popular standard for privacy-preserving data collection. Numerous LDP protocols have been proposed in the literature which differ in how they provide higher utility in different settings. Yet, few have engaged in analyzing the privacy relationships of these protocols under varying settings, and consequently, it is non-trivial to select which LDP protocol is best to use in a newly emerging application. In this paper, we present an adversarial approach to protocol analysis and selection and make three original contributions. First, we introduce a Bayesian adversary to analyze the privacy relationships of LDP protocols under varying settings. We show that different protocols have substantially different responses to the attack effectiveness of the Bayesian adversary, measured in terms of Adversarial Success Rate (ASR). Second, we provide a formal and empirical analysis on a set of privacy and utility-critical factors, including encoding parameters, privacy budget, data domain, adversarial knowledge, and statistical distribution. We show that different settings of these factors have significant effects on the ASRs of LDP protocols, and no protocol provides consistently low ASR across all settings. Third, we design and develop LDPLens, a prototype implementation of our proposed framework. Given a data collection scenario with various factors and constraints, LDPLens enables optimized selection of a desirable LDP protocol for the given scenario. We evaluate the effectiveness of LDPLens using three case studies with real-world datasets. Results show that LDPLens recommends a different protocol in each case study, and the protocol recommended by LDPLens can yield up to 1.5–2 fold reduction in utility loss, ASR or privacy budget compared to a randomly selected protocol. Mehmet Emre Gursoy, Ling Liu 0001, Ka-Ho Chow 0001, Stacey Truex, Wenqi Wei 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Secure and Utility-Aware Data Collection with Condensed Local Differential PrivacyabstractLocal Differential Privacy (LDP) is popularly used in practice for privacy-preserving data collection. Although existing LDP protocols offer high utility for large user populations (100,000 or more users), they perform poorly in scenarios with small user populations (such as those in the cybersecurity domain) and lack perturbation mechanisms that are effective for both ordinal and non-ordinal item sequences while protecting sequence length and content simultaneously. In this paper, we address the small user population problem by introducing the concept of Condensed Local Differential Privacy (CLDP) as a specialization of LDP, and develop a suite of CLDP protocols that offer desirable statistical utility while preserving privacy. Our protocols support different types of client data, ranging from ordinal data types in finite metric spaces (numeric malware infection statistics), to non-ordinal items (OS versions, transaction categories), and to sequences of ordinal and non-ordinal items. Extensive experiments are conducted on multiple datasets, including datasets that are an order of magnitude smaller than those used in existing approaches, which show that proposed CLDP protocols yield high utility. Furthermore, case studies with Symantec datasets demonstrate that our protocols accurately support key cybersecurity-focused tasks of detecting ransomware outbreaks, identifying targeted and vulnerable OSs, and inspecting suspicious activities on infected machines. Mehmet Emre Gursoy, Acar Tamersoy, Stacey Truex, Wenqi Wei 0001, Ling Liu 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Demystifying Membership Inference Attacks in Machine Learning as a ServiceabstractMembership inference attacks seek to infer membership of individual training instances of a model to which an adversary has black-box access through a machine learning-as-a-service API. In providing an in-depth characterization of membership privacy risks against machine learning models, this paper presents a comprehensive study towards demystifying membership inference attacks from two complimentary perspectives. First, we provide a generalized formulation of the development of a black-box membership inference attack model. Second, we characterize the importance of model choice on model vulnerability through a systematic evaluation of a variety of machine learning models and model combinations using multiple datasets. Through formal analysis and empirical evidence from extensive experimentation, we characterize under what conditions a model may be vulnerable to such black-box membership inference attacks. We show that membership inference vulnerability is data-driven and corresponding attack models are largely transferable. Though different model types display different vulnerabilities to membership inference, so do different datasets. Our empirical results additionally show that (1) using the type of target model under attack within the attack model may not increase attack effectiveness and (2) collaborative learning exposes vulnerabilities to membership inference risks when the adversary is a participant. We also discuss countermeasure and mitigation strategies. Stacey Truex, Ling Liu 0001, Mehmet Emre Gursoy, Lei Yu 0002, Wenqi Wei 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | Understanding Object Detection Through an Adversarial Lens
Ka-Ho Chow 0001, Ling Liu 0001, Mehmet Emre Gursoy, Stacey Truex, Wenqi Wei 0001, Yanzhao Wu 0001 |
ESORICS (2) | 3 |
| 2020 | Data Poisoning Attacks Against Federated Learning Systems
Vale Tolpegin, Stacey Truex, Mehmet Emre Gursoy, Ling Liu 0001 |
ESORICS (1) | 3 |
| 2020 | A Framework for Evaluating Client Privacy Leakages in Federated Learning
Wenqi Wei 0001, Ling Liu 0001, Margaret L. Loper, Ka-Ho Chow 0001, Mehmet Emre Gursoy, Stacey Truex, Yanzhao Wu 0001 |
ESORICS (1) | 5 |
| 2020 | Sensitivity Analysis for Non-Interactive Differential Privacy: Bounds and Efficient AlgorithmsabstractDifferential privacy (DP) has gained significant attention lately as the state of the art in privacy protection. It achieves privacy by adding noise to query answers. We study the problem of privately and accurately answering a set of statistical range queries in batch mode (i.e., under non-interactive DP). The noise magnitude in DP depends directly on the sensitivity of a query set, and calculating sensitivity was proven to be NP-hard. Therefore, efficiently bounding the sensitivity of a given query set is still an open research problem. In this work, we propose upper bounds on sensitivity that are tighter than those in previous work. We also propose a formulation to exactly calculate sensitivity for a set of COUNT queries. However, it is impractical to implement these bounds without sophisticated methods. We therefore introduce methods that build a graph model G based on a query set Q, such that implementing the aforementioned bounds can be achieved by solving two well-known clique problems on G. We make use of the literature in solving these clique problems to realize our bounds efficiently. Experimental results show that for query sets with a few hundred queries, it takes only a few seconds to obtain results. Ali Inan, Mehmet Emre Gursoy, Yücel Saygin |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | Known Sample Attacks on Relation Preserving Data TransformationsabstractMany data mining applications such as clustering and k-NN search rely on distances and relations in the data. Thus, distance preserving transformations, which perturb the data but retain records' distances, have emerged as a prominent privacy protection method. In this paper, we present a novel attack on a generalized form of distance preserving transformations, called relation preserving transformations. Our attack exploits not the exact distances between data, but the relationships between the distances. We show that an attacker with few known samples (4 to 10) and direct access to relations can retrieve unknown data records with more than 95 percent precision. In addition, experiments demonstrate that simple methods of noise addition or perturbation are not sufficient to prevent our attack, as they decrease precision by only 10 percent. Emre Kaplan, Mehmet Emre Gursoy, Mehmet Ercan Nergiz, Yücel Saygin |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2019 | Classification of Driving Behavior Events Utilizing Kinematic Classification and Machine Learning for Down Sampled Time Series DataabstractThe proliferation of connected cars globally has the potential to produce torrents of Big Data that will enable improvements in driver safety, new location based services, improvements in vehicle quality, and optimized vehicle designs. One aspect of connected car data involves driving behavior data and its use for Usage Based Insurance (UBI). UBI has become one of the most widely used applications of driving behavior data. Currently, the transmission and processing of high frequency driving behavior data from the connected car to the cloud is limited by wireless data costs and in-vehicle hardware complexity. To alleviate these issues, we detail the development of a machine learning framework utilizing a kinematic classification methodology applied to down sampled time series vehicle data sets for accurate imputation of driving behavior events in UBI applications. The down-sampled data, consisting of 5 second frames with data fields of timestamp, vehicle speed, and vehicle acceleration is classified into unique kinematic clusters to standardize any driving behavior data distributions. Subsequently, machine learning is used to impute harsh driving events for each 5 second frame in select kinematic clusters. This novel machine learning methodology reduced data set sizes by 75%, utilized a limited set of five attributes, and achieved an average precision and recall of 84.5% and 63.5% for two distinct connected car data sets with 1/5 Hz down-sampled data. Vikram Krishnamurthy, Kusha Nezafati, Juhyun Bae, Mehmet Emre Gursoy, Mian Zhong, Vikrant Singh |
IEEE BigData | 4 |
| 2019 | Deep Neural Network Ensembles Against Deception: Ensemble Diversity, Accuracy and RobustnessabstractEnsemble learning is a methodology that integrates multiple DNN learners for improving prediction performance of individual learners. Diversity is greater when the errors of the ensemble prediction is more uniformly distributed. Greater diversity is highly correlated with the increase in ensemble accuracy. Another attractive property of diversity optimized ensemble learning is its robustness against deception: an adversarial perturbation attack can mislead one DNN model to misclassify but may not fool other ensemble DNN members consistently. In this paper we first give an overview of the concept of ensemble diversity and examine the three types of ensemble diversity in the context of DNN classifiers. We then describe a set of ensemble diversity measures, a suite of algorithms for creating diversity ensembles and for performing ensemble consensus (voted or learned) for generating high accuracy ensemble output by strategically combining outputs of individual members. This paper concludes with a discussion on a set of open issues in quantifying ensemble diversity for robust deep learning. Ling Liu 0001, Wenqi Wei 0001, Ka-Ho Chow 0001, Margaret L. Loper, Mehmet Emre Gursoy, Stacey Truex, Yanzhao Wu 0001 |
MASS | 5 |
| 2019 | Differentially Private Model Publishing for Deep LearningabstractDeep learning techniques based on neural networks have shown significant success in a wide range of AI tasks. Large-scale training datasets are one of the critical factors for their success. However, when the training datasets are crowdsourced from individuals and contain sensitive information, the model parameters may encode private information and bear the risks of privacy leakage. The recent growing trend of the sharing and publishing of pre-trained models further aggravates such privacy risks. To tackle this problem, we propose a differentially private approach for training neural networks. Our approach includes several new techniques for optimizing both privacy loss and model accuracy. We employ a generalization of differential privacy called concentrated differential privacy(CDP), with both a formal and refined privacy loss analysis on two different data batching methods. We implement a dynamic privacy budget allocator over the course of training to improve model accuracy. Extensive experiments demonstrate that our approach effectively improves privacy loss accounting, training efficiency and model quality under a given privacy budget. Lei Yu 0002, Ling Liu 0001, Calton Pu, Mehmet Emre Gursoy, Stacey Truex |
IEEE Symposium on Security and Privacy | 4 |
| 2019 | Differentially Private and Utility Preserving Publication of Trajectory DataabstractThe universal popularity of GPS-enabled mobile devices and traffic navigation services has fueled the growth of trajectory data, as evidenced by Uber Movement and NYC taxi data release. Although trajectory data can generate valuable insights and value-added services for many, publishing this data while respecting mobile users' privacy has been a long-standing challenge. In this paper, we present DP-Star, a methodical framework for publishing trajectory data with differential privacy guarantee as well as high utility preservation. DP-Star relies on a novel combination of several components. First, DP-Star's normalization algorithm uses the Minimum Description Length metric to summarize raw trajectories using their representative points, thereby achieving a desirable trade-off between the preciseness and conciseness of their information content. Second, DP-Star constructs a density-aware grid which ensures spatial densities can be preserved despite the noise added to satisfy differential privacy. Third, DP-Star preserves the correlations between trajectories' end points through a private trip distribution, and intermediate points through a private Markov mobility model. Finally, DP-Star estimates users' trip lengths using a median length estimation method, and generates synthetic trajectories that preserve both differential privacy and high utility. Our experimental comparison shows that DP-Star significantly outperforms existing approaches in terms of trajectory utility and accuracy. Mehmet Emre Gursoy, Ling Liu 0001, Stacey Truex, Lei Yu 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2018 | PrivacyZone: A Novel Approach to Protecting Location Privacy of Mobile UsersabstractWhile location-based services and applications are increasing in popularity, there are growing concerns over users' location privacy. Although there exist general purpose mobile permission systems and cloaking techniques, these techniques suffer from several problems when applied to continuous location and GPS access, as they are often rigid, coarse-grained, not sufficiently personalizable, and unaware of road network semantics. This paper proposes PrivacyZone, a novel system for constructing personalized fine-grained privacy quarantine regions and protecting users' privacy within these regions. PrivacyZone allows users to seamlessly enter their privacy specifications under spatial, temporal, and semantic customization. Novel challenges arise from having to enforce privacy zones for large volume and variety of users with frequent location updates. We show that naive privacy zone processing techniques are inefficient and cause excessive energy consumption. We therefore develop advanced processing techniques based on the concept of safe hibernation. We empirically evaluate our techniques to demonstrate their trade-offs with respect to hibernation time, computation effort, and network bandwidth usage. Our results show that PrivacyZone is efficient, scalable, and flexible, while preserving users' location privacy. Emre Yigitoglu, Mehmet Emre Gursoy, Ling Liu 0001, Margaret L. Loper, Bhuvan Bamba, Kisung Lee |
IEEE BigData | 2 |
| 2018 | Utility-Aware Synthesis of Differentially Private and Attack-Resilient Location TracesabstractAs mobile devices and location-based services become increasingly ubiquitous, the privacy of mobile users' location traces continues to be a major concern. Traditional privacy solutions rely on perturbing each position in a user's trace and replacing it with a fake location. However, recent studies have shown that such point-based perturbation of locations is susceptible to inference attacks and suffers from serious utility losses, because it disregards the moving trajectory and continuity in full location traces. In this paper, we argue that privacy-preserving synthesis of complete location traces can be an effective solution to this problem. We present AdaTrace, a scalable location trace synthesizer with three novel features: provable statistical privacy, deterministic attack resilience, and strong utility preservation. AdaTrace builds a generative model from a given set of real traces through a four-phase synthesis process consisting of feature extraction, synopsis learning, privacy and utility preserving noise injection, and generation of differentially private synthetic location traces. The output traces crafted by AdaTrace preserve utility-critical information existing in real traces, and are robust against known location trace attacks. We validate the effectiveness of AdaTrace by comparing it with three state of the art approaches (ngram, DPT, and SGLT) using real location trace datasets (Geolife and Taxi) as well as a simulated dataset of 50,000 vehicles in Oldenburg, Germany. AdaTrace offers up to 3-fold improvement in trajectory utility, and is orders of magnitude faster than previous work, while preserving differential privacy and attack resilience. Mehmet Emre Gursoy, Ling Liu 0001, Stacey Truex, Lei Yu 0002, Wenqi Wei 0001 |
CCS | 1 |
| 2018 | Location disclosure risks of releasing trajectory distances
Emre Kaplan, Mehmet Emre Gursoy, Mehmet Ercan Nergiz, Yücel Saygin |
Data Knowl. Eng. | 2 |
| 2017 | Differentially private nearest neighbor classification
Mehmet Emre Gursoy, Ali Inan, Mehmet Ercan Nergiz, Yücel Saygin |
Data Min. Knowl. Discov. | 1 |
| 2016 | Graph-based modelling of query sets for differential privacyabstractDifferential privacy has gained attention from the community as the mechanism for privacy protection. Significant effort has focused on its application to data analysis, where statistical queries are submitted in batch and answers to these queries are perturbed with noise. The magnitude of this noise depends on the privacy parameter ϵ and the sensitivity of the query set. However, computing the sensitivity is known to be NP-hard. Ali Inan, Mehmet Emre Gursoy, Emir Esmerdag, Yücel Saygin |
SSDBM | 2 |
| 2016 | Privacy-Preserving Publishing of Hierarchical DataabstractMany applications today rely on storage and management of semi-structured information, for example, XML databases and document-oriented databases. These data often have to be shared with untrusted third parties, which makes individuals’ privacy a fundamental problem. In this article, we propose anonymization techniques for privacy-preserving publishing of hierarchical data. We show that the problem of anonymizing hierarchical data poses unique challenges that cannot be readily solved by existing mechanisms. We extend two standards for privacy protection in tabular data ( k -anonymity and ℓ-diversity) and apply them to hierarchical data. We present utility-aware algorithms that enforce these definitions of privacy using generalizations and suppressions of data values. To evaluate our algorithms and their heuristics, we experiment on synthetic and real datasets obtained from two universities. Our experiments show that we significantly outperform related methods that provide comparable privacy guarantees. Ismet Ozalp, Mehmet Emre Gursoy, Mehmet Ercan Nergiz, Yücel Saygin |
ACM Trans. Priv. Secur. | 2 |
| 2015 | Scalable reactive vehicle-to-vehicle congestion avoidance mechanismabstractThe increasing popularity and acceptance of VANETs will make the deployment of autonomous vehicles easier and faster since the VANET will reduce dependence on expensive sensors. Many useful applications will be possible with the usage of VANETs, which will improve the safety and quality of trips for the owners of these vehicles. One of these applications is the avoidance of traffic congestion by smart dynamic rerouting. For scalability, current cloud-based solutions, like Google Maps traffic, update congestion levels after a time interval rather than providing real-time measurements. In this paper, we introduce a vehicle-to-vehicle congestion avoidance mechanism, which detects real-time congestion levels and reroutes vehicles accordingly to minimize their trip times. Our system is highly distributed and is, therefore, not subjected to the limitations of centralized congestion avoidance mechanisms. We show via simulation that our system can significantly decrease the trip times of vehicles as well as the average car density on the map. Our proposed system, with its checkpoint and offline path generation approaches, is more responsive to local congestion level changes and computationally less complex for least congested route calculations than state-of-the-art congestion avoidance mechanisms. Mevlut Turker Garip, Mehmet Emre Gursoy, Peter L. Reiher, Mario Gerla |
CCNC | 2 |