Luca Bonomi

dblp:117/9414 · DBLP profile ↗
← Back
25ranked-venue papers
14as first author
10since 2021 · last 2025
0000-0002-5751-1341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 13 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Quality and Robustness of Generative Watermarking on Biomedical Data
abstract
Sharing biomedical data for research can advance methodology development and improve patient care. However, data ownership, integrity, and authenticity are among the key concerns for broad biomedical data sharing. Data watermarking, a technique that embeds a message in the data, has been applied in medical imaging to address those concerns. Specifically, the majority of medical image watermarking methods embed watermarks in the frequency domains to achieve minimal data distortions. However, such methods exhibit low robustness, where perturbing the watermarked data may challenge watermark detection. In this work, we study the state-of-the-art watermarking methods based on generative models and examine their applicability to biomedical data. Those methods aim to embed watermarks in the latent space of generative models for improved robustness. We evaluate the quality of watermarked data and adopt a wide range of image perturbations to quantify the robustness of watermark detection. Our results with two real biomedical data sets show that generative watermarking is more robust than the classic approach. While pre-trained generative models may be adopted to watermark biomedical data, there might be a gap in data quality and future work could conduct comprehensive, task-specific evaluations.
Liyue Fan, Luca Bonomi
IJCNN2
2023 Private Continuous Survival Analysis with Distributed Multi-Site Data
abstract
Effective disease surveillance systems require large-scale epidemiological data to improve health outcomes and quality of care for the general population. As data may be limited within a single site, multi-site data (e.g., from a number of local/regional health systems) need to be considered. Leveraging distributed data across multiple sites for epidemiological analysis poses significant challenges. Due to the sensitive nature of epidemiological data, it is imperative to design distributed solutions that provide strong privacy protections. Current privacy solutions often assume a central site, which is responsible for aggregating the distributed data and applying privacy protection before sharing the results (e.g., aggregation via secure primitives and differential privacy for sharing aggregate results). However, identifying such a central site may be difficult in practice and relying on a central site may introduce potential vulnerabilities (e.g., single point of failure). Furthermore, to support clinical interventions and inform policy decisions in a timely manner, epidemiological analysis need to reflect dynamic changes in the data. Yet, existing distributed privacy-protecting approaches were largely designed for static data (e.g., one-time data sharing) and cannot fulfill dynamic data requirements. In this work, we propose a privacy-protecting approach that supports the sharing of dynamic epidemiological analysis and provides strong privacy protection in a decentralized manner. We apply our solution in continuous survival analysis using the Kaplan-Meier estimation model while providing differential privacy protection. Our evaluations on a real dataset containing COVID-19 cases show that our method provides highly usable results.
Luca Bonomi, Marilyn Lionts, Liyue Fan
IEEE Big Data1
2023 Enabling Health Data Sharing with Fine-Grained Privacy
abstract
Sharing health data is vital in advancing medical research and transforming knowledge into clinical practice. Meanwhile, protecting the privacy of data contributors is of paramount importance. To that end, several privacy approaches have been proposed to protect individual data contributors in data sharing, including data anonymization and data synthesis techniques. These approaches have shown promising results in providing privacy protection at the dataset level. In this work, we study the privacy challenges in enabling fine-grained privacy in health data sharing. Our work is motivated by recent research findings, in which patients and healthcare providers may have different privacy preferences and policies that need to be addressed. Specifically, we propose a novel and effective privacy solution that enables data curators (e.g., healthcare providers) to protect sensitive data elements while preserving data usefulness. Our solution builds on randomized techniques to provide rigorous privacy protection for sensitive elements and leverages graphical models to mitigate privacy leakage due to dependent elements. To enhance the usefulness of the shared data, our randomized mechanism incorporates domain knowledge to preserve semantic similarity and adopts a block-structured design to minimize utility loss. Evaluations with real-world health data demonstrate the effectiveness of our approach and the usefulness of the shared data for health applications.
Luca Bonomi, Sepand Gousheh, Liyue Fan
CIKM1
2023 Hide Your Distance: Privacy Risks and Protection in Spatial Accessibility Analysis
abstract
Measuring spatial accessibility to healthcare resources and facilities has long been an important problem in public health. For example, during disease outbreaks, sharing spatial accessibility data such as individual travel distances to health facilities is vital to policy making and designing effective interventions. However, sharing these data may raise privacy concerns, as information about individual data contributors (e.g., health status and residential address) may be disclosed. In this work, we investigate those unintended information leakage in spatial accessibility analysis. Specifically, we are interested in understanding whether sharing data for spatial accessibility computations may disclose individual participation (i.e., membership inference) and personal identifiable information (i.e., address inference). Furthermore, we propose two provably private algorithms that mitigate those privacy risks. The evaluation is conducted with real population and healthcare facilities data from Mecklenburg county, NC and Nashville, TN. Compared to state-of-the-art privacy practices, our methods effectively reduce the risks of membership and address disclosure, while providing useful data for spatial accessibility analysis.
Liyue Fan, Luca Bonomi
SIGSPATIAL/GIS2
2022 Training Deep Learning Models Privately in Survival Studies
Liyue Fan, Luca Bonomi
AMIA2
2022 Sharing personal ECG time-series data privately
abstract
OBJECTIVE: Emerging technologies (eg, wearable devices) have made it possible to collect data directly from individuals (eg, time-series), providing new insights on the health and well-being of individual patients. Broadening the access to these data would facilitate the integration with existing data sources (eg, clinical and genomic data) and advance medical research. Compared to traditional health data, these data are collected directly from individuals, are highly unique and provide fine-grained information, posing new privacy challenges. In this work, we study the applicability of a novel privacy model to enable individual-level time-series data sharing while maintaining the usability for data analytics. METHODS AND MATERIALS: We propose a privacy-protecting method for sharing individual-level electrocardiography (ECG) time-series data, which leverages dimensional reduction technique and random sampling to achieve provable privacy protection. We show that our solution provides strong privacy protection against an informed adversarial model while enabling useful aggregate-level analysis. RESULTS: We conduct our evaluations on 2 real-world ECG datasets. Our empirical results show that the privacy risk is significantly reduced after sanitization while the data usability is retained for a variety of clinical tasks (eg, predictive modeling and clustering). DISCUSSION: Our study investigates the privacy risk in sharing individual-level ECG time-series data. We demonstrate that individual-level data can be highly unique, requiring new privacy solutions to protect data contributors. CONCLUSION: The results suggest our proposed privacy-protection method provides strong privacy protections while preserving the usefulness of the data.
Luca Bonomi, Zeyun Wu, Liyue Fan
J. Am. Medical Informatics Assoc.1
2022 VERTICOX: Vertically Distributed Cox Proportional Hazards Model Using the Alternating Direction Method of Multipliers
abstract
The Cox proportional hazards model is a popular semi-parametric model for survival analysis. In this paper, we aim at developing a federated algorithm for the Cox proportional hazards model over vertically partitioned data (i.e., data from the same patient are stored at different institutions). We propose a novel algorithm, namely VERTICOX, to obtain the global model parameters in a distributed fashion based on the Alternating Direction Method of Multipliers (ADMM) framework. The proposed model computes intermediary statistics and exchanges them to calculate the global model without collecting individual patient-level data. We demonstrate that our algorithm achieves equivalent accuracy for the estimation of model parameters and statistics to that of its centralized realization. The proposed algorithm converges linearly under the ADMM framework. Its computational complexity and communication costs are polynomially and linearly associated with the number of subjects, respectively. Experimental results show that VERTICOX can achieve accurate model parameter estimation to support federated survival analysis over vertically distributed data by saving bandwidth and avoiding exchange of information about individual patients. The source code for VERTICOX is available at: https://github.com/daiwenrui/VERTICOX.
Wenrui Dai, Xiaoqian Jiang, Luca Bonomi, Yong Li 0033, Hongkai Xiong, Lucila Ohno-Machado
IEEE Trans. Knowl. Data Eng.3
2021 Collecting and Sharing ECG Data with Privacy Protection
Liyue Fan, Zeyun Wu, Luca Bonomi
AMIA3
2021 Privacy-Preserving Federated Biomedical Data Analysis
Xiaoqian Jiang, Luca Bonomi, Jaideep Vaidya, Li Xiong 0001, Lucila Ohno-Machado
AMIA2
2021 Noise-tolerant similarity search in temporal medical data
Luca Bonomi, Liyue Fan, Xiaoqian Jiang
J. Biomed. Informatics1
2020 Protecting patient privacy in survival analyses
abstract
OBJECTIVE: Survival analysis is the cornerstone of many healthcare applications in which the "survival" probability (eg, time free from a certain disease, time to death) of a group of patients is computed to guide clinical decisions. It is widely used in biomedical research and healthcare applications. However, frequent sharing of exact survival curves may reveal information about the individual patients, as an adversary may infer the presence of a person of interest as a participant of a study or of a particular group. Therefore, it is imperative to develop methods to protect patient privacy in survival analysis. MATERIALS AND METHODS: We develop a framework based on the formal model of differential privacy, which provides provable privacy protection against a knowledgeable adversary. We show the performance of privacy-protecting solutions for the widely used Kaplan-Meier nonparametric survival model. RESULTS: We empirically evaluated the usefulness of our privacy-protecting framework and the reduced privacy risk for a popular epidemiology dataset and a synthetic dataset. Results show that our methods significantly reduce the privacy risk when compared with their nonprivate counterparts, while retaining the utility of the survival curves. DISCUSSION: The proposed framework demonstrates the feasibility of conducting privacy-protecting survival analyses. We discuss future research directions to further enhance the usefulness of our proposed solutions in biomedical research applications. CONCLUSION: The results suggest that our proposed privacy-protection methods provide strong privacy protections while preserving the usefulness of survival analyses.
Luca Bonomi, Xiaoqian Jiang, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.1
2019 Protecting Patient Privacy in Survival Analyses
Luca Bonomi, Xiaoqian Jiang, Lucila Ohno-Machado
AMIA1
2019 VERTICOX: Vertically Distributed Cox Proportional Hazards Model
Xiaoqian Jiang, Luca Bonomi, Lucila Ohno-Machado
AMIA2
2018 Differentially Private Sanitization of ECG Time Series
Liyue Fan, Luca Bonomi
AMIA2
2018 Optimal group route query: Finding itinerary for group of users in spatial databases
Liyue Fan, Luca Bonomi, Cyrus Shahabi, Li Xiong 0001
GeoInformatica2
2018 Patient ranking with temporally annotated data
Luca Bonomi, Xiaoqian Jiang
J. Biomed. Informatics1
2017 A Mortality Study for ICU Patients Using Bursty Medical Events
abstract
The study of patients in Intensive Care Units (ICUs) is a crucial task in critical care research which has significant implications both in identifying clinical risk factors and defining institutional guidances. The mortality study of ICU patients is of particular interest because it provides useful indications to healthcare institutions for improving patients experience, internal policies, and procedures (e.g. allocation of resources). To this end, many research works have been focused on the length of stay (LOS) for ICU patients as a feature for studying the mortality. In this work, we propose a novel mortality study based on the notion of burstiness, where the temporal information of patients longitudinal data is taken into consideration. The burstiness of temporal data is a popular measure in network analysis and time-series anomaly detection, where high values of burstiness indicate presence of rapidly occurring events in short time periods (i.e. burst). Our intuition is that these bursts may relate to possible complications in the patient's medical condition and hence provide indications on the mortality. Compared to the LOS, the burstiness parameter captures the temporality of the medical events providing information about the overall dynamic of the patients condition. To the best of our knowledge, we are the first to apply the burstiness measure in the clinical research domain. Our preliminary results on a real dataset show that patients with high values of burstiness tend to have higher mortality rate compared to patients with more regular medical events. Overall, our study shows promising results and provides useful insights for developing predictive models on temporal data and advancing modern critical care medicine.
Luca Bonomi, Xiaoqian Jiang
ICDE1
2017 Multi-user Itinerary Planning for Optimal Group Preference
Liyue Fan, Luca Bonomi, Cyrus Shahabi, Li Xiong 0001
SSTD2
2016 Linking patients with non-PHI data
Luca Bonomi, Xiaoqian Jiang
AMIA1
2016 An Information-Theoretic Approach to Individual Sequential Data Sanitization
abstract
Fine-grained, personal data has been largely, continuously generated nowadays, such as location check-ins, web histories, physical activities, etc. Those data sequences are typically shared with untrusted parties for data analysis and promotional services. However, the individually-generated sequential data contains behavior patterns and may disclose sensitive information if not properly sanitized. Furthermore, the utility of the released sequence can be adversely affected by sanitization techniques. In this paper, we study the problem of individual sequence data sanitization with minimum utility loss, given user-specified sensitive patterns. We propose a privacy notion based on information theory and sanitize sequence data via generalization. We show the optimization problem is hard and develop two efficient heuristic solutions. Extensive experimental evaluations are conducted on real-world datasets and the results demonstrate the efficiency and effectiveness of our solutions.
Luca Bonomi, Liyue Fan, Hongxia Jin
WSDM1
2014 Monitoring web browsing behavior with differential privacy
abstract
Monitoring web browsing behavior has benefited many data mining applications, such as top-K discovery and anomaly detection. However, releasing private user data to the greater public would concern web users about their privacy, especially after the incident of AOL search log release where anonymization was not correctly done. In this paper, we adopt differential privacy, a strong, provable privacy definition, and show that differentially private aggregates of web browsing activities can be released in real-time while preserving the utility of shared data. Our proposed algorithms utilize the rich correlation of the time series of aggregated data and adopt a state-space approach to estimate the underlying, true aggregates from the perturbed values by the differential privacy mechanism. We evaluate our algorithms with real-world web browsing data. Utility evaluations with three metrics demonstrate that the quality of the private, released data by our solutions closely resembles that of the original, unperturbed aggregates.
Liyue Fan, Luca Bonomi, Li Xiong 0001, Vaidy S. Sunderam
WWW2
2013 A two-phase algorithm for mining sequential patterns with differential privacy
abstract
Frequent sequential pattern mining is a central task in many fields such as biology and finance. However, release of these patterns is raising increasing concerns on individual privacy. In this paper, we study the sequential pattern mining problem under the differential privacy framework which provides formal and provable guarantees of privacy. Due to the nature of the differential privacy mechanism which perturbs the frequency results with noise, and the high dimensionality of the pattern space, this mining problem is particularly challenging. In this work, we propose a novel two-phase algorithm for mining both prefixes and substring patterns. In the first phase, our approach takes advantage of the statistical properties of the data to construct a model-based prefix tree which is used to mine prefixes and a candidate set of substring patterns. The frequency of the substring patterns is further refined in the successive phase where we employ a novel transformation of the original data to reduce the perturbation noise. Extensive experiment results using real datasets showed that our approach is effective for mining both substring and prefix patterns in comparison to the state-of-the-art solutions.
Luca Bonomi, Li Xiong 0001
CIKM1
2013 LinkIT: privacy preserving record linkage and integration via transformations
abstract
We propose to demonstrate an open-source tool, LinkIT, for privacy preserving record Linkage and Integration via data Transformations. LinkIT implements novel algorithms that support data transformations for linking sensitive attributes, and is designed to work with our previously developed tool, FRIL (Fine-grained Record Integration and Linkage), to provide a complete record linkage solution. LinkIT can be also used as a stand-alone secure transformation tool to link string records. The system uses a novel embedding technique based on frequent variable length grams mined from original records with differential privacy, and utilizes a personalized threshold for performing linkage in the embedded space. Compared to the state-of-the-art secure transformation method [16], LinkIT guarantees stronger privacy with better scalability while achieving comparable utility results.
Luca Bonomi, Li Xiong 0001, James J. Lu
SIGMOD Conference1
2013 Mining Frequent Patterns with Differential Privacy
abstract
The mining of frequent patterns is a fundamental component in many data mining tasks. A considerable amount of research on this problem has led to a wide series of efficient and scalable algorithms for mining frequent patterns. However, releasing these patterns is posing concerns on the privacy of the users participating in the data. Indeed the information from the patterns can be linked with a large amount of data available from other sources creating opportunities for adversaries to break the individual privacy of the users and disclose sensitive information. In this proposal, we study the mining of frequent patterns in a privacy preserving setting. We first investigate the difference between sequential and itemset patterns, and second we extend the definition of patterns by considering the absence and presence of noise in the data. This leads us in distinguishing the patterns between exact and noisy. For exact patterns, we describe two novel mining techniques that we previously developed. The first approach has been applied in a privacy preserving record linkage setting, where our solution is used to mine frequent patterns which are employed in a secure transformation procedure to link records that are similar. The second approach improves the mining utility results using a two-phase strategy which allows to effectively mine frequent substrings as well as prefixes patterns. For noisy patterns, first we formally define the patterns according to the type of noise and second we provide a set of potential applications that require the mining of these patterns. We conclude the paper by stating the challenges in this new setting and possible future research directions.
Luca Bonomi
Proc. VLDB Endow.1
2012 Frequent grams based embedding for privacy preserving record linkage
abstract
In this paper, we study the problem of privacy preserving record linkage which aims to perform record linkage without revealing anything about the non-linked records. We propose a new secure embedding strategy based on frequent variable length grams which allows record linkage on the embedded space. The frequent grams used for constructing the embedding base are mined from the original database under the framework of differential privacy. Compared with the state-of-the-art secure matching schema [15], our approach provides formal, provable privacy guarantees and achieves better scalability while providing comparable utility.
Luca Bonomi, Li Xiong 0001, Rui Chen 0012, Benjamin C. M. Fung
CIKM1