VLDB 2026 Research / reviewers in the wild / expert
Fabian Prasser
dblp:94/7923 · also Fabian Praßer
· DBLP profile ↗
18ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0003-3172-3095ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data infrastructure for integrating clinical data in the large-scale international ORCHESTRA cohort: from data import to federated analysisabstractLarge-scale international collaborations are increasingly managing large volumes of sensitive health data for research purposes. The use of such infrastructure requires fulfilling various technical and organizational requirements to ensure usability and security. Before data scientists can access the data, it must be imported into the infrastructure with high data security requirements. For federated analysis workflows, additional criteria, such as data harmonization and prevention of individual patient information disclosure, must also be met. This paper outlines the key components of the data infrastructure implemented in the European research project ORCHESTRA and elaborates on the methods that support federated analysis workflows in a heterogeneous legal environment. Special attention is given to data security, interoperability, and usability, from which data scientists and researchers would benefit. We demonstrate the usability of data infrastructure on a federated analysis and machine learning use cases on remote datasets which satisfies the previously mentioned requirements. Furthermore, we propose organizational measures to optimize the process, reducing the time between a data access request and granting access. Miroslav Puskaric, Hammam Abu Attieh, Fabian Prasser, Roy Gusinow, Chiara Dellacasa, Elisa Rossi, Juan Mata Naranjo, Lorenzo Maria Canziani, Anna Górska, Jan Hasenauer |
IEEE Big Data | 3 |
| 2023 | Generating evidence on privacy outcomes to inform privacy risk management: A way forward?
Daniel Strech, Tamarinde Haven, Vince I. Madai, Thierry Meurers, Fabian Prasser |
J. Biomed. Informatics | 5 |
| 2022 | Open tools for quantitative anonymization of tabular phenotype data: literature reviewabstractPrecision medicine relies on molecular and systems biology methods as well as bidirectional association studies of phenotypes and (high-throughput) genomic data. However, the integrated use of such data often faces obstacles, especially in regards to data protection. An important prerequisite for research data processing is usually informed consent. But collecting consent is not always feasible, in particular when data are to be analyzed retrospectively. For phenotype data, anonymization, i.e. the altering of data in such a way that individuals cannot be identified, can provide an alternative. Several re-identification attacks have shown that this is a complex task and that simply removing directly identifying attributes such as names is usually not enough. More formal approaches are needed that use mathematical models to quantify risks and guide their reduction. Due to the complexity of these techniques, it is challenging and not advisable to implement them from scratch. Open software libraries and tools can provide a robust alternative. However, also the range of available anonymization tools is heterogeneous and obtaining an overview of their strengths and weaknesses is difficult due to the complexity of the problem space. We therefore performed a systematic review of open anonymization tools for structured phenotype data described in the literature between 1990 and 2021. Through a two-step eligibility assessment process, we selected 13 tools for an in-depth analysis. By comparing the supported anonymization techniques and further aspects, such as maturity, we derive recommendations for tools to use for anonymizing phenotype datasets with different properties. Anna Christine Haber, Ulrich Sax, Fabian Prasser |
Briefings Bioinform. | 3 |
| 2022 | EasySMPC: a simple but powerful no-code tool for practical secure multiparty computationabstractBACKGROUND: Modern biomedical research is data-driven and relies heavily on the re-use and sharing of data. Biomedical data, however, is subject to strict data protection requirements. Due to the complexity of the data required and the scale of data use, obtaining informed consent is often infeasible. Other methods, such as anonymization or federation, in turn have their own limitations. Secure multi-party computation (SMPC) is a cryptographic technology for distributed calculations, which brings formally provable security and privacy guarantees and can be used to implement a wide-range of analytical approaches. As a relatively new technology, SMPC is still rarely used in real-world biomedical data sharing activities due to several barriers, including its technical complexity and lack of usability. RESULTS: To overcome these barriers, we have developed the tool EasySMPC, which is implemented in Java as a cross-platform, stand-alone desktop application provided as open-source software. The tool makes use of the SMPC method Arithmetic Secret Sharing, which allows to securely sum up pre-defined sets of variables among different parties in two rounds of communication (input sharing and output reconstruction) and integrates this method into a graphical user interface. No additional software services need to be set up or configured, as EasySMPC uses the most widespread digital communication channel available: e-mails. No cryptographic keys need to be exchanged between the parties and e-mails are exchanged automatically by the software. To demonstrate the practicability of our solution, we evaluated its performance in a wide range of data sharing scenarios. The results of our evaluation show that our approach is scalable (summing up 10,000 variables between 20 parties takes less than 300 s) and that the number of participants is the essential factor. CONCLUSIONS: We have developed an easy-to-use "no-code solution" for performing secure joint calculations on biomedical data using SMPC protocols, which is suitable for use by scientists without IT expertise and which has no special infrastructure requirements. We believe that innovative approaches to data sharing with SMPC are needed to foster the translation of complex protocols into practice. Felix Nikolaus Wirth, Tobias Kussel, Armin Müller, Kay Hamacher, Fabian Prasser |
BMC Bioinform. | 5 |
| 2020 | Improving Data Quality in Medical Research: A Monitoring Architecture for Clinical and Translational Data WarehousesabstractClinical and translational data warehouses are important infrastructure building blocks for modern data-driven approaches in medical research. These analytics-oriented databases have been designed to integrate heterogeneous biomedical datasets from different sources and to support use cases such as cohort selection and ad-hoc data analyses. However, the lack of clear definitions of source data and controlled data collection procedures often raises concerns about the quality of data provided in such environments and, consequently, about the evidence level of related findings. To address these problems, we present an architecture that helps to monitor data quality issues when importing data into warehousing solutions using ETL (Extraction, Transformation, Load) processes. Our approach provides software developers with an API (Application Programming Interface) for logging detailed and structured information about data quality issues encountered. This information can then be displayed in dynamic dashboards, the evolution of data quality can be monitored over time, and quality issues can be traced back to their source. Our architecture supports several well-known data quality dimensions, addressing conformance, completeness, and plausibility. We present an open-source implementation, which is compatible with common clinical and translational data warehousing platforms, such as i2b2 and tranSMART, and which can be used in conjunction with many ETL environments. Helmut Spengler, Ingrid Gatz, Florian Kohlmayer, Klaus A. Kuhn, Fabian Prasser |
CBMS | 5 |
| 2020 | SCOR: A secure international informatics infrastructure to investigate COVID-19abstractGlobal pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale. Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux |
J. Am. Medical Informatics Assoc. | 19 |
| 2020 | Flexible data anonymization using ARX - Current status and challenges aheadabstractSummary The race for innovation has turned into a race for data. Rapid developments of new technologies, especially in the field of artificial intelligence, are accompanied by new ways of accessing, integrating, and analyzing sensitive personal data. Examples include financial transactions, social network activities, location traces, and medical records. As a consequence, adequate and careful privacy management has become a significant challenge. New data protection regulations, for example in the EU and China, are direct responses to these developments. Data anonymization is an important building block of data protection concepts, as it allows to reduce privacy risks by altering data. The development of anonymization tools involves significant challenges, however. For instance, the effectiveness of different anonymization techniques depends on context, and thus tools need to support a large set of methods to ensure that the usefulness of data is not overly affected by risk‐reducing transformations. In spite of these requirements, existing solutions typically only support a small set of methods. In this work, we describe how we have extended an open source data anonymization tool to support almost arbitrary combinations of a wide range of techniques in a scalable manner. We then review the spectrum of methods supported and discuss their compatibility within the novel framework. The results of an extensive experimental comparison show that our approach outperforms related solutions in terms of scalability and output data quality—while supporting a much broader range of techniques. Finally, we discuss practical experiences with ARX and present remaining issues and challenges ahead. Fabian Prasser, Johanna Eicher, Helmut Spengler, Raffael Bild, Klaus A. Kuhn |
Softw. Pract. Exp. | 1 |
| 2019 | Experiences from the National Demonstrator Study within the German Medical Informatics Initiative
Thomas Ganslandt, Jannik Schaaf, Josef Schepers, Holger Storf, Felix Balzer, Silke Haferkamp, Robert Lodahl, Fabian Prasser, Ulrich Sax, Holger Stenzhorn, Hans-Ulrich Prokosch, Martin Boeker |
AMIA | 8 |
| 2018 | SafePub: A Truthful Data Anonymization Algorithm With Strong Privacy GuaranteesabstractAbstract Methods for privacy-preserving data publishing and analysis trade off privacy risks for individuals against the quality of output data. In this article, we present a data publishing algorithm that satisfies the differential privacy model. The transformations performed are truthful, which means that the algorithm does not perturb input data or generate synthetic output data. Instead, records are randomly drawn from the input dataset and the uniqueness of their features is reduced. This also offers an intuitive notion of privacy protection. Moreover, the approach is generic, as it can be parameterized with different objective functions to optimize its output towards different applications. We show this by integrating six well-known data quality models. We present an extensive analytical and experimental evaluation and a comparison with prior work. The results show that our algorithm is the first practical implementation of the described approach and that it can be used with reasonable privacy parameters resulting in high degrees of protection. Moreover, when parameterizing the generic method with an objective function quantifying the suitability of data for building statistical classifiers, we measured prediction accuracies that compare very well with results obtained using state-of-the-art differentially private classification algorithms. Raffael Bild, Klaus A. Kuhn, Fabian Prasser |
Proc. Priv. Enhancing Technol. | 3 |
| 2018 | A Scalable and Pragmatic Method for the Safe Sharing of High-Quality Health DataabstractThe sharing of sensitive personal health data is an important aspect of biomedical research. Methods of data de-identification are often used in this process to trade the granularity of data off against privacy risks. However, traditional approaches, such as HIPAA safe harbor or -anonymization, often fail to provide data with sufficient quality. Alternatively, data can be de-identified only to a degree which still allows us to use it as required, e.g., to carry out specific analyses. Controlled environments, which restrict the ways recipients can interact with the data, can then be used to cope with residual risks. The contributions of this article are twofold. First, we present a method for implementing controlled data sharing environments and analyze its privacy properties. Second, we present a de-identification method which is specifically suited for sanitizing health data which is to be shared in such environments. Traditional de-identification methods control the uniqueness of records in a dataset. The basic idea of our approach is to reduce the probability that a record in a dataset has characteristics which are unique within the underlying population. As the characteristics of the population are typically not known, we have implemented a pragmatic solution in which properties of the population are modeled with statistical methods. We have further developed an accompanying process for evaluating and validating the degree of protection provided. The results of an extensive experimental evaluation show that our approach enables the safe sharing of high-quality data and that it is highly scalable. Fabian Prasser, Florian Kohlmayer, Helmut Spengler, Klaus A. Kuhn |
IEEE J. Biomed. Health Informatics | 1 |
| 2017 | An Open Source Tool for Game Theoretic Health Data De-Identification
Fabian Prasser, James Gaupp, Zhiyu Wan, Weiyi Xia, Yevgeniy Vorobeychik, Murat Kantarcioglu, Klaus A. Kuhn, Bradley A. Malin |
AMIA | 1 |
| 2017 | A Tool for Optimizing De-identified Health Data for Use in Statistical ClassificationabstractWhen individual-level health data is shared in biomedical research the privacy of patients and probands must be protected. This is typically achieved with methods of data de-identification, which transform data in such a way that formal guarantees about the degree of protection from re-identification can be provided. In the process it is important to minimize loss of information to ensure that the resulting data is useful. A typical use case is the creation of predictive models for knowledge discovery and decision support, e.g. to infer diagnoses or to predict outcomes of therapies. A variety of methods have been developed which can be used to build robust statistical classifiers from de-identified data. However, they have not been tuned for practical use and they have not been implemented into mature software tools. To bridge this gap, we have extended ARX, an open source anonymization tool for health data, with several new features. We have implemented a method for optimizing the suitability of de-identified data for building statistical classifiers and a method for assessing the performance of classifiers built from de-identified data. All methods are accessible via a comprehensive graphical user interface. We have used our implementation to create logistic regression models from a patient discharge dataset for predicting the costs of hospital stays. The results show that our approach enables the creation of privacy-preserving classifiers with optimal prediction accuracy. Fabian Prasser, Johanna Eicher, Raffael Bild, Helmut Spengler, Klaus A. Kuhn |
CBMS | 1 |
| 2015 | The cost of quality: Implementing generalization and suppression for anonymizing biomedical data with minimal information lossabstractOBJECTIVE: With the ARX data anonymization tool structured biomedical data can be de-identified using syntactic privacy models, such as k-anonymity. Data is transformed with two methods: (a) generalization of attribute values, followed by (b) suppression of data records. The former method results in data that is well suited for analyses by epidemiologists, while the latter method significantly reduces loss of information. Our tool uses an optimal anonymization algorithm that maximizes output utility according to a given measure. To achieve scalability, existing optimal anonymization algorithms exclude parts of the search space by predicting the outcome of data transformations regarding privacy and utility without explicitly applying them to the input dataset. These optimizations cannot be used if data is transformed with generalization and suppression. As optimal data utility and scalability are important for anonymizing biomedical data, we had to develop a novel method. METHODS: In this article, we first confirm experimentally that combining generalization with suppression significantly increases data utility. Next, we proof that, within this coding model, the outcome of data transformations regarding privacy and utility cannot be predicted. As a consequence, existing algorithms fail to deliver optimal data utility. We confirm this finding experimentally. The limitation of previous work can be overcome at the cost of increased computational complexity. However, scalability is important for anonymizing data with user feedback. Consequently, we identify properties of datasets that may be predicted in our context and propose a novel and efficient algorithm. Finally, we evaluate our solution with multiple datasets and privacy models. RESULTS: This work presents the first thorough investigation of which properties of datasets can be predicted when data is anonymized with generalization and suppression. Our novel approach adopts existing optimization strategies to our context and combines different search methods. The experiments show that our method is able to efficiently solve a broad spectrum of anonymization problems. CONCLUSION: Our work shows that implementing syntactic privacy models is challenging and that existing algorithms are not well suited for anonymizing data with transformation models which are more complex than generalization alone. As such models have been recommended for use in the biomedical domain, our results are of general relevance for de-identifying structured biomedical data. Florian Kohlmayer, Fabian Prasser, Klaus A. Kuhn |
J. Biomed. Informatics | 2 |
| 2014 | ARX - A Comprehensive Tool for Anonymizing Biomedical Data
Fabian Prasser, Florian Kohlmayer, Ronald R. Lautenschläger, Klaus A. Kuhn |
AMIA | 1 |
| 2014 | A Benchmark of Globally-Optimal Anonymization Methods for Biomedical DataabstractCollaboration and data sharing have become core elements of biomedical research. At the same time, there is a growing understanding of privacy threats related to data sharing, especially when sensitive data from distributed sources become available for linkage. Statistical disclosure control comprises well-known data anonymization techniques that allow the protection of data by introducing fuzziness. To protect datasets from different types of threats, different privacy criteria are commonly implemented. Data anonymization is an important measure, but it is computationally complex, and it can significantly reduce the expressiveness of data. To attenuate these problems, a number of algorithms has been proposed, which aim at increasing data quality or improving efficiency. Previous evaluations of such algorithms lack a systematic approach, as they focus on specific algorithms, specific privacy criteria, and specific runtime environments. Therefore, it is difficult for decision makers to decide which algorithm is best suited for their requirements. As a first step towards a comprehensive and systematic evaluation of anonymity algorithms, we report on our ongoing efforts for providing an open source benchmark. In this contribution, we focus on optimal algorithms utilizing global recoding with full-domain generalization. We present a systematic evaluation of domain-specific algorithms and generic search methods for a broad set of privacy criteria, including k-anonymity, l-diversity, t-closeness and d-presence, and their use in multiple real-world datasets. Our results show that there is no single solution fitting all needs, and that generic search methods can outperform highly specialized algorithms. Fabian Prasser, Florian Kohlmayer, Klaus A. Kuhn |
CBMS | 1 |
| 2014 | A flexible approach to distributed data anonymization
Florian Kohlmayer, Fabian Prasser, Claudia Eckert 0001, Klaus A. Kuhn |
J. Biomed. Informatics | 2 |
| 2012 | Highly efficient optimal k-anonymity for biomedical datasetsabstractK-anonymization is a wide-spread technique for the de-identification of biomedical datasets. To not render the data useless for further analysis it is often important to find an optimal solution to the k-anonymity problem, i.e., a transformation with minimum information loss. As performance is often a key requirement this paper describes an efficient implementation of a k-anonymization algorithm which is especially suitable for biomedical datasets. Although our basic implementation already offers excellent performance we present several further optimizations and show that these yield an additional speedup of up to a factor offive even for large datasets. Florian Kohlmayer, Fabian Prasser, Claudia Eckert 0001, Alfons Kemper, Klaus A. Kuhn |
CBMS | 2 |
| 2012 | Efficient distributed query processing for autonomous RDF databasesabstractThe inherent flexibility of the RDF data model has led to its notable adoption in many domains, especially in the area of life-sciences. Some of these domains have an emerging need to access data integrated from various distributed sources of information. It is not always possible to implement this by simply loading all data into one central RDF store. For example, in the context of inter-institutional collaboration for drug development and clinical research participants often want to maintain control over their local databases. Alternatively, distributed query processing techniques can be utilized to evaluate queries by accessing the remote data sources only on demand and in conformance with local authorization models. In this paper we present an efficient approach to distributed query processing for large autonomous RDF databases. The groundwork is laid by a comprehensive RDF-specific schema- and instance-level synopsis. We present an optimizer that is able to utilize this synopsis to generate compact execution plans by precisely determining, at compile-time, those sources that are relevant to a query. Furthermore we present a tightly integrated query engine that is able to further reduce the volume of intermediate results at run-time. An extensive evaluation shows that our approach improves query execution times by up to two and transferred data volumes by up to three orders of magnitude compared to a naïve implementation. Fabian Prasser, Alfons Kemper, Klaus A. Kuhn |
EDBT | 1 |