EDBT 2026 Demo / reviewers in the wild / expert
Markus Hittmeir
dblp:203/4428
· DBLP profile ↗
11ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0002-3363-6270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Distance-based linkage of personal microbiome records for identification and its privacy implicationsabstractDue to its high potential for analysis in clinical settings, research on the human microbiome has been flourishing for several years. As an increasing amount of data on the microbiome is gathered and stored, analysing the temporal and individual stability of microbiome readings, and the succeeding privacy risks, has gained importance. In 2015, Franzosa et al. demonstrated the feasibility of matching and linking individuals in microbiome-based datasets from the Human Microbiome Project, which could lead to re-identification of individuals, and thus poses privacy implications for microbiome study designs. Their technique is based on the construction of body site-specific metagenomic codes that maintain a certain stability over time. In this paper, we establish a distance-based technique for personal microbiome identification, which is combined with a solution for avoiding spurious, false positive matches. In a direct comparison with the approach from Franzosa et al., which assumes that information is available as microbial records, rather than at the more detailed (but less likely to be shared) nucleic acid level, our method improves upon the identification results on most of the considered datasets. Our main finding is an increase of the average percentage of true positive identifications of 30% on the widely studied microbiome of the gastrointestinal tract. While we particularly recommend our method for application on the gut microbiome, we also observed substantial identification success on other body sites. Our results demonstrate the potential of privacy threats in microbiome data gathering, storage, sharing, and analysis, and thus underline the need for solutions to protect the microbiome as personal and sensitive medical data. We also show that the method is robust to various hyper-parameter settings. Based on our observations, we further identify challenges in personal microbiome identification research, specifically, the scarcity of benchmark data and associated data analysis tasks. Based on our experience, we propose solutions for a more systematic and comparable evaluation, considering also aspects of costs entailed with applying privacy-preserving methods. Rudolf Mayer, Markus Hittmeir, Andreas Ekelhart |
Comput. Secur. | 2 |
| 2023 | k-Anonymity on Metagenomic Features in Microbiome DatabasesabstractThe human microbiome is increasingly subject to extensive research, due to its relations to health, diet, exercise and illness. While ever more microbiome data is gathered and stored, recent works have demonstrated the threat of individual re-identification based on matching samples taken at different points in time, by matching metagenomic features extracted from microbiome readings. The individual and temporal stability of the microbiome varies for different body sites and is particularly pronounced for readings from the gastrointestinal tract. To meet the resulting need for privacy-protecting solutions, we adapt the well-known concept of k-anonymity and make it suitable for application to microbiome datasets. In particular, our approach for establishing k-anonymity is based on micro-aggregation.Our evaluation uses ten datasets containing samples of gut microbiomes, and analyzes the decreased privacy risk on the anonymised dataset as well as the incurred information loss. The analysis demonstrates the suitability of our approach for the protection of sensitive microbiome data. Rudolf Mayer, Alicja Karlowicz, Markus Hittmeir |
ARES | 3 |
| 2022 | Distance-based Techniques for Personal Microbiome Identification✱abstractDue to its high potential for analysis in clinical settings, research on the human microbiome has been flourishing for several years. As an increasing amount of data on the microbiome is gathered and stored, analysing the temporal and individual stability of microbiome readings and the ensuing privacy risks has gained importance. In 2015, Franzosa et al. demonstrated the feasibility of microbiome-based identifiability on datasets from the Human Microbiome Project, thus posing privacy implications for microbiome study designs. Their technique is based on the construction of body site-specific metagenomic codes that maintain a certain stability over time. Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart |
ARES | 1 |
| 2022 | Efficient Bayesian Network Construction for Increased Privacy on Synthetic DataabstractThe use of synthetic data is a widely acknowledged privacy-preserving measure that reduces identity and attribute disclosure risks in micro-data. The idea is to learn the statistical properties of an original dataset, store this information in a model, and then use this model to generate artificial samples and build a synthetic dataset that resembles the original. One of the many different approaches of synthetization tools relies on describing the original dataset by using a Bayesian network. This method is implemented in the open-source tool DataSynthesizer and has proven particularly suitable for datasets with a small to moderate number of attributes. In this paper, we will substitute the greedy algorithm used for learning the Bayesian network by a substantially faster genetic algorithm. In addition, our goal is to protect particularly sensitive attributes by decreasing specific correlations in the synthetic data that may reveal personal information. We will thus show how to customize the network structures for specific machine learning tasks. Our experiments demonstrate that this technique allows to further decrease the disclosure risks and, hence, add to the applicability of synthetic data as technique for privacy preservation. Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart |
IEEE Big Data | 1 |
| 2022 | Utility and Privacy Assessment of Synthetic Microbiome Data
Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart |
DBSec | 1 |
| 2021 | RandRunner: Distributed Randomness from Trapdoor VDFs with Strong Uniqueness
Philipp Schindler, Aljosha Judmayer, Markus Hittmeir, Nicholas Stifter, Edgar R. Weippl |
NDSS | 3 |
| 2020 | A Baseline for Attribute Disclosure Risk in Synthetic DataabstractThe generation of synthetic data is widely considered as viable method for alleviating privacy concerns and for reducing identification and attribute disclosure risk in micro-data. The records in a synthetic dataset are artificially created and thus do not directly relate to individuals in the original data in terms of a 1-to-1 correspondence. As a result, inferences about said individuals appear to be infeasible and, simultaneously, the utility of the data may be kept at a high level. In this paper, we challenge this belief by interpreting the standard attacker model for attribute disclosure as classification problem. We show how disclosure risk measures presented in recent publications may be compared to or even be reformulated as machine learning classification models. Our overall goal is to empirically analyze attribute disclosure risk in synthetic data and to discuss its close relationship to data utility. Moreover, we improve the baseline for attribute disclosure risk from the attacker's perspective by applying variants of the RadiusNearestNeighbor and the EnsembleVote classifier. Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart |
CODASPY | 1 |
| 2020 | Privacy-Preserving Anomaly Detection Using Synthetic Data
Rudolf Mayer, Markus Hittmeir, Andreas Ekelhart |
DBSec | 2 |
| 2020 | Deterministic Integer Factorization with Oracles for Euler's Totient FunctionabstractIn this paper, we construct deterministic factorization algorithms for natural numbers N under the assumption that the prime power decomposition of Euler’s totient function φ( N) is known. Their runtime complexities depend on the number ω( N) of distinct prime divisors of N, and we present efficient methods for relatively small values of ω( N) as well as for its large values. One of our main goals is to establish an asymptotic expression with explicit remainder term O( x/ A) for the number of positive integers N ≤ x composed of s distinct prime factors that can be factored nontrivially in deterministic time t = t( x), provided that the prime power decomposition of φ( N) is known. We obtain it for A = A( x) = x 1– ɛ , where ɛ = ɛ( s) > 0 is sufficiently small and t = t( x) is a polynomial in log x of degree d = d( ɛ). An analogous bound is deduced under the assumption of the oracle providing the decomposition of orders of elements in [Formula: see text]. Markus Hittmeir, Jacek Pomykala |
Fundam. Informaticae | 1 |
| 2019 | On the Utility of Synthetic Data: An Empirical Evaluation on Machine Learning TasksabstractWith the recent advances and increasing activities in data mining and analysis, the protection of the privacy of individuals is crucial. Several approaches address this concern, from techniques like data anonymisation to secure, non-disclosive computation, all of which have their specific strengths and weaknesses, depending on the specific requirements. A slightly different approach is the generation of synthetic data, which tries to preserve the overall properties and characteristics of the original data without revealing information about actual individual data samples. The promise is that, for most purposes, models trained on the synthetic data instead of the real data do not show a significant loss of performance. In this paper, we give an overview on currently available approaches for synthetic data generation, and empirically evaluate the utility of the generated synthetic data by testing them on a number of supervised machine learning tasks on several publicly available datasets. Markus Hittmeir, Andreas Ekelhart, Rudolf Mayer |
ARES | 1 |
| 2019 | Utility and Privacy Assessments of Synthetic Data for Regression TasksabstractWith ever increasing capacity for collecting, storing, and processing of data, there is also a high demand for intelligent data analysis methods. While there have been impressive advances in machine learning and similar domains in recent years, this also gives rise to concerns regarding the protection of personal and otherwise sensitive data, especially if it is to be analysed by third parties. Besides anonymisation, which becomes challenging with high dimensional data, one approach for privacy-preserving data mining lies in the usage of synthetic data, which comes with the promise of protecting the users' data and producing analysis results close to those achieved by using real data. In this paper, we analyse a number of different approaches for creating synthetic data, and study the utility of the created datasets for regression tasks, i.e. the prediction of a numeric value. We further investigate the similarity of real and synthetic data samples. Finally, we contribute to privacy assessments and measurements of the risk of attribute disclosure on synthetic data by extending an approach developed for categorical data. Markus Hittmeir, Andreas Ekelhart, Rudolf Mayer |
IEEE BigData | 1 |