Markus Hittmeir

dblp:203/4428 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0002-3363-6270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2024 Distance-based linkage of personal microbiome records for identification and its privacy implications
abstract
Due to its high potential for analysis in clinical settings, research on the human microbiome has been flourishing for several years. As an increasing amount of data on the microbiome is gathered and stored, analysing the temporal and individual stability of microbiome readings, and the succeeding privacy risks, has gained importance. In 2015, Franzosa et al. demonstrated the feasibility of matching and linking individuals in microbiome-based datasets from the Human Microbiome Project, which could lead to re-identification of individuals, and thus poses privacy implications for microbiome study designs. Their technique is based on the construction of body site-specific metagenomic codes that maintain a certain stability over time. In this paper, we establish a distance-based technique for personal microbiome identification, which is combined with a solution for avoiding spurious, false positive matches. In a direct comparison with the approach from Franzosa et al., which assumes that information is available as microbial records, rather than at the more detailed (but less likely to be shared) nucleic acid level, our method improves upon the identification results on most of the considered datasets. Our main finding is an increase of the average percentage of true positive identifications of 30% on the widely studied microbiome of the gastrointestinal tract. While we particularly recommend our method for application on the gut microbiome, we also observed substantial identification success on other body sites. Our results demonstrate the potential of privacy threats in microbiome data gathering, storage, sharing, and analysis, and thus underline the need for solutions to protect the microbiome as personal and sensitive medical data. We also show that the method is robust to various hyper-parameter settings. Based on our observations, we further identify challenges in personal microbiome identification research, specifically, the scarcity of benchmark data and associated data analysis tasks. Based on our experience, we propose solutions for a more systematic and comparable evaluation, considering also aspects of costs entailed with applying privacy-preserving methods.
Rudolf Mayer, Markus Hittmeir, Andreas Ekelhart
Comput. Secur.2
2023 k-Anonymity on Metagenomic Features in Microbiome Databases
abstract
The human microbiome is increasingly subject to extensive research, due to its relations to health, diet, exercise and illness. While ever more microbiome data is gathered and stored, recent works have demonstrated the threat of individual re-identification based on matching samples taken at different points in time, by matching metagenomic features extracted from microbiome readings. The individual and temporal stability of the microbiome varies for different body sites and is particularly pronounced for readings from the gastrointestinal tract. To meet the resulting need for privacy-protecting solutions, we adapt the well-known concept of k-anonymity and make it suitable for application to microbiome datasets. In particular, our approach for establishing k-anonymity is based on micro-aggregation.Our evaluation uses ten datasets containing samples of gut microbiomes, and analyzes the decreased privacy risk on the anonymised dataset as well as the incurred information loss. The analysis demonstrates the suitability of our approach for the protection of sensitive microbiome data.
Rudolf Mayer, Alicja Karlowicz, Markus Hittmeir
ARES3
2022 Distance-based Techniques for Personal Microbiome Identification✱
abstract
Due to its high potential for analysis in clinical settings, research on the human microbiome has been flourishing for several years. As an increasing amount of data on the microbiome is gathered and stored, analysing the temporal and individual stability of microbiome readings and the ensuing privacy risks has gained importance. In 2015, Franzosa et al. demonstrated the feasibility of microbiome-based identifiability on datasets from the Human Microbiome Project, thus posing privacy implications for microbiome study designs. Their technique is based on the construction of body site-specific metagenomic codes that maintain a certain stability over time.
Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart
ARES1
2022 Efficient Bayesian Network Construction for Increased Privacy on Synthetic Data
abstract
The use of synthetic data is a widely acknowledged privacy-preserving measure that reduces identity and attribute disclosure risks in micro-data. The idea is to learn the statistical properties of an original dataset, store this information in a model, and then use this model to generate artificial samples and build a synthetic dataset that resembles the original. One of the many different approaches of synthetization tools relies on describing the original dataset by using a Bayesian network. This method is implemented in the open-source tool DataSynthesizer and has proven particularly suitable for datasets with a small to moderate number of attributes. In this paper, we will substitute the greedy algorithm used for learning the Bayesian network by a substantially faster genetic algorithm. In addition, our goal is to protect particularly sensitive attributes by decreasing specific correlations in the synthetic data that may reveal personal information. We will thus show how to customize the network structures for specific machine learning tasks. Our experiments demonstrate that this technique allows to further decrease the disclosure risks and, hence, add to the applicability of synthetic data as technique for privacy preservation.
Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart
IEEE Big Data1
2022 Utility and Privacy Assessment of Synthetic Microbiome Data
Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart
DBSec1
2021 RandRunner: Distributed Randomness from Trapdoor VDFs with Strong Uniqueness
Philipp Schindler, Aljosha Judmayer, Markus Hittmeir, Nicholas Stifter, Edgar R. Weippl
NDSS3
2020 A Baseline for Attribute Disclosure Risk in Synthetic Data
abstract
The generation of synthetic data is widely considered as viable method for alleviating privacy concerns and for reducing identification and attribute disclosure risk in micro-data. The records in a synthetic dataset are artificially created and thus do not directly relate to individuals in the original data in terms of a 1-to-1 correspondence. As a result, inferences about said individuals appear to be infeasible and, simultaneously, the utility of the data may be kept at a high level. In this paper, we challenge this belief by interpreting the standard attacker model for attribute disclosure as classification problem. We show how disclosure risk measures presented in recent publications may be compared to or even be reformulated as machine learning classification models. Our overall goal is to empirically analyze attribute disclosure risk in synthetic data and to discuss its close relationship to data utility. Moreover, we improve the baseline for attribute disclosure risk from the attacker's perspective by applying variants of the RadiusNearestNeighbor and the EnsembleVote classifier.
Markus Hittmeir, Rudolf Mayer, Andreas Ekelhart
CODASPY1
2020 Privacy-Preserving Anomaly Detection Using Synthetic Data
Rudolf Mayer, Markus Hittmeir, Andreas Ekelhart
DBSec2
2020 Deterministic Integer Factorization with Oracles for Euler's Totient Function
abstract
In this paper, we construct deterministic factorization algorithms for natural numbers N under the assumption that the prime power decomposition of Euler’s totient function φ( N) is known. Their runtime complexities depend on the number ω( N) of distinct prime divisors of N, and we present efficient methods for relatively small values of ω( N) as well as for its large values. One of our main goals is to establish an asymptotic expression with explicit remainder term O( x/ A) for the number of positive integers N ≤ x composed of s distinct prime factors that can be factored nontrivially in deterministic time t = t( x), provided that the prime power decomposition of φ( N) is known. We obtain it for A = A( x) = x 1– ɛ , where ɛ = ɛ( s) > 0 is sufficiently small and t = t( x) is a polynomial in log x of degree d = d( ɛ). An analogous bound is deduced under the assumption of the oracle providing the decomposition of orders of elements in [Formula: see text].
Markus Hittmeir, Jacek Pomykala
Fundam. Informaticae1
2019 On the Utility of Synthetic Data: An Empirical Evaluation on Machine Learning Tasks
abstract
With the recent advances and increasing activities in data mining and analysis, the protection of the privacy of individuals is crucial. Several approaches address this concern, from techniques like data anonymisation to secure, non-disclosive computation, all of which have their specific strengths and weaknesses, depending on the specific requirements. A slightly different approach is the generation of synthetic data, which tries to preserve the overall properties and characteristics of the original data without revealing information about actual individual data samples. The promise is that, for most purposes, models trained on the synthetic data instead of the real data do not show a significant loss of performance. In this paper, we give an overview on currently available approaches for synthetic data generation, and empirically evaluate the utility of the generated synthetic data by testing them on a number of supervised machine learning tasks on several publicly available datasets.
Markus Hittmeir, Andreas Ekelhart, Rudolf Mayer
ARES1
2019 Utility and Privacy Assessments of Synthetic Data for Regression Tasks
abstract
With ever increasing capacity for collecting, storing, and processing of data, there is also a high demand for intelligent data analysis methods. While there have been impressive advances in machine learning and similar domains in recent years, this also gives rise to concerns regarding the protection of personal and otherwise sensitive data, especially if it is to be analysed by third parties. Besides anonymisation, which becomes challenging with high dimensional data, one approach for privacy-preserving data mining lies in the usage of synthetic data, which comes with the promise of protecting the users' data and producing analysis results close to those achieved by using real data. In this paper, we analyse a number of different approaches for creating synthetic data, and study the utility of the created datasets for regression tasks, i.e. the prediction of a numeric value. We further investigate the similarity of real and synthetic data samples. Finally, we contribute to privacy assessments and measurements of the risk of attribute disclosure on synthetic data by extending an approach developed for categorical data.
Markus Hittmeir, Andreas Ekelhart, Rudolf Mayer
IEEE BigData1