Erik Buchmann

dblp:11/226 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0009-5874-4313ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 1 first-authorArtificial intelligence and machine learning · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Helping Johnny Make Sense of Privacy Policies with LLMs
abstract
Understanding and engaging with privacy policies is crucial for online privacy, yet these documents remain notoriously complex and difficult to navigate. We present PRISMe, an interactive browser extension that combines LLM-based policy assessment with a dashboard and customizable chat interface, enabling users to skim quick overviews or explore policy details in depth while browsing. We conduct a user study (N=22) with participants of diverse privacy knowledge to investigate how users interpret the tool’s explanations and how it shapes their engagement with privacy policies, identifying distinct interaction patterns. Participants valued the clear overviews and conversational depth, but flagged some issues, particularly adversarial robustness and hallucination risks. Thus, we investigate how a retrieval-augmented generation (RAG) approach can alleviate issues by re-running the chat queries from the study. Our findings surface design challenges as well as technical trade-offs, contributing actionable insights for developing future user-centered, trustworthy privacy policy analysis tools.
Vincent Freiberger, Arthur Fleig, Erik Buchmann
CHI3
2026 Predicting User Perception based on Stimuli-Independent Saccade Transitions
abstract
Scanpaths and eye movements provide insight into how users perceive and interact with digital products. However, most studies assess user states using stimulus-dependent metrics, like fixations or areas of interest (AOIs). This paper examines whether stimulus-independent saccadic transitions — gaze movements not tied to predefined stimulus elements — carry predictive information about user states and experiences. To do so, eye-tracking data and perceived usability and UX ratings from a study with 121 participants interacting with websites were analyzed. Saccadic transitions were extracted from the scanpaths and analyzed using machine learning models to identify transition patterns predictive of the user ratings. Results show that models predict perceived usability and UX most accurately when saccade transitions are grouped into the eight inter-cardinal directions and further differentiated by median saccade length. This demonstrates that even brief, often-overlooked gaze shifts within the stimulus might provide valuable insight into how users perceive websites.
Fabian Engl, Erik Buchmann, Jürgen Mottok
ETRA2
2026 Constructing Machine Learning Features from Eye Movement Metrics: Feature Engineering Techniques for HCI Research
abstract
Eye-tracking is increasingly used in human–computer interaction (HCI) research to measure visual attention and user interaction. More recently, established eye-tracking metrics have been used as inputs to machine learning models. However, the transformation of traditional eye movement metrics into machine-learning-compatible features is rarely discussed in the literature. This limits the reproducibility and interpretability of how different eye-tracking metrics influence classification outcomes, particularly for researchers new to data-driven modeling. This paper provides a survey of how established third- and fourth-order scanpath metrics can be systematically converted into features suitable for classification and regression models. Building on existing metric categorizations, common gaze metrics are linked to concrete feature conversion techniques, including discretization, matrix unfolding, n-gram analysis, and scanpath segmentation.
Fabian Engl, Jürgen Mottok, Erik Buchmann
ETRA3
2026 Can Knowledge of Demographics and Privacy Parameters Break Location Privacy?
Maja Schneider, Charini Nanayakkara, Peter Christen, Erik Buchmann, Erhard Rahm
ICISSP (1)4
2023 Is Homomorphic Encryption Feasible for Smart Mobility?
abstract
Smart mobility is a promising approach to meet urban transport needs in an environmentally and and user-friendly way.Smart mobility computes itineraries with multiple means of transportation, e.g., trams, rental bikes or electric scooters, according to customer preferences.A mobility platform cares for reservations, connecting transports, invoicing and billing.This requires sharing sensible personal data with multiple parties, and puts data privacy at risk.In this paper, we investigate if fully homomorphic encryption (FHE) can be applied in practice to mitigate such privacy issues.FHE allows to calculate on encrypted data, without having to decrypt it first.We implemented three typical distributed computations in a smart mobility scenario with SEAL, a recent programming library for FHE.With this implementation, we have measured memory consumption and execution times for three variants of distributed transactions, that are representative for a wide range of smart mobility tasks.Our evaluation shows, that FHE is indeed applicable to smart mobility: With today's processing capabilities, state-of-the-art FHE increases a smart mobility transaction by about 100 milliseconds and less than 3 microcents.
Anika Hannemann, Erik Buchmann
FedCSIS2
2021 Securing Orchestrated Containers with BSI Module SYS.1.6
abstract
Orchestrated container virtualization, such as Docker/Kubernetes, is an attractive option to transfer complex IT ecosystems into the cloud. However, this is associated with new challenges for IT security. Containers store sensitive data with the code. The orchestration decides at run-time which containers are executed on which host. Application code is obtained as images from external sources at run-time. Typically, the operator of the cloud is not the owner of the data. Therefore, the configuration of the orchestration is critical, and an attractive target for attackers. A prominent option to secure IT infrastructures is to use security guidelines from agencies, such as Germany’s Federal Office for Information Security. In this work, we analyze the module ”SYS.1.6 Container” from this agency. We want to find out how suitable this module is to secure a typical Kubernetes scenario. Our scenario is a classical 3-tier architecture with front end, business logic and databaseback end. We show that with orchestration, the protection needs for the entire Kubernetes cluster in terms of confidentiality, integrity and availability automatically become ”high” as soon as a sensitive data object is processed or stored in any container. Our analysis has shown that the SYS.1.6 module is generally suitable. However, we have identified three additional threats. Two of them could be exploited automatically, as soon as a respective vulnerability in Docker/Kubernetes appears.
Christoph Haar, Erik Buchmann
ICISSP2
2021 Text Mining for Standardized Quality Criteria of Natural-Language IT-Requirements
abstract
Without a precise specification, an IT project might not remain on time and on budget constraints, or it might lead to a different outcome than desired. A number of established standards define how requirements must be written to avoid such issues. This paper describes our ongoing work to derive a comprehensive set of standardized criteria that IT-requirements must meet in accordance with IEEE 1233-1996 and ISO/IEC/IEEE 29148-2011. We also use a text-mining approach to identify IT-requirements that violate these standards. Our preliminary results are promising: In our biased dataset, we can use text features that are easy to compute, to filter out requirements that do not comply with the standards. Our beneficiaries are auditors, developers, Scrum teams, customers and other stakeholders whose projects are highly dependent on extensive IT-requirements specification.
Erik Buchmann, Serda Hauser
RE1
2019 Fane: A Firewall Appliance For The Smart Home
abstract
With the advent of the Internet of Things (IoT), many domestic devices have been equipped with information technology.By connecting IoT devices with each other and with the Internet, Smart Home installations exist that allow the automation of complex household tasks.A popular example is Google Nest that controls cooling, heating and home security.However, Smart Home users are tempted to neglect that such IoT devices pose IT-Security risks.Examples like the Mirai malware have already shown that insecure IoT devices can be used for large-scale network attacks.Thus, it is important to adapt security approaches to Smart Home installations.In this paper, we introduce FANE, our concept for a Firewall AppliaNcE for Smart Home installations.FANE makes a few realistic assumptions on the network segmentation and the communication profile of IoT devices.This allows FANE to learn firewall rules automatically.Our prototypical implementation indicates that FANE can secure a wide range of IoT devices without requiring network-security expertise from the Smart Home user.
Christoph Haar, Erik Buchmann
FedCSIS2
2019 Deriving Workflow Privacy Patterns from Legal Documents
abstract
The General Data Protection Regulation (GDPR) has strengthened the importance of data privacy and protection for enterprises offering their services in the EU.An important part of intensified efforts towards better privacy protection is enterprise workflow (re)design.In particular, the GDPR as strengthen the imperative to apply the privacy by design principle when (re)designing workflows.A conforming and promising approach is to model privacy relevant workflow fragments as Workflow Privacy Patterns (WPPs).Such WPPs allow to specify abstract templates for recurring data-privacy problems in workflows.Thus, WPPs are intended to support workflow engineers, auditors and privacy officers by providing pre-validated patterns that comply with existing data privacy regulations.However, it is unclear yet how to obtain WPPs systematically with an appropriate level of detail.In this paper, we introduce our approach to derive WPPs from legal texts and similar normative regulations.We propose a structure of a WPP, which we derive from pattern approaches from other research areas.We also introduce a framework that allows to design WPPs which make legal regulations accessible for persons who do not possess in-depth legal expertise.We have applied our approach to different articles of the GDPR, and we have obtained evidence that we can transfer legal text into a structured WPP representation.If a workflow correctly implements a WPP that has been designed that way, the workflow automatically complies to the respective fragment of the underlying legal text.
Marcin Robak, Erik Buchmann
FedCSIS2
2017 An evaluation of combinations of lossy compression and change-detection approaches for time-series data
Gregor Hollmig, Matthias Horne, Simon Leimkühler, Frederik Schöll, Carsten Strunk, Adrian Englhardt, Pavel Efros, Erik Buchmann, Klemens Böhm
Inf. Syst.8
2016 Identifying defective nodes in wireless sensor networks
Christopher Oßner, Erik Buchmann, Klemens Böhm
Distributed Parallel Databases2
2015 How to quantify the impact of lossy transformations on change detection
abstract
To ease the proliferation of big data, it frequently is transformed, be it by compression, be it by anonymization. Such transformations however modify characteristics of the data, such as changes in the case of time series. Changes however are important for subsequent analyses. The impact of those modifications depends on the application scenario, and quantifying it is far from trivial. This is because a transformation can shift or modify existing changes or introduce new ones. In this paper, we propose MILTON, a flexible and robust Measure for quantifying the Impact of Lossy Transformations on subsequent change detectiON. MILTON is applicable to any lossy transformation technique on time-series data and to any general-purpose change-detection approach. We have evaluated it with three real-world use cases. Our evaluation shows that MILTON allows to quantify the impact of lossy transformations and to choose the best one from a class of transformation techniques for a given application scenario.
Pavel Efros, Erik Buchmann, Adrian Englhardt, Klemens Böhm
SSDBM2
2015 Individual privacy constraints on time-series data
Fabian Laforet, Erik Buchmann, Klemens Böhm
Inf. Syst.2
2015 Efficient and secure exact-match queries in outsourced databases
Clemens Heidinger, Klemens Böhm, Erik Buchmann, Martin Spoo
World Wide Web3
2014 FACTS: A Framework for Anonymity towards Comparability, Transparency, and Sharing - Exploratory Paper
Clemens Heidinger, Klemens Böhm, Erik Buchmann
CAiSE3
2013 Privacy through Uncertainty in Location-Based Services
abstract
Location-Based Services (LBS) are becoming more prevalent. While there are many benefits, there are also real privacy risks. People are unwilling to give up the benefits - but can we reduce privacy risks without giving up on LBS entirely? This paper explores the possibility of introducing uncertainty into location information when using an LBS, so as to reduce privacy risk while maintaining good quality of service. This paper also explores the current uses of uncertainty information in a selection of mobile applications.
Shawn Merrill, Nilgun Basalp, Joachim Biskup, Erik Buchmann, Chris Clifton, Bart Kuijpers, Walied Othman, Erkay Savas
MDM (2)4
2013 Re-identification of Smart Meter data
Erik Buchmann, Klemens Böhm, Thorben Burghardt, Stephan Kessler
Pers. Ubiquitous Comput.1
2011 Query processing in sensor networks
Erik Buchmann, Nesime Tatbul, Mario A. Nascimento
Distributed Parallel Databases1
2010 Finding misplaced items in retail by clustering RFID data
abstract
In retail, products are organized according to layout plans, so-called planograms. Compliance to planograms is important, since good product placement can significantly increase sales. Currently, retailers are about to implement RFID installations consisting of smart shelves and RFID-tagged items to support in-store logistics and processes. In principle, they can also use these installations to implement planogram compliance verification: Each antenna is supposed to detect all tagged items in one location of the planogram. But due to physical constraints, RFID tags can be identified by more than one RFID antenna. Thus, one cannot decide if an item carrying such a tag complies with the planogram. We propose a new method called RPCV which checks planogram compliance on large databases of items. It is based on the observation that the number of times an antenna identifies each item of a certain product type roughly follows a normal distribution. RPCV represents each item as a two-dimensional vector containing the number of readings both by the right antenna and by wrong ones according to the planogram. It clusters this data, separately for each product type. A cluster then is a set of correctly placed items or of misplaced ones. RPCV produces one order of magnitude less wrong predictions than current state of the art, and it requires less data to yield good predictions. A study with RFID-equipped goods and smart shelves shows that our approach is effective in realistic scenarios.
Leonardo Weiss Ferreira Chaves, Erik Buchmann, Klemens Böhm
EDBT2
2010 Energy-efficient processing of spatio-temporal queries in wireless sensor networks
abstract
Research on Moving Object Databases (MOD) has resulted in sophisticated query mechanisms for moving objects and regions. Wireless Sensor Networks (WSN) support a wide range of applications that track or monitor moving objects. However, applying the concepts of MOD to WSN is difficult: While MOD tend to require precise object positions, the information acquired in WSN may be incomplete or inaccurate. This may be because of limited detection ranges, node failures or detection mechanisms that only determine if an object is in the vicinity of a node, but not its exact position. In this paper, we study the processing of spatiotemporal queries in WSN. First, we adapt the models used in MOD to WSN while keeping their semantical depth. Second, we propose two approaches for processing such queries in WSN in-network instead of collecting all data at the base station. Our experimental evaluations using simulation as well as a Sun SPOT deployment show that our measures reduce communication by up to 89%, compared to collecting all information at the base station.
Markus Bestehorn, Klemens Böhm, Erik Buchmann, Stephan Kessler
GIS3
2010 Processing continuous join queries in sensor networks: a filtering approach
abstract
While join processing in wireless sensor networks has received a lot of attention recently, current solutions do not work well for continuous queries. In those networks however, continuous queries are the rule. To minimize the communication costs of join processing, it is important to not ship non-joining tuples. In order to know which tuples do not join, prior work has proposed a precomputation step. For continuous queries however, repeating the precomputation for each execution is unnecessary and leaves aside that data tends to be temporally correlated. In this paper, we present a filtering approach for the processing of continuous join queries. We propose to keep the filters and to maintain them. The problems are determining the sizes of the filters and deciding which filters to update. Simplistic approaches result in bad performance. We show how to compute solutions that are optimal. Experiments on real-world sensor data indicate that our method performs close to a theoretical optimum and consistently outperforms state-of-the-art join approaches.
Mirco Stern, Klemens Böhm, Erik Buchmann
SIGMOD Conference3
2010 Deriving Spatio-temporal Query Results in Sensor Networks
Markus Bestehorn, Klemens Böhm, Patrick Erik Bradley, Erik Buchmann
SSDBM4
2010 Fault-tolerant query processing in structured P2P-systems
Markus Bestehorn, Christian von der Weth, Erik Buchmann, Klemens Böhm
Distributed Parallel Databases3
2009 Towards materialized view selection for distributed databases
abstract
Materialized views (MV) can significantly improve the query performance of relational databases. In this paper, we con-sider MVs to optimize complex scenarios where many het-erogeneous nodes with different resource constraints (e.g., CPU, IO and network bandwidth) query and update nu-merous tables on different nodes. Such problems are typical for large enterprises, e.g., global retailers storing thousands of relations on hundreds of nodes at different subsidiaries. Choosing which views to materialize in a distributed, com-plex scenario is NP-hard. Furthermore, the solution space is huge, and the large number of input factors results in non-monotonic cost models. This prohibits the straightfor-ward use of brute-force algorithms, greedy approaches or proposals from organic computing. For the same reason, all solutions for choosing MVs we are aware of do not consider either distributed settings or update costs. In this paper we describe an algorithmic framework which restricts the sets of considered MVs so that a genetic algo-rithm can be applied. In order to let the genetic algorithm converge quickly, we generate initial populations based on knowledge on database tuning, and devise a selection func-tion which restricts the solution space by taking the simi-larity of MV configurations into account. We evaluate our approach both with artificial settings and a real-world RFID scenario from retail. For a small setting consisting of 24 ta-bles distributed over 9 nodes, an exhaustive search needs 10 hours processing time. Our approach derives a compa-rable set of MVs within 30 seconds. Our approach scales well: Within 15 minutes it chooses a set of MVs for a real-world scenario consisting of 1,000 relations, 400 hosts, and a workload of 3,000 queries and updates. 1.
Leonardo Weiss Ferreira Chaves, Erik Buchmann, Fabian Hueske, Klemens Böhm
EDBT2
2009 Towards Efficient Processing of General-Purpose Joins in Sensor Networks
abstract
Join processing in wireless sensor networks is difficult: As the tuples can be arbitrarily distributed within the network, matching pairs of tuples is communication intensive and costly in terms of energy. Current solutions only work well with specific placements of the nodes and/or make restrictive assumptions. In this paper, we present SENS-Join, an efficient general-purpose join method for sensor networks. To obtain efficiency, SENS-Join does not ship tuples that do not join, based on a filtering step. Our main contribution is the design of this filtering step which is highly efficient in order not to exhaust the potential savings. We demonstrate the performance of SENS-Join experimentally: The overall energy consumption can be reduced by more than 80%, as compared to the state-of-the-art approach. The per node energy consumption of the most loaded nodes can be reduced by more than an order of magnitude.
Mirco Stern, Erik Buchmann, Klemens Böhm
ICDE2
2009 A Wavelet Transform for Efficient Consolidation of Sensor Relations with Quality Guarantees
abstract
Answering queries with a low selectivity in wireless sensor networks is a challenging problem. A simple tree-based data collection is communication-intensive and costly in terms of energy. Prior work has addressed the problem by approximating query results based on models of sensor readings. This cuts communication effort if the accuracy requirements are loose, e.g., if the temperature is required within ±0.5°C. For more accuracy, the models need frequent updates, and the communication costs quickly increase. In addition, sophisticated models incur substantial training costs. We propose a query-processing scheme that efficiently consolidates sensor data based on wavelet synopses. The difficulty is that the synopsis has to be constructed incrementally during data collection to ensure efficiency. Our core contribution is to show how to distribute the construction of wavelet synopses in sensor networks. In addition, our approach provides strict error guarantees. We evaluate our distributed wavelet compaction on real-world and on synthetic sensor data. Our solution reduces communication costs by more than a factor of five compared to state-of-the-art approaches. Further, our error guarantees for which efficient data consolidation is possible are better than theirs by more than an order of magnitude.
Mirco Stern, Erik Buchmann, Klemens Böhm
Proc. VLDB Endow.2
2008 Collaborative Search and User Privacy: How Can They Be Reconciled?
Thorben Burghardt, Erik Buchmann, Klemens Böhm, Chris Clifton
CollaborateCom2
2008 Tagmark: reliable estimations of RFID tags for business processes
abstract
Radio Frequency Identification (RFID) promises optimization of commodity flows in all industry segments. But due to physical constraints, RFID technology cannot detect all RFID tags from an assembly of items. This poses problems when integrating RFID data with enterprise-backend systems for tasks like inventory management or shelf replenishment. In this paper we propose the TagMark method to accomplish this integration. TagMark targets at a retailer scenario, where it estimates the number of tagged items from samples like the sales history or the tags read by smart shelves. The problem is challenging because most existing estimation methods depend on assumptions that do not hold in typical RFID applications, e.g., static item sets, simple random samples, or the availability of samples with user-defined sizes. TagMark adapts mark-recapture-methods in order to provide guarantees for the accuracy of the estimation and bounds for the sample sizes. It can be implemented as a database extension, allowing seamless integration into existing enterprise backend systems. A study with RFID-equipped goods acknowledges that our approach is effective in realistic scenarios, and database experiments with up to 1,000,000 items confirm that it can be efficiently implemented. Finally, we explore a broad range of extreme conditions that might stress TagMark, including a thief who knows the location of unread items.
Leonardo Weiss Ferreira Chaves, Erik Buchmann, Klemens Böhm
KDD2
2008 Discovering the Scope of Privacy Needs in Collaborative Search
abstract
Collaborative search engines (CSE) are an upcoming trend in WWW search. CSE let knowledge workers concert their efforts and support user collaboration. However, search terms and links clicked that are shared among users reveal their interests, habits, social relations and intentions. Thus, CSE might put the privacy of the users at risk. In this paper, we describe our first steps towards discovering the scope of privacy needs in CSE. We identify common components of CSE, and we describe typical use cases and user groups. Based on these information, we explore the range of privacy threats that might arise from query and link sharing. Furthermore, we outline a conceptual framework to explore the privacy needs of CSE users. Finally, we describe two findings from preliminary study results: First, our participants were less concerned about what providers might learn, but wanted to restrict information disclosed to people in their social network. Second, we have identified a new class of reciprocal privacy preferences, which allow or prohibit information disclose depending on the behavior of others.
Thorben Burghardt, Erik Buchmann, Klemens Böhm
Web Intelligence2
2007 Free riding-aware forwarding in Content-Addressable Networks
Klemens Böhm, Erik Buchmann
VLDB J.2
2005 Piggyback Meta-Data Propagation in Distributed Hash Tables
Erik Buchmann, Sven Apel, Gunter Saake
WEBIST1
2004 How to Run Experiments with Large Peer-to-Peer Data Structure
abstract
Summary form only given. Distributed hash tables (DHT) promise to administer huge sets of (key, value)-pairs under high workloads. DHT currently are a hot topic of research in various disciplines of computer science. Experimental results that are convincing require evaluations with large DHT (i.e., more than 100,000 nodes). However, many studies confine themselves to (less convincing) experimental examinations with much fewer nodes. Information on how to run experiments with DHT with many nodes is not available. Based on experience gained with a DHT implementation of our own, this article describes how to carry out such experiments successfully. The infrastructure used is a cluster of 32 commodity workstations. We start by compiling requirements regarding such experiments. We then identify the various bottlenecks that may be the result of a naive implementation, and we describe their negative effects. We propose various countermeasures, e.g., an experiment clock, and a component that maintains persistent network connections between cluster nodes. The features proposed are beneficial: A naive experimental setup allows for 10,000 peers maximum and a total of 20 operations per second, a sophisticated one following our proposal for 1,000,000 peers and 150 operations per second. Furthermore, we say why experimental results gained in such a way are meaningful in many situations.
Erik Buchmann, Klemens Böhm
IPDPS1