Rüdiger Kapitza

dblp:27/2697 · DBLP profile ↗
← Back
74ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-8116-7763ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 24 · 5 since 2021Software engineering, systems software and programming languages · 20 · 1 first-author · 10 since 2021Systems, architecture and hardware · 16 · 1 first-author · 3 since 2021Computer networks · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reaktor: Go as a First-Class Citizen for Client-Side Web Development Using WebAssembly
Arne Vogel, Moritz Constantin Tietze, Rüdiger Kapitza
ICWE3
2026 Wasm-WCET: Worst-Case Execution-Time Analysis of WebAssembly Modules on Updatable Resource-Constrained Embedded Devices
Maximilian Seidler, Martin Michelis, Peter Wägemann, Rüdiger Kapitza
RTAS4
2025 TEE-Assisted Recovery and Upgrades for Long-Running BFT Services
Ines Messadi, Markus Elias Gerber, Tobias Distler, Rüdiger Kapitza
ARES (2)4
2025 A Cloudy View on Trust Relationships of CVMs: How Confidential Virtual Machines are Falling Short in Public Cloud
abstract
Confidential computing in the public cloud intends to safeguard workload privacy while outsourcing infrastructure management to a cloud provider. This is achieved by executing customer workloads within so called Trusted Execution Environments, such as Confidential Virtual Machines (CVMs), which protect them from unauthorised access by cloud administrators and privileged system software. At the core of confidential computing lies remote attestation—a mechanism that enables workload owners to verify the initial state of their workload and authenticate the underlying hardware. This paper critically examines the confidential computing offerings of market-leading cloud providers to assess whether they genuinely adhere to its core principles. We develop a taxonomy based on carefully selected criteria to systematically evaluate these offerings, enabling us to analyse the components responsible for remote attestation, the evidence provided at each stage, the extent of cloud provider influence and whether this undermines the threat model of confidential computing. Specifically, we investigate how CVMs are deployed in public cloud infrastructures, the extent to which customers can request and verify attestation evidence, and their ability to define and enforce configuration and attestation requirements. This analysis provides insight into whether confidential computing guarantees—namely confidentiality and integrity—are genuinely upheld. Our findings reveal that major cloud providers retain control over critical parts of the trusted software stack and, in some cases, intervene in the standard remote attestation process. This directly contradicts their claims of delivering confidential computing, as the model fundamentally excludes the cloud provider from the set of trusted entities.
Jana Eisoldt, Anna Galanou, Andrey Ruzhanskiy, Nils Küchenmeister, Yewgenij Baburkin, Tianxiang Dai, Ivan Gudymenko, Stefan Köpsell, Rüdiger Kapitza
ACSAC9
2025 Full Trust Alchemist: Reforging Attestation for Cloud-based Confidential Workloads
abstract
Although confidential virtual machines (CVMs) offer strong isolation in untrusted cloud environments, their attestation mechanisms are restricted to static boot-time measurements. This means they cannot capture the detailed post-boot state necessary for real-world deployments. Modern workloads demand context-specific trust decisions that vary across verifiers, operational stages and workloads, like software supply chains or cloud-native workload deployments.
Anna Galanou, Florian Lubitz, Hajeong Jeon, Christof Fetzer, Rüdiger Kapitza
Middleware5
2025 WasmEye: Language- and Platform-independent Anomaly Detection for WebAssembly
abstract
With WebAssembly, you can write and run code in various languages on almost any platform. This has already led to its versatile use in IoT, edge, and cloud environments. With this broad use in complex distributed environments, additional control and management support, such as anomaly detection, to ensure reliable and secure execution will become key to WebAssembly's future success. We want to detect anomalies in WebAssembly modules to protect the system from bugs or malicious code. However, current anomaly detection solutions do not adhere to the WebAssembly philosophy by ignoring platform and source language independence or by being limited to specific types of attacks.
Arne Vogel, Timothee Glörfeld, Alexander Szekely-Schenker, Thomas Trenner, Rene Ermler, Rüdiger Kapitza
Middleware6
2025 Practical Whole-System Persistence
abstract
Sudden power outages remain one of the biggest threats to losing data, disrupting systems and causing financial damages. Whole system persistence (WSP) has previously been proposed as a solution to mitigate such threats through the use of non-volatile main memory (NVRAM). However, it missed out on external device state persistence and the NVRAM technology used at the time was expensive and limited with regard to their scalability. Today's NVRAM technologies are both more affordable and offer much higher storage capacities, but are typically slower than DRAM.
Dustin T. Nguyen, Oliver Giersch, Thomas Preisner, Jonathan Krebs, Henriette Herzog, Rüdiger Kapitza, Jörg Nolte, Timo Hönig, Wolfgang Schröder-Preikschat
SYSTOR6
2024 VeriFence: Lightweight and Precise Spectre Defenses for Untrusted Linux Kernel Extensions
abstract
High-performance IO demands low-overhead communication between user- and kernel space. This demand can no longer be fulfilled by traditional system calls. Linux's extended Berkeley Packet Filter (BPF) avoids user-/kernel transitions by just-in-time compiling user-provided bytecode and executing it in kernel mode with near-native speed. To still isolate BPF programs from the kernel, they are statically analyzed for memory- and type-safety, which imposes some restrictions but allows for good expressiveness and high performance. However, to mitigate the Spectre vulnerabilities disclosed in 2018, defenses which reject potentially-dangerous programs had to be deployed. We find that this affects 31 % to 54 % of programs in a dataset with 844 real-world BPF programs from popular open-source projects. To solve this, users are forced to disable the defenses to continue using the programs, which puts the entire system at risk.
Luis Gerhorst, Henriette Herzog, Peter Wägemann, Maximilian Ott, Rüdiger Kapitza, Timo Hönig
RAID5
2023 SoK: Scalability Techniques for BFT Consensus
abstract
With the advancement of blockchain systems, many recent research works have proposed distributed ledger technology (DLT) that employs Byzantine fault-tolerant (BFT) consensus protocols to decide which block to append next to the ledger. Notably, BFT consensus can offer high performance, energy efficiency, and provable correctness properties, and it is thus considered a promising building block for creating highly resilient and performant blockchain infrastructures. Yet, a major ongoing challenge is to make BFT consensus applicable to large-scale environments. A large body of recent work addresses this challenge by developing novel ideas to improve the scalability of BFT consensus, thus opening the path for a new generation of BFT protocols tailored to the needs of blockchain. In this survey, we create a systematization of knowledge about the novel scalability-enhancing techniques that state-of-the-art BFT consensus protocols use. For our comparison, we closely analyze the efforts, assumptions, and trade-offs these protocols make.
Christian Berger 0006, Signe Rüsch, Arne Vogel, Kai Bleeke, Leander Jehl, Hans P. Reiser, Rüdiger Kapitza
ICBC7
2023 Trustworthy confidential virtual machines for the masses
abstract
Confidential computing alleviates the concerns of distrustful customers by removing the cloud provider from their trusted computing base and resolves their disincentive to migrate their workloads to the cloud. This is facilitated by new hardware extensions, like AMD's SEV Secure Nested Paging (SEV-SNP), which can run a whole virtual machine with confidentiality and integrity protection against a potentially malicious hypervisor owned by an untrusted cloud provider. However, the assurance of such protection to either the service providers deploying sensitive workloads or the end-users passing sensitive data to services requires sending proof to the interested parties. Service providers can retrieve such proof by performing remote attestation while end-users have typically no means to acquire this proof or validate its correctness and therefore have to rely on the trustworthiness of the service providers.
Anna Galanou, Khushboo Bindlish, Luca Preibsch, Yvonne-Anne Pignolet, Christof Fetzer, Rüdiger Kapitza
Middleware6
2023 ContractBox: Realizing accountable data sharing on the edge using a small scale blockchain
abstract
The utilization of IoT devices is becoming omnipresent in industrial settings. However, adoption in more rural areas still poses challenges. Especially when processing data on the edge, privacy, data integrity, accountability and data ownership pose challenges since the devices may be easily accessed and manipulated. We present ContractBox, a system that provides accountable and trusted data sharing based on a publisher–subscriber system, as well as trusted computing on the edge. ContractBox uses a Trusted Execution Environment to guarantee the confidentiality and integrity of clients’ data and code. Furthermore, it provides security by using WebAssembly to execute the smart contracts in their own sandboxed environment. This design protects the host as well as co-located smart contracts from misbehaving executions. Lastly, it ensures the immutability and accountability of the published data by storing it in a blockchain. We show that ContractBox can process several thousand publications per second with various payloads and host multiple smart contract runtimes on a single edge device. ContractBox also achieves up to 35 times higher throughput than a comparable deployment of Hyperledger Fabric.
Lennart Almstedt, Kai Bleeke, Mohammad Mahhouk, Leander Jehl, Rüdiger Kapitza, Lars C. Wolf
Comput. Networks5
2022 ZugChain: Blockchain-Based Juridical Data Recording in Railway Systems
abstract
In modern trains, a juridical recording unit logs events that occur during operation. This data is used to reconstruct the exact chain of events in case of failures and crashes. To ensure data recovery after an accident, the recorder is hardened against physical damage and secured against tampering; however, it is a single proprietary device and by no means indestructible.This paper presents ZugChain, a distributed, blockchain-based juridical recording unit that opportunistically utilizes on-train hardware. ZugChain offers high reliability via replication and tamper-resistance due to the nature of blockchains. It implements a permissioned blockchain based on a Byzantine fault-tolerant agreement protocol suitable for diverse communication systems. To utilize the logged data for advanced services, e. g., predictive maintenance, ZugChain securely and continuously exports traces to private data centers. We demonstrate ZugChain's feasibility with an implementation running on real train hardware, where we show that ZugChain orders data within 14 ms using at maximum 15 % of the total available shared CPU resources, thus fulfilling requirements of juridical recorders.
Signe Rüsch, Kai Bleeke, Ines Messadi, Andreas Krampf, Katharina Olze, Susanne Stahnke, Robert Schmid, Lukas Pirl, Roland Kittel, Andreas Polze, Marquart Franz, Leander Jehl, Rüdiger Kapitza
DSN15
2022 MATEE: multimodal attestation for trusted execution environments
abstract
Confidential computing services enable users to run their workloads in Trusted Execution Environments (TEEs) leveraging secure hardware like Intel SGX, and verify them by performing remote attestation. This process offers necessary proof for the integrity of users' software and the authenticity of the hardware, signed by a hardware-specific attestation key. Recent side-channel attacks have successfully retrieved such keys, enabling attackers to forge the attestation data and thereby undermining users' trust in their TEE. If the attestation proof is bound to a second hardware root of trust impervious to side-channel attacks, then the remote attestation process can maintain its security guarantees.
Anna Galanou, Franz Gregor, Rüdiger Kapitza, Christof Fetzer
Middleware3
2022 SplitBFT: Improving Byzantine Fault Tolerance Safety Using Trusted Compartments
abstract
Byzantine fault-tolerant agreement (BFT) in a partially synchronous system usually requires 3f + 1 nodes to tolerate f faulty replicas. Due to their high throughput and finality property, BFT algorithms build the core of recent permissioned blockchains. As a complex and resource-demanding infrastructure, multiple cloud providers have started offering Blockchain-as-a-Service. This eases the deployment of permissioned blockchains but places the cloud provider in a central controlling position, thereby questioning blockchains' fault tolerance and decentralization properties and their underlying BFT algorithm. This paper presents SplitBFT, a new way to utilize trusted execution technology (TEEs), such as Intel SGX, to harden the safety and confidentiality guarantees of BFT systems, thereby strengthening the trust in could-based deployments of permissioned blockchains. Deviating from standard assumptions, SplitBFT acknowledges that code protected by trusted execution may fail. We address this by splitting and isolating the core logic of BFT protocols into multiple compartments resulting in a more resilient architecture. We apply SplitBFT to the traditional practical byzantine fault tolerance algorithm (PBFT) and evaluate it using SGX. Our results show that SplitBFT adds only a reasonable overhead compared to the non-compartmentalized variant.
Ines Messadi, Markus Horst Becker, Kai Bleeke, Leander Jehl, Sonia Ben Mokhtar, Rüdiger Kapitza
Middleware6
2022 EventChain: a blockchain framework for secure, privacy-preserving event verification
abstract
The number of fake news written by bots or malicious actors on social media is rising. One cause is the ability of users to post anything, at any place, at any time. This offers great flexibility, but it also poses the risk that users share misinformation. One type includes events that allegedly have occurred in a location, without the reporting user necessarily having been present. A user may add her location to posts to appear as an eyewitness and thus more trustworthy; however, basing that trust on commonly used and easily faked GPS locations is not reasonable.
Signe Rüsch, Michael Behlendorf, Markus Horst Becker, René Kudlek, Hesham Hosney Elsayed Mohamed, Felix Schoenitz, Leander Jehl, Rüdiger Kapitza
Middleware8
2021 Precursor: a fast, client-centric and trusted key-value store using RDMA and Intel SGX
abstract
As offered by the Intel Software Guard Extensions (SGX), trusted execution enables confidentiality and integrity for off-site deployed services. Thereby, securing key-value stores has received particular attention, as they are a building block for many complex applications to speed-up request processing. Initially, the developers' main design challenge has been to address the performance barriers of SGX. Besides, we identified the integration of a SGX-secured key-value store with recent network technologies, especially RDMA, as an essential emerging requirement. RDMA allows fast direct access to remote memory at high bandwidth. As SGX-protected memory cannot be directly accessed over the network, a fast exchange between the main and trusted memory must be enabled. More importantly, SGX-protected services can be expected to be CPU-bound as a result of the vast number of cryptographic operations required to transfer and store data securely.
Ines Messadi, Shivananda Neumann, Nico Weichbrodt, Lennart Almstedt, Mohammad Mahhouk, Rüdiger Kapitza
Middleware6
2021 sgx-dl: dynamic loading and hot-patching for secure applications: experience paper
abstract
Trusted execution as offered by Intel's Software Guard Extensions (SGX) is considered as an enabler to protect the integrity and confidentiality of stateful workloads such as key-value stores and databases in untrusted environments. These systems are typically long running and require extension mechanisms built on top of dynamic loading as well as hot-patching to avoid downtimes and apply security updates faster. However, such essential mechanisms are currently neglected or even missing in combination with trusted execution.
Nico Weichbrodt, Joshua Heinemann, Lennart Almstedt, Pierre-Louis Aublin, Rüdiger Kapitza
Middleware5
2019 Bursting: Increasing Energy Efficiency of Erasure-Coded Data in Animal-Borne Sensor Networks
Björn Cassens, Markus Hartmann, Thorsten Nowak, Niklas Duda, Jörn Thielecke, Alexander Koelpin, Rüdiger Kapitza
EWSN7
2019 AccTEE: A WebAssembly-based Two-way Sandbox for Trusted Resource Accounting
abstract
Remote computation has numerous use cases such as cloud computing, client-side web applications or volunteer computing. Typically, these computations are executed inside a sandboxed environment for two reasons: first, to isolate the execution in order to protect the host environment from unauthorised access, and second to control and restrict resource usage. Often, there is mutual distrust between entities providing the code and the ones executing it, owing to concerns over three potential problems: (i) loss of control over code and data by the providing entity, (ii) uncertainty of the integrity of the execution environment for customers, and (iii) a missing mutually trusted accounting of resource usage.
David Goltzsche, Manuel Nieke, Thomas Knauth, Rüdiger Kapitza
Middleware4
2019 Trusted Computing Meets Blockchain: Rollback Attacks and a Solution for Hyperledger Fabric
abstract
A smart contract on a blockchain cannot keep a secret because its data is replicated on all nodes in a network. To remedy this problem, it has been suggested combining blockchains with trusted execution environments (TEEs), such as Intel SGX, for executing applications that demand confidentiality. As a consequence, untrusted blockchain nodes cannot get access to the data and computations inside the TEE. This paper first explores issues that arise from the combination of TEEs with blockchains: Smart contracts executed inside TEEs are susceptible to rollback attacks, which should be prevented to maintain confidentiality for the application. However, in blockchains with non-final consensus protocols, such as the proof-of-work in Ethereum and others, the contract execution must handle rollbacks by design. This implies that TEEs for securing smart-contract execution cannot be directly used for such blockchains; this approach works only when the consensus decisions are final. Second, this work introduces an architecture and a prototype for smart-contract execution within Intel SGX for Hyperledger Fabric, a prominent enterprise blockchain platform. Our system resolves additional difficulties posed by the specific execute-order-validate architecture of Fabric, prevents rollback attacks on TEE-based execution as far as possible, and minimizes the trusted computing base. For increasing security, our design encapsulates each application on the blockchain within its own enclave that shields it from the host system. An evaluation shows that the overhead of moving the execution into SGX is within 10%-20% for a sealed-bid auction application.
Marcus Brandenburger, Christian Cachin, Rüdiger Kapitza, Alessandro Sorniotti
SRDS3
2019 Bloxy: Providing Transparent and Generic BFT-Based Ordering Services for Blockchains
abstract
With the wide-spread use of blockchain technology, Byzantine fault-tolerant (BFT) protocols are explored as a means to achieve consensus on which transactions should be processed next. BFT protocols are not a one-size-fits-all solution: they should be chosen according to the blockchain's use case, which can range from supply chain management to decentralised storage, requiring specialisation e.g. regarding throughput, latency, or level of decentralisation. Previously, consensus protocols were usually hardcoded into the blockchain infrastructure and could not be exchanged, therefore inhibiting flexible use of an otherwise generic blockchain infrastructure. Hyperledger Fabric claims to provide modular consensus and support for crash-fault and Byzantine fault tolerant protocols. However, integrating a BFT protocol has shown that Fabric's architecture is currently not well-suited for this fault model as it requires substantial changes and thereby breaks Fabric's modularity. This also has to be repeated for each integrated BFT protocol. In this paper, we present Bloxy, a blockchain-aware trusted proxy running on the replica that encapsulates all BFT client functionality. Bloxy enables transparent access to generic BFT frameworks and preserves Fabric's modularity even for the Byzantine fault model. It runs inside a trusted execution environment based on Intel's Software Guard Extensions. Bloxy offers blockchain-specific communication mechanisms as well as short-term block storage to handle crashes or disconnects to ensure that all nodes receive block updates. We implemented two Bloxy-based ordering services based on PBFT and the hybrid BFT protocol Hybster. Our evaluation shows that our approach increases throughput by up to 71% compared to directly integrated BFT protocols.
Signe Rüsch, Kai Bleeke, Rüdiger Kapitza
SRDS3
2019 Trust more, serverless
abstract
The increasingly popular and novel Function-as-a-Service (FaaS) clouds allow users the deployment of single functions. Compared to Infrastructure-as-a-Service or Platform-as-a-Service, this enables providers even more aggressive and rigorous resource sharing and liberates customers from tedious maintenance tasks. However, as a crucial factor of cloud adoption, FaaS clouds need to provide security and privacy guarantees in order to allow sensitive data processing.
Stefan Brenner, Rüdiger Kapitza
SYSTOR2
2018 EndBox: Scalable Middlebox Functions Using Client-Side Trusted Execution
abstract
Many organisations enhance the performance, security, and functionality of their managed networks by deploying middleboxes centrally as part of their core network. While this simplifies maintenance, it also increases cost because middlebox hardware must scale with the number of clients. A promising alternative is to outsource middlebox functions to the clients themselves, thus leveraging their CPU resources. Such an approach, however, raises security challenges for critical middlebox functions such as firewalls and intrusion detection systems. We describe EndBox, a system that securely executes middlebox functions on client machines at the network edge. Its design combines a virtual private network (VPN) with middlebox functions that are hardware-protected by a trusted execution environment (TEE), as offered by Intel's Software Guard Extensions (SGX). By maintaining VPN connection endpoints inside SGX enclaves, EndBox ensures that all client traffic, including encrypted communication, is processed by the middlebox. Despite its decentralised model, EndBox's middlebox functions remain maintainable: they are centrally controlled and can be updated efficiently. We demonstrate EndBox with two scenarios involving (i) a large company; and (ii) an Internet service provider that both need to protect their network and connected clients. We evaluate EndBox by comparing it to centralised deployments of common middlebox functions, such as load balancing, intrusion detection, firewalling, and DDoS prevention. We show that EndBox achieves up to 3.8x higher throughput and scales linearly with the number of clients.
David Goltzsche, Signe Rüsch, Manuel Nieke, Sébastien Vaucher, Nico Weichbrodt, Valerio Schiavoni, Pierre-Louis Aublin, Paolo Costa, Christof Fetzer, Pascal Felber, Peter R. Pietzuch, Rüdiger Kapitza
DSN12
2018 Troxy: Transparent Access to Byzantine Fault-Tolerant Systems
abstract
Various protocols and architectures have been proposed to make Byzantine fault tolerance (BFT) increasingly practical. However, the deployment of such systems requires dedicated client-side functionality. This is necessary as clients have to connect to multiple replicas and perform majority voting over the received replies to outvote faulty responses. Deploying custom client-side code is cumbersome, and often not an option, especially in open heterogeneous systems and for well-established protocols (e.g., HTTP and IMAP) where diverse client-side implementations co-exist. We propose Troxy, a system which relocates the BFT-specific client-side functionality to the server side, thereby making BFT transparent to legacy clients. To achieve this, Troxy relies on a trusted subsystem built upon hardware protection enabled by Intel SGX. Additionally, Troxy reduces the replication cost of BFT for read-heavy workloads by offering an actively maintained cache that supports trustworthy read operations while preserving the consistency guarantees offered by the underlying BFT protocol. A prototype of Troxy has been built and evaluated, and results indicate that using Troxy (1) leads to at most 43% performance loss with small ordered messages in a local network environment, while (2) improves throughput by 130% with read-heavy workloads in a simulated wide-area network.
Nico Weichbrodt, Johannes Behl, Pierre-Louis Aublin, Tobias Distler, Rüdiger Kapitza
DSN6
2018 STANlite - A Database Engine for Secure Data Processing at Rack-Scale Level
abstract
Intel's novel Software Guard eXtensions (SGX) enable secure and trusted execution of services, thereby paving the way to outsource sensitive data processing to external data centers. While SGX promises trusted execution close to native speed, frequent I/O operations and memory usage beyond a hardware-dependent threshold of currently 92 MiB result in substantial performance degradation. For memory-intensive workloads such as key-value stores and databases these penalties can be prohibitively high. We present STANlite - an in-memory database engine for SGX-enabled secure data processing in rack-scale environments. STANlite performs efficient user-level paging, whenever a database workload requires more space than the performance-friendly in-memory state size. Furthermore, STANlite smartly combines the properties of Remote Direct Memory Access (RDMA) and SGX to reduce the overhead of network-based I/O operations. While SGX usually provides confidentiality and integrity at the same time, STANlite enables a purely integrity preserving data management mode for additional performance. Finally, STANlite features a small trusted computing base and is memory-efficient, as it extends SQLite, a database for embedded use. We evaluated STANlite in terms of query response time. It outperforms a vanilla SGX-based SQLite version by 1.79x for microbenchmarks and 2.44x for TPC-C.
Vasily A. Sartakov, Nico Weichbrodt, Sebastian Krieter, Thomas Leich, Rüdiger Kapitza
IC2E5
2018 CYCLOSA: Decentralizing Private Web Search through SGX-Based Browser Extensions
abstract
By regularly querying Web search engines, users (unconsciously) disclose large amounts of their personal data as part of their search queries, among which some might reveal sensitive information (e.g. health issues, sexual, political or religious preferences). Several solutions exist to allow users querying search engines while improving privacy protection. However, these solutions suffer from a number of limitations: some are subject to user re-identification attacks, while others lack scalability or are unable to provide accurate results. This paper presents CYCLOSA, a secure, scalable and accurate private Web search solution. CYCLOSA improves security by relying on trusted execution environments (TEEs) as provided by Intel SGX. Further, CYCLOSA proposes a novel adaptive privacy protection solution that reduces the risk of user re-identification. CYCLOSA sends fake queries to the search engine and dynamically adapts their count according to the sensitivity of the user query. In addition, CYCLOSA meets scalability as it is fully decentralized, spreading the load for distributing fake queries among other nodes. Finally, CYCLOSA achieves accuracy of Web search as it handles the real query and the fake queries separately, in contrast to other existing solutions that mix fake and real query results.
Rafael Pires 0001, David Goltzsche, Sonia Ben Mokhtar, Sara Bouchenak, Antoine Boutet, Pascal Felber, Rüdiger Kapitza, Marcelo Pasin, Valerio Schiavoni
ICDCS7
2018 EActors: Fast and flexible trusted computing using SGX
abstract
Novel trusted execution support, as offered by Intel's Software Guard eXtensions (SGX), embeds seamlessly into user space applications by establishing regions of encrypted memory, called enclaves. Enclaves comprise code and data that is executed under special protection of the CPU and can only be accessed via an enclave defined interface. To facilitate the usability of this new system abstraction, Intel offers a software development kit (SGX SDK). While the SDK eases the use of SGX, it misses appropriate programming support for inter-enclave interaction, and demands to hardcode the exact use of trusted execution into applications, which restricts flexibility.
Vasily A. Sartakov, Stefan Brenner, Sonia Ben Mokhtar, Sara Bouchenak, Gaël Thomas 0001, Rüdiger Kapitza
Middleware6
2018 sgx-perf: A Performance Analysis Tool for Intel SGX Enclaves
abstract
Novel trusted execution technologies such as Intel's Software Guard Extensions (SGX) are considered a cure to many security risks in clouds. This is achieved by offering trusted execution contexts, so called enclaves, that enable confidentiality and integrity protection of code and data even from privileged software and physical attacks. To utilise this new abstraction, Intel offers a dedicated Software Development Kit (SDK). While it is already used to build numerous applications, understanding the performance implications of SGX and the offered programming support is still in its infancy. This inevitably leads to time-consuming trial-and-error testing and poses the risk of poor performance.
Nico Weichbrodt, Pierre-Louis Aublin, Rüdiger Kapitza
Middleware3
2018 Hybrid Fault-Tolerant Consensus in Asynchronous and Wireless Embedded Systems
abstract
Byzantine fault-tolerant (BFT) consensus in an asynchronous system can only tolerate up to floor[(n-1)/3] faulty processes in a group of n processes. This is quite a strict limit in certain application scenarios, for example a group consisting of only 3 processes. In order to break through this limit, we can leverage a hybrid fault model, in which a subset of the system is enhanced and cannot be arbitrarily faulty except for crashing. Based on this model, we propose a randomized binary consensus algorithm that executes in complete asynchrony, rather than in partial synchrony required by deterministic algorithms. It can tolerate up to floor[(n-1)/2] Byzantine faulty processes as long as the trusted subsystem in each process is not compromised, and terminates with a probability of one. The algorithm is resilient against a strong adversary, i. e. the adversary is able to inspect the state of the whole system, manipulate the delay of every message and process, and then adjust its faulty behaviour during execution. From a practical point of view, the algorithm is lightweight and has little dependency on lower level protocols or communication primitives. We evaluate the algorithm and the results show that it performs promisingly in a testbed consisting of up to 10 embedded devices connected via an ad hoc wireless network.
Wenbo Xu 0002, Signe Rüsch, Rüdiger Kapitza
OPODIS4
2018 RATCHETA: Memory-Bounded Hybrid Byzantine Consensus for Cooperative Embedded Systems
abstract
Cooperative autonomous systems gain increasing popularity nowadays. Most of these systems demand for high fault-resilience, otherwise a single faulty node could render the whole system useless. This essentially calls for a Byzantine fault-tolerant consensus. However, in such algorithms typically only (n-1)/3 faulty nodes can be tolerated in a group of n nodes and the message complexity is high. Even worse, systems with only 3 nodes are too small to even tolerate a single Byzantine node. In this work we present a novel consensus algorithm, RATCHETA. On the one hand it increases the maximum tolerable faulty nodes to (n-1)/2 and lowers the message complexity. This is achieved by assuming a hybrid fault model, which features the use of a small trusted subsystem that hosts a pair of monotonic counters for message authentication to prevent equivocation. Moreover, it can ensure an upper bound of the memory usage and message size, which is not addressed by most other hybrid consensus algorithms. On the other hand RATCHETA is tailored for wireless embedded systems. It uses multicast to reduce the communication overhead, and it does not rely on any packet loss detection or retransmission mechanisms. We implemented RATCHETA with its trusted subsystem built on top of ARM TrustZone. Our experimental results show that RATCHETA can tolerate both Byzantine faults and a certain amount of omission faults. With 20% message omissions, a 10- node group needs less than 1 second on average to reach a consensus. If 4 nodes out of 10 become Byzantine, the consensus latency is only about 1-3.6 seconds even under rough network conditions.
Wenbo Xu 0002, Rüdiger Kapitza
SRDS2
2018 Dependable Non-Volatile Memory
abstract
Recent advances in persistent memory (PM) enable fast, byte-addressable main memory that maintains its state across power cycling events. To survive power outages and prevent inconsistent application state, current approaches introduce persistent logs and require expensive cache flushes. Thus, these solutions cause a performance penalty of up to 10x for write operations on PM. With respect to wear-out effects, and a significantly lower write performance compared to read operations, we identify this as a major flaw that impacts performance and lifetime of PM. In addition, most PM technologies are susceptible to soft-errors that cause corrupted data, which implies a high risk of a permanently inconsistent system state.
Arthur Martens, Rouven Scholz, Phil Lindow, Niklas Lehnfeld, Marc A. Kastner 0001, Rüdiger Kapitza
SYSTOR6
2017 Secure Cloud Micro Services Using Intel SGX
Stefan Brenner, Tobias Hundt, Giovanni Mazzeo, Rüdiger Kapitza
DAIS4
2017 Rollback and Forking Detection for Trusted Execution Environments Using Lightweight Collective Memory
abstract
Novel hardware-aided trusted execution environments, as provided by Intel's Software Guard Extensions (SGX), enable to execute applications in a secure context that enforces confidentiality and integrity of the application state even when the host system is misbehaving. While this paves the way towards secure and trustworthy cloud computing, essential system support to protect persistent application state against rollback and forking attacks is missing. In this paper we present LCM - a lightweight protocol to establish a collective memory amongst all clients of a remote application to detect integrity and consistency violations. LCM enables the detection of rollback attacks against the remote application, enforces the consistency notion of fork-linearizability and notifies clients about operation stability. The protocol exploits the trusted execution environment, complements it with simple client-side operations, and maintains only small, constant storage at the clients. This simplifies the solution compared to previous approaches, where the clients had to verify all operations initiated by other clients. We have implemented LCM and demonstrated its advantages with a key-value store application. The evaluation shows that it introduces low network and computation overhead, in particular, a LCM-protected key-value store achieves 0.72x - 0.98x of an SGX-secured key-value store throughput.
Marcus Brandenburger, Christian Cachin, Matthias Lorenz, Rüdiger Kapitza
DSN4
2017 Hybrids on Steroids: SGX-Based High Performance BFT
abstract
With the advent of trusted execution environments provided by recent general purpose processors, a class of replication protocols has become more attractive than ever: Protocols based on a hybrid fault model are able to tolerate arbitrary faults yet reduce the costs significantly compared to their traditional Byzantine relatives by employing a small subsystem trusted to only fail by crashing. Unfortunately, existing proposals have their own price: We are not aware of any hybrid protocol that is backed by a comprehensive formal specification, complicating the reasoning about correctness and implications. Moreover, current protocols of that class have to be performed largely sequentially. Hence, they are not well-prepared for just the modern multi-core processors that bring their very own fault model to a broad audience. In this paper, we present Hybster, a new hybrid state-machine replication protocol that is highly parallelizable and specified formally. With over 1 million operations per second using only four cores, the evaluation of our Intel SGX-based prototype implementation shows that Hybster makes hybrid state-machine replication a viable option even for today's very demanding critical services.
Johannes Behl, Tobias Distler, Rüdiger Kapitza
EuroSys3
2017 Automated Encounter Detection for Animal-Borne Sensor Nodes
Björn Cassens, Simon Ripperger, Martin Hierold, Frieder Mayer, Rüdiger Kapitza
EWSN5
2017 Multi-site Synchronous VM Replication for Persistent Systems with Asymmetric Read/Write Latencies
abstract
Novel non-volatile memory is considered as a future replacement for conventional main memory. While besides persistence non-volatile memory technologies promise higher storage density and lower power demand, they also possess an asymmetry between fast read and slow write access times. The latter can be in the order of 2x to even 20x. While persistent main memory per se asks for novel system software support, the asymmetry between read and write access times needs to be addressed for building efficient solutions. This paper addresses virtual machine replication for high availability when non-volatile memory is put in use. So far asynchronous VM replication that stops a virtual machine, performs a local copy of dirty pages, resumes the virtual machine, before the data is transferred to a back-up site, is the state of the art. In contrast, we propose a multi-site zero-copy synchronous VM replication, which utilizes Remote Direct Memory Access and is tolerant to different read/write latencies. We demonstrate that for future hardware settings synchronous VM replication provides a performance increase of up to 27% compared to current best practices.
Vasily A. Sartakov, Rüdiger Kapitza
PRDC2
2017 Glamdring: Automatic Application Partitioning for Intel SGX
Joshua Lind, Christian Priebe, Divya Muthukumaran, Dan O'Keeffe, Pierre-Louis Aublin, Florian Kelbert, Tobias Reiher, David Goltzsche, David M. Eyers, Rüdiger Kapitza, Christof Fetzer, Peter R. Pietzuch
USENIX ATC10
2017 Telling Your Secrets without Page Faults: Stealthy Page Table-Based Attacks on Enclaved Execution
Jo Van Bulck, Nico Weichbrodt, Rüdiger Kapitza, Frank Piessens, Raoul Strackx
USENIX Security Symposium3
2016 BFT-Dep: Automatic Deployment of Byzantine Fault-Tolerant Services in PaaS Cloud
abstract
Cloud computing has been a massive trend over the recent years and eased the deployment of scalable distributed applications. While initially renting out virtual machines was the predominant form of cloud computing, nowadays Platform as a Service (PaaS) solutions are emerging. The main advantages of the latter are a faster and easier application deployment as well as built-in support for horizontal scalability. However, when it comes to services with more demanding dependability requirements, currently provided deployment mechanisms quickly become insufficient, forcing cloud customers to fall back into manual, self-made strategies. In this paper, we present BFT-Dep , a framework for deploying Byzantine fault-tolerant (BFT) services in a PaaS cloud automatically. BFT-Dep leverages the existing PaaS functionality to address specific deployment requirements and provides tailored support to set up and manage replicated services. It flexibly integrates BFT protocols as an independent service layer, thereby alleviating the complexity that the deployment of such systems entails. An initial prototype of BFT-Dep has been implemented on top of the open-source PaaS platform OpenShift and a first evaluation shows its practicability.
Rüdiger Kapitza
DAIS2
2016 AsyncShock: Exploiting Synchronisation Bugs in Intel SGX Enclaves
Nico Weichbrodt, Anil Kurmus, Peter R. Pietzuch, Rüdiger Kapitza
ESORICS (1)4
2016 The Neverending Runtime: Using new Technologies for Ultra-Low Power Applications with an Unlimited Runtime
Björn Cassens, Arthur Martens, Rüdiger Kapitza
EWSN3
2016 SecureKeeper: Confidential ZooKeeper using Intel SGX
Stefan Brenner, Colin Wulf, David Goltzsche, Nico Weichbrodt, Matthias Lorenz, Christof Fetzer, Peter R. Pietzuch, Rüdiger Kapitza
Middleware8
2016 SCONE: Secure Linux Containers with Intel SGX
Sergei Arnautov, Bohdan Trach, Franz Gregor, Thomas Knauth, André Martin, Christian Priebe, Joshua Lind, Divya Muthukumaran, Dan O'Keeffe, Mark Stillwell, David Goltzsche, David M. Eyers, Rüdiger Kapitza, Peter R. Pietzuch, Christof Fetzer
OSDI13
2016 Resource-Efficient Byzantine Fault Tolerance
abstract
One of the main reasons why Byzantine fault-tolerant (BFT) systems are currently not widely used lies in their high resource consumption:$3f+1$replicas are required to tolerate only$f$faults. Recent works have been able to reduce the minimum number of replicas to$2f+1$by relying on trusted subsystems that prevent a faulty replica from making conflicting statements to other replicas without being detected. Nevertheless, having been designed with the focus on fault handling, during normal-case operation these systems still use more resources than actually necessary to make progress in the absence of faults. This paper presentsResource-efficient Byzantine Fault Tolerance(ReBFT), an approach that minimizes the resource usage of a BFT system during normal-case operation by keeping$f$replicas in a passive mode. In contrast to active replicas, passive replicas neither participate in the agreement protocol nor execute client requests; instead, they are brought up to speed by verified state updates provided by active replicas. In case of suspected or detected faults, passive replicas are activated in a consistent manner. To underline the flexibility of our approach, we applyReBFTto two existing BFT systems: PBFT and MinBFT.
Tobias Distler, Christian Cachin, Rüdiger Kapitza
IEEE Trans. Computers3
2016 Monitoring Bats in the Wild: On Using Erasure Codes for Energy-Efficient Wireless Sensor Networks
abstract
We explore the advantages of using Erasure Codes (ECs) in a very challenging sensor networking scenario, namely, monitoring and tracking bats in the wild. The mobile bat nodes collect contact information that needs to be transmitted to stationary base stations whenever they are in communication range. We are particularly interested in improving the overall communication reliability of the wireless communication. The mobile nodes are capable of storing a few 100kB of data and to exchange contact information in aggregated form. Due to the continuous flight of the bats and the forest environment, the wireless channel quality varies quickly and, thus, the communication is in general assumed to be highly unreliable. Given the very strict energy constraints of the mobile node and the inherently asymmetric channels, conventional techniques such as full data replication or Automatic Repeat Request to improve the communication reliability are prohibitive. In this work, we investigate the tradeoff between reliability achieved and the cost in form of additional transmissions, that is, the additional energy costs. Our energy measurements on a real platform combined with larger-scale simulation of the wireless communication clearly indicate the advantages of using ECs in our scenario. The results are also applicable in other configurations when unreliable communication channels meet tight energy budgets.
Falko Dressler, Margit Mutschlechner, Rüdiger Kapitza, Simon Ripperger, Christopher Eibel, Benedict Herzog, Timo Hönig, Wolfgang Schröder-Preikschat
ACM Trans. Sens. Networks4
2015 Worst-Case Energy Consumption Analysis for Energy-Constrained Embedded Systems
abstract
The fact that energy is a scarce resource in many embedded real-time systems creates the need for energy-aware task schedulers, which not only guarantee timing constraints but also consider energy consumption. Unfortunately, existing approaches to analyze the worst-case execution time (WCET) of a task usually cannot be directly applied to determine its worst-case energy consumption (WCEC) due to execution time and energy consumption not being closely correlated on many state-of-the-art processors. Instead, a WCEC analyzer must take into account the particular energy characteristics of a target platform. In this paper, we present 0g, a comprehensive approach to WCEC analysis that combines different techniques to speed up the analysis and to improve results. If detailed knowledge about the energy costs of instructions on the target platform is available, our tool is able to compute upper bounds for the WCEC by statically analyzing the program code. Otherwise, a novel approach allows 0g to determine the WCEC by measurement after having identified a set of suitable program inputs based on an auxiliary energy model, which specifies the energy consumption of instructions in relation to each other. Our experiments for three target platforms show that 0g provides precise WCEC estimates.
Peter Wägemann, Tobias Distler, Timo Hönig, Heiko Janker, Rüdiger Kapitza, Wolfgang Schröder-Preikschat
ECRTS5
2015 Consensus-Oriented Parallelization: How to Earn Your First Million
abstract
Consensus protocols employed in Byzantine fault-tolerant systems are notoriously compute intensive. Unfortunately, the traditional approach to execute instances of such protocols in a pipelined fashion is not well suited for modern multi-core processors and fundamentally restricts the overall performance of systems based on them. To solve this problem, we present the consensus-oriented parallelization (COP) scheme, which disentangles consecutive consensus instances and executes them in parallel by independent pipelines; or to put it in the terminology of our main target, today's processors: COP is the introduction of superscalarity to the field of consensus protocols. In doing so, COP achieves 2.4 million operations per second on commodity server hardware, a factor of 6 compared to a contemporary pipelined approach measured on the same code base and a factor of over 20 compared to the highest throughput numbers published for such systems so far. More important, however, is: COP provides up to 3 times as much throughput on a single core than its competitors and it can make use of additional cores where other approaches are confined by the slowest stage in their pipeline. This enables Byzantine fault tolerance for the emerging market of extremely demanding transactional systems and gives more room for conventional deployments to increase their quality of service.
Johannes Behl, Tobias Distler, Rüdiger Kapitza
Middleware3
2015 Temporality a NVRAM-based Virtualization Platform
abstract
Power failures in data centers and Cloud Computing infrastructures can cause loss of data and impact revenue. Existing best practice such as persistent logging and checkpointing add overhead during operation and increase recovery time. Other solutions like the use of an uninterruptable power supply incur additional costs and are maintenance-intensive. Novel persistent main memory, i.e. memory that retains stored data without an external source of power, firstly prevents data loss in case of a power outage, secondly reduces the time for a system reboot and thirdly enables to continue operation at full-speed after a recovery. Yet new architectures and programming models are required to utilize persistent main memory. We present Temporality a virtualization layer that runs virtual machines in persistent memory and offers virtual persistent memory. It can be used as a basis for future Cloud platforms to allow applications the utilization of persistent memory without any changes. It provides safety of volatile data, significantly decreases overall recovery time and prevents subsequent performance degradation.
Vasily A. Sartakov, Arthur Martens, Rüdiger Kapitza
SRDS3
2014 Adaptive and Scalable High Availability for Infrastructure Clouds
Stefan Brenner, Benjamin Garbers, Rüdiger Kapitza
DAIS3
2014 Quantifiable Run-Time Kernel Attack Surface Reduction
Anil Kurmus, Sergej Dechand, Rüdiger Kapitza
DIMVA3
2014 Crosscheck: Hardening Replicated Multithreaded Services
abstract
State-machine replication has received widespread attention for the provisioning of highly available services in data centers. However, current production systems focus on tolerating crash faults only and prominent service outages caused by state corruptions have indicated that this is a risky strategy. In the future, state corruptions due to transient faults (such as bit flips) become even more likely, caused by ongoing hardware trends regarding the shrinking of structure sizes and reduction of operating voltages. In this paper we present Crosscheck, an approach to tolerate arbitrary state corruption (ASC) in the context of fault-tolerant replication of multithreaded services. Crosscheck is able to detect silent data corruptions ahead of execution, and by crosschecking state changes with co-executing replicas, even ASCs can be detected. Finally, fault tolerance is achieved by a fine-grained recovery using fault-free replicas. Our implementation is transparent to the application by utilizing fine-grained software-hardening mechanisms using aspect-oriented programming. To validate Crosscheck we present a replicated multithreaded key-value store that is resilient to state corruptions.
Arthur Martens, Christoph Borchert, Tobias Oliver Geissler, Daniel Lohmann, Olaf Spinczyk, Rüdiger Kapitza
DSN6
2014 NV-Hypervisor: Hypervisor-Based Persistence for Virtual Machines
abstract
Power outages and subsequent recovery are major causes of service downtimes. This issue is amplified by the ongoing trend of steadily growing in-memory state of Internet-based services which increases the risk of data loss and extends recovery time. Protective measures against power outages, such as uninterruptible power supply are expensive, maintenance-intensive and often fragile. With the advent of non-volatile random-access memory (NVRAM) provided by commodity servers, there is a scalable, less costly and robust alternative to recover from power outages and other failures. However, as of today, off-the-shelf software is not ready for benefiting from NVRAM. We present NV-Hyper visor a lightweight hyper visor extension that transparently provides persistence for virtual machines. NV-Hyper visor paves the way for utilizing NVRAM in virtualized environments (i.e., infrastructure-as-a-service clouds) and protects stateful services such as key-value stores and databases from data loss and time-consuming recovery.
Vasily A. Sartakov, Rüdiger Kapitza
DSN2
2014 Effectiveness of Fault Detection Mechanisms in Static and Dynamic Operating System Designs
abstract
Developers of embedded (real-time) systems can choose from a variety of operating systems. While some embedded operating systems provide very flexible APIs, e.g., a POSIX-compliant interface for run-time management, others have a completely static structure, which is generated at compile time by utilizing detailed application knowledge. A prominent example for the latter class from the domain of automotive operating systems is OSEK/OS and its successor AUTOSAR/OS. As we have shown in previous work, the design of the operating system has a strong impact on its vulnerability for system failure caused by hardware faults. This observation is gaining importance, because there is an ongoing trend towards low-power and low-cost, yet less reliable, hardware. This work quantifies the difference in vulnerability for soft errors in main memory of a flexible (dynamic) operating systems (eCos) and a static system (CiAO), which has an OSEK-compliant structure. We also analyze the additional degree of robustness that is achieved by hardening an operating system with software-based and hardware-based fault-tolerance measures and the corresponding costs. Covering this design space gives developers a better chance for good design decisions with respect to the trade-off between fault tolerance, resource consumption, and interface convenience. Our results indicate that with a combination of hardware- and software-based fault-tolerance measures, silent data corruptions in both operating systems can be reduced to below one percent (compared to eCos). However, the analyzed fault-tolerance mechanisms are expensive for the dynamic system, whereas the statically designed operating system can be hardened at much lower price.
Martin Hoffmann 0001, Christoph Borchert, Christian Dietrich 0001, Horst Schirmeier, Rüdiger Kapitza, Olaf Spinczyk, Daniel Lohmann
ISORC5
2014 DrySim: simulation-aided deployment-specific tailoring of mote-class WSN software
abstract
Despite intensive research in the field of mote-class Wireless Sensor Networks in recent years, real-life deployments are still challenging and systems are prone to failures. This can typically be attributed to fragile hardware or misbehaving software. Issues caused by software, often induced by the inherent constraints of resources, can be countered using simulations. However the simulation results often do not reflect those of the specific deployment. We suggest analyzing the actual environment conditions of a deployed network and map them to a simulator. Then, based on simulations, software and parameters can be tailored to the specific deployment. We developed two tool chains, RealSim and DryRun, and compared results from simulation runs to those acquired from two different testbeds using Tmote Sky nodes. This was done in two campaigns, each altering 2 configuration parameters from the hardware to the application layer. The presented data is based on over 1100 experiments, respectively over 270 h, on real hardware and almost 7000 simulations. The close relation of simulation and real measurements shows that our DrySim approach is feasible.
Moritz Strübe, Florian Lukas, Rüdiger Kapitza
MSWiM4
2013 Bandwidth Prediction in the Face of Asymmetry
Sven Schober, Stefan Brenner, Rüdiger Kapitza, Franz J. Hauck
DAIS3
2013 Attack Surface Metrics and Automated Compile-Time OS Kernel Tailoring
Anil Kurmus, Reinhard Tartler, Daniela Dorneanu, Bernhard Heinloth, Valentin Rothberg, Andreas Ziegler 0002, Wolfgang Schröder-Preikschat, Daniel Lohmann, Rüdiger Kapitza
NDSS9
2013 Distributed applications and interoperable systems (Extended papers from DAIS'10)
abstract
This special issue contains selected papers from the 10th IFIP International Conference on Distributed Applications and Interoperable Systems (DAIS ’10), held in Amsterdam in the Netherlands on June 7–9 in 2010. For the 10th anniversary, the conference was organized for the first time in cooperation with ACM SIGSOFT and ACM SIGAPP. Furthermore, the program co-chairs Frank Eliassen and Rüdiger Kapitza initiated this special issue. The presented works address various aspects of distributed applications, including their design, implementation and operation, the supporting middleware, appropriate software engineering methodologies and tools. The selection of papers was compiled by inviting 8 out of the 17 full conference papers that were originally presented at the conference. In a three-stage review process 5 out of the 8 papers were finally accepted for this special issue. Walraven et al. 1 propose a coordination architecture for flexible and policy-driven composition of cross-organizational features in distributed service systems. The underlying approach of this architecture is to specify the features and their composition at a higher level that abstracts the internal implementation mechanisms of the organizations involved. Felber et al. 2 present a collaborative search companion system, CoFeed, that collects user search queries and that considers feedback to build user-centric and document-centric profiling information. Over time, the system constructs ranked collections of elements that maintain the required information diversity and enhance the user search experience by presenting additional results tailored to the user's interest space. Zaplata et al. 3 propose a concept for the development of future context-aware applications based on the novel approach of structured context prediction. As a framework, this approach allows integration of domain-specific knowledge and facilitates the application, combination and implementation of suitable prediction methods. Romero et al. 4 focus on home environments and propose a middleware solution, called DigiHome, that applies the Service Component Architecture (SCA) component model to integrate data and events generated by heterogeneous devices in such environments. DigiHome exploits the extensibility of SCA to incorporate the representational state transfer architectural style and, in this way, leverages on the integration of multiscale systems of systems (from wireless sensor networks to the Internet). Additionally, the platform applies complex event processing technology that detects application-specific situations. Careton et al. 5 extend the ambient-oriented programming paradigm to program radio frequency identification (RFID) applications. They considered RFID tags as intermittently connected mutable proxy objects hosted on mobile-distributed computing devices and detail their prototype implementation. We thank the editors of Software – Practice and Experience for hosting this special issue, the authors for their excellent work of extending and improving their original conference publications and especially the DAIS program committee members who served for the special issue for taking the time and providing high value feedback during the three review rounds.
Rüdiger Kapitza
Softw. Pract. Exp.1
2012 CheapBFT: resource-efficient byzantine fault tolerance
abstract
One of the main reasons why Byzantine fault-tolerant (BFT) systems are not widely used lies in their high resource consumption: 3f+1 replicas are necessary to tolerate only f faults. Recent works have been able to reduce the minimum number of replicas to 2f+1 by relying on a trusted subsystem that prevents a replica from making conflicting statements to other replicas without being detected. Nevertheless, having been designed with the focus on fault handling, these systems still employ a majority of replicas during normal-case operation for seemingly redundant work. Furthermore, the trusted subsystems available trade off performance for security; that is, they either achieve high throughput or they come with a small trusted computing base.
Rüdiger Kapitza, Johannes Behl, Christian Cachin, Tobias Distler, Simon Kuhnle, Seyed Vahid Mohammadi, Wolfgang Schröder-Preikschat, Klaus Stengel
EuroSys1
2012 DQMP: A Decentralized Protocol to Enforce Global Quotas in Cloud Environments
Johannes Behl, Tobias Distler, Rüdiger Kapitza
SSS3
2011 Providing Context-Aware Adaptations Based on a Semantic Model
Guido Söldner, Rüdiger Kapitza, René Meier 0001
DAIS2
2011 Increasing performance in byzantine fault-tolerant systems with on-demand replica consistency
abstract
Traditional agreement-based Byzantine fault-tolerant (BFT) systems process all requests on all replicas to ensure consistency. In addition to the overhead for BFT protocol and state-machine replication, this practice degrades performance and prevents throughput scalability. In this paper, we propose an extension to existing BFT architectures that increases performance for the default number of replicas by optimizing the resource utilization of their execution stages.
Tobias Distler, Rüdiger Kapitza
EuroSys2
2011 SPARE: Replicas on Hold
Tobias Distler, Ivan Popov, Wolfgang Schröder-Preikschat, Hans P. Reiser, Rüdiger Kapitza
NDSS5
2011 Revisiting Fault-Injection Experiment-Platform Architectures
abstract
Many years of research on dependable, fault-tolerant software systems yielded a myriad of tool implementations for vulnerability analysis and experimental validation of resilience measures. Trace recording and fault injection are among the core functionalities these tools provide for hardware debuggers or system simulators, partially including some means to automate larger experiment campaigns. We argue that current fault-injection tools are too highly specialized for specific hardware devices or simulators, and are developed in poorly modularized implementations impeding evolution and maintenance. In this article, we present a novel design approach for a fault-injection infrastructure that allows experimenting researchers to switch simulator or hardware back ends with little effort, fosters experiment code reuse, and retains a high level of maintainability.
Horst Schirmeier, Martin Hoffmann 0001, Rüdiger Kapitza, Daniel Lohmann, Olaf Spinczyk
PRDC3
2010 Stateful Mobile Modules for Sensor Networks
Moritz Strübe, Rüdiger Kapitza, Klaus Stengel, Michael Daum 0001, Falko Dressler
DCOSS2
2010 Dynamic operator replacement in sensor networks
abstract
We present an integrated approach for supporting in-network sensor data processing in dynamic and heterogeneous sensor networks. The concept relies on data stream processing techniques that define and optimize the distribution of queries and their operators. We anticipate a high degree of dynamics in the network, which can for example be expected in the case of wildlife monitoring applications. The distribution of operators to individual nodes demands system level capabilities not available in current sensor node operating systems. In particular, we present a system for seamless and on demand operator migration between sensor nodes. Our framework, which we implemented for Contiki running on TelosB nodes, supports stateful module migration including selected parts of the code and data sections.
Moritz Strübe, Michael Daum 0001, Rüdiger Kapitza, Felix Jesús Villanueva, Falko Dressler
MASS3
2009 Recoverable Class Loaders for a Fast Restart of Java Applications
Vladimir Nikolov, Rüdiger Kapitza, Franz J. Hauck
Mob. Networks Appl.2
2008 Using Object Replication for Building a Dependable Version Control System
Rüdiger Kapitza, Hans P. Reiser
DAIS1
2008 Adaptive Web Service Migration
Holger Schmidt 0002, Rüdiger Kapitza, Franz J. Hauck, Hans P. Reiser
DAIS2
2008 Multithreading Strategies for Replicated Objects
Jörg Domaschka, Thomas Bestfleisch, Franz J. Hauck, Hans P. Reiser, Rüdiger Kapitza
Middleware5
2007 A Generic Infrastructure for Decentralised Dynamic Loading of Platform-Specific Code
Rüdiger Kapitza, Holger Schmidt 0002, Udo Bartlang, Franz J. Hauck
DAIS1
2007 Parallel State Transfer in Object Replication Systems
Rüdiger Kapitza, Thomas Zeman, Franz J. Hauck, Hans P. Reiser
DAIS1
2007 Hypervisor-Based Efficient Proactive Recovery
abstract
Proactive recovery is a promising approach for building fault and intrusion tolerant systems that tolerate an arbitrary number of faults during system lifetime. This paper investigates the benefits that a virtualization-based replication infrastructure can offer for implementing proactive recovery. Our approach uses the hypervisor to initialize a new replica in parallel to normal system execution and thus minimizes the time in which a proactive reboot interferes with system operation. As a consequence, the system maintains an equivalent degree of system availability without requiring more replicas than a traditional replication system. Furthermore, having the old replica available on the same physical host as the rejuvenated replica helps to optimize state transfer. The problem of remote transfer is reduced to remote validation of the state in the frequent case when the local replica has not been corrupted.
Hans P. Reiser, Rüdiger Kapitza
SRDS2
2006 Fault-Tolerant Replication Based on Fragmented Objects
Hans P. Reiser, Rüdiger Kapitza, Jörg Domaschka, Franz J. Hauck
DAIS2
2006 Consistent Replication of Multithreaded Distributed Objects
abstract
Determinism is mandatory for replicating distributed objects with strict consistency guarantees. Multithreaded execution of method invocations is a source of nondeterminism, but helps to improve performance and avoids deadlocks that nested invocations can cause in a single-threaded execution model. This paper contributes a novel algorithm for deterministic thread scheduling based on the interception of synchronisation statements. It assumes that shared data are protected by mutexes and client requests are sent to all replicas in total order; requests are executed concurrently as long as they do not issue potentially conflicting synchronisation operations. No additional communication is required for granting locks in a consistent order in all replicas. In addition to reentrant mutex locks, the algorithm supports condition variables and time-bounded wait operations. An experimental evaluation shows that, in some typical usage patterns of distributed objects, the algorithm is superior to other existing approaches
Hans P. Reiser, Jörg Domaschka, Franz J. Hauck, Rüdiger Kapitza, Wolfgang Schröder-Preikschat
SRDS4