EDBT 2026 Demo / reviewers in the wild / expert
Christof Fetzer
dblp:f/ChristofFetzer
· DBLP profile ↗
175ranked-venue papers
31as first author
24since 2021 · last 2026
0000-0001-8240-5420ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 81 · 17 first-author · 2 since 2021Security and privacy · 55 · 15 first-author · 9 since 2021Software engineering, systems software and programming languages · 32 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Databases, data management, data science and information retrieval · 9Artificial intelligence and machine learning · 2Computer networks · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Confidential Key Management as a Service: Enhancing Availability and Isolation in Key Protection
Huyen Tran Ngoc Nhat, Christof Fetzer |
SECRYPT (1) | 2 |
| 2025 | Understanding the Latency-Security Tradeoff: TEE-based Confidential Computing for Streaming WorkloadsabstractDistributed streaming platforms such as Pravega, Kafka, and Pulsar are widely used for high-throughput, low-latency data processing. As these platforms increasingly handle sensitive data, ensuring data confidentiality and integrity becomes critical. Trusted Execution Environments (TEEs) offer secure computations that can be used on client-side processing, but their impact on performance must be carefully assessed. This study evaluates the write latency of Pravega clients running in TEEs compared to those in standard (non-secured) environments. We found that under typical workloads, TEE-based clients experience approximately 50% higher latency due to the overhead of secure executions. However, when data rates exceed 976 MB/s, the Pravega broker reaches its throughput limit, causing latency to spike for standard clients. In contrast, TEE-based clients exhibit more stable latency under these high-throughput conditions. These findings can be helpful for data architects, as systems highlight a trade-off: while latency may increase, the impact could be acceptable in certain scenarios given the enhanced security benefits. Alan Cueva Mora, K. P. N. Jayasena, Robert Krahn, Enrique Chirivella-Perez, Christof Fetzer |
ICNP | 5 |
| 2025 | Full Trust Alchemist: Reforging Attestation for Cloud-based Confidential WorkloadsabstractAlthough confidential virtual machines (CVMs) offer strong isolation in untrusted cloud environments, their attestation mechanisms are restricted to static boot-time measurements. This means they cannot capture the detailed post-boot state necessary for real-world deployments. Modern workloads demand context-specific trust decisions that vary across verifiers, operational stages and workloads, like software supply chains or cloud-native workload deployments. Anna Galanou, Florian Lubitz, Hajeong Jeon, Christof Fetzer, Rüdiger Kapitza |
Middleware | 4 |
| 2025 | SCONE Confidential Computing Environment: Protecting Applications Against Powerful Adversaries (Invited Talk)abstractOur objective is to protect the code, data, and keys of applications against all users with access to the computer systems. In some domains (e.g., healthcare domain), this must be guaranteed, even if the application is not entirely correct. To simplify the adoption of confidential computing, SCONE transforms cloud-native applications into confidential cloud-native applications running on vanilla Kubernetes clusters. The applications can run on Intel SGX, Intel TDX, and AMD SEV. In the near future, SCONE will also support confidential GPUs. The confidentiality, integrity, and consistency of an application’s data and keys are guaranteed by always keeping the data encrypted, i.e., at rest, in transit, and in use. This enables us to add a protection layer around applications to prevent data loss caused be bugs and backdoors in the application code. Christof Fetzer |
OPODIS | 1 |
| 2024 | CRISP: Confidentiality, Rollback, and Integrity Storage Protection for Confidential Cloud-Native ComputingabstractTrusted execution environments (TEEs) protect the integrity and confidentiality of running code and its associated data. Nevertheless, TEEs' integrity protection does not extend to the state saved on disk. Furthermore, modern cloud-native applications heavily rely on orchestration (e.g., through systems such as Kubernetes) and, thus, have their services frequently restarted. During restarts, attackers can revert the state of confidential services to a previous version that may aid their malicious intent. This paper presents CRISP, a rollback protection mechanism that uses an existing runtime for Intel SGX and transparently prevents rollback. Our approach can constrain the attack window to a fixed and short period or give developers the tools to avoid the vulnerability window altogether. Finally, experiments show that applying CRISP in a critical stateful cloud-native application may incur a resource increase but only a minor performance penalty. Ardhi Putra Pratama Hartono, Andrey Brito, Christof Fetzer |
CLOUD | 3 |
| 2024 | Traceability and Accountability by Construction
Julius Wenzel, Maximilian A. Köhl, Sarah Sterz, Hanwei Zhang 0001, Andreas Schmidt 0003, Christof Fetzer, Holger Hermanns |
ISoLA (4) | 6 |
| 2024 | A Comprehensive Study on the Impact of Vulnerable Dependencies on Open-Source SoftwareabstractOpen-source libraries are widely used by software developers to speed up the development of products, however, they can introduce security vulnerabilities, leading to incidents like Log4Shell. With the expanding usage of open-source libraries, it becomes even more imperative to comprehend and address these dependency vulnerabilities. The use of Software Composition Analysis (SCA) tools does greatly help here as they provide a deep insight on what dependencies are used in a project, enhancing the security and integrity in the software supply chain. In order to learn how wide spread vulnerabilities are and how quickly they are being fixed, we conducted a study on over 1k open-source software projects with about 50k releases comprising several languages such as Java, Python, Rust, Go, Ruby, PHP, and JavaScript. Our objective is to investigate the severity, persistence, and distribution of these vulnerabilities, as well as their correlation with project metrics such as team and contributors size, activity and release cycles. In order to perform such analysis, we crawled over 1k projects from github including their version history ranging from 2013 to 2023 using VODA, our SCA tool. Using our approach, we can provide information such as library versions, dependency depth, and known vulnerabilities, and how they evolved over the software development cycle. Being larger and more diverse than datasets used in earlier works and studies, ours provides better insights and generalizability of the gained results. The data collected answers several research questions about the dependency depth and the average time a vulnerability persists. Among other findings, we observed that for most programming languages, vulnerable dependencies are transitive, and a critical vulnerability persists in average for over a year before being fixed. The results furthermore emphasize the importance of managing dependencies, performing timely updates, and suggests types of vulnerabilities that can be fixed faster. Shree Hari Bittugondanahalli Indra Kumar, Lilia Rodrigues Sampaio, André Martin, Andrey Brito, Christof Fetzer |
ISSRE | 5 |
| 2024 | Invited Paper: Using Signed Formulas for Online Certification
Julius Wenzel, Andreas Berg, Christof Fetzer |
SSS | 3 |
| 2023 | Triad: Trusted Timestamps in Untrusted EnvironmentsabstractWe aim to provide trusted time measurement mechanisms to applications and cloud infrastructure deployed in environments that could harbor potential adversaries, including the hardware infrastructure provider. Despite Trusted Execution Environments (TEEs) providing multiple security functionalities, timestamps from the Operating System are not covered. Nevertheless, some services require time for validating permissions or ordering events. To address that need, we introduce Triad, a trusted timestamp dispatcher of time readings. The solution provides trusted timestamps enforced by mutually supportive enclave-based clock servers that create a continuous trusted timeline. We leverage enclave properties such as forced exits and CPU-based counters to mitigate attacks on the server’s timestamp counters. Triad produces trusted, confidential, monotonically-increasing timestamps with bounded error and desirable, nontrivial properties. Our implementation relies on Intel SGX and SCONE, allowing transparent usage. We evaluate Triad’s error and behavior in multiple dimensions. Gabriel Fernandez 0001, Andrey Brito, Christof Fetzer |
CloudCom | 3 |
| 2023 | Trustworthy confidential virtual machines for the massesabstractConfidential computing alleviates the concerns of distrustful customers by removing the cloud provider from their trusted computing base and resolves their disincentive to migrate their workloads to the cloud. This is facilitated by new hardware extensions, like AMD's SEV Secure Nested Paging (SEV-SNP), which can run a whole virtual machine with confidentiality and integrity protection against a potentially malicious hypervisor owned by an untrusted cloud provider. However, the assurance of such protection to either the service providers deploying sensitive workloads or the end-users passing sensitive data to services requires sending proof to the interested parties. Service providers can retrieve such proof by performing remote attestation while end-users have typically no means to acquire this proof or validate its correctness and therefore have to rely on the trustworthiness of the service providers. Anna Galanou, Khushboo Bindlish, Luca Preibsch, Yvonne-Anne Pignolet, Christof Fetzer, Rüdiger Kapitza |
Middleware | 5 |
| 2023 | SinClave: Hardware-assisted Singletons for TEEsabstractFor trusted execution environments (TEEs), remote attestation permits establishing trust in software executed on a remote host. It requires that the measurement of a remote TEE is both complete and fresh: We need to measure all aspects that might determine the behavior of an application, and this measurement has to be reasonably fresh. Performing measurements only at the start of a TEE simplifies the attestation but enables "reuse" attacks of enclaves. We demonstrate how to perform such reuse attacks for different TEE frameworks. We also show how to address this issue by enforcing freshness -- through the concept of a singleton enclave -- and completeness of the measurements. Completeness of measurements is not trivial since the secrets provisioned to an enclave and the content of the filesystem can both affect the behavior of the software, i.e., can be used to mount reuse attacks. We present mechanisms to include measurements of these two components in the remote attestation. Our evaluation based on real-world applications shows that our approach incurs only negligible overhead ranging from 1.03% to 13.2%. Franz Gregor, Robert Krahn, Do Le Quoc, Christof Fetzer |
Middleware | 4 |
| 2023 | Confidential computing and related technologies: a critical reviewabstractAbstract This research critically reviews the definition of confidential computing (CC) and the security comparison of CC with other related technologies by the Confidential Computing Consortium (CCC). We demonstrate that the definitions by CCC are ambiguous, incomplete and even conflicting. We also demonstrate that the security comparison of CC with other technologies is neither scientific nor fair. We highlight the issues in the definitions and comparisons and provide initial recommendations for fixing the issues. These recommendations are the first step towards more precise definitions and reliable comparisons in the future. Muhammad Usama Sardar, Christof Fetzer |
Cybersecur. | 2 |
| 2023 | Capacity planning for dependable services
Rasha Faqeh, André Martin, Valerio Schiavoni, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
Theor. Comput. Sci. | 6 |
| 2022 | Revizor: testing black-box CPUs against speculation contractsabstractSpeculative vulnerabilities such as Spectre and Meltdown expose speculative execution state that can be exploited to leak information across security domains via side-channels. Such vulnerabilities often stay undetected for a long time as we lack the tools for systematic testing of CPUs to find them. Oleksii Oleksenko, Christof Fetzer, Boris Köpf, Mark Silberstein |
ASPLOS | 2 |
| 2022 | MATEE: multimodal attestation for trusted execution environmentsabstractConfidential computing services enable users to run their workloads in Trusted Execution Environments (TEEs) leveraging secure hardware like Intel SGX, and verify them by performing remote attestation. This process offers necessary proof for the integrity of users' software and the authenticity of the hardware, signed by a hardware-specific attestation key. Recent side-channel attacks have successfully retrieved such keys, enabling attackers to forge the attestation data and thereby undermining users' trust in their TEE. If the attestation proof is bound to a second hardware root of trust impervious to side-channel attacks, then the remote attestation process can maintain its security guarantees. Anna Galanou, Franz Gregor, Rüdiger Kapitza, Christof Fetzer |
Middleware | 4 |
| 2022 | Capacity Planning for Dependable Services
Rasha Faqeh, André Martin, Valerio Schiavoni, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
SSS | 6 |
| 2022 | A Sorted Datalog Hammer for Supervisor Verification Conditions Modulo Simple Linear ArithmeticabstractAbstract In a previous paper, we have shown that clause sets belonging to the Horn Bernays-Schönfinkel fragment over simple linear real arithmetic (HBS(SLR)) can be translated into HBS clause sets over a finite set of first-order constants. The translation preserves validity and satisfiability and it is still applicable if we extend our input with positive universally or existentially quantified verification conditions (conjectures). We call this translation a Datalog hammer. The combination of its implementation in SPASS-SPL with the Datalog reasoner VLog establishes an effective way of deciding verification conditions in the Horn fragment. We verify supervisor code for two examples: a lane change assistant in a car and an electronic control unit of a supercharged combustion engine. In this paper, we improve our Datalog hammer in several ways: we generalize it to mixed real-integer arithmetic and finite first-order sorts; we extend the class of acceptable inequalities beyond variable bounds and positively grounded inequalities; and we significantly reduce the size of the hammer output by a soft typing discipline. We call the result the sorted Datalog hammer. It not only allows us to handle more complex supervisor code and to model already considered supervisor code more concisely, but it also improves our performance on real world benchmark examples. Finally, we replace the before file-based interface between SPASS-SPL and VLog by a close coupling resulting in a single executable binary. Martin Bromberger, Irina Dragoste, Rasha Faqeh, Christof Fetzer, Larry González, Markus Krötzsch, Maximilian Marx 0001, Harish K. Murali, Christoph Weidenbach |
TACAS (1) | 4 |
| 2022 | SGXTuner: Performance Enhancement of Intel SGX Applications Via Stochastic OptimizationabstractIntelSGXhas started to be widely adopted. Cloud providers (Microsoft Azure, IBM Cloud, Alibaba Cloud) are offering new solutions, implementingdata-in-useprotection via SGX. A major challenge faced by both academia and industry is providing transparent SGX support to legacy applications. The approach with the highest consensus is linking the target software with SGX-extendedlibclibraries. Unfortunately, the increased security entails a dramatic performance penalty, which is mainly due to the intrinsic overhead of context switches, and the limited size of protected memory. Performance optimization is non-trivial since it depends on key parameters whose manual tuning is a very long process. We present the architecture of an automated tool, calledSGXTuner, which is able to find the best setting of SGX-extendedlibclibrary parameters, by iteratively adjusting such parameters based on continuous monitoring of performance data. The tool is — to a large extent — algorithm agnostic. We decided to base the current implementation on a particular type of stochastic optimization algorithm, specificallySimulated Annealing. A massive experimental campaign was conducted on a relevant case study. Three client-server applications —Memcached,Redis, andApache— were compiled with SCONE'ssgx-musland tuned for best performance. Results demonstrate the effectiveness ofSGXTuner. Giovanni Mazzeo, Sergei Arnautov, Christof Fetzer, Luigi Romano |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | TRIGLAV: Remote Attestation of the Virtual Machine's Runtime Integrity in Public CloudsabstractTrust is of paramount concern for tenants to deploy their security-sensitive services in the cloud. The integrity of virtual machines (VMs) in which these services are deployed needs to be ensured even in the presence of powerful adversaries with administrative access to the cloud. Traditional approaches for solving this challenge leverage trusted computing techniques, e.g., vTPM, or hardware CPU extensions, e.g., AMD SEV. But, they are vulnerable to powerful adversaries, or they provide only load time (not runtime) integrity measurements of VMs. We propose TRIGLAV, a protocol allowing tenants to establish and maintain trust in VM runtime integrity of software and its configuration. TRIGLAV is transparent to the VM configuration and setup. It performs an implicit attestation of VMs during a secure login and binds the VM integrity state with the secure connection. Our prototype's evaluation shows that TRIGLAV is practical and incurs low performance overhead (< 6%). Wojciech Ozga, Do Le Quoc, Christof Fetzer |
CLOUD | 3 |
| 2021 | Perun: Confidential Multi-stakeholder Machine Learning Framework with Hardware Acceleration Support
Wojciech Ozga, Do Le Quoc, Christof Fetzer |
DBSec | 3 |
| 2021 | ADAM-CS: Advanced Asynchronous Monotonic Counter ServiceabstractTrusted execution environments (TEEs) offer the technological breakthrough to allow several applications to be deployed and executed over untrusted public cloud environments. Although TEEs (e. g., Intel SGX, ARM TrustZone, AMD SEV) provide several mechanisms to ensure confidentiality and integrity of data and code, they do not offer freshness out of the box, a critical aspect yet often overlooked, for instance, to protect against rollback attacks. Monotonic counters are a popular way to detect rollbacks, as their counter values cannot be decremented. However, counter increments are slow (i.e., 10thof milliseconds), making their use impractical for distributed services and applications processing thousands of transactions simultaneously, for which an order of magnitude improvement is needed. ADAM-CS is an asynchronous monotonic counter service to protect such high-traffic applications against rollback attacks. Leveraging a set of distributed monotonic counters and specific algorithms, ADAM-CS minimizes the maximum vulnerability window (MVW), i.e., the amount of transactions an adversary could successfully rollback. Thanks to its asynchronous nature, ADAM-CS supports thousands of increments per second without introducing additional latency in the transactions performed by applications. Our measurements indicate that we can keep the MVW well below 10ms while supporting a throughput of more than 21K requests/s when using eight counters. André Martin, Cong Lian, Franz Gregor, Robert Krahn, Valerio Schiavoni, Pascal Felber, Christof Fetzer |
DSN | 7 |
| 2021 | Credentials as a Service Providing Self Sovereign Identity as a Cloud Service Using Trusted Execution EnvironmentsabstractWith increasing digitization, more and more people use their identification credentials for accessing online services; which increases concern for data privacy. To ensure user's privacy, alternate credential management schemes must be adopted. Self-Sovereign Identity (SSI) is a form of credential management where users are in charge of their credentials. Privacy-critical data is stored at the user's end and they can choose to do selective disclosure of minimal required information to access services. Currently, SSI solutions are not being widely adopted by service providers and the ecosystem is fragmented. One of the reasons for the lack of adoption is the need for maintaining private infrastructure for credential issuance, as critical user information is to be handled during credential issuance. To cater to this, we present a solution that enables the service providers to run their credential issuers on public cloud - a so-called Credentials as a Service (CaaS). CaaS issuers run inside Trusted Execution Environments (TEE) enabling credential issuers to ensure user's privacy while enjoying the flexibility of the pay-per-use cloud model. CaaS can pave the way for making SSI credentials ubiquitous in identity management solutions. Hira Siddiqui, Mujtaba Idrees, Ivan Gudymenko, Do Le Quoc, Christof Fetzer |
IC2E | 5 |
| 2021 | BROFY: Towards Essential Integrity Protection for MicroservicesabstractTrusted computing has emerged as one of the main components in a critical microservice application. A powerful adversary such as the cloud provider could harm its integrity by altering the application's code, behavior, and memory. Numerous attempts to preserve application integrity have been made, especially using Trusted Execution Environments (TEE). However, recent studies show that a CPU bitflip, which both adversary or faulty hardware can trigger, may invalidate its integrity despite being executed inside TEE. In the form of Silent Data Corruption (SDC), this bitflip may come undetected and shamble the trust built in a distributed system. We present BROFY, a toolchain that makes the program reliably perform correct computation inside the Intel SGX enclave that already provides code and memory integrity protection out-of-the-box. BROFY is compatible with multiple programming languages, needs no specific requirements or changes on the codebase, and offers a configurable trade-off between recovery ability and performance. We tested BROFY against actual bitflips by undervolting CPU, and our results show a significant decrease in irrecoverable failure rate from 96.7% to 0.5%, with a 100% detection rate inside an SGX enclave. Our experiment shows that programs armored by BROFY, compared to native execution, have 84% overhead on average based on the computation-intensive Starbench benchmark and only 3% overhead on a multithreaded HTTP server application written in C. Ardhi Putra Pratama Hartono, Christof Fetzer |
SRDS | 2 |
| 2021 | Active replication for latency-sensitive stream processing in Apache FlinkabstractStream processing frameworks allow processing massive amounts of data shortly after it is produced, and enable a fast reaction to events in scenarios such as data center monitoring, smart transportation, or telecommunication networks. Many scenarios depend on the fast and reliable processing of incoming data, requiring low end-to-end latencies from the ingest of a new event to the corresponding output. The occurrence of faults jeopardizes these guarantees: Currently-leading high-availability solutions for stream processing such as Spark Streaming or Apache Flink's implement passive replication through snapshotting, requiring a stop-the-world operation to recover from a failure. Active replication, while incurring higher deployment costs, can overcome these limitations and allow to mask the impact of faults and match stringent end-to-end latency requirements. We present the design, implementation, and evaluation of active replication in the popular Apache Flink platform. Our study explores two alternative designs, a leader-based approach leveraging external services (Kafka and ZooKeeper) and a leaderless implementation leveraging a novel deterministic merging algorithm. Our evaluation using a series of microbenchmarks and a SaaS cloud monitoring scenario on a 37-server cluster show that the actively-replicated Flink can fully mask the impact of faults on end-to-end latency. Guillaume Rosinosky, Florian Schmidt 0009, Oleh Bodunov, Christof Fetzer, André Martin, Etienne Rivière |
SRDS | 4 |
| 2020 | Vallum-Med: Protecting Medical Data in Cloud EnvironmentsabstractDespite the many advantages of cloud computing, keeping information in such an environment increases the risk of cyber attacks, as well as the possibility of unauthorized access by cloud provider employees. Another critical concern is privacy protection, since depending on data access control, confidential information may be exposed even through authorized access. To solve these issues we have previously proposed Vallum, a platform that leverages Intel SGX protection to ensure the security, confidentiality, and integrity of data at rest and during processing. It also provides tools for privacy protection, following policies set by the data owner. In this demo we present Vallum-Med, an application of Vallum for the protection of medical patient personal data, including imaging results of their cardiac examinations. We will demonstrate that this system fully supports cloud protection of such sensitive data as well as the definition of privacy policies and ensuring that all results of queries are compliant to these policies. All processing, data storage and network traffic are protected using SCONE, a docker container-based technology for seamlessly incorporating SGX protection for applications, which provides a fully encrypted memory environment. Ronny Peterson, Altigran S. da Silva, Christof Fetzer, André Martin, Ignacio Blanquer |
CIKM | 4 |
| 2020 | T-Lease: a trusted lease primitive for distributed systemsabstractA lease is an important primitive for building distributed protocols, and it is ubiquitously employed in distributed systems. However, the scope of the classic lease abstraction is restricted to the trusted computing infrastructure. Unfortunately, this important primitive cannot be employed in the untrusted computing infrastructure because the trusted execution environments (TEEs) do not provide a trusted time source. In the untrusted environment, an adversary can easily manipulate the system clock to violate the correctness properties of lease-based systems. Bohdan Trach, Rasha Faqeh, Oleksii Oleksenko, Wojciech Ozga, Pramod Bhatotia, Christof Fetzer |
SoCC | 6 |
| 2020 | LEGaTO: Low-Energy, Secure, and Resilient Toolset for Heterogeneous ComputingabstractThe LEGaTO project leverages task-based programming models to provide a software ecosystem for Made in-Europe heterogeneous hardware composed of CPUs, GPUs, FPGAs and dataflow engines. The aim is to attain one order of magnitude energy savings from the edge to the converged cloud/HPC, balanced with the security and resilience challenges. LEGaTO is an ongoing three-year EU H2020 project started in December 2017. Behzad Salami 0001, Konstantinos Parasyris, Adrián Cristal, Osman S. Unsal, Xavier Martorell, Raúl de la Cruz, Leonardo Arturo Bautista-Gomez, Daniel A. Jiménez, Carlos Álvarez 0001, Seyed Saber Nabavi Larimi, Sergi Madonar, Miquel Pericàs, Pedro Trancoso, Mustafa Abdul Jabbar, Jing Chen 0038, Pirah Noor Soomro, Madhavan Manivannan, Micha vor dem Berge, Stefan Krupop, Frank Klawonn, Al Mekhlafi, Sigrun May, Tobias Becker, Georgi Gaydadjiev, Hans Salomonsson, Devdatt P. Dubhashi, Oron Port, Yoav Etsion, Do Le Quoc, Christof Fetzer, Martin Kaiser, Nils Kucza, Jens Hagemeyer, René Griessl, Lennart Tigges, Kevin Mika, A. Hüffmeier, Marcelo Pasin, Valerio Schiavoni, Isabelly Rocha, Christian Göttel, Pascal Felber |
DATE | 31 |
| 2020 | Towards Formalization of Enhanced Privacy ID (EPID)-based Remote Attestation in Intel SGXabstractVulnerabilities in privileged software layers have been exploited with severe consequences. Recently, Trusted Execution Environments (TEEs) based technologies have emerged as a promising approach since they claim strong confidentiality and integrity guarantees regardless of the trustworthiness of the underlying system software. In this paper, we consider one of the most prominent TEE technologies, referred to as Intel Software Guard Extensions (SGX). Despite many formal approaches, there is still a lack of formal proof of some critical processes of Intel SGX, such as remote attestation. To fill this gap, we propose a fully automated, rigorous, and sound formal approach to specify and verify the Enhanced Privacy ID (EPID)-based remote attestation in Intel SGX under the assumption that there are no side-channel attacks and no vulnerabilities inside the enclave. The evaluation indicates that the confidentiality of attestation keys is preserved against a Dolev-Yao adversary in this technology. We also present a few of the many inconsistencies found in the existing literature on Intel SGX attestation during formal specification. Muhammad Usama Sardar, Do Le Quoc, Christof Fetzer |
DSD | 3 |
| 2020 | Trust Management as a Service: Enabling Trusted Execution in the Face of Byzantine StakeholdersabstractTrust is arguably the most important challenge for critical services both deployed as well as accessed remotely over the network. These systems are exposed to a wide diversity of threats, ranging from bugs to exploits, active attacks, rogue operators, or simply careless administrators. To protect such applications, one needs to guarantee that they are properly configured and securely provisioned with the "secrets" (e.g., encryption keys) necessary to preserve not only the confidentiality, integrity and freshness of their data but also their code. Furthermore, these secrets should not be kept under the control of a single stakeholder—which might be compromised and would represent a single point of failure—and they must be protected across software versions in the sense that attackers cannot get access to them via malicious updates. Traditional approaches for solving these challenges often use ad hoc techniques and ultimately rely on a hardware security module (HSM) as root of trust. We propose a more powerful and generic approach to trust management that instead relies on trusted execution environments (TEEs) and a set of stakeholders as root of trust. Our system, PALÆMON, can operate as a managed service deployed in an untrusted environment, i.e., one can delegate its operations to an untrusted cloud provider with the guarantee that data will remain confidential despite not trusting any individual human (even with root access) nor system software. PALÆMON addresses in a secure, efficient and cost-effective way five main challenges faced when developing trusted networked applications and services. Our evaluation on a range of benchmarks and real applications shows that PALÆMON performs efficiently and can protect secrets of services without any change to their source code. Franz Gregor, Wojciech Ozga, Sébastien Vaucher, Rafael Pires 0001, Do Le Quoc, Sergei Arnautov, André Martin, Valerio Schiavoni, Pascal Felber, Christof Fetzer |
DSN | 10 |
| 2020 | Formal Foundations for Intel SGX Data Center Attestation Primitives
Muhammad Usama Sardar, Rasha Faqeh, Christof Fetzer |
ICFEM | 3 |
| 2020 | Towards Dynamic Dependable Systems Through Evidence-Based Continuous Certification
Rasha Faqeh, Christof Fetzer, Holger Hermanns, Jörg Hoffmann 0001, Michaela Klauck, Maximilian A. Köhl, Marcel Steinmetz, Christoph Weidenbach |
ISoLA (2) | 2 |
| 2020 | TEEMon: A continuous performance monitoring framework for TEEsabstractTrusted Execution Environments (TEEs), such as Intel Software Guard eXtensions (SGX), are considered as a promising approach to resolve security challenges in clouds. TEEs protect the confidentiality and integrity of application code and data even against privileged attackers with root and physical access by providing an isolated secure memory area, i.e., enclaves. The security guarantees are provided by the CPU, thus even if system software is compromised, the attacker can never access the enclave's content. While this approach ensures strong security guarantees for applications, it also introduces a considerable runtime overhead in part by the limited availability of protected memory (enclave page cache). Currently, only a limited number of performance measurement tools for TEE-based applications exist and none offer performance monitoring and analysis during runtime. Robert Krahn, Donald Dragoti, Franz Gregor, Do Le Quoc, Valerio Schiavoni, Pascal Felber, Clenimar Souza, Andrey Brito, Christof Fetzer |
Middleware | 9 |
| 2020 | A practical approach for updating an integrity-enforced operating systemabstractTrusted computing defines how to securely measure, store, and verify the integrity of software controlling a computer. One of the major challenge that make them hard to be applied in practice is the issue with software updates. Specifically, an operating system update causes the integrity violation because it changes the well-known initial state trusted by remote verifiers, such as integrity monitoring systems. Consequently, the integrity monitoring of remote computers becomes unreliable due to the high amount of false positives. Wojciech Ozga, Do Le Quoc, Christof Fetzer |
Middleware | 3 |
| 2020 | secureTF: A Secure TensorFlow FrameworkabstractData-driven intelligent applications in modern online services have become ubiquitous. These applications are usually hosted in the untrusted cloud computing infrastructure. This poses significant security risks since these applications rely on applying machine learning algorithms on large datasets which may contain private and sensitive information. Do Le Quoc, Franz Gregor, Sergei Arnautov, Roland Kunkel, Pramod Bhatotia, Christof Fetzer |
Middleware | 6 |
| 2020 | SpecFuzz: Bringing Spectre-type vulnerabilities to the surface
Oleksii Oleksenko, Bohdan Trach, Mark Silberstein, Christof Fetzer |
USENIX Security Symposium | 4 |
| 2020 | Federated and secure cloud services for building medical image classifiers on an intercontinental infrastructure
Ignacio Blanquer, Francisco Vilar Brasileiro, Andrey Brito, Amanda Calatrava, Christof Fetzer, Flavio Figueiredo, Ronny Petterson Guimarães, Leandro Bezerra Marinho, Wagner Meira Jr., Altigran S. da Silva, Angel Alberich-Bayarri, Eduardo Camacho-Ramos, Ana Jimenez-Pastor, Antonio Luiz L. Ribeiro, Bruno Ramos Nascimento |
Future Gener. Comput. Syst. | 6 |
| 2019 | Vallum: Privacy, Confidentiality and Access Controlfor Sensitive Data in Cloud EnvironmentsabstractManaging sensitive data in shared environments such as public clouds is an enduring challenge. While several approaches exist to protect data at rest such as end-to-end encryption, there exist only a few solutions such as homomorphic encryption that offer secure data processing. Unfortunately, these solutions cannot be used in practice as they incur non-negligible run-time overheads and security risks. Moreover, as the majority of data management systems were designed to operate in private cloud environments, which are under the control of the data owner, they often lack appropriate mechanisms for access control as well as privacy assurance. In this paper we propose Vallum, a data access and protection layer that closes these gaps while enabling users to operate data management systems in shared environments such securely as public clouds. Vallum utilizes Intel SGX and remote attestation to ensure confidentiality and integrity of the data being stored and processed. Furthermore, it provides access protection and privacy assurance through a extensible architecture. Our performance evaluation indicates that the overhead introduced by Vallum makes it viable to be deployed in cloud infrastructures. Ronny Peterson, Altigran S. da Silva, Gabriel Fernandez 0001, André Martin, Christof Fetzer, Andrey Brito |
CloudCom | 6 |
| 2019 | TEE-Perf: A Profiler for Trusted Execution EnvironmentsabstractWe introduce TEE-PERF, an architecture-and platform-independent performance measurement tool for trusted execution environments (TEEs). More specifically, TEE-PERF supports method-level profiling for unmodified multithreaded applications, without relying on any architecture-specific hardware features (e.g. Intel VTune Amplifier), or without requiring platform-dependent kernel features (e.g. Linux perf). Moreover, TEE-PERF provides accurate profiling measurements since it traces the entire process execution without employing instruction pointer sampling. Thus, TEE-PERF does not suffer from sampling frequency bias, which can occur with threads scheduled to align to the sampling frequency. We have implemented TEE-P ERF with an easy to use interface, and integrated it with Flame Graphs to visualize the performance bottlenecks. We have evaluated TEE-PERF based on the Phoenix multithreaded benchmark suite and real-world applications (RocksDB, SPDK, etc.), and compared it with Linux perf. Our experimental evaluation shows that TEE-PERF incurs low profiling overheads, while providing accurate profile measurements to identify and optimize the application bottlenecks in the context of TEEs. TEE-PERF is publicly available. Maurice Bailleu, Donald Dragoti, Pramod Bhatotia, Christof Fetzer |
DSN | 4 |
| 2019 | SPEICHER: Securing LSM-based Key-Value Stores using Shielded Execution
Maurice Bailleu, Jörg Thalheim, Pramod Bhatotia, Christof Fetzer, Michio Honda, Kapil Vaswani |
FAST | 4 |
| 2019 | Clemmys: towards secure remote execution in FaaSabstractWe introduce Clemmys, a security-first serverless platform that ensures confidentiality and integrity of users' functions and data as they are processed on untrusted cloud premises, while keeping the cost of protection low. We provide a design for hardening FaaS platforms with Intel SGX---a hardware-based shielded execution technology. We explain the protocol that our system uses to ensure confidentiality and integrity of data, and integrity of function chains. To overcome performance and latency issues that are inherent in SGX applications, we apply several SGX-specific optimizations to the runtime system: we use SGXv2 to speed up the enclave startup and perform batch EPC augmentation. To evaluate our approach, we implement our design over Apache Open-Whisk, a popular serverless platform. Lastly, we show that Clemmys achieved same throughput and similar latency as native Apache OpenWhisk, while allowing it to withstand several new attack vectors. Bohdan Trach, Oleksii Oleksenko, Franz Gregor, Pramod Bhatotia, Christof Fetzer |
SYSTOR | 5 |
| 2019 | CoSMIX: A Compiler-based System for Secure Memory Instrumentation and Execution in Enclaves
Meni Orenbach, Yan Michalevsky, Christof Fetzer, Mark Silberstein |
USENIX ATC | 3 |
| 2019 | SGX-PySpark: Secure Distributed Data AnalyticsabstractData analytics is central to modern online services, particularly those data-driven. Often this entails the processing of large-scale datasets which may contain private, personal and sensitive information relating to individuals and organisations. Particular challenges arise where cloud is used to store and process the sensitive data. In such settings, security and privacy concerns become paramount, as the cloud provider is trusted to guarantee the security of the services they offer, including data confidentiality. Therefore, the issue this work tackles is “How to securely perform data analytics in a public cloud?” Do Le Quoc, Franz Gregor, Jatinder Singh, Christof Fetzer |
WWW | 4 |
| 2018 | LEGaTO: towards energy-efficient, secure, fault-tolerant toolset for heterogeneous computingabstractLEGaTO is a three-year EU H2020 project which started in December 2017. The LEGaTO project will leverage task-based programming models to provide a software ecosystem for Made-in-Europe heterogeneous hardware composed of CPUs, GPUs, FPGAs and dataflow engines. The aim is to attain one order of magnitude energy savings from the edge to the converged cloud/HPC. Adrián Cristal, Osman S. Unsal, Xavier Martorell, Raúl de la Cruz, Leonardo Arturo Bautista-Gomez, Daniel Jiménez-González, Carlos Álvarez 0001, Behzad Salami 0001, Sergi Madonar, Miquel Pericàs, Pedro Trancoso, Micha vor dem Berge, Gunnar Billung-Meyer, Stefan Krupop, Wolfgang Christmann, Frank Klawonn, Amani Mihklafi, Tobias Becker, Georgi Gaydadjiev, Hans Salomonsson, Devdatt P. Dubhashi, Oron Port, Yoav Etsion, Vesna Nowack, Christof Fetzer, Jens Hagemeyer, Thorsten Jungeblut, Nils Kucza, Martin Kaiser, Mario Porrmann, Marcelo Pasin, Valerio Schiavoni, Isabelly Rocha, Christian Göttel, Pascal Felber |
CF | 26 |
| 2018 | ApproxJoin: Approximate Distributed JoinsabstractA distributed join is a fundamental operation for processing massive datasets in parallel. Unfortunately, computing an equi-join over such datasets is very resource-intensive, even when done in parallel. Given this cost, the equi-join operator becomes a natural candidate for optimization using approximation techniques, which allow users to trade accuracy for latency. Finding the right approximation technique for joins, however, is a challenging task. Sampling, in particular, cannot be directly used in joins; naïvely performing a join over a sample of the dataset will not preserve statistical properties of the query result. Do Le Quoc, Istemi Ekin Akkus, Pramod Bhatotia, Spyros Blanas, Ruichuan Chen, Christof Fetzer, Thorsten Strufe |
SoCC | 6 |
| 2018 | EndBox: Scalable Middlebox Functions Using Client-Side Trusted ExecutionabstractMany organisations enhance the performance, security, and functionality of their managed networks by deploying middleboxes centrally as part of their core network. While this simplifies maintenance, it also increases cost because middlebox hardware must scale with the number of clients. A promising alternative is to outsource middlebox functions to the clients themselves, thus leveraging their CPU resources. Such an approach, however, raises security challenges for critical middlebox functions such as firewalls and intrusion detection systems. We describe EndBox, a system that securely executes middlebox functions on client machines at the network edge. Its design combines a virtual private network (VPN) with middlebox functions that are hardware-protected by a trusted execution environment (TEE), as offered by Intel's Software Guard Extensions (SGX). By maintaining VPN connection endpoints inside SGX enclaves, EndBox ensures that all client traffic, including encrypted communication, is processed by the middlebox. Despite its decentralised model, EndBox's middlebox functions remain maintainable: they are centrally controlled and can be updated efficiently. We demonstrate EndBox with two scenarios involving (i) a large company; and (ii) an Internet service provider that both need to protect their network and connected clients. We evaluate EndBox by comparing it to centralised deployments of common middlebox functions, such as load balancing, intrusion detection, firewalling, and DDoS prevention. We show that EndBox achieves up to 3.8x higher throughput and scales linearly with the number of clients. David Goltzsche, Signe Rüsch, Manuel Nieke, Sébastien Vaucher, Nico Weichbrodt, Valerio Schiavoni, Pierre-Louis Aublin, Paolo Costa, Christof Fetzer, Pascal Felber, Peter R. Pietzuch, Rüdiger Kapitza |
DSN | 9 |
| 2018 | LibSEAL: revealing service integrity violations using trusted executionabstractUsers of online services such as messaging, code hosting and collaborative document editing expect the services to uphold the integrity of their data. Despite providers' best efforts, data corruption still occurs, but at present service integrity violations are excluded from SLAs. For providers to include such violations as part of SLAs, the competing requirements of clients and providers must be satisfied. Clients need the ability to independently identify and prove service integrity violations to claim compensation. At the same time, providers must be able to refute spurious claims. Pierre-Louis Aublin, Florian Kelbert, Dan O'Keeffe, Divya Muthukumaran, Christian Priebe, Joshua Lind, Robert Krahn, Christof Fetzer, David M. Eyers, Peter R. Pietzuch |
EuroSys | 8 |
| 2018 | Pesos: policy enhanced secure object storeabstractThird-party storage services pose the risk of integrity and confidentiality violations as the current storage policy enforcement mechanisms are spread across many layers in the system stack. To mitigate these security vulnerabilities, we present the design and implementation of Pesos, a Policy Enhanced Secure Object Store (Pesos) for untrusted third-party storage providers. Pesos allows clients to specify per-object security policies, concisely and separately from the storage stack, and enforces these policies by securely mediating the I/O in the persistence layer through a single unified enforcement layer. More broadly, Pesos exposes a rich set of storage policies ensuring the integrity, confidentiality, and access accounting for data storage through a declarative policy language. Robert Krahn, Bohdan Trach, Anjo Vahldiek-Oberwagner, Thomas Knauth, Pramod Bhatotia, Christof Fetzer |
EuroSys | 6 |
| 2018 | SGX-Aware Container Orchestration for Heterogeneous ClustersabstractContainers are becoming the de facto standard to package and deploy applications and micro-services in the cloud. Several cloud providers (e.g., Amazon, Google, Microsoft) begin to offer native support on their infrastructure by integrating container orchestration tools within their cloud offering. At the same time, the security guarantees that containers offer to applications remain questionable. Customers still need to trust their cloud provider with respect to data and code integrity. The recent introduction by Intel of Software Guard Extensions (SGX) into the mass market offers an alternative to developers, who can now execute their code in a hardware-secured environment without trusting the cloud provider. This paper provides insights regarding the support of SGX inside Kubernetes, an industry-standard container orchestrator. We present our contributions across the whole stack supporting execution of SGX-enabled containers. We provide details regarding the architecture of the scheduler and its monitoring framework, the underlying operating system support and the required kernel driver extensions. We evaluate our complete implementation on a private cluster using the real-world Google Borg traces. Our experiments highlight the performance trade-offs that will be encountered when deploying SGX-enabled micro-services in the cloud. Sébastien Vaucher, Rafael Pires 0001, Pascal Felber, Marcelo Pasin, Valerio Schiavoni, Christof Fetzer |
ICDCS | 6 |
| 2018 | PubSub-SGX: Exploiting Trusted Execution Environments for Privacy-Preserving Publish/Subscribe SystemsabstractThis paper presents PUBSUB-SGX, a content-based publish-subscribe system that exploits trusted execution environments (TEEs), such as Intel SGX, to guarantee confidentiality and integrity of data as well as anonymity and privacy of publishers and subscribers. We describe the technical details of our Python implementation, as well as the required system support introduced to deploy our system in a container-based runtime. Our evaluation results show that our approach is sound, while at the same time highlighting the performance and scalability trade-offs. In particular, by supporting just-in-time compilation inside of TEEs, Python programs inside of TEEs are in general faster than when executed natively using standard CPython. Sergei Arnautov, Andrey Brito, Pascal Felber, Christof Fetzer, Franz Gregor, Robert Krahn, Wojciech Ozga, André Martin, Valerio Schiavoni, Marcus Tenorio, Nikolaus Thummel |
SRDS | 4 |
| 2018 | Varys: Protecting SGX Enclaves from Practical Side-Channel Attacks
Oleksii Oleksenko, Bohdan Trach, Robert Krahn, Mark Silberstein, Christof Fetzer |
USENIX ATC | 5 |
| 2017 | Integrating Reactive Cloud Applications in SERECAabstractA consolidated trend in designing cloud-based applications is to make use of a reactive microservice architecture, which allows to divide an application in several well-partitioned software units with specific responsibilities. Such an architecture perfectly fits in cloud environments, ensuring a number of advantages (i.e., high availability and scalability, ease of deployment and development). However, the new way of designing cloud applications introduces challenging security threats. Besides the difficulty in monitoring security of the overall distributed application, an important aspect of concern relates to the risk of break the chain of trust established among the different microservices belonging to the application. That is, a compromised single microservice may bring down the other related ones. Christof Fetzer, Giovanni Mazzeo, John Oliver, Luigi Romano, Martijn Verburg |
ARES | 1 |
| 2017 | SecureCloud: Secure big data processing in untrusted cloudsabstractWe present the SecureCloud EU Horizon 2020 project, whose goal is to enable new big data applications that use sensitive data in the cloud without compromising data security and privacy. For this, SecureCloud designs and develops a layered architecture that allows for (i) the secure creation and deployment of secure micro-services; (ii) the secure integration of individual micro-services to full-fledged big data applications; and (iii) the secure execution of these applications within untrusted cloud environments. To provide security guarantees, SecureCloud leverages novel security mechanisms present in recent commodity CPUs, in particular, Intel's Software Guard Extensions (SGX). SecureCloud applies this architecture to big data applications in the context of smart grids. We describe the SecureCloud approach, initial results, and considered use cases. Florian Kelbert, Franz Gregor, Rafael Pires 0001, Stefan Köpsell, Marcelo Pasin, Aurelien Havet, Valerio Schiavoni, Pascal Felber, Christof Fetzer, Peter R. Pietzuch |
DATE | 9 |
| 2017 | Fex: A Software Systems EvaluatorabstractSoftware systems research relies on experimental evaluation to assess the effectiveness of newly developed solutions. However, the existing evaluation frameworks are rigid (do not allow creation of new experiments), often simplistic (may not reveal issues that appear in real-world applications), and can be inconsistent (do not guarantee reproducibility of experiments across platforms). This paper presents Fex, a software systems evaluation framework that addresses these limitations. Fex is extensible (can be easily extended with custom experiment types), practical (supports composition of different benchmark suites and real-world applications), and reproducible (it is built on container technology to guarantee the same software stack across platforms). We show that Fex achieves these design goals with minimal end-user effort - for instance, adding Nginx web-server to evaluation requires only 160 LoC. Going forward, we discuss the architecture of the framework, explain its interface, show common usage scenarios, and evaluate the efforts for writing various custom extensions. Oleksii Oleksenko, Dmitrii Kuvaiskii, Pramod Bhatotia, Christof Fetzer |
DSN | 4 |
| 2017 | SGXBOUNDS: Memory Safety for Shielded ExecutionabstractShielded execution based on Intel SGX provides strong security guarantees for legacy applications running on untrusted platforms. However, memory safety attacks such as Heartbleed can render the confidentiality and integrity properties of shielded execution completely ineffective. To prevent these attacks, the state-of-the-art memory-safety approaches can be used in the context of shielded execution. Dmitrii Kuvaiskii, Oleksii Oleksenko, Sergei Arnautov, Bohdan Trach, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
EuroSys | 7 |
| 2017 | GENPACK: A Generational Scheduler for Cloud Data CentersabstractCloud data centers largely rely on virtualization to provision resources and host services across their infrastructure. The scheduling problem has been widely studied and is well understood when the resource requirements and the expected lifetime of services are known beforehand. In contrast, when workloads are not known in advance, effective scheduling of services, and more generally system containers, becomes much more complex. In this paper, we propose GENPACK, a framework for system containers scheduling in cloud data centers that leverages principles from generational garbage collection (GC). It combines runtime monitoring of system containers to learn their requirements and properties, and a scheduler that manages different generations of servers. The population of these generations may vary over time depending on the global load, hence they are subject to being shut down when idle to save energy. We implemented GENPACK and tested it in a dedicated data center, showing that it can be up to 23% more energy-efficient that SWARM's built-in scheduling policies on a real-world trace. Aurelien Havet, Valerio Schiavoni, Pascal Felber, Maxime Colmant, Romain Rouvoy, Christof Fetzer |
IC2E | 6 |
| 2017 | FFQ: A Fast Single-Producer/Multiple-Consumer Concurrent FIFO QueueabstractWith the spreading of multi-core architectures, operating systems and applications are becoming increasingly more concurrent and their scalability is often limited by the primitives used to synchronize the different hardware threads. In this paper, we address the problem of how to optimize the throughput of a system with multiple producer and consumer threads. Such applications typically synchronize their threads via multi-producer/multi-consumer FIFO queues, but existing solutions have poor scalability, as we could observe when designing a secure application framework that requires high-throughput communication between many concurrent threads. In our target system, however, the items enqueued by different producers do not necessarily need to be FIFO ordered. Hence, we propose a fast FIFO queue, FFQ, that aims at maximizing throughput by specializing the algorithm for single-producer/multiple-consumer settings: each producer has its own queue from which multiple consumers can concurrently dequeue. Furthermore, while we provide a wait-free interface for producers, we limit ourselves to lock-free consumers to eliminate the need for helping. We also propose a multi-producer variant to show which synchronization operations we were able to remove by focusing on a single producer variant. Our evaluation analyses the performance using micro-benchmarks and compares our results with other state-of-the-art solutions: FFQ exhibits excellent performance and scalability. Sergei Arnautov, Pascal Felber, Christof Fetzer, Bohdan Trach |
IPDPS | 3 |
| 2017 | StreamApprox: approximate computing for stream analyticsabstractApproximate computing aims for efficient execution of workflows where an approximate output is sufficient instead of the exact output. The idea behind approximate computing is to compute over a representative sample instead of the entire input dataset. Thus, approximate computing --- based on the chosen sample size --- can make a systematic trade-off between the output accuracy and computation efficiency. Do Le Quoc, Ruichuan Chen, Pramod Bhatotia, Christof Fetzer, Volker Hilt, Thorsten Strufe |
Middleware | 4 |
| 2017 | Sieve: actionable insights from monitored metrics in distributed systemsabstractMajor cloud computing operators provide powerful monitoring tools to understand the current (and prior) state of the distributed systems deployed in their infrastructure. While such tools provide a detailed monitoring mechanism at scale, they also pose a significant challenge for the application developers/operators to transform the huge space of monitored metrics into useful insights. These insights are essential to build effective management tools for improving the efficiency, resiliency, and dependability of distributed systems. Jörg Thalheim, Antonio Rodrigues, Istemi Ekin Akkus, Pramod Bhatotia, Ruichuan Chen, Bimal Viswanath, Lei Jiao 0002, Christof Fetzer |
Middleware | 8 |
| 2017 | Glamdring: Automatic Application Partitioning for Intel SGX
Joshua Lind, Christian Priebe, Divya Muthukumaran, Dan O'Keeffe, Pierre-Louis Aublin, Florian Kelbert, Tobias Reiher, David Goltzsche, David M. Eyers, Rüdiger Kapitza, Christof Fetzer, Peter R. Pietzuch |
USENIX ATC | 11 |
| 2017 | PrivApprox: Privacy-Preserving Stream Analytics
Do Le Quoc, Martin Beck, Pramod Bhatotia, Ruichuan Chen, Christof Fetzer, Thorsten Strufe |
USENIX ATC | 5 |
| 2016 | Energy minimization at all layers of the data center: The ParaDIME project
Oscar Palomar, Santhosh Kumar Rethinagiri, Gulay Yalcin, J. Rubén Titos Gil, Pablo Prieto, Emma Torrella, Osman S. Unsal, Adrián Cristal, Pascal Felber, Anita Sobe, Yaroslav Hayduk, Mascha Kurpicz, Christof Fetzer, Thomas Knauth, Malte Schneegaß, Jens Struckmeier, Dragomir Milojevic |
DATE | 13 |
| 2016 | ELZAR: Triple Modular Redundancy Using Intel AVX (Practical Experience Report)abstractInstruction-Level Redundancy (ILR) is a well-known approach to tolerate transient CPU faults. It replicates instructions in a program and inserts periodic checks to detect and correct CPU faults using majority voting, which essentially requires three copies of each instruction and leads to high performance overheads. As SIMD technology can operate simultaneously on several copies of the data, it appears to be a good candidate for decreasing these overheads. To verify this hypothesis, we propose ELZAR, a compiler framework that transforms unmodified multithreaded applications to support triple modular redundancy using Intel AVX extensions for vectorization. Our experience with several benchmark suites and real-world case-studies yields mixed results: while SIMD may be beneficial for some workloads, e.g., CPU-intensive ones with many floating-point operations, it exposes higher overhead than ILR in many applications we tested. Dmitrii Kuvaiskii, Oleksii Oleksenko, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
DSN | 5 |
| 2016 | HAFT: hardware-assisted fault toleranceabstractTransient hardware faults during the execution of a program can cause data corruptions. We present HAFT, a fault tolerance technique using hardware extensions of commodity CPUs to protect unmodified multithreaded applications against such corruptions. HAFT utilizes instruction-level redundancy for fault detection and hardware transactional memory for fault recovery. We evaluated HAFT with Phoenix and PARSEC benchmarks. The observed normalized runtime is 2x, with 98.9% of the injected data corruptions being detected and 91.2% being corrected. To demonstrate the effectiveness of HAFT, we applied it to real-world case studies including Memcached, Apache, and SQLite. Dmitrii Kuvaiskii, Rasha Faqeh, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
EuroSys | 5 |
| 2016 | INSPECTOR: Data Provenance Using Intel Processor Trace (PT)abstractData provenance strives for explaining how the computation was performed by recording a trace of the execution. The provenance trace is useful across a wide-range of workflows to improve the dependability, security, and efficiency of software systems. In this paper, we present Inspector, a POSIX-compliant data provenance library for shared-memory multithreaded programs. The Inspector library is completely transparent and easy to use: it can be used as a replacement for the pthreads library by a simple exchange of libraries linked, without even recompiling the application code. To achieve this result, we present a parallel provenance algorithm that records control, data, and schedule dependencies using a Concurrent Provenance Graph (CPG). We implemented our algorithm to operate at the compiled binary code level by leveraging a combination of OS-specific mechanisms, and recently released Intel PT ISA extensions as part of the Broadwell micro-architecture. Our evaluation on a multicore platform using applications from multithreaded benchmarks suites (PARSEC and Phoenix) shows reasonable provenance overheads for a majority of applications. Lastly, we briefly describe three case-studies where the generic interface exported by Inspector is being used to improve the dependability, security, and efficiency of systems. The Inspector library is publicly available for further use in a wide range of other provenance workflows. Jörg Thalheim, Pramod Bhatotia, Christof Fetzer |
ICDCS | 3 |
| 2016 | Quality-driven disorder handling for m-way sliding window stream joinsabstractSliding window join is one of the most important operators for stream applications. To produce high quality join results, a stream processing system must deal with the ubiquitous disorder within input streams which is caused by network delay, parallel processing, etc. Disorder handling involves an inevitable tradeoff between the latency and the quality of produced join results. To meet different requirements of stream applications, it is desirable to provide a user-configurable result-latency vs. result-quality tradeoff. Existing disorder handling approaches either do not provide such configurability, or support only user-specified latency constraints. In this work, we advocate the idea of quality-driven disorder handling, and propose a buffer-based disorder handling approach for sliding window joins, which minimizes sizes of input-sorting buffers, thus the result latency, while respecting user-specified result-quality requirements. The core of our approach is an analytical model which directly captures the relationship between sizes of input buffers and the produced result quality. Our approach is generic. It supports m-way sliding window joins with arbitrary join conditions. Experiments on real-world and synthetic datasets show that, compared to the state of the art, our approach can reduce the result latency incurred by disorder handling by up to 95% while providing the same level of result quality. Yuanzhen Ji, Anisoara Nica, Zbigniew Jerzak, Gregor Hackenbroich, Christof Fetzer |
ICDE | 6 |
| 2016 | Compliance, Functional Safety and Fault Detection by Formal Methods
Christof Fetzer, Christoph Weidenbach, Patrick Wischnewski |
ISoLA (2) | 1 |
| 2016 | SecureKeeper: Confidential ZooKeeper using Intel SGX
Stefan Brenner, Colin Wulf, David Goltzsche, Nico Weichbrodt, Matthias Lorenz, Christof Fetzer, Peter R. Pietzuch, Rüdiger Kapitza |
Middleware | 6 |
| 2016 | Secure Content-Based Routing Using Intel Software Guard Extensions
Rafael Pires 0001, Marcelo Pasin, Pascal Felber, Christof Fetzer |
Middleware | 4 |
| 2016 | SCONE: Secure Linux Containers with Intel SGX
Sergei Arnautov, Bohdan Trach, Franz Gregor, Thomas Knauth, André Martin, Christian Priebe, Joshua Lind, Divya Muthukumaran, Dan O'Keeffe, Mark Stillwell, David Goltzsche, David M. Eyers, Rüdiger Kapitza, Peter R. Pietzuch, Christof Fetzer |
OSDI | 15 |
| 2016 | IncApprox: A Data Analytics System for Incremental Approximate ComputingabstractIncremental and approximate computations are increasingly being adopted for data analytics to achieve low-latency execution and efficient utilization of computing resources. Incremental computation updates the output incrementally instead of re-computing everything from scratch for successive runs of a job with input changes. Approximate computation returns an approximate output for a job instead of the exact output. Both paradigms rely on computing over a subset of data items instead of computing over the entire dataset, but they differ in their means for skipping parts of the computation. Incremental computing relies on the memoization of intermediate results of sub-computations, and reusing these memoized results across jobs. Approximate computing relies on representative sampling of the entire dataset to compute over a subset of data items. Dhanya R. Krishnan, Do Le Quoc, Pramod Bhatotia, Christof Fetzer, Rodrigo Rodrigues 0001 |
WWW | 4 |
| 2015 | Scalable Network Traffic Classification Using Distributed Support Vector MachinesabstractInternet traffic has increased dramatically in recent years due to the popularization of the Internet and the appearance of wireless Internet mobile devices such as smart-phones and tablets. The explosive growth of Internet traffic has introduced a practical example that demonstrates the concept of Big Data. Accurate identification and classification of large network traffic data plays an important role in network management including capacity planning, network forensics, QoS and intrusion detection. However, the state-of-the-art solutions, which rely on a dedicated server, are not scalable for analyzing high volume network traffic data. In this paper, we implement a distributed Support Vector Machines (SVMs) framework for classifying network traffic using Hadoop, an open-source distributed computing framework for Big Data processing. We design a global parameter store that maintains the global shared parameters between SVM training nodes. The distributed SVMs have been deployed on a 20 node cluster to analyze real network traffic trace. The results demonstrate that with 19 Mapper nodes the system is around 30% faster than Cloud SVM solution and outperforms the standalone SVM with nearly 9 times faster in training process and 15 times in the classifying process. In addition, the distributed SVMs architecture is designed to analyze large scale datasets. Therefore, it can be used not only for processing network traffic dataset, but also other large scale datasets such as Web data. Do Le Quoc, Valerio D'Alessandro, Byungchul Park, Luigi Romano, Christof Fetzer |
CLOUD | 5 |
| 2015 | UniCrawl: A Practical Geographically Distributed Web CrawlerabstractAs the wealth of information available on the web keeps growing, being able to harvest massive amounts of data has become a major challenge. Web crawlers are the core components to retrieve such vast collections of publicly available data. The key limiting factor of any crawler architecture is however its large infrastructure cost. To reduce this cost, and in particular the high upfront investments, we present in this paper a geo-distributed crawler solution, UniCrawl. UniCrawl orchestrates several geographically distributed sites. Each site operates an independent crawler and relies on well-established techniques for fetching and parsing the content of the web. UniCrawl splits the crawled domain space across the sites and federates their storage and computing resources, while minimizing thee inter-site communication cost. To assess our design choices, we evaluate UniCrawl in a controlled environment using the ClueWeb12 dataset, and in the wild when deployed over several remote locations. We conducted several experiments over 3 sites spread across Germany. When compared to a centralized architecture with a crawler simply stretched over several locations, UniCrawl shows a performance improvement of 93.6% in terms of network bandwidth consumption, and a speedup factor of 1.75. Do Le Quoc, Christof Fetzer, Pascal Felber, Etienne Rivière, Valerio Schiavoni, Pierre Sutra |
CLOUD | 2 |
| 2015 | EHadoop: Network I/O Aware Scheduler for Elastic MapReduce ClusterabstractOver the last few years the usage of cloud computing dramatically increased. Many data analytics platforms run on the cloud. Such systems characterized by large data transfer among VMs. The network isolation between cloud users in modern data centers is not as good as CPU and memory isolation [20]. The weak isolation leads to unpredictable performance of the inter data center network. Moreover, with the raise of popularity of cloud computing the competition between providers get tougher, which leads to prices decrease. Some users decide to perform data-analytics in a cross-cloud fashion [11], which requires data transfer over WAN. It is known that WAN provides lower than LAN performance. We show that saturated network can greatly impact MapReduce job's task completion time. It results in higher costs for the user, because according to the pay-as-you-go model the user pays for the time resources being used. In this work we present EHadoop network I/O aware scheduler for elastic MapReduce cluster which performs online job profiling and schedules tasks based on available network bandwidth. The evaluation results show that EHadoop allows to avoid network contention and does not increase MapReduce task completion time with network bandwidth degradation. Lenar Yazdanov, Maxim Gorbunov, Christof Fetzer |
CLOUD | 3 |
| 2015 | Online parameter optimization for elastic data stream processingabstractElastic scaling allows data stream processing systems to dynamically scale in and out to react to workload changes. As a consequence, unexpected load peaks can be handled and the extent of the overprovisioning can be reduced. However, the strategies used for elastic scaling of such systems need to be tuned manually by the user. This is an error prone and cumbersome task, because it requires a detailed knowledge of the underlying system and workload characteristics. In addition, the resulting quality of service for a specific scaling strategy is unknown a priori and can be measured only during runtime. Thomas Heinze 0001, Lars Roediger, Andreas Meister 0001, Yuanzhen Ji, Zbigniew Jerzak, Christof Fetzer |
SoCC | 6 |
| 2015 | Resiliency-aware Data Compression for In-memory Database SystemsabstractNowadays, database systems pursuit a main memory-centric architecture, where the entire business-related
data is stored and processed in a compressed form in main memory. In this case, the performance gain is
massive because database operations can benefit from its higher bandwidth and lower latency. However,
current main memory-centric database systems utilize general-purpose error detection and correction solutions
to address the emerging problem of increasing dynamic error rate of main memory. The costs of these generalpurpose
methods dramatically increases with increasing error rates. To reduce these costs, we have to exploit
context knowledge of database systems for resiliency. Therefore, we introduce our vision of resiliency-aware
data compression in this paper, where we want to exploit the benefits of both fields in an integrated approach
with low performance and memory overhead. In detail, we present and evaluate a first approach using AN
encoding and two different compression schemes to show the potentials and challenges of our vision. Till Kolditz, Dirk Habich, Patrick Damme, Wolfgang Lehner, Dmitrii Kuvaiskii, Oleksii Oleksenko, Christof Fetzer |
DATA | 7 |
| 2015 | Δ-Encoding: Practical Encoded ProcessingabstractTransient and permanent errors in memory and CPUs occur with alarming frequency. Although most of these errors are masked at the hardware level or result in crashes, a non-negligible number of them leads to Silent Data Corruptions (SDCs), i.e., incorrect results of computations. Safety-critical programs require a very high level of confidence that such faults are detected and not propagated to the outside. Unfortunately, state-of-the-art fault detection techniques generally assume a limited Single Event Upset fault model, concentrating only on transient faults.We present Δ-encoding: a software-only approach to detect hardware faults with very high probability. Δ-encoding makes no assumptions on the rate and type of faults. Our approach combines AN codes and duplicated instructions to harden programs against transient and permanent hardware errors. Our evaluation shows that Δ-encoding detects 99.997% of all injected errors with performance slowdown of 2 - 4 times. Dmitrii Kuvaiskii, Christof Fetzer |
DSN | 2 |
| 2015 | User-Constraint and Self-Adaptive Fault Tolerance for Event Stream Processing SystemsabstractEvent Stream Processing (ESP) Systems are currently enabling a renaissance in the data processing area as they provide results at low latency compared to the traditional MapReduce approach. Although the majority of ESP systems offer some form of fault tolerance to their users, the provided fault tolerance scheme is often not tailored to the application at hand. For example, active replication is well suited for critical applications where unresponsiveness due to a background recovery process is not acceptable. However, for other classes of applications without such tight constraints, the use of passive replication, based on checkpoints and logging, is a better choice as it can save a significant amount of resources compared to active replication. In this paper, we present StreamMine3G, a fault tolerant and elastic ESP system which employs several fault tolerance schemes, such as passive and active replication as well as intermediate alternatives such as active and passive standby. In order to free the user from the burden of choosing the correct scheme for the application at hand, StreamMine3G is equipped with a fault-tolerance controller that transitions between the employed schemes during runtime in response to the evolution of the given workload and the user's provided constraints (recovery time and semantics, i.e., gap or precise). Our evaluation shows that the overall resource footprint for fault tolerance can be considerably reduced using our adaptive approach without consequences to the recovery time. André Martin, Tiaraju Smaneoto, Tobias Dietze, Andrey Brito, Christof Fetzer |
DSN | 5 |
| 2015 | VeCycle: Recycling VM Checkpoints for Faster MigrationsabstractVirtual machine migration is a useful and widely used workload management technique. However, the overhead of moving gigabytes of data across machines, racks, or even data centers limits its applicability. According to a recent study by IBM [7], the number of distinct servers visited by a migrating VM is small; often just two. By storing a checkpoint on each server, a subsequent incoming migration of the same VM must transfer less data over the network. Our analysis shows that for short migration intervals of 2 hours on average 50% to 70% of the checkpoint can be reused. For longer migration intervals of up to 24 hours still between 20% to 50% can be reused. In addition, we compared different methods to reduce the migration traffic. We find that content-based redundancy elimination consistently achieves better results than relying on dirty page tracking alone. Sometimes the difference is only a few percent, but can reach up to 50% and more. Our empirical measurements with a QEMU-based prototype confirm the reduction in migration traffic and time. Thomas Knauth, Christof Fetzer |
Middleware | 2 |
| 2015 | Scalable Error Isolation for Distributed Systems
Diogo Behrens, Marco Serafini, Flavio Paiva Junqueira, Sergei Arnautov, Christof Fetzer |
NSDI | 5 |
| 2015 | Quality-Driven Continuous Query Execution over Out-of-Order Data StreamsabstractExecuting continuous queries over out-of-order data streams, where tuples are not ordered according to timestamps, is challenging; because high result accuracy and low result latency are two conflicting performance metrics. Although many applications allow trading exact query results for lower latency, they still expect the produced results to meet a certain quality requirement. However, none of existing disorder handling approaches have considered minimizing the result latency while meeting user-specified requirements on the quality of query results. Yuanzhen Ji, Hongjin Zhou, Zbigniew Jerzak, Anisoara Nica, Gregor Hackenbroich, Christof Fetzer |
SIGMOD Conference | 6 |
| 2015 | ControlFreak: Signature Chaining to Counter Control Flow AttacksabstractMany modern embedded systems use networks to communicate. This increases the attack surface: the adversary does not need to have physical access to the system and can launch remote attacks. By exploiting software bugs, the attacker might be able to change the behavior of a program. Security violations in safety-critical systems are particularly dangerous since they might lead to catastrophic results. Hence, safety-critical software requires additional protection. We present an approach to detect and prevent control flow attacks. Such attacks maliciously modify program's control flow to achieve the desired behavior. We develop ControlFreak, a hardware watchdog to monitor program execution and to prevent illegal control flow transitions. The watchdog employs chained signatures to detect any modification of the instruction stream and any illegal jump in the program even if signatures are maliciously modified. Sergei Arnautov, Christof Fetzer |
SRDS | 2 |
| 2014 | PowerCass: Energy Efficient, Consistent Hashing Based Storage for Micro Clouds Based InfrastructureabstractConsistent hash based storage systems are used in many real world applications for which energy is one of the main cost factors. However, these systems are typically designed and deployed without any mechanisms to save energy at times of low demand. We present an energy conserving implementation of a consistent hashing based key-value store, called PowerCass, based on Apache's Cassandra. In PowerCass, nodes are divided into three groups: active, dormant, and sleepy. Nodes in the active group store cover all the data and running continuously. Dormant nodes are only powered during peak activity time and for replica synchronization. Sleepy nodes are offline almost all the time except for replica synchronization and exceptional peak loads. With this simple and elegant approach we are able to reduce the energy consumption by up to 66% compared to the unmodified key-value store Cassandra. Frezewd Lemma Tena, Thomas Knauth, Christof Fetzer |
IEEE CLOUD | 3 |
| 2014 | Lightweight Automatic Resource Scaling for Multi-tier Web ApplicationsabstractDynamic resource scaling is a key property of cloud computing. Users can acquire or release required capacity for their applications on-the-fly. The most widely used and practical approach for dynamic scaling based on predefined policies (rules). For example, IaaS providers such as RightScale asks application owners to manually set the scaling rules. This task assumes, that the user has an expertise knowledge about the application being run on the cloud. However, it is not always true. In this paper we propose a lightweight adaptive multi-tier scaling framework VscalerLight, which learns scaling policy online. Our framework performs fine-grained vertical resource scaling of multi-tier web application. We present the design and implementation of VscalerLight. We evaluate the framework against widely used RUBiS benchmark. Results show that the application under control of VscalerLight guarantees 95th percentile response time specified in SLA. Lenar Yazdanov, Christof Fetzer |
IEEE CLOUD | 2 |
| 2014 | ParaDIME: Parallel Distributed Infrastructure for Minimization of EnergyabstractDramatic environmental and economic impact of the ever increasing power and energy consumption of modern computing devices in data centers is now a critical challenge. On one hand, designers use technology scaling as one of the methods to face the phenomenon called dark silicon (only segments of a chip function concurrently due to power restrictions). On the other hand, designers use extreme-scale systems such as teradevices to meet the performance needs of their applications which in turn increases the power consumption of the platform. In order to overcome these challenges, we need novel computing paradigms that address energy efficiency. One of the promising solutions is to incorporate parallel distributed methodologies at different abstraction levels. The FP7 project ParaDIME focuses on this objective to provide different distributed methodologies (software-hardware techniques) at different abstraction levels to attack the power-wall problem. In particular, the ParaDIME framework will utilize: circuit and architecture operation below safe voltage limits for drastic energy savings, specialized energy-aware computing accelerators, heterogeneous computing, energy-aware runtime, approximate computing and power-aware message passing. The major outcome of the project will be a processor architecture for a heterogeneous distributed system that utilizes future device characteristics for drastic energy savings. Wherever possible, ParaDIME will adopt multidisciplinary techniques, such as hardware support for message passing, runtime energy optimization utilizing new hardware energy performance counters, use of accelerators for error recovery from sub-safe voltage operation, and approximate computing through annotated code. Furthermore, we will establish and investigate the theoretical limits of energy savings at the device, circuit, architecture, runtime and programming model levels of the computing stack, as well as quantify the actual energy savings achieved by the ParaDIME approach for the complete computing stack with the real environment. Santhosh Kumar Rethinagiri, Oscar Palomar, Anita Sobe, Thomas Knauth, Wojciech M. Barczynski, Gulay Yalcin, Yaroslav Hayduk, Adrián Cristal, Osman S. Unsal, Pascal Felber, Christof Fetzer, Julien Ryckaert, Gina Alioto |
DSD | 11 |
| 2014 | Elastic Scaling of a High-Throughput Content-Based Publish/Subscribe EngineabstractPublish/subscribe (pub/sub) infrastructures running as a service on cloud environments offer simplicity and flexibility for composing distributed applications. Provisioning them appropriately is however challenging. The amount of stored subscriptions and incoming publications varies over time, and the computational cost depends on the nature of the applications and in particular on the filtering operation they require (e.g., content-based vs. topic-based, encrypted vs. non-encrypted filtering). The ability to elastically adapt the amount of resources required to sustain given throughput and delay requirements is key to achieving cost-effectiveness for a pub/sub service running in a cloud environment. In this paper, we present the design and evaluation of an elastic content-based pub/sub system: E-STREAMHUB. Specific contributions of this paper include: (1) a mechanism for dynamic scaling, both out and in, of stateful and stateless pub/sub operators, (2) a local and global elasticity policy enforcer maintaining high system utilization and stable end-to-end latencies, and (3) an evaluation using real-world tick workload from the Frankfurt Stock Exchange and encrypted content-based filtering. Raphaël Barazzutti, Thomas Heinze 0001, André Martin, Emanuel Onica, Pascal Felber, Christof Fetzer, Zbigniew Jerzak, Marcelo Pasin, Etienne Rivière |
ICDCS | 6 |
| 2014 | Combining Error Detection and Transactional Memory for Energy-Efficient Computing below Safe Operation MarginsabstractThe power envelope has become a major issue for the design of computer systems. One way of reducing energy consumption is to downscale the voltage of microprocessors. However, this does not come without costs. By decreasing the voltage, the likelihood of failures increases drastically and without mechanisms for reliability, the systems would not operate any more. For reliability we need (1) error detection and (2) error recovery mechanisms. We provide in this paper a first study investigating the combination of different error detection mechanisms with transactional memory, with the objective to improve energy efficiency. According to our evaluation, using reliability schemes combined with transactional memory for error recovery reduces energy by 54% while providing a reliability level of 100%. Gulay Yalcin, Anita Sobe, Derin Harmanci, Alexey Voronin, Jons-Tobias Wamhoff, Pascal Felber, Osman S. Unsal, Adrián Cristal, Christof Fetzer |
PDP | 9 |
| 2014 | DeTrans: Deterministic and Parallel execution of TransactionsabstractDeterministic execution of a multithreaded application guarantees the same output as long as the application runs with the same input parameters. Determinism helps a programmer to test and debug an application and to provide fault-tolerance in the systems based on replicas. Additionally, Transactional Memory (TM) greatly simplifies development of multithreaded applications where applications use transactions (instead of locks) as a concurrency control mechanism to synchronize accesses to shared memory. However, deterministic systems proposed so far are not TM-aware. They violate the main properties of TM (atomicity, consistency and isolation of transactions), and execute TM applications incorrectly. In this paper, we present DeTrans, a runtime system for deterministic execution of multithreaded TM applications. DeTrans executes TM applications deterministically, it executes nontransactional code serially in round-robin order, and transactional code in parallel. Also, we show how DeTrans works with both eager and lazy software TM. We compare DeTrans with Dthreads, a state-of-the-art deterministic execution system. Unlike Dthreads, DeTrans does not use memory protection hardware nor facilities of the underlying operating system (OS) to execute multithreaded applications deterministically. DeTrans uses properties of software TM to ensure deterministic execution. We evaluate DeTrans using the STAMP benchmark suite and we compare DeTrans and Dthreads performance costs. DeTrans incurs less overhead because threads execute in the same address space without any OS system calls overhead. According to our results, DeTrans is 3.99x, 3.39x, 2.44x faster on average than Dthreads for 2, 4 and 8 threads, respectively. Vesna Smiljkovic, Srdjan Stipic, Christof Fetzer, Osman S. Unsal, Adrián Cristal, Mateo Valero |
SBAC-PAD | 3 |
| 2014 | Chained Signatures for Secure Program ExecutionabstractEvery somewhat complex computer system contains bugs. As it is nearly impossible to fix all bugs in the software stack, the only alternative remains is to make the system secure accepting the fact that software is vulnerable. In this work, a hardware monitor is proposed that checks the correctness of program execution using chained signatures. Sergei Arnautov, Christof Fetzer |
SRDS | 2 |
| 2014 | HardPaxos: Replication Hardened against Hardware ErrorsabstractState Machine Replication (SMR) is a common technique to make services fault-tolerant. Practical SMR systems tolerate process crashes, but no hardware errors such as bit flips. Still, hardware errors can cause major service outages, and their rate is expected to increase in the future. Current approaches either incur a high overhead by hardening large parts of the system in software, or increase the cost of ownership by introducing additional hardware components. This work presents HardPaxos, an atomic broadcast algorithm for SMR that enables services to tolerate hardware errors, while incurring little performance and state overhead. HardPaxos requires no additional hardware and has only a small part of its functionality hardened using a combination of AN-encoding and duplicated execution. Our evaluation shows a throughput overhead of at most 5% for typical payload sizes. Moreover, fault injection experiments show that our hardening decreases the number of undetected errors from 15% to 0.02%. Diogo Behrens, Dmitrii Kuvaiskii, Christof Fetzer |
SRDS | 3 |
| 2014 | Practical Encoded ProcessingabstractEmbedded distributed systems are becoming increasingly complex and interconnected. Some of the challenges in building such systems are safety, i.e., the ability to operate correctly even in the face of arbitrary hardware errors, and security, i.e., the ability to withstand hacker attacks. In this paper, an approach to improve both safety and security for embedded distributed systems with low performance overhead is proposed. Preliminary results indicate that applications hardened using the proposed technique have less than 2x performance overhead and fault coverage of 99.9% (assuming no control flow faults). Dmitrii Kuvaiskii, Christof Fetzer |
SRDS | 2 |
| 2014 | DreamServer: Truly On-Demand Cloud ServicesabstractToday's cloud offerings, while promising flexibility, fail to deliver this flexibility to lower-end services with frequent, minute-long idle times. We present DreamServer, an architecture and combination of technologies to deploy virtualized services just-in-time: virtualized web applications are suspended when idle and resurrected only when the next request arrives. We demonstrate that stateful VM resume can be accomplished in less than one second for select applications. Thomas Knauth, Christof Fetzer |
SYSTOR | 2 |
| 2014 | The TURBO Diaries: Application-controlled Frequency Scaling Explained
Jons-Tobias Wamhoff, Stephan Diestelhorst, Christof Fetzer, Patrick Marlier, Pascal Felber, David Dice |
USENIX ATC | 3 |
| 2013 | Improving Wide-Area Replication Performance through Informed Leader Election and Overlay ConstructionabstractReplication is an important building block to achieve high availability in the presence of failures. Until recently, wide-area replication with strong consistency guarantees was regarded as impractical due to performance constraints. We investigate how informed leader election combined with a network overlay can improve the performance of distributed consensus, which is at the heart of every replicated data store. Leader election and overlay construction are particularly relevant when replicating data at global scale where network links exhibit diverse performance characteristics. We propose to incorporate knowledge about the link quality and network overlay topology into the leader election algorithm. In particular, we show how optimizing only for a quorum, instead of all replicas, we can increase replication throughput or decrease the request latency. Our measurements show a throughput increase of 1.5x when optimizing for throughput of all replicas and a 3x improvement when the throughput is optimized only for a quorum. Syed Kewaan Ejaz, Diogo Behrens, Thomas Knauth, Christof Fetzer |
IEEE CLOUD | 4 |
| 2013 | VScaler: Autonomic Virtual Machine ScalingabstractRecent research results in cloud community found that cloud users increasingly force providers to shift from fixed bundle instance types(e.g. Amazon instances) to flexible bundles and shrinked billing cycles. This means that cloud applications can dynamically provision the used amount of resources in a more fine-grained fashion. This observation calls for approaches which are able to automatically implement fine granular VM resource allocation with respect to user-provided SLAs. In this work we propose VScaler, a framework which implements autonomic resource allocation using a novel approach to reinforcement learning. Lenar Yazdanov, Christof Fetzer |
IEEE CLOUD | 2 |
| 2013 | StreamMine3G OneClick - Deploy and Monitor ESP Applications with a Single ClickabstractIn this paper, we present StreamMine3G One Click, a web based application relieving users from the burden of the complex installation and setup process of Event Stream Processing (ESP) systems such as StreamMine3G in cloud environments. Using StreamMine3G One Click, users just upload their application logic and choose their favorite cloud provider such as Amazon AWS. StreamMine3G One Click performs an automatic installation process ensuring a correct setup of StreamMine3G and dependent services such as Zookeeper in the cloud. Furthermore, StreamMine3GOneClick offers a convenient graphical monitoring interface giving users an instantaneous insight of the performance to identify bottlenecks as well as to easily trace down bugs in their applications. Andrey Brito, André Martin, Christof Fetzer, Isabelly Rocha, Telles Nobrega |
ICPP | 3 |
| 2013 | dsync: Efficient Block-wise Synchronization of Multi-Gigabyte Binary Data
Thomas Knauth, Christof Fetzer |
LISA | 2 |
| 2013 | FastLane: improving performance of software transactional memory for low thread countsabstractSoftware transactional memory (STM) can lead to scalable implementations of concurrent programs, as the relative performance of an application increases with the number of threads that support it. However, the absolute performance is typically impaired by the overheads of transaction management and instrumented accesses to shared memory. This often leads STM-based programs with low thread counts to perform worse than a sequential, non-instrumented version of the same application. Jons-Tobias Wamhoff, Christof Fetzer, Pascal Felber, Etienne Rivière, Gilles Muller |
PPoPP | 2 |
| 2013 | Brief announcement: between all and nothing - versatile aborts in hardware transactional memoryabstractHardware Transactional Memory (HTM) implementations are becoming available in commercial, off-the-shelf components. While generally comparable, some implementations deviate from the strict all-or-nothing property of pure Transactional Memory. We analyse these deviations and find that with small modifications, they can be used to accelerate and simplify both transactional and non-transactional programming constructs. At the heart of our extensions we enable access to the transaction's full register state in the abort handler in an existing HTM without extending the architectural register state. Access to the full register state enables applications in both transactional and non-transactional parallel programming: hybrid transactional memory; transactional escape actions; transactional suspend/resume; and alert-on-update. Stephan Diestelhorst, Martin Nowack, Michael F. Spear, Christof Fetzer |
SPAA | 4 |
| 2013 | Transactional Encoding for Tolerating Transient Hardware Errors
Jons-Tobias Wamhoff, Mario Schwalbe, Rasha Faqeh, Christof Fetzer, Pascal Felber |
SSS | 4 |
| 2012 | Energy-aware scheduling for infrastructure cloudsabstractMore and more data centers are built, consuming ever more kilo watts of energy. Over the years, energy has become a dominant cost factor for data center operators. Utilizing low-power idle modes is an immediate remedy to reduce data center power consumption. We use simulation to quantify the difference in energy consumption caused exclusively by virtual machine schedulers. Besides demonstrating the inefficiency of wide-spread default schedulers, we present our own optimized scheduler. Using a range of realistic simulation scenarios, our customized scheduler OptSched reduces cumulative machine uptime by up to 60.1%. We evaluate the effect of data center composition, run time distribution, virtual machine sizes, and batch requests on cumulative machine uptime. IaaS administrators can use our results to quickly assess possible reductions in machine uptime and, hence, untapped energy saving potential. Thomas Knauth, Christof Fetzer |
CloudCom | 2 |
| 2012 | Fault-tolerant complex event processing using customizable state machine-based operatorsabstractModern Complex Event Processing (CEP) systems often need an high degree of customization in order to implement required application logic. The use of declarative languages, such as CQL, often leads to complicated and hard to maintain application code. In this demo, we show how state machine-based CEP operators help to cope with these problems. State machine-based CEP operators allow for a high flexibility as well as a re-usability of application logic components. A major benefit of the presented solution is its easy integration with existing streaming engines, which we demonstrate using StreamMine, a highly parallel and faulttolerant streaming engine prototype. In this demo we show: (1) how state machine-based operators allow for an easy definition of custom, reusable CEP operators, (2) how resulting state machines can be easily combined with existing faulttolerance techniques within StreamMine and (3) how the resulting CEP applications can be tested in a cost efficient way. Thomas Heinze 0001, Zbigniew Jerzak, André Martin, Lenar Yazdanov, Christof Fetzer |
EDBT | 5 |
| 2012 | Infrastructure Provisioning for Scalable Content-Based Routing: Framework and AnalysisabstractContent-based publish/subscribe is an attractive paradigm for designing large-scale systems, as it decouples producers of information from consumers. This provides extensive flexibility for applications, which can use a modular architecture. Using this architecture, each participant expresses its interest in events by means of filters on the content of those events instead of using pre-established communication channels. However, matching events against filters has a non-negligible processing cost. Scaling the infrastructure with the number of users or events requires appropriate provisioning of resources for each of the operations involved: routing and filtering. In this paper, we propose and describe a generic, modular, and scalable infrastructure for supporting high-performance content-based publish/subscribe. We analyze its properties and show how it dynamically scales in a realistic setting. Our results provide valuable insights into the design and deployment of scalable content-based routing infrastructures. Raphaël Barazzutti, Pascal Felber, Hugues Mercier, Emanuel Onica, Jean-Francois Pineau, Etienne Rivière, Christof Fetzer |
NCA | 7 |
| 2012 | Brief Announcement: Fast Travellers: Infrastructure-Independent Deadlock Resolution in Resource-restricted Distributed Systems
Sebastian Ertel, Christof Fetzer, Michael J. Beckerle |
DISC | 2 |
| 2011 | Scaling Non-elastic Applications Using Virtual MachinesabstractHardware virtualization is a cost effective mean to reduce the number of physical machines (PMs) required to handle computational tasks. Virtualization also guarantees high levels of isolation (performance and security wise) between virtual machines running on the same physical hardware. Besides enabling consolidation of workloads, virtual machine (VM) technology also offers an application independent way of shifting workloads between physical machines. Live migration, i.e., shifting workloads without explicitly stopping the virtual machine, is particularly attractive because of the minimal impact on virtual machine and hence service availability. We explore the use of live migration to scale non-elastic (i.e., static runtime configuration) applications dynamically. Virtual machines thus provide an application agnostic way to dynamic scalability, and open new venues for minimizing the physical resource usage in a data center. We will show that virtualization technology in connection with the live migration capabilities of modern hyper visors can be used to scale non-elastic application in a generic way. Some problems still present in current virtualization techniques with respect to live migration will also be highlighted. Thomas Knauth, Christof Fetzer |
IEEE CLOUD | 2 |
| 2011 | Scalable and Low-Latency Data Processing with Stream MapReduceabstractWe present StreamMapReduce, a data processing approach that combines ideas from the popular MapReduce paradigm and recent developments in Event Stream Processing. We adopted the simple and scalable programming model of MapReduce and added continuous, low-latency data processing capabilities previously found only in Event Stream Processing systems. This combination leads to a system that is efficient and scalable, but at the same time, simple from the user's point of view. For latency-critical applications, our system allows a hundred-fold improvement in response time. Notwithstanding, when throughput is considered, our system offers a ten-fold per node throughput increase in comparison to Hadoop. As a result, we show that our approach addresses classes of applications that are not supported by any other existing system and that the MapReduce paradigm is indeed suitable for scalable processing of real-time data streams. Andrey Brito, André Martin, Thomas Knauth, Stephan Creutz, Diogo Becker de Brum, Stefan Weigert, Christof Fetzer |
CloudCom | 7 |
| 2011 | Boundless memory allocations for memory safety and high availabilityabstractSpatial memory errors (like buffer overflows) are still a major threat for applications written in C. Most recent work focuses on memory safety - when a memory error is detected at runtime, the application is aborted. Our goal is not only to increase the memory safety of applications but also to increase the application's availability. Therefore, we need to tolerate spatial memory errors at runtime. We have implemented a compiler extension, Boundless, that automatically adds the tolerance feature to C applications at compile time. We show that this can increase the availability of applications. Our measurements also indicate that Boundless has a lower performance overhead than SoftBound, a state-of-the-art approach to detect spatial memory errors. Our performance gains result from a novel way to represent pointers. Nevertheless, Boundless is compatible with existing C code. Additionally, Boundless provides a trade-off to reduce the runtime overhead even further: We introduce vulnerability specific patching for spatial memory errors to tolerate only known vulnerabilities. Vulnerability specific patching has an even lower runtime overhead than full tolerance. Marc Brünink, Martin Süßkraut, Christof Fetzer |
DSN | 3 |
| 2011 | Aaron: An adaptable execution environmentabstractSoftware bugs and hardware errors are the largest contributors to downtime, and can be permanent (e.g. deterministic memory violations, broken memory modules) or transient (e.g. race conditions, bitflips). Although a large variety of dependability mechanisms exist, only few are used in practice. The existing techniques do not prevail for several reasons: (1) the introduced performance overhead is often not negligible, (2) the gained coverage is not sufficient, and (3) users cannot control and adapt the mechanism. Aaron tackles these challenges by detecting hardware and software errors using automatically diversified software components. It uses these software variants only if CPU spare cycles are present in the system. In this way, Aaron increases fault coverage without incurring a perceivable performance penalty. Our evaluation shows that Aaron provides the same throughput as an execution of the original application while checking a large percentage of requests - whenever load permits. Marc Brünink, André Schmitt, Thomas Knauth, Martin Süßkraut, Ute Schiffel, Stephan Creutz, Christof Fetzer |
DSN | 7 |
| 2011 | Low-Overhead Fault Tolerance for High-Throughput Data Processing SystemsabstractThe MapReduce programming paradigm proved to be a useful approach for building highly scalable data processing systems. One important reason for its success is simplicity, including the fault tolerance mechanisms. However, this simplicity comes at a price: efficiency. MapReduce's fault tolerance scheme stores too much intermediate information on disk. This inefficiency negatively affects job completion time. Furthermore, this inefficiency in particular forbids the application of MapReduce in near real-time scenarios where jobs need to produce results quickly. In this paper, we discuss an alternative fault tolerance scheme that is inspired by virtual synchrony. The key feature of our approach is a low-overhead deterministic execution. Deterministic execution reduces the amount of persistently stored information. In addition, because persisting intermediate results are no longer required for fault tolerance, we use more efficient communication techniques that considerably improve job completion time and throughput. Our contribution is twofold: (i) we enable the use of MapReduce for jobs ranging from seconds to a few tens of seconds, satisfying these deadlines even in the case of failures, (ii) we considerably reduce the fault tolerance overhead and as such the overhead of MapReduce in general. Our modifications are transparent to the application. André Martin, Thomas Knauth, Stephan Creutz, Diogo Becker de Brum, Stefan Weigert, Christof Fetzer, Andrey Brito |
ICDCS | 6 |
| 2011 | Community-based Analysis of Netflow for Early Detection of Security Incidents
Stefan Weigert, Matti A. Hiltunen, Christof Fetzer |
LISA | 3 |
| 2011 | Optimizing hybrid transactional memory: the importance of nonspeculative operationsabstractTransactional memory (TM) is a speculative shared-memory synchronization mechanism used to speed up concurrent programs. Most current TM implementations are software-based (STM) and incur noticeable overheads for each transactional memory access. Hardware TM proposals (HTM) address this issue but typically suffer from other restrictions such as limits on the number of data locations that can be accessed in a transaction.In this paper, we present several new hybrid TM algorithms that can execute HTM and STM transactions concurrently and can thus provide good performance over a large spectrum of workloads. The algorithms exploit the ability of some HTMs to have both speculative and nonspeculative (nontransactional) memory accesses within a transaction to decrease the transactions' runtime overhead, abort rates, and hardware capacity requirements. We evaluate implementations of these algorithms based on AMD's Advanced Synchronization Facility, an x86 instruction set extension proposal that has been shown to provide a sound basis for HTM. Torvald Riegel, Patrick Marlier, Martin Nowack, Pascal Felber, Christof Fetzer |
SPAA | 5 |
| 2011 | Active Replication at (Almost) No CostabstractMapReduce has become a popular programming paradigm in the domain of batch processing systems. Its simplicity allows applications to be highly scalable and to be easily deployed on large clusters. More recently, the MapReduce approach has been also applied to Event Stream Processing (ESP) systems. This approach, which we call StreamMapReduce, enabled many novel applications that require both scalability and low latency. Another recent trend is to move distributed applications to public clouds such as Amazon EC2 rather than running and maintaining private data centers. Most cloud providers charge their customers on an hourly basis rather than on CPU cycles consumed. However, many applications, especially those that process online data, need to limit their CPU utilization to conservative levels (often as low as $50\%$) to be able to accommodate natural and sudden load variations without causing unacceptable deterioration in responsiveness. In this paper, we present a new fault tolerance approach based on active replication for StreamMapReduce systems. This approach is cost effective for cloud consumers as well as cloud providers. Cost effectiveness is achieved by fully utilizing the acquired computational resources without performance degradation and by reducing the need for additional nodes dedicated to fault tolerance. André Martin, Christof Fetzer, Andrey Brito |
SRDS | 2 |
| 2011 | Resiliency-Aware Data Management
Matthias Boehm 0001, Wolfgang Lehner, Christof Fetzer |
Proc. VLDB Endow. | 3 |
| 2010 | Prospect: a compiler framework for speculative parallelizationabstractMaking efficient use of modern multi-core and future many-core CPUs is a major challenge. We describe a new compiler-based platform, Prospect, that supports the parallelization of sequential applications. The underlying approach is a generalization of an existing approach to parallelize runtime checks. The basic idea is to generate two variants of the application: (1) a fast variant having bare bone functionality, and (2) a slow variant with extra functionality. The fast variant is executed sequentially. Its execution is divided into epochs. Martin Süßkraut, Thomas Knauth, Stefan Weigert, Ute Schiffel, Martin Meinhold, Christof Fetzer |
CGO | 6 |
| 2010 | Evaluation of AMD's advanced synchronization facility within a complete transactional memory stackabstractAMD's Advanced Synchronization Facility (ASF) is an x86 instruction set extension proposal intended to simplify and speed up the synchronization of concurrent programs. In this paper, we report our experiences using ASF for implementing transactional memory. We have extended a C/C++ compiler to support language-level transactions and generate code that takes advantage of ASF. We use a software fallback mechanism for transactions that cannot be committed within ASF (e.g., because of hardware capacity limitations). Our evaluation uses a cycle-accurate x86 simulator that we have extended with ASF support. Building a complete ASF-based software stack allows us to evaluate the performance gains that a user-level program can obtain from ASF. Our measurements on a wide range of benchmarks indicate that the overheads traditionally associated with software transactional memories can be significantly reduced with the help of ASF. David Christie, Jae-Woong Chung, Stephan Diestelhorst, Michael Hohmuth, Martin Pohlack, Christof Fetzer, Martin Nowack, Torvald Riegel, Pascal Felber, Patrick Marlier, Etienne Rivière |
EuroSys | 6 |
| 2010 | ANB- and ANBDmem-Encoding: Detecting Hardware Errors in Software
Ute Schiffel, André Schmitt, Martin Süßkraut, Christof Fetzer |
SAFECOMP | 4 |
| 2010 | RobuSTM: A Robust Software Transactional Memory
Jons-Tobias Wamhoff, Torvald Riegel, Christof Fetzer, Pascal Felber |
SSS | 3 |
| 2010 | Brief Announcement: Hybrid Time-Based Transactional Memory
Pascal Felber, Christof Fetzer, Patrick Marlier, Martin Nowack, Torvald Riegel |
DISC | 2 |
| 2010 | Extensible transactional memory testbed
Derin Harmanci, Vincent Gramoli, Pascal Felber, Christof Fetzer |
J. Parallel Distributed Comput. | 4 |
| 2010 | Time-Based Software Transactional MemoryabstractSoftware transactional memory (STM) is a concurrency control mechanism that is widely considered to be easier to use by programmers than other mechanisms such as locking. The first generations of STMs have either relied on visible read designs, which simplify conflict detection while pessimistically ensuring a consistent view of shared data to the application, or optimistic invisible read designs that are significantly more efficient but require incremental validation to preserve consistency, at a cost that increases quadratically with the number of objects read in a transaction. Most of the recent designs now use a “time-based” (or “time stamp-based”) approach to still benefit from the performance advantage of invisible reads without incurring the quadratic overhead of incremental validation. In this paper, we give an overview of the time-based STM approach and discuss its benefits and limitations. We formally introduce the first time-based STM algorithm, the Lazy Snapshot Algorithm (LSA). We study its semantics and the impact of its design parameters, notably multiversioning and dynamic snapshot extension. We compare it against other classical designs and we demonstrate that its performance is highly competitive, both for obstruction-free and lock-based STM designs. Pascal Felber, Christof Fetzer, Patrick Marlier, Torvald Riegel |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2009 | Fifth Workshop on Hot Topics in System Dependability (HotDep 2009)abstractThe FifthWorkshop on Hot Topics in System Dependability (HotDep'09) brings forth cutting-edge research ideas in fault tolerance, reliability and systems. This year's edition of the workshop will feature a total of 10 presentations of original research on wide array of topics, including cloud computing, storage, program analysis, operating systems, replication protocols, or failure prediction. Christof Fetzer, Rodrigo Rodrigues 0001 |
DSN | 1 |
| 2009 | Minimizing Latency in Fault-Tolerant Distributed Stream Processing SystemsabstractEvent stream processing (ESP) applications target the real-time processing of huge amounts of data. Events traverse a graph of stream processing operators where the information of interest is extracted. As these applications gain popularity, the requirements for scalability, availability, and dependability increase. In terms of dependability and availability, many applications require a precise recovery, i.e., a guarantee that the outputs during and after a recovery would be the same as if the failure that triggered recovery had never occurred. Existing solutions for precise recovery induce prohibitive latency costs, either by requiring continuous checkpoint or logging (in a passive replication approach) or perfect synchronization between replicas executing the same operations (in an active replication approach). We introduce a novel technique to guarantee precise recovery for ESP applications while minimizing the latency costs as compared to traditional approaches. The technique minimizes latencies via speculative execution in a distributed system. In terms of scalability, the key component of our approach is a modified software transactional memory that provides not only the speculation capabilities but also optimistic parallelization for costly operations. Andrey Brito, Christof Fetzer, Pascal Felber |
ICDCS | 2 |
| 2009 | TM-Stream: An STM framework for distributed event stream processingabstractWe extend DSTM2 with a combination of two techniques: First, we applied speculative dependencies between transactions, as first introduced in. Specifically, transactions may read data of earlier transactions that have completed their execution, but are not yet committed. This is the case, for instance, when transactions have to commit in a certain order and must wait for the completion of earlier transactions to detect possible conflicts (e.g., in stream processing systems). Second, we expand speculation to distributed settings, by allowing not yet committed transactions to trigger execution of other speculative transactions on a remote machine. We use a simple notification mechanism to commit or abort remote speculative transactions once the outcome of all the transactions they depend on is known. In this paper we describe our extensions to the DSTM2 framework to enable distributed speculation and evaluate their performance on a simple distributed application. Heiko Sturzrehm, Pascal Felber, Christof Fetzer |
IPDPS | 3 |
| 2009 | Towards Improved Overlay Simulation Using Realistic TopologiesabstractSimulation of distributed applications and overlay networks is challenging. Often the results generated in simulation do not match experimental results. Distributed testbeds like Planet-Lab help to bridge the gap, but they do not offer enough nodes to do an Internet scale evaluation. In this paper we use a tool called TopDNS for generating realistic topologies for simulations, using the Planet-Lab to collect measurement data. We show, that simulation results may differ significantly from earlier results using synthesized topologies. We provide a data analysis to explain the observed results and to provide a better understanding of latency between hosts in certain DNS name spaces. Gert Pfeifer, Ryan C. Spring, Christof Fetzer |
NCA | 3 |
| 2009 | AN-Encoding Compiler: Building Safety-Critical Systems with Commodity Hardware
Christof Fetzer, Ute Schiffel, Martin Süßkraut |
SAFECOMP | 1 |
| 2009 | Multithreading-Enabled Active Replication for Event Stream Processing OperatorsabstractEvent stream processing (ESP) systems are very popular in monitoring applications. Algorithmic trading, network monitoring and sensor networks are good examples of applications that rely upon ESP systems. As these systems become larger and more widely deployed, they have to answer increasingly stronger requirements that are often difficult to satisfy. Fault-tolerance is a good example of such a non-trivial requirement. Making ESP operators fault-tolerant can add considerable performance overhead to the application. In this paper, we focus on active replication as an approach to provide fault-tolerance to ESP operators. More precisely, we address the performance costs of active replication for operators in distributed ESP applications.We use a speculation mechanism based on software transactional memory (STM) to achieve the following goals: (i) enable replicas to make progress using optimistic delivery; (ii) enable early forwarding of speculative computation results; (iii) enable active replication of multi-threaded operators using transactional executions. Experimental evaluation shows that, using this combination of mechanisms, one can implement highly efficient fault-tolerant ESP operators. Andrey Brito, Christof Fetzer, Pascal Felber |
SRDS | 2 |
| 2009 | Speculation for Parallelizing Runtime Checks
Martin Süßkraut, Stefan Weigert, Ute Schiffel, Thomas Knauth, Martin Nowack, Diogo Becker de Brum, Christof Fetzer |
SSS | 7 |
| 2008 | Dependable Embedded Systems Special Day Panel: Issues and Challenges in Dependable Embedded SystemsabstractThe paper presents a panel discussion on the issues and challenges in dependable embedded system from both the academic and industrial perspectives. The panelists are Jacob Abraham from the University of Texas at Austin-USA, Stefan Poledna from TTTech-Austria, Avi Mendelson from Intel-Israel, and Subhasish Mitra from Stanford University-USA. Neeraj Suri, Christof Fetzer, Jacob A. Abraham, Stefan Poledna, Avi Mendelson, Subhasish Mitra |
DATE | 2 |
| 2008 | Enhanced server fault-tolerance for improved user experienceabstractInteractive applications such as email, calendar, and maps are migrating from local desktop machines to data centers due to the many advantages offered by such a computing environment. Furthermore, this trend is creating a marked increase in the deployment of servers at data centers. To ride the price/performance curves for CPU, memory and other hardware, inexpensive commodity machines are the most cost effective choices for a data center. However, due to low availability numbers of these machines, the probability of server failures is relatively high. Server failures can in turn cause service outages, degrade user experience and eventually result in lost revenue for businesses. We propose a TCP splice-based Web server architecture that seamlessly tolerates both Web proxy and backend server failures. The client TCP connection and sessions are preserved, and failover to alternate servers in case of server failures is fast and client transparent. The architecture provides support for both deterministic and non-deterministic server applications. A prototype of this architecture has been implemented in Linux, and the paper presents detailed performance results for a PHP-based Webmail application deployed over this architecture. Manish Marwah, Shivakant Mishra, Christof Fetzer |
DSN | 3 |
| 2008 | Switchblade: enforcing dynamic personalized system call modelsabstractSystem call interposition is a common approach to restrict the power of applications and to detect code injections. It enforces a model that describes what system calls and/or what sequences thereof are permitted. However, there exist various issues like concurrency vulnerabilities and incomplete models that restrict the power of system call interposition approaches. We present a new system, SwitchBlade, that uses randomized and personalized fine-grained system call models to increase the probability of detecting code injections. However, using a fine-grain system call model, we cannot exclude the possibility that the model is violated during normal program executions. To cope with false positives, SwitchBlade uses on-demand taint analysis to update a system call model during runtime. Christof Fetzer, Martin Süßkraut |
EuroSys | 1 |
| 2008 | Dynamic performance tuning of word-based software transactional memoryabstractThe current generation of software transactional memories has the advantage of being simple and efficient. Nevertheless, there are several parameters that affect the performance of a transactional memory, for example the locality of the application and the cache line size of the processor. In this paper, we investigate dynamic tuning mechanisms on a new time-based software transactional memory implementation. We study in extensive measurements the performance of our implementation and exhibit the benefits of dynamic tuning. We compare our results with TL2, which is currently one of the fastest word-based software transactional memories. Pascal Felber, Christof Fetzer, Torvald Riegel |
PPoPP | 2 |
| 2008 | Automatic data partitioning in software transactional memoriesabstractWe investigate to which extent data partitioning can help improve the performance of software transactional memory (STM). Our main idea is that the access patterns of the various data structures of an application might be sufficiently different so that it would be beneficial to tune the behavior of the STM for individual data partitions. We evaluate our approach using standard transactional memory benchmarks. We show that these applications contain partitions with different characteristics and, despite the runtime overhead introduced by partition tracking and dynamic tuning, that partitioning provides significant performance improvements. Torvald Riegel, Christof Fetzer, Pascal Felber |
SPAA | 2 |
| 2008 | Adaptive Internal Clock SynchronizationabstractExisting clock synchronization algorithms assume a bounded clock reading error. This, in turn, results in an inflexible design that typically requires node crashes whenever the given bound might be violated. We propose a novel, adaptive internal clock synchronization algorithm which allows to compute the deviation between the clocks during runtime. The computed deviation can be propagated to the application layer to allow it to adapt its behavior according to the current clock deviation. The contributions of this paper are: (1) a new specification of a relaxed clock synchronization problem, and (2) a new clock synchronization algorithm with a novel approach to dealing with crash failures. Zbigniew Jerzak, Robert Fach, Christof Fetzer |
SRDS | 3 |
| 2007 | DSN 2007 WorkshopsabstractWorkshops at DSN provide a forum for a group of participants (typically 20 to 50 in size) to interact and exchange opinions on topics related to any of the many facets of dependable systems and networks. We welcome participation by professionals from a range of diverse backgrounds who can contribute to advancing the technology and understanding of the workshop subject. Christof Fetzer |
DSN | 1 |
| 2007 | Robustness and Security Hardening of COTS Software LibrariesabstractCOTS components, like software libraries, can be used to reduce the development effort. Unfortunately, many COTS components have been developed without a focus on robust- ness and security. We propose a novel approach to harden software libraries to improve their robustness and security. Our approach is automated, general and extensible and consists of the following stages. First, we use a static analysis to prepare and guide the following fault injection. In the dynamic analysis stage, fault injection experiments execute the library functions with both usual and extreme input values. The experiments are used to derive and verify one protection hypothesis per function (for instance, function foo fails if argument 1 is a NULL pointer). In the hardening stage, a protection wrapper is generated from these hypothesis to reject unrobust input values of library functions. We evaluate our approach by hardening a library used by Apache (a web server). Martin Süßkraut, Christof Fetzer |
DSN | 2 |
| 2007 | Topic 8 Distributed Systems and Algorithms
Luís E. T. Rodrigues, Achour Mostéfaoui, Christof Fetzer, Philippas Tsigas |
Euro-Par | 3 |
| 2007 | Fail-Aware Publish/SubscribeabstractIn this paper we present a wide area distributed system using a content-based publish/subscribe communication middleware which can deterministically detect and report failures with respect to timely message delivery and message omission. Our approach does not require external clock synchronization nor does it impose any constraints on the publish/subscribe middleware. We show that our system performs better and is safer than when using NTP for external clock synchronization. We provide a proof of concept implementation and present results of experiments carried out in the PlanetLab environment. Zbigniew Jerzak, Robert Fach, Christof Fetzer |
NCA | 3 |
| 2007 | Exploiting Host Name Locality for Reduced Stretch P2P RoutingabstractStructured P2P networks are a promising alternative for engineering new distributed services and for replacing existing distributed services like DNS. Providing competitive performance with traditional distributed services is however very difficult because existing services like DNS are highly tuned using a combination of caching and localized communication. Typically, P2P systems use randomized host IDs which destroys any locality that might have been inherent in the IP addresses or the names of the hosts. In this way, P2P communication can result in a high stretch. We propose a locality preserving structured P2P system that supports efficient local communication and low stretch. While this system was optimized for resolving domain names, it will also provide a low stretch to other applications and it can be combined with existing replication schemes to optimize the response times even further. Gert Pfeifer, Christof Fetzer, Thomas Hohnstein |
NCA | 2 |
| 2007 | From causal to z-linearizable transactional memoryabstractThe current generation of time-based transactional memories (TMs) has the advantage of being simple and efficient, and providing strong linearizability semantics. Linearizability matches well the goal of TM to simplify the design and implementation of concurrent applications. However, long transactions can have a much lower likelihood of committing than smaller transactions because of the strict ordering constraints imposed by linearizability. We investigate the use of weaker semantics for TM and introduce a new consistency criterion that we call z-linearizability. By combining properties of linearizability and serializability, z-linearizability provides a good trade-off between strong semantics and good practical performance even for long transactions. Torvald Riegel, Christof Fetzer, Heiko Sturzrehm, Pascal Felber |
PODC | 2 |
| 2007 | Software Encoded Processing: Building Dependable Systems with Commodity Hardware
Ute Schiffel, Christof Fetzer |
SAFECOMP | 2 |
| 2007 | Time-based transactional memory with scalable time basesabstractTime-based transactional memories use time to reason about the consistency of data accessed by transactions and about the order in which transactions commit. They avoid the large read overhead of transactional memories that always check consistency when a new object is accessed, while still guaranteeing consistency at all times--in contrast to transactional memories that only check consistency on transaction commit. Current implementations of time-based transactional memories use a single global clock that is incremented by the commit operation for each update transaction that commits. In large systems with frequent commits, the contention on this global counter can thus become a major bottleneck. We present a scalable replacement for this global counter and describe how the Lazy Snapshot Algorithm (LSA), which forms the basis for our LSA-STM time-based software transactional memory, has to be changed to support these new time bases. In particular, we show how the global counter can be replaced (1) by an external or physical clock that can be accessed efficiently, and (2) by multiple synchronized physical clocks. Torvald Riegel, Christof Fetzer, Pascal Felber |
SPAA | 2 |
| 2006 | Student ForumabstractThe Student Forum will provide an opportunity for students working in the area of dependable computing to present and discuss their research objectives, approaches and preliminary results. The Forum is centered around a conference track during which the selected "student research papers" are presented. Christof Fetzer |
DSN | 1 |
| 2006 | Leader Election in the Timed Finite Average Response Time ModelabstractLeader election is one of the fundamental problems in distributed systems. A leader is a correct process that can be used to coordinate the work of a set of processes. An algorithm has to implement two properties to solve the leader election problem: (1) safety, (2) liveness. In this work we show that the stabilization property is not necessary for the leader election problem. We do this by examine the ability to solve leader election in the FAR model. The FAR model does neither assume the existence of an upper bound on the communication or computation delays nor that the system stabilizes. Instead it assumes that the system is in a certain balance: computation is not infinitely fast, the communication subsystem has rudimentary congestion control and the average response time between two correct processes is finite. Our contribution is twofold: (1) we show that leader election is not solvable in the pure FAR model and (2) that it becomes solvable with local clocks with a bounded drift rate Christof Fetzer, Martin Süßkraut |
PRDC | 1 |
| 2006 | Fault-tolerant and scalable TCP splice and web server architectureabstractThis paper describes three enhancements to the TCP splicing mechanism: (1) Enable a TCP connection to be simultaneously spliced through multiple machines for higher scalability; (2) Make a spliced connection fault-tolerant to proxy failures; and (3) Provide flexibility of splitting a TCP splice between a proxy and a backend server for further increasing the scalability of a Web server system. A Web server architecture based on this enhanced TCP splicing is proposed. This architecture provides a highly scalable, seamless service to the users with minimal disruption during server failures. In addition to the traditional Web services in which users download Web pages, multimedia files and other types of data from a Web server, the proposed architecture supports newly emerging Web services that are highly interactive, and involve relatively longer, stateful client-server sessions. A prototype of this architecture has been implemented as a Linux 2.6 kernel module, and the paper presents important performance results measured from this implementation Manish Marwah, Shivakant Mishra, Christof Fetzer |
SRDS | 3 |
| 2006 | A Lazy Snapshot Algorithm with Eager Validation
Torvald Riegel, Pascal Felber, Christof Fetzer |
DISC | 3 |
| 2005 | A System Demonstration of ST-TCPabstractST-TCP (server fault-tolerant TCP) is an extension of TCP to tolerate TCP server failures. Server fault tolerance is provided by using an active-backup server that keeps track of the state of a TCP connection. The backup server takes over the TCP connection if the primary server fails. This take-over is fast, seamless, and completely transparent to the client. This paper provides a system demonstration of a new ST-TCP prototype. The new prototype incorporates a performance-enhanced architecture and addresses application failure scenarios. Five experiments using this prototype are proposed to demonstrate the following useful features: (1) client-transparent, seamless failover; (2) insignificant performance overhead of ST-TCP during failure-free periods; and (3) an ability to tolerate all single crash failures at the hardware and operating system levels, and, most crash failures at the application level. Manish Marwah, Shivakant Mishra, Christof Fetzer |
DSN | 3 |
| 2005 | On the Possibility of Consensus in Asynchronous Systems with Finite Average Response TimesabstractIt has long been known that the consensus problem cannot be solved deterministically in completely asynchronous distributed systems, i.e., systems (1) without assumptions on communication delays and relative speed of processes and (2) without access to real-time clocks. In this paper, we define a new asynchronous system model. Instead of assuming reliable channels with finite transmission delays, stubborn channels with a finite average response time was assumed (if neither the sender nor the receiver crashes), and it is assumed that there exists some unknown physical bound on how fast an integer can be incremented. Note that there is no limit on how slow a program can be executed or how fast other statements can be executed. Also, there exists no upper or lower bound on the transmission delay of messages or the relative speed of processes. The are no additional assumptions about clocks, failure detectors, etc. that would aid in solving consensus either. It is shown that consensus can nevertheless be solved deterministically in this asynchronous system model Christof Fetzer, Ulrich Schmid 0001, Martin Süßkraut |
ICDCS | 1 |
| 2004 | Brief announcement: on the possibility of consensus in asynchronous systems with finite average response timesabstractNo abstract available. Christof Fetzer, Ulrich Schmid 0001 |
PODC | 1 |
| 2004 | Automatic Detection and Masking of Nonatomic Exception HandlingabstractThe development of robust software is a difficult undertaking and is becoming increasingly more important as applications grow larger and more complex. Although modern programming languages such as C++ and Java provide sophisticated exception handling mechanisms to detect and correct runtime error conditions, exception handling code must still be programmed with care to preserve application consistency. In particular, exception handling is only effective if the premature termination of a method due to an exception does not leave an object in an inconsistent state. We address this issue by introducing the notion of failure atomicity in the context of exceptions. We propose practical techniques to automatically detect and mask the nonatomic exception handling situations encountered during program execution. These techniques can be applied to applications written in various programming languages that support exceptions. We perform experimental evaluation on both C++ and Java applications to demonstrate the effectiveness of our techniques and measure the overhead that they introduce. Christof Fetzer, Pascal Felber, Karin Högstedt |
IEEE Trans. Software Eng. | 1 |
| 2003 | Automatic Detection and Masking of Non-Atomic Exception Handling
Christof Fetzer, Karin Högstedt, Pascal Felber |
DSN | 1 |
| 2003 | HEALERS: A Toolkit for Enhancing the Robustness and Security of Existing ApplicationsabstractHEALERS is a practical, high-performance toolkit that can enhance the robustness and security of existing applications. For any shared library, it can find all functions defined in that library and automatically derives properties for those functions. Through automated faultinjection experiments, it can detect arguments that cause the library to crash and derive safe argument types for each function. The toolkit can prevent heap and stack buffer overflows that are a common cause of security breaches. The nice feature of the HEALERS approach is that it can protect existing applications without access to the source code. Christof Fetzer |
DSN | 1 |
| 2003 | TCP Server Fault Tolerance Using Connection Migration to a Backup ServerabstractThis paper describes the design, implementation, and performance evaluation of ST-TCP (Server fault-Tolerant TCP), which is an extension of TCP to tolerate TCP server failures. This is done by using an active backup server that keeps track of the state of the TCP connection and takes over the TCP connection whenever the primary fails. This migration of the TCP connection to the backup is completely transparent to the client. Because no changes are required on the client machine, any TCP client can access a ST-TCP server. The performance overhead of ST-TCP over standard TCP is minimal, and during normal operation its behavior is the same as that of a regular TCP. In addition, ST-TCP provides a fast and seamless failover whenever the primary server fails. This is verified by a prototype implementation of ST-TCP in the Linux operating system, and experiments with a number of simulated applications which have different communication characteristics. Manish Marwah, Shivakant Mishra, Christof Fetzer |
DSN | 3 |
| 2003 | Elastic Vector TimeabstractIn recent years there has been an increasing demand to build "soft" real-time applications on top of asynchronous distributed systems. Designing and implementing such applications is a non-trivial task and application designers are often faced with the need to circumvent impossibility results. In this paper we discuss how to ensure that actions are executed in the correct order even in the face of failures. We propose a novel time base and a new synchronization mechanism for the design of distributed "soft" real-time applications. We demonstrate (1) how this time base can be used to enforce an externally consistent ordering, and (2) how it permits to circumvent impossibility results by sketching how to solve the leader election and perfect failure detection problem. Christof Fetzer, Michel Raynal |
ICDCS | 1 |
| 2003 | Randomized Asynchronous Consensus with Imperfect CommunicationsabstractWe introduce a novel hybrid failure model, which facilitates an accurate and detailed analysis of round-based synchronous, partially synchronous and asynchronous distributed algorithms under both process and link failures. Granting every process in the system up to f/sub /spl lscr// send and receive link failures (with f/sub /spl lscr///sup a/ arbitrary faulty ones among those) in every round, without being considered faulty, we show that the well-known randomized Byzantine agreement algorithm of (Srikanth & Toueg 1987) needs just n /spl ges/ 4f/sub /spl lscr// + 2ff/sub /spl lscr///sup a/+ 3f/sub a/ + 1 processes for coping with f/sub a/ Byzantine faulty processes. The probability of disagreement after R iterations is only 2/sup -R/, which is the same as in the FLP model and thus much smaller than the lower bound 0(1/R) known for synchronous systems with lossy links. Moreover, we show that 2-stubborn links are sufficient for this algorithm. Hence, contrasting widespread belief, a perfect communications subsystem is not required for efficiently solving randomized Byzantine agreement. Ulrich Schmid 0001, Christof Fetzer |
SRDS | 2 |
| 2003 | Fail-Awareness: An Approach to Construct Fail-Safe Systems
Christof Fetzer, Flaviu Cristian |
Real Time Syst. | 1 |
| 2003 | Perfect Failure Detection in Timed Asynchronous SystemsabstractPerfect failure detectors can correctly decide whether a computer is crashed. However, it is impossible to implement a perfect failure detector in purely asynchronous systems. We show how to enforce perfect failure detection in timed asynchronous systems with hardware watchdogs. The two main system model assumptions are: 1) each computer can measure time intervals with a known maximum error and 2) each computer has a watchdog that crashes the computer unless the watchdog is periodically updated. We have implemented a system that satisfies both assumptions using a combination of off-the-shelf software and hardware. To implement a perfect failure detector for process crash failures, we show that, in some systems, a hardware watchdog is actually not necessary. Christof Fetzer |
IEEE Trans. Computers | 1 |
| 2002 | An Automated Approach to Increasing the Robustness of C LibrariesabstractAs our reliance on computers increases, so does the need for robust software. Previous studies have shown that many C libraries exhibit robustness problems due to exceptional inputs. This paper describes the HEALERS system that uses an automated approach to increasing the robustness of C libraries without source code access. The system extracts the C type information for a shared library using header files and manual pages. Then it generates for each global function a fault-injector to determine a "robust " argument type for each argument. Based on this information and optionally, some manual editing, the system generates a robustness wrapper that performs careful argument checking before invoking C library functions. A robustness evaluation using Ballista tests has shown that our wrapper can prevent crash, hang, and abort failures. Moreover the wrapper generation process is highly automated and can easily adapt to new library releases. Christof Fetzer |
DSN | 1 |
| 2002 | A Flexible Generator Architecture for Improving Software DependabilityabstractImproving the dependability of computer systems is increasingly important as more and more of our lives depend on the availability of such systems. Wrapping dynamic link libraries is an effective approach for improving the reliability and security of computer software without source code access. We describe a flexible framework to generate a rich set of software wrappers for shared libraries. We describe the architecture of the wrapper generator, the problems of how to generate wrappers efficiently, and our solutions to these problems. Based on a set of properties declared for a function, the generator can create a variety of wrappers to suit the diverse requirements of application programs. Performance measurements indicate that the overhead of the generated wrappers is small. Christof Fetzer |
ISSRE | 1 |
| 2002 | The Timewheel Group Communication SystemabstractDescribes the timewheel group communication system, which has been designed for a timed asynchronous distributed system model. All protocols in the timewheel group communication system have been designed to be fail-aware in the sense that a process can detect, at any point in time, whether any of its properties is violated. Although these protocols have been designed to operate in an asynchronous distributed computing environment, they provide timeliness properties. The timewheel group communication system provides nine group communication semantics that a user can dynamically choose from while broadcasting an update. This system provides high throughput, fast delivery and stability times, uses a small number of messages per update broadcast, and evenly distributes the processing load among group members. Shivakant Mishra, Christof Fetzer, Flaviu Cristian |
IEEE Trans. Computers | 2 |
| 2001 | Enforcing Perfect Failure DetectionabstractPerfect failure detectors can correctly decide whether a computer is crashed. However it is impossible to implement a perfect failure detector in purely asynchronous systems. We show how to enforce perfect failure detection in timed distributed systems with hardware watchdogs. The two main system model assumptions are: each computer can measure time intervals with a known maximum error; and each computer has a watchdog that crashes the computer unless the watchdog is periodically updated. We have implemented a system that satisfies both assumptions using a combination of off-the-shelf software and hardware. Christof Fetzer |
ICDCS | 1 |
| 2001 | Tapping TCP StreamsabstractProviding transparent replication of servers has been a major goal in the fault tolerance community. Transparent replication is particularly challenging for highly nondeterministic applications, such as the ones that use multithreading. For such applications, keeping replicas in a consistent state becomes non-trivial. One way to deal with the non-determinism is to use a leader/follower approach. In this paper we describe the design and performance of a TCP tapping mechanism we implemented. This mechanism was designed to improve the efficiency of leader/follower replication. We argue that TCP tapping can address a major efficiency bottleneck of leader/follower replication. Maxim Orgiyan, Christof Fetzer |
NCA | 2 |
| 2001 | Rejuvenation and Failure Detection in Partitionable SystemsabstractCertain gateways (e.g., some cable or DSL modems) are known to have low reliability and low availability. Most failures of these devices can however be "fixed" by rejuvenating the device after a failure has been detected. Such a detection based rejuvenation strategy permits increasing the availability of these gateways. In the considered scenario, rejuvenation is non-trivial since a failure of such a gateway will leave it partitioned away from the network. In particular, network operators that want to rejuvenate these gateways are in a different network partition, and can therefore not initiate a remote rejuvenation. In this paper we propose a failure detection based rejuvenation service and a remote detection service. The rejuvenation service detects and fixes "soft" failures automatically (in one partition), and the detection service detects (in another partition) all rejuvenations exactly once, within a bounded amount of time, even when the gateway is rejuvenated consecutively. The detection service also allows the detection of "hard" failures, and filtering of notifications of soft failures. Christof Fetzer, Karin Högstedt |
PRDC | 1 |
| 2001 | An Adaptive Failure Detection ProtocolabstractThe detection of process failures is a crucial problem system designers have to cope with in order to build fault-tolerant distributed platforms. Unfortunately, it is impossible to distinguish with certainty a crashed process from a very slow process in a purely asynchronous distributed system. This prevents some problems from being solved in such systems. That is why failure detector oracles have been introduced to circumvent these impossibility results. The paper presents a relatively simple protocol that allows a process to "monitor" another process, and consequently to detect its crash. This protocol relies as much as possible on application messages to do this monitoring. Different from previous process crash detection protocols, it uses control messages only when no application message is sent by the monitoring process to the observed process. When the underlying system satisfies the partial synchrony assumption, it actually implements an eventually perfect failure detector (i.e., a failure detector of the class usually denoted OP). Moreover if the average observed transmission delay is finite and the upper layer application terminates within a bounded number of steps for any failure detector in OP after the failure detector becomes "perfect", then, when run with the proposed protocol, it also terminates correctly. These properties make the protocol inexpensive, implementable, and powerful. The paper also describes performance measurements of an implementation of the protocol. Christof Fetzer, Michel Raynal, Frédéric Tronel |
PRDC | 1 |
| 2001 | Detecting Heap Smashing Attacks through Fault Containment WrappersabstractBuffer overflow attacks are a major cause of security breaches in modern operating systems. Not only are overflows of buffers on the stack a security threat, overflows of buffers kept on the heap can be too. A malicious user might be able to hijack the control flow of a root-privileged program if the user can initiate an overflow of a buffer on the heap when this overflow overwrites a function pointer stored on the heap. The paper presents a fault-containment wrapper which provides effective and efficient protection against heap buffer overflows caused by C library functions. The wrapper intercepts every function call to the C library that can write to the heap and performs careful boundary checks before it calls the original function. This method is transparent to existing programs and does not require source code modification or recompilation. Experimental results on Linux machines indicate that the performance overhead is small. Christof Fetzer |
SRDS | 1 |
| 2000 | he Timely Computing Base: Timely Actions in the Presence of Uncertain TimelinessabstractReal-time behavior is specified in compliance with timeliness requirements, which in essence calls for synchronous system models. However systems often rely on unpredictable and unreliable infrastructures, that suggest the use of asynchronous models. Several models have been proposed to address this issue. We propose an architectural construct that takes a generic approach to the problem of programming in the presence of uncertain timeliness. We assume the existence of a component, capable of executing timing functions, which helps applications with varying degrees of synchrony to behave reliably despite the occurrence of timing failures. We call this component the Timely Computing Base, TCB. This paper describes the TCB architecture and model, and discusses the application programming interface for accessing the TCB services. The implementation of the TCB services uses fail-awareness techniques to increase the coverage of TCB properties. Paulo Veríssimo, António Casimiro, Christof Fetzer |
DSN | 3 |
| 2000 | Enforcing synchronous system properties on top of timed systemsabstractA synchronous system model is a simple yet powerful distributed system model that reduces the complexity of the design and implementation of dependable distributed applications. However, a late message arrival or a missed deadline violates the properties of a completely synchronous system. Therefore, an application that depends upon these properties might violate its safety and timeliness properties due to a late message or a missed deadline. In this paper, we propose a family of protocols that enforce the synchronous system properties. These protocols transform performance and omission failures that cannot be masked into crash failures. The protocols are designed to be correct for any number of performance and omission failures: they run on top of timed systems extended by hardware watchdogs. The described approach is targeted towards "nearly synchronous systems", i.e., systems in which the probability of performance and omission failures is low but not negligible. Christof Fetzer |
PRDC | 1 |
| 1999 | The Timed Asynchronous Distributed System ModelabstractWe propose a formal definition for the timed asynchronous distributed system model. We present extensive measurements of actual message and process scheduling delays and hardware clock drifts. These measurements confirm that this model adequately describes current distributed systems such as a network of workstations. We also give an explanation of why practically needed services, such as consensus or leader election, which are not implementable in the time-free model, are implementable in the timed asynchronous system model. Flaviu Cristian, Christof Fetzer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1999 | A Highly Available Local Leader Election ServiceabstractWe define the highly available local leader election problem (G. LeLann, 1977), a generalization of the leader election problem for partitionable systems. We propose a protocol that solves the problem efficiently and give some performance measurements of our implementation. The local leader election service has been proven useful in the design and implementation of several fail-aware services for partitionable systems. Christof Fetzer, Flaviu Cristian |
IEEE Trans. Software Eng. | 1 |
| 1998 | The Message Classification ModelabstractWe propose a new system model for asynchronous distributed systems that we call the message classification model.Motivation for this model is its ability 1) to support a restricted but useful form of "communication by time" by classifying messages as either "slow" or Lifast" but without incorporating neither real-time clocks nor "time-outs", and 2) to describe transient and permanent network partitions.The message classification model allows the definition of different classes of classification schemes.To show that the model is indeed useful, we show how one can solve the consensus and the election problem for a certain class of message classification schemes. IntroductionWe introduce a new model that we call the message classification model (MCM) to describe asynchronous distributed systems.The goal of this model is to achieve a similar generality as the FLP model [ll] (in the sense that almost all distributed computing systems running an appropriate software layer can be described by the MCM) while still allowing to solve interesting problems like the consensus or the election problem.Another goal of this model is to enable the description of permanent and transient network partitions.A distributed system consists of a set of processes that can communicate with each other by exchanging messages.An asynchronous (distributed) system is a distributed system in which it is not possible to place bounds on communication delays, clock drift, or relative speeds of processes.We use the phrase distributed computing system to refer to a set of computers connected by a network.A system model is an abstract description of the properties of a distributed system, e.g. it specifies if processes have access to real-time clocks *This research was supported Christof Fetzer |
PODC | 1 |
| 1997 | A Fail-Awar Membership ServiceabstractWe propose a new protocol that can be used to implement a partitionable membership service for timed asynchronous systems. The protocol is fail-aware in the sense that a process p knows at all times if its approximation of the set of processes in its partition is up-to-date or out-of-date. The protocol minimizes wrong suspicions of processes by giving processes a second chance to stay in the membership before they are removed. Our measurements show that the exclusion of live processes is rare and the crash detection times are good. The protocol guarantees that the memberships of two partitions never overlap. Christof Fetzer, Flaviu Cristian |
SRDS | 1 |
| 1997 | Integrating External and Internal Clock Synchronization
Christof Fetzer, Flaviu Cristian |
Real Time Syst. | 1 |
| 1996 | Fail-Awareness in Timed Asynchronous SystemsabstractWe address the problem of the impossibdity of implementing synchronous fault-tolerant service specifications in asynchronous distributed systems.We introduce a method for weakening a synchronous service specification so that it becomes implementable in "timed" asynchronous systems, that q This research was partially sponsored by a grant from the Air Force Office of Scientific Research Fe fmkioo to meted@d/bcrd copies of cll or pert of W:s rnetericl for perm-mel or clcssroarn use is grcnted withcwt &c provided Urct the copies not IIU& or dktdwtrd fw protit or commerc"ml q dvmtege, the.c~yrigbt rdce, the title of the publkxtion q nd ite dcte q ppecr, q nd aottce u given tbct copyright is by pcrmisrkM of tbe ACM, inc.To copy othcnviee, to republicb, to poA 00 aetvers or to rdetribute to Iietcj requires epccitic pcrmhioo cndhr f-.PODC'%, Pttiladelphis PA, USA O l% ACM &SgT$)l.~~%/OS..$3.50 asynchronous distributed systems in which processes have access to local hardware clocks.Hardware clocks and the notion of "performance failures" are essential for our approach.This work is therefore based on the timed asynchronous system model[11] Christof Fetzer, Flaviu Cristian |
PODC | 1 |
| 1996 | Fail-Aware Failure DetectorsabstractIn existing asynchronous distributed systems it is impossible to implement failure detectors which are perfect, i.e. they only suspect crashed processes and eventually suspect all crashed processes. Some recent research has however proposed that any "reasonable" failure detector for solving the election problem must be perfect. We address this problem by introducing two new classes of fail-aware failure detectors that are (1) implementable in existing asynchronous distributed systems, (2) not necessarily perfect, and (3) can be used to solve the election problem. In particular we show that there exists a fail-aware failure detector that allows to solve the election problem and which is strictly weaker than a perfect failure detector. Christof Fetzer, Flaviu Cristian |
SRDS | 1 |
| 1995 | Fault-Tolerant External Clock SynchronizationabstractWe address the problem of how to integrate fault-tolerant internal and external clock synchronization. We propose a new algorithm which provides both external and internal clock synchronization for as long as no more than F reference time servers out of a total of 2F+1 are faulty. When the number of faulty reference time servers exceeds F, the algorithm degrades to a fault-tolerant internal clock synchronization algorithm. We prove that at least 2F+1 reference time servers are necessary for achieving external clock synchronization when up to F reference time servers can suffer arbitrary failures, thus our algorithm provides maximum fault-tolerance. The algorithm is also optimal in another sense: we show that the maximum deviation between reference time and the clocks of nonreference time servers is minimal. Flaviu Cristian, Christof Fetzer |
ICDCS | 2 |
| 1995 | Lower Bounds for Convergence Function Based Clock SynchronizationabstractArticle Free Access Share on Lower bounds for convergence function based clock synchronization Authors: Christof Fetzer Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CA Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CAView Profile , Flaviu Cristian Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CA Department of Computer Science & Engineering, University of California, San Diego, La Jolla, CAView Profile Authors Info & Claims PODC '95: Proceedings of the fourteenth annual ACM symposium on Principles of distributed computingAugust 1995 Pages 137–143https://doi.org/10.1145/224964.224980Online:20 August 1995Publication History 10citation288DownloadsMetricsTotal Citations10Total Downloads288Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Christof Fetzer, Flaviu Cristian |
PODC | 1 |
| 1994 | Probabilistic Internal Clock SynchronizationabstractWe propose an improved probabilistic method for reading remote clocks in systems subject to unbounded communication delays and use this method to design a fault-tolerant probabilistic internal clock synchronization protocol. This protocol masks clock reading failures and arbitrary failures of processes. Because of probabilistic reading, our protocol achieves better synchronization precisions than those achievable by previously known deterministic algorithms. Another advantage of the proposed protocol is that it uses a linear, instead of quadratic, number of messages, and that message exchanges are staggered in time instead of all happening in narrow synchronization intervals. The drift rate of the synchronized clocks is optimal.> Flaviu Cristian, Christof Fetzer |
SRDS | 2 |