EDBT 2026 Demo / reviewers in the wild / expert
Valerio Schiavoni
dblp:15/2850
· DBLP profile ↗
95ranked-venue papers
2as first author
45since 2021 · last 2026
0000-0003-1493-6603ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 32 · 16 since 2021Systems, architecture and hardware · 28 · 1 first-author · 14 since 2021Software engineering, systems software and programming languages · 18 · 1 first-author · 9 since 2021Computer networks · 5 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DroidHunter: A Robust Vision-Based Detection Against Hidden Android MalwareabstractDue to their large popularity, Android smartphones are often targeted by malware attacks. Several strategies exist to detect malware code. However, we show that they are insufficient when dealing with obfuscation techniques. DroidHunter is our novel method for detecting Android malwares. DroidHunter leverages opcodes and their parameters, transforming those into RGB images and specific encoding techniques. The generated images are then used to train two different classification models based on support vector machines, convolutional neural networks, and a vision-based transformer. We evaluate DroidHunter on several datasets with up to 476,937 APKs from multiple sources. With detection rates from 98.65% to 99.94%, DroidHunter overcomes nine state-of-the-art malware detection techniques, including Drebin, MaMadroid, DexRay. Moreover, DroidHunter demonstrates strong resilience against hidden malware with detection rates up to 98.98%, and shows robustness on newly emerging threats, achieving an AUT of 0.89 on recent malware samples. We release our code to the research community, with instructions to reproduce our evaluation available at: https://zenodo.org/doi/10.5281/zenodo.10977166. Victoire Nganfang, Simon Queyrut, Yérom-David Bromberg, Valerio Schiavoni, Djob Mvondo, Kengne Tchendji Vianney |
AsiaCCS | 4 |
| 2026 | An Automated IoT-Based Infrastructure for Real-Time Soil Water Deficit Prediction - (Use-Case Paper)
Michèle Fischer, Hugo Delottier, Qi Tang 0004, Oliver Schilling, Valerio Schiavoni, Philip Brunner |
DAIS | 5 |
| 2026 | Heterogeneous Application Orchestration in Cyber-Physical Systems
Mehmet Cihan Sakman, Valerio Schiavoni, Ronny Seiger, Olaf Zimmermann, Josef Spillner |
DAIS | 2 |
| 2026 | Graph-Matrix Model for Data Storage Systems
Quentin Voiret, Bertrand Ducourthial, Pascal Felber, Valerio Schiavoni |
DAIS | 4 |
| 2026 | TriHaRd: Higher Resilience for TEE Trusted TimeabstractAccurately measuring time passing is critical for many applications. However, in Trusted Execution Environments (TEEs) such as Intel SGX, the time source is outside the Trusted Computing Base: a malicious host can manipulate the TEE’s notion of time, jumping in time or affecting perceived time speed. Previous work (Triad) proposes protocols for TEEs to maintain a trustworthy time source by building a cluster of TEEs that collaborate with each other and with a remote Time Authority to maintain a continuous notion of passing time. However, such approaches still allow an attacker to control the operating system and arbitrarily manipulate their own TEE’s perceived clock speed. An attacker can even propagate faster passage of time to honest machines participating in Triad’s trusted time protocol, causing them to skip to timestamps arbitrarily far in the future. We propose TriHaRd, a TEE trusted time protocol achieving high resilience against clock speed and offset manipulations, notably through Byzantine-resilient clock updates and consistency checks. We empirically show that TriHaRd mitigates known attacks against Triad. This repository contains the source code, as well as deployment and analysis scripts, for the "TriHaRd: Higher Resilience for TEE Trusted Time" paper, accepted for publication at the INFOCOM'26 conference. Matthieu Bettinger, Sonia Ben Mokhtar, Pascal Felber, Etienne Rivière, Valerio Schiavoni, Anthony Simonet |
INFOCOM | 5 |
| 2026 | A comprehensive performance evaluation of TEEs for confidential DNA alignment
Lorenzo Brescia, Iacopo Colonnelli, Robert Birke, Valerio Schiavoni, Pascal Felber, Marco Aldinucci |
Future Gener. Comput. Syst. | 4 |
| 2025 | ConfBench: A Tool for Easy Evaluation of Confidential Virtual MachinesabstractEnsuring the security and confidentiality of cloud computing workloads is essential. To this end, major cloud providers offer computing instances based on trusted execution environments (TEEs) to support confidential computing in virtual machines. TEEs are hardware-based shielded environments building on technologies available today, such as Intel TDX or AMD SEV-SNP or that will soon be, as with ARM CCA.To lower the barriers to experimenting with these technologies for researchers and practitioners, we developed ConfBench, a tool for easy evaluation of confidential virtual machines. ConfBench supports both cloud-native workloads (Function-as-a-Service) and classic applications. ConfBench facilitates the management of the full lifecycle of such workloads, from their deployment to the gathering of performance metrics, taking into account the specifics of TEE-enabled confidential virtual machines. We use ConfBench to collect execution overhead measurements for different VM-enabled TEEs (Intel TDX and AMD SEV-SNP) through extensive experiments. We also showcase how ConfBench’s architecture allows for validating also simulation-based TEEs, reporting preliminary results with ARM CCA. We highlight the intrinsic overheads of such confidential VMs by conducting stress tests against machine learning inference tasks, DBMS and native-OS operations benchmarking, as well as by evaluating the costs of attestation operations required in the context of confidential computing. The results indicate generally tenable overheads with modern TEEs, with exceptions mainly from I/O-intensive tasks, especially with TDX. ConfBench’s multi-language support for FaaS workloads also lets us gain insights into differences stemming from varying complexities behind language runtimes. We release ConfBench to the research community and provide instructions to reproduce our experiments. Andrea De Murtas, Daniele Cono D'Elia, Giuseppe Antonio Di Luna, Pascal Felber, Leonardo Querzoni, Valerio Schiavoni |
DSN | 6 |
| 2025 | PhishingHook: Catching Phishing Ethereum Smart Contracts leveraging EVM OpcodesabstractThe Ethereum Virtual Machine (EVM) is a decentralized computing engine. It enables the Ethereum blockchain to execute smart contracts and decentralized applications (dApps). The increasing adoption of Ethereum sparked the rise of phishing activities. Phishing attacks often target users through deceptive means, e.g., fake websites, wallet scams, or malicious smart contracts, aiming to steal sensitive information or funds. A timely detection of phishing activities in the EVM is therefore crucial to preserve the user trust and network integrity. Some state-of-the art approaches to phishing detection in smart contracts rely on the online analysis of transactions and their traces. However, replaying transactions often exposes sensitive user data and interactions, with several security concerns. In this work, we present PhishingHook, a framework that applies machine learning techniques to detect phishing activities in smart contracts by directly analyzing the contract’s bytecode and its constituent opcodes. We evaluate the efficacy of such techniques in identifying malicious patterns, suspicious function calls, or anomalous behaviors within the contract’s code itself before it is deployed or interacted with. We experimentally compare 16 techniques, belonging to four main categories (Histogram Similarity Classifiers, Vision Models, Language Models and Vulnerability Detection Models), using 7,000 real-world malware smart contracts. Our results demonstrate the efficiency of PhishingHook in performing phishing classification systems, with about 90% average accuracy among all the models. We support experimental reproducibility, and we release our code and datasets to the research community. Pasquale De Rosa, Simon Queyrut, Yérom-David Bromberg, Pascal Felber, Valerio Schiavoni |
DSN | 5 |
| 2025 | On Real-Time Guarantees in Intel SGX and TDXabstractTrusted execution environments (TEE) represent a major technological breakthrough that provide strong confidentiality and integrity guarantees for code and data running on potentially vulnerable or untrustworthy computing systems, such as cloud, edge, embedded, mobile, or even blockchain systems. However, the performance overhead associated with TEEs still poses a limitation on the extent to which real-time (RT) sensitive applications can benefit from this technology, e.g., to run on untrusted third-party infrastructures. This work investigates various TEE-based architectures spanning from process-based to virtual-machine-based implementations, for securing RT applications. It offers in addition an in-depth evaluation of these architectures, providing insights into how various TEE deployments influence the temporal compute and communication guarantees of RT systems. Peterson Yuhala, Christian Göttel, Jämes Ménétrey, Valerio Schiavoni, David Kozhaya, Pascal Felber |
ECRTS | 4 |
| 2025 | Blockchain Energy Consumption: Unveiling the Impact of Network Topologies
Vincenzo P. Di Perna, Valerio Schiavoni, Francesco Fabris, Marco Bernardo 0001 |
ICBC | 2 |
| 2025 | IM-PIR: In-Memory Private Information RetrievalabstractPrivate information retrieval (PIR) is a cryptographic primitive that allows a client to securely query one or multiple servers without revealing their specific interests. In spite of their strong security guarantees, current PIR constructions are computationally costly. Specifically, most PIR implementations are memory-bound due to the need to scan extensive databases (in the order of GB), making them inherently constrained by the limited memory bandwidth in traditional processor-centric computing architectures. Processing-in-memory (PIM) is an emerging computing paradigm that augments memory with compute capabilities, addressing the memory bandwidth bottleneck while simultaneously providing extensive parallelism. Recent research has demonstrated PIM's potential to significantly improve performance across a range of data-intensive workloads, including graph processing, genome analysis, and machine learning. Mpoki Mwaisela, Peterson Yuhala, Pascal Felber, Valerio Schiavoni |
Middleware | 4 |
| 2025 | Where to Place Your TEE? In Search of a Censorship-Resilient Design for Rollup Sequencers
Andrei Arusoaie, Claudiu-Nicu Barbieru, Oana-Otilia Captarencu, Pascal Felber, Corentin Libert, Emanuel Onica, Etienne Rivière, Valerio Schiavoni, Peterson Yuhala |
OPODIS | 8 |
| 2025 | Speeding Up the Development for the Computing Continuum with WebAssemblyabstractWebAssembly is a lightweight binary format that achieves portable software across heterogeneous compute nodes in the computing continuum. One of its main drawbacks is the limited availability of pre-packaged software, in comparison to larger ecosystems, as it is instead the case with container registries or programming language package repositories. A transpilation approach might close this gap. However, it requires overcoming several technical issue dealing with native code execution. This work assesses the current transpilation technologies and the reduction of deployment friction. We report about the state of technology and contribute three advancements: (i) a compilation pipeline for Python applications using componentize-py and wasi-wheels, enabling native extension support; (ii) exploitation of hardware accelerators in Wasm via WASI-nn, and (iii) seamless cross-language composition using WebAssembly Interface Types (WIT). We demonstrate the effectiveness of our result with two real-world applications in a 3-zones continuum, highlighting faster development without compromising software dependability. We support experimental reproducibility and release our prototype and dataset as open-source. Mehmet Cihan Sakman, Josef Spillner, Valerio Schiavoni |
PRDC | 3 |
| 2025 | Kollaps: Decentralized and Efficient Network Emulation for Large-Scale SystemsabstractThe performance and behavior of distributed systems is highly influenced by network properties, latency, bandwidth, packet loss, and jitter. When developing a distributed system, questions like, “how sensitive is the application’s performance to network latency and bandwidth?” commonly arise. Answering these questions systematically and in a reproducible manner is very hard due to the variability and lack of control over the network. Moreover, state-of-the-art approaches are focused exclusively on the control plane, lack support for network dynamics or do not scale beyond a single machine or small cluster, which further aggravates this problem. Kollaps is a distributed, scalable, and efficient network emulator addressing these limitations by hinging on two observations. First, from an application’s perspective, what matters are the emergent end-to-end properties (e. g., latency, bandwidth, jitter) rather than the internal state of the routers and switches leading to those properties. Second, this model is amenable to decentralized management, allowing the emulation to scale with the number of machines required by the application. This premise allows for building a simpler, dynamic emulation model that does not require maintaining the full network state. Kollaps is agnostic of the application language and transport protocol, scales to thousands of application nodes, and is accurate when compared against a bare-metal deployment or state-of-the-art approaches that emulate the full network state. We use Kollaps to accurately reproduce results from the literature and predict the behavior of complex unmodified distributed systems under different network dynamics. Sebastião Amaro, Miguel Matos, Valerio Schiavoni |
IEEE Trans. Netw. | 3 |
| 2024 | On the Cost of Model-Serving Frameworks: An Experimental EvaluationabstractIn machine learning (ML), the inference phase is the process of applying pre-trained models to new, unseen data with the objective of making predictions. During the inference phase, end-users interact with ML services to gain insights, recommendations, or actions based on the input data. For this reason, serving strategies are nowadays crucial for deploying and managing models in production environments effectively. These strategies ensure that models are available, scalable, reliable, and performant for real-world applications, such as time series forecasting, image classification, natural language processing, and so on. In this paper, we evaluate the performances of five widely-used model serving frameworks (TensorFlow Serving, TorchServe, MLServer, MLflow, and BentoML) under four different scenarios (malware detection, cryptocoin prices forecasting, image classification, and sentiment analysis). We demonstrate that TensorFlow Serving is able to outperform all the other frameworks in serving deep learning (DL) models. Moreover, we show that DL-specific frameworks (TensorFlow Serving and TorchServe) display significantly lower latencies than the three general-purpose ML frameworks (BentoML, MLFlow, and MLServer). Pasquale De Rosa, Yérom-David Bromberg, Pascal Felber, Djob Mvondo, Valerio Schiavoni |
IC2E | 5 |
| 2024 | P4ce: Consensus over RDMA at Line SpeedabstractP4ce is the first replication protocol that exhibits the same latency and requires the same network capacity as sending data to a single server. P4ce builds upon previous RDMA-based consensus protocols. They achieve consensus with a single network round-trip, but with a reduced network throughput. P4ce also achieves consensus with a single round-trip, but without degrading throughput by decoupling the consensus decisions from the RDMA communications. The decision part of the consensus protocol runs on a commodity server, but the communication part of P4ce is fully implemented on a programmable switch, which replicates data and aggregates the acknowledgements in the network, avoiding the throughput bottleneck at the leader. Although simple in its principle, the implementation of P4ce raises many challenging issues, notably caused by the complexity of RDMA and the underlying network protocols, the intricacies of packet rewriting during replication and aggregation, and the restricted set of operations that can be implemented at wire speed in the programmable switch. We implemented P4ce and deployed it on a commercially-available Intel Tofino switch, achieving up to 4x better through-put and better latency than state-of-the-art consensus protocols. Rémi Dulong, Nathan Felber, Pascal Felber, Gilles Hopin, Baptiste Lepers, Valerio Schiavoni, Gaël Thomas 0001, Sébastien Vaucher |
ICDCS | 6 |
| 2024 | Evaluating the Potential of In-Memory Processing to Accelerate Homomorphic Encryption: Practical Experience ReportabstractThe widespread adoption of cloud-based solutions introduces privacy and security concerns. Techniques such as homomorphic encryption (HE) mitigate this problem by allowing computation over encrypted data without the need for decryption. However, the high computational and memory overhead associated with the underlying cryptographic operations has hindered the practicality of HE-based solutions. While a significant amount of research has focused on reducing computational overhead by utilizing hardware accelerators like GPUs and FPGAs, there has been relatively little emphasis on addressing HE memory overhead. Processing in-memory (PIM) presents a promising solution to this problem by bringing computation closer to data, thereby reducing the overhead resulting from processor-memory data movements. In this work, we evaluate the potential of a PIM architecture from UPMEM for accelerating HE operations. Firstly, we focus on PIM-based acceleration for polynomial operations, which underpin HE algorithms. Subsequently, we conduct a case study analysis by integrating PIM into two popular and open-source HE libraries, OpenFHE and HElib. Our study concludes with key findings and takeaways gained from the practical application of HE operations using PIM, providing valuable insights for those interested in adopting this technology. Mpoki Mwaisela, Joel Hari, Peterson Yuhala, Jämes Ménétrey, Pascal Felber, Valerio Schiavoni |
SRDS | 6 |
| 2024 | CLUES: Collusive Theft of Conditional Generative Adversarial NetworksabstractConditional Generative Adversarial Networks (cGANs) are increasingly popular web-based synthesis services accessed through a query API, e.g., cGANs generate a cat image based on a “cat” query. However, cGAN-based synthesizers can be stolen via adversaries' queries, i.e., model thieves. The prevailing adversarial assumption is that thieves act independently: they query the deployed cGAN (i.e., the victim), and train a stolen cGAN using the images obtained from the victim. A popular anti-theft defense consists in throttling down the number of queries from any given user. We consider a more realistic adversarial scenario: model thieves collude to query the victim, and then train the stolen cGAN. Clues is a new collusive model stealing framework, enabling thieves to bypass throttle-based defenses and steal cGANs more efficiently than through individual efforts. Thieves collect queried images and train a stolen cGAN in a federated manner. We evaluate Clues on three image datasets, e.g., MNIST, FashionMNIST and CelebA. We experimentally show the scalability of the proposed attack strategies against the number of thieves and the queried images, the impact of a classical noise-based defense, a passive watermarking defense and a JPEG-based countermeasure. Our evaluation shows that such a collusive stealing strategy gets close to 4 units of Frechet Inception Distance from a victim model. Our code is readily available to the research community: https://zenodo.org/records/10224340. Simon Queyrut, Valerio Schiavoni, Lydia Y. Chen, Pascal Felber, Robert Birke |
SRDS | 2 |
| 2024 | BlindexTEE: A Blind Index Approach Towards TEE-Supported End-to-End Encrypted DBMS
Louis Vialar, Jämes Ménétrey, Valerio Schiavoni, Pascal Felber |
SSS | 3 |
| 2024 | Reliable IoT analytics at scaleabstractSocieties and legislations are moving towards automated decision-making based on measured data in safety-critical environments. Over the next years, density and frequency of measurements will increase to generate more insights and get a more solid basis for decisions, including through redundant low-cost sensor deployments. The resulting data characteristics lead to large-scale system design in which small input data errors may lead to severe cascading problems including ultimately wrong decisions. To ensure internal data consistency to mitigate this risk in such IoT environments, fast-paced data fusion and consensus among redundant measurements need to be achieved. In this context, we introduce history-aware sensor fusion powered by accurate voting with clustering as a promising approach to achieve fast and informed consensus, which can converge to the output up to 4X faster than the state of the art history-based voting. Leveraging three case studies, we investigate different voting schemes and show how this approach can improve data accuracy by up to 30% and performance by up to 12% compared to state-of-the-art sensor fusion approaches. We furthermore contribute a specification format for easily deploying our methods in practice and use it to develop a pilot implementation. Panagiotis Gkikopoulos, Peter G. Kropf, Valerio Schiavoni, Josef Spillner |
J. Parallel Distributed Comput. | 3 |
| 2024 | A Comprehensive Trusted Runtime for WebAssembly With Intel SGXabstractIn real-world scenarios, trusted execution environments (TEEs) frequently host applications that lack the trust of the infrastructure provider, as well as data owners who have specifically outsourced their data for remote processing. We presentTwine, a trusted runtime for running WebAssembly-compiled applications within TEEs, establishing a two-way sandbox.Twineleverages memory safety guarantees of WebAssembly (Wasm) and abstracts the complexity of TEEs, empowering the execution of legacy and language-agnostic applications. It extends the standard WebAssembly system interface (WASI), providing controlled OS services, focusing on I/O. Additionally, through built-in TEE mechanisms,Twinedelivers attestation capabilities to ensure the integrity of the runtime and the OS services supplied to the application. We evaluate its performance using general-purpose benchmarks and real-world applications, showing it compares on par with state-of-the-art solutions. A case study involving fintech companyCredorareveals thatTwinecan be deployed in production with reasonable performance trade-offs, ranging from a 0.7× slowdown to a 1.17× speedup compared to native run time. Finally, we identify performance improvement through library optimisation, showcasing one such adjustment that leads up to$4.1\times$speedup.Twineis open-source and has been upstreamed into the original Wasm runtime, WAMR. Jämes Ménétrey, Marcelo Pasin, Pascal Felber, Valerio Schiavoni, Giovanni Mazzeo, Arne Hollum, Darshan Vaydia |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | SGX Switchless Calls Made ConfiglessabstractIntel's software guard extensions (SGX) provide hardware enclaves to guarantee confidentiality and integrity for sensitive code and data. However, systems leveraging such security mechanisms must often pay high performance overheads. A major source of this overhead is SGX enclave transitions which induce expensive cross-enclave context switches. The Intel SGX SDK mitigates this with a switchless call mechanism for transitionless cross-enclave calls using worker threads. Intel's SGX switchless call implementation improves performance but provides limited flexibility: developers need to statically fix the system configuration at build time, which is error-prone and misconfigurations lead to performance degradations and waste of CPU resources. ZC-Switchless is a configless and efficient technique to drive the execution of SGX switchless calls. Its dynamic approach optimises the total switchless worker threads at runtime to minimise CPU waste. The experimental evaluation shows that ZC-Switchless obviates the performance penalty of misconfigured switchless systems while minimising CPU waste. Peterson Yuhala, Michael Paper, Timothée Zerbib, Pascal Felber, Valerio Schiavoni, Alain Tchana |
DSN | 5 |
| 2023 | Mitigating Adversarial Attacks in Federated Learning with Trusted Execution EnvironmentsabstractThe main premise of federated learning (FL) is that machine learning model updates are computed locally to preserve user data privacy. This approach avoids by design user data to ever leave the perimeter of their device. Once the updates aggregated, the model is broadcast to all nodes in the federation. However, without proper defenses, compromised nodes can probe the model inside their local memory in search for adversarial examples, which can lead to dangerous real-world scenarios. For instance, in image-based applications, adversarial examples consist of images slightly perturbed to the human eye getting misclassified by the local model. These adversarial images are then later presented to a victim node's counterpart model to replay the attack. Typical examples harness dissemination strategies such as altered traffic signs (patch attacks) no longer recognized by autonomous vehicles or seemingly unaltered samples that poison the local dataset of the FL scheme to undermine its robustness. PELTA is a novel shielding mechanism leveraging Trusted Execution Environments (TEEs) that reduce the ability of attackers to craft adversarial samples. PELTA masks inside the TEE the first part of the back-propagation chain rule, typically exploited by attackers to craft the malicious samples. We evaluate PELTA on state-of-the-art accurate models using three well-established datasets: CIFAR-10, CIFAR-100 and ImageNet. We show the effectiveness of PELTA in mitigating six white-box state-of-the-art adversarial attacks, such as Projected Gradient Descent, Momentum Iterative Method, Auto Projected Gradient Descent, the Carlini & Wagner attack. In particular, PELTA constitutes the first attempt at defending an ensemble model against the Self-Attention Gradient attack to the best of our knowledge. Our code is available to the research community at https://github.com/queyrusi/Pelta Simon Queyrut, Valerio Schiavoni, Pascal Felber |
ICDCS | 2 |
| 2023 | Characterizing Distributed Machine Learning Workloads on Apache Spark: (Experimentation and Deployment Paper)abstractDistributed machine learning (DML) environments are widely used in many application domains to build decision-making systems. However, the complexity of these environments is overwhelming for novice users. On the one hand, data scientists are more familiar with hyper-parameter tuning and typically lack an understanding of the trade-offs and challenges of parameterizing DML platforms to achieve good performance. On the other hand, system administrators focus on tuning distributed platforms, unaware of the possible implications of the platform on the quality of the learning models. To shed light on such parameter configuration interplay, we run multiple DML workloads on the widely used Apache Spark distributed platform, leveraging 13 popular learning methods and 6 real-world datasets on two distinct clusters. We collect and perform an in-depth analysis of workload execution traces to compare the efficiency of different configuration strategies. We consider tuning only hyper-parameters, tuning only platform parameters, and jointly tuning both hyper-parameters and platform parameters. We publicly release our collected traces and derive key takeaways on DML workloads. Counter-intuitively, platform parameters have a higher impact on the model quality than hyper-parameters. More generally, we show that multi-level parameter configuration can provide better results in terms of model quality and execution time while also optimizing resource costs. Yasmine Djebrouni, Isabelly Rocha, Sara Bouchenak, Lydia Y. Chen, Pascal Felber, Vania Marangozova-Martin, Valerio Schiavoni |
Middleware | 7 |
| 2023 | SecV: Secure Code Partitioning via Multi-Language Secure ValuesabstractTrusted execution environments like Intel SGX provide enclaves, which offer strong security guarantees for applications. Running entire applications inside enclaves is possible, but this approach leads to a large trusted computing base (TCB). As such, various tools have been developed to partition programs written in languages such as C or Java into trusted and untrusted parts, which are run in and out of enclaves respectively. However, those tools depend on language-specific taint-analysis and partitioning techniques. They cannot be reused for other languages and there is thus a need for tools that transcend this language barrier. Peterson Yuhala, Pascal Felber, Hugo Guiroux, Jean-Pierre Lozi, Alain Tchana, Valerio Schiavoni, Gaël Thomas 0001 |
Middleware | 6 |
| 2023 | A Holistic Approach for Trustworthy Distributed Systems with WebAssembly and TEEsabstractEthereum is the dominant blockchain ecosystem capable of executing Turing-complete smart contracts. Rollups gained significant traction as the primary layer 2 (L2) solution meant to bring horizontal scalability to the main Ethereum network (L1). A core component of any rollup is the sequencer, which creates new L2 blocks to be submitted in rollup batches to L1. In most of the current rollup architectures, this component is centralised. As a result, these designs are prone to inconspicuous censorship practices by the sequencer. Trusted execution environments (TEEs) can guarantee the integrity of various sequencer components, which is instrumental in addressing censorship. However, the reaction of the system design to censorship attempts depends on where a TEE is integrated and which components it protects. In particular, this reaction is limited in the case of a monolithic TEE-protected sequencer design. Proposer-Builder Separation (PBS) is a non-monolithic paradigm adopted on L1, which separates the production of blocks from proposing them for inclusion in the blockchain. Recently, PBS has been considered for integration with L2 sequencers, with an impact on alleviating censorship. In this paper, we explore the design space of TEE-integrating PBS and non-PBS sequencer variants. First, we introduce a formal framework for the censorship actions that captures the specificity of the L2 sequencer. Then, we analyse to what extent the different designs address these censorship actions. Our main contribution is a novel design variation that allows for a precise observation of censored transactions. In the presence of TEEs, in a PBS setting, we demonstrate this precise observability, which is necessary to enable resilience to censorship. Jämes Ménétrey, Aeneas Grüter, Peterson Yuhala, Julius Oeftiger, Pascal Felber, Marcelo Pasin, Valerio Schiavoni |
OPODIS | 7 |
| 2023 | TL4x: Buffered Durable Transactions on Disk as Fast as in MemoryabstractThe arrival of persistent memory devices to consumer market has revived the interest in transactional durable algorithms. Persistent memory (PM) is touted as having two attributes that distinguish it from other storage technologies: byte-addressability and fast transactional persistence. Gal Assa, Andreia Correia, Pedro Ramalhete, Valerio Schiavoni, Pascal Felber |
PPoPP | 4 |
| 2023 | Capacity planning for dependable services
Rasha Faqeh, André Martin, Valerio Schiavoni, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
Theor. Comput. Sci. | 3 |
| 2022 | Attestation Mechanisms for Trusted Execution Environments DemystifiedabstractAttestation is a fundamental building block to establish trust over software systems. When used in conjunction with trusted execution environments, it guarantees the genuineness of the code executed against powerful attackers and threats, paving the way for adoption in several sensitive application domains. This paper reviews remote attestation principles and explains how the modern and industrially well-established trusted execution environments Intel SGX, Arm TrustZone and AMD SEV, as well as emerging RISC-V solutions, leverage these mechanisms. Jämes Ménétrey, Christian Göttel, Anum Khurshid, Marcelo Pasin, Pascal Felber, Valerio Schiavoni, Shahid Raza |
DAIS | 6 |
| 2022 | Understanding Cryptocoins Trends Correlations
Pasquale De Rosa, Valerio Schiavoni |
DAIS | 2 |
| 2022 | VEDLIoT: Very Efficient Deep Learning in IoTabstractThe VEDLIoT project targets the development of energy-efficient Deep Learning for distributed AIoT applications. A holistic approach is used to optimize algorithms while also dealing with safety and security challenges. The approach is based on a modular and scalable cognitive IoT hardware platform. Using modular microserver technology enables the user to configure the hardware to satisfy a wide range of applications. VEDLIoT offers a complete design flow for Next-Generation IoT devices required for collaboratively solving complex Deep Learning applications across distributed systems. The methods are tested on various use-cases ranging from Smart Home to Automotive and Industrial IoT appliances. VEDLIoT is an H2020 EU project which started in November 2020. It is currently in an intermediate stage with the first results available. Martin Kaiser, René Griessl, Nils Kucza, Carola Haumann, Lennart Tigges, Kevin Mika, Jens Hagemeyer, Florian Porrmann, Ulrich Rückert 0001, Micha vor dem Berge, Stefan Krupop, Mario Porrmann, Marco Tassemeier, Pedro Trancoso, Fareed Qararyah, Stavroula Zouzoula, António Casimiro, Alysson Neves Bessani, José Cecílio, Stefan Andersson, Oliver Brunnegård, Olof Eriksson, Roland Weiss 0001, Franz Meierhöfer, Hans Salomonsson, Elaheh Malekzadeh, Daniel Ödman, Anum Khurshid, Pascal Felber, Marcelo Pasin, Valerio Schiavoni, Jämes Ménétrey, Karol Gugala, Piotr Zierhoffer, Eric Knauss, Hans-Martin Heyn |
DATE | 31 |
| 2022 | WaTZ: A Trusted WebAssembly Runtime Environment with Remote Attestation for TrustZoneabstractWebAssembly (Wasm) is a novel low-level bytecode format that swiftly gained popularity for its efficiency, versatility and security, with near-native performance. Besides, trusted execution environments (TEEs) shield critical software assets against compromised infrastructures. However, TEEs do not guarantee the code to be trustworthy or that it was not tampered with. Instead, one relies on remote attestation to assess the code before execution. This paper describes WaTZ, which is (i) an efficient and secure runtime for trusted execution of Wasm code for Arm’s TrustZone TEE, and (ii) a lightweight remote attestation system optimised for Wasm applications running in TrustZone, as it lacks built-in mechanisms for attestation. The remote attestation protocol is formally verified using a state-of-the-art analyser and model checker. Our extensive evaluation of Arm-based hardware uses synthetic and real-world benchmarks, illustrating typical tasks IoT devices achieve. WaTZ’s execution speed is on par with Wasm runtimes in the normal world and reaches roughly half the speed of native execution, which is compensated by the additional security guarantees and the inter-operability offered by Wasm. WaTZ is open-source and available on GitHub along with instructions to reproduce our experiments. Jämes Ménétrey, Marcelo Pasin, Pascal Felber, Valerio Schiavoni |
ICDCS | 4 |
| 2022 | Shielding federated learning systems against inference attacks with ARM TrustZoneabstractFederated Learning (FL) opens new perspectives for training machine learning models while keeping personal data on the users premises. Specifically, in FL, models are trained on the users' devices and only model updates (i.e., gradients) are sent to a central server for aggregation purposes. However, the long list of inference attacks that leak private data from gradients, published in the recent years, have emphasized the need of devising effective protection mechanisms to incentivize the adoption of FL at scale. While there exist solutions to mitigate these attacks on the server side, little has been done to protect users from attacks performed on the client side. In this context, the use of Trusted Execution Environments (TEEs) on the client side are among the most proposing solutions. However, existing frameworks (e.g., DarkneTZ) require statically putting a large portion of the machine learning model into the TEE to effectively protect against complex attacks or a combination of attacks. We present GradSec, a solution that allows protecting in a TEE only sensitive layers of a machine learning model, either statically or dynamically, hence reducing both the Trusted Computing Base (TCB) size and the overall training time by up to 30% and 56%, respectively compared to state-of-the-art competitors. Aghiles Ait Messaoud, Sonia Ben Mokhtar, Vlad Nitu, Valerio Schiavoni |
Middleware | 4 |
| 2022 | EdgeTune: Inference-Aware Multi-Parameter TuningabstractDeep Neural Networks (DNNs) have demonstrated impressive performance on many machine-learning tasks such as image recognition and language modeling, and are becoming prevalent even on mobile platforms. Despite so, designing neural architectures still remains a manual, time-consuming process that requires profound domain knowledge. Recently, Parameter Tuning Servers have gathered the attention o industry and academia. Those systems allow users from all domains to automatically achieve the desired model accuracy for their applications. However, although the entire process of tuning and training models is performed solely to be deployed for inference, state-of-the-art approaches typically ignore system-oriented and inference-related objectives such as runtime, memory usage, and power consumption. This is a challenging problem: besides adding one more dimension to an already complex problem, the information about edge devices available to the user is rarely known or complete. To accommodate all these objectives together, it is crucial for tuning system to take a holistic approach to parameter tuning and consider all levels of parameters simultaneously into account. We present EdgeTune, a novel inference-aware parameter tuning server. It considers the tuning of parameters in all levels backed by an optimization function capturing multiple objectives. Our approach relies on inference estimated metrics collected from our emulation server running asynchronously from the main tuning process. The latter can then leverage the inference performance while still tuning the model. We propose a novel one-fold tuning algorithm that employs the principle of multi-fidelity and simultaneously explores multiple tuning budgets, which the prior art can only handle as suboptimal case of single type of budget. EdgeTune outputs inference recommendations to the user while improving tuning time and energy by at least 18\% and 53\% when compared to the baseline. Isabelly Rocha, Pascal Felber, Valerio Schiavoni, Lydia Y. Chen |
Middleware | 3 |
| 2022 | Capacity Planning for Dependable Services
Rasha Faqeh, André Martin, Valerio Schiavoni, Pramod Bhatotia, Pascal Felber, Christof Fetzer |
SSS | 3 |
| 2022 | Decentralised Data Quality Control in Ground Truth Production for Autonomic DecisionsabstractAutonomic decision-making based on rules and metrics is inevitably on the rise in distributed software systems. Often, the metrics are acquired from system observations such as static checks and runtime traces. To avoid bias propagation and hence reduce wrong decisions in increasingly autonomous systems due to poor observation data quality, multiple independent observers can exchange their findings and produce a majority-accepted, complete and outlier-cleaned ground truth in the form of consensus-supported metrics. In this work, we motivate the growing importance of metrics for informed and autonomic decisions in clouds and other distributed systems, present reasons for diverging observations, and describe a federated approach to produce ground truth with data-centric consensus voting for more reliable decision making processes. We validate the system design with experiments in the area of cloud software artefact observations and highlight benefits for reproducible distributed system behaviour. Panagiotis Gkikopoulos, Valerio Schiavoni, Josef Spillner |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | KeVlar-Tz: A Secure Cache for ArmTrustZone - (Practical Experience Report)
Oscar Benedito, Ricard Delgado-Gonzalo, Valerio Schiavoni |
DAIS | 3 |
| 2021 | Analysis and Improvement of Heterogeneous Hardware Support in Docker Images
Panagiotis Gkikopoulos, Valerio Schiavoni, Josef Spillner |
DAIS | 2 |
| 2021 | NVCache: A Plug-and-Play NVMM-based I/O Booster for Legacy SystemsabstractThis paper introduces NVCACHE, an approach that uses a non-volatile main memory (NVMM) as a write cache to improve the write performance of legacy applications. We compare NVCACHE against file systems tailored for NVMM (Ext4-DAX and NOVA) and with I/O-heavy applications (SQLite, RocksDB). Our evaluation shows that NVCACHE reaches the performance level of the existing state-of-the-art systems for NVMM, but without their limitations: NVCACHE does not limit the size of the stored data to the size of the NVMM, and works transparently with unmodified legacy applications, providing additional persistence guarantees even when their source code is not available. Rémi Dulong, Rafael Pires 0001, Andreia Correia, Valerio Schiavoni, Pedro Ramalhete, Pascal Felber, Gaël Thomas 0001 |
DSN | 4 |
| 2021 | ADAM-CS: Advanced Asynchronous Monotonic Counter ServiceabstractTrusted execution environments (TEEs) offer the technological breakthrough to allow several applications to be deployed and executed over untrusted public cloud environments. Although TEEs (e. g., Intel SGX, ARM TrustZone, AMD SEV) provide several mechanisms to ensure confidentiality and integrity of data and code, they do not offer freshness out of the box, a critical aspect yet often overlooked, for instance, to protect against rollback attacks. Monotonic counters are a popular way to detect rollbacks, as their counter values cannot be decremented. However, counter increments are slow (i.e., 10thof milliseconds), making their use impractical for distributed services and applications processing thousands of transactions simultaneously, for which an order of magnitude improvement is needed. ADAM-CS is an asynchronous monotonic counter service to protect such high-traffic applications against rollback attacks. Leveraging a set of distributed monotonic counters and specific algorithms, ADAM-CS minimizes the maximum vulnerability window (MVW), i.e., the amount of transactions an adversary could successfully rollback. Thanks to its asynchronous nature, ADAM-CS supports thousands of increments per second without introducing additional latency in the transactions performed by applications. Our measurements indicate that we can keep the MVW well below 10ms while supporting a throughput of more than 21K requests/s when using eight counters. André Martin, Cong Lian, Franz Gregor, Robert Krahn, Valerio Schiavoni, Pascal Felber, Christof Fetzer |
DSN | 5 |
| 2021 | Plinius: Secure and Persistent Machine Learning Model TrainingabstractWith the increasing popularity of cloud based machine learning (ML) techniques there comes a need for privacy and integrity guarantees for ML data. In addition, the significant scalability challenges faced by DRAM coupled with the high access-times of secondary storage represent a huge performance bottleneck for ML systems. While solutions exist to tackle the security aspect, performance remains an issue. Persistent memory (PM) is resilient to power loss (unlike DRAM), provides fast and fine-granular access to memory (unlike disk storage) and has latency and bandwidth close to DRAM (in the order of ns and GB/s, respectively). We present PLINIUS, a ML framework using Intel SGX enclaves for secure training of ML models and PM for fault tolerance guarantees. PLINIUS uses a novel mirroring mechanism to create and maintain (i) encrypted mirror copies of ML models on PM, and (ii) encrypted training data in byte-addressable PM, for near-instantaneous data recovery after a system failure. Compared to disk-based checkpointing systems, PLINIUS is 3.2× and 3.7× faster respectively for saving and restoring models on real PM hardware, achieving robust and secure ML model training in SGX enclaves. Peterson Yuhala, Pascal Felber, Valerio Schiavoni, Alain Tchana |
DSN | 3 |
| 2021 | Twine: An Embedded Trusted Runtime for WebAssemblyabstractWebAssembly is an Increasingly popular lightweight binary instruction format, which can be efficiently embedded and sandboxed. Languages like C, C++, Rust, Go, and many others can be compiled into WebAssembly. This paper describes Twine, a WebAssembly trusted runtime designed to execute unmodified, language-independent applications. We leverage Intel SGX to build the runtime environment without dealing with language-specific, complex APIs. While SGX hardware provides secure execution within the processor, Twine provides a secure, sandboxed software runtime nested within an SGX enclave, featuring a WebAssembly system interface (WASI) for compatibility with unmodified WebAssembly applications. We evaluate Twine with a large set of general-purpose benchmarks and real-world applications. In particular, we used Twine to implement a secure, trusted version of SQLite, a well-known full-fledged embeddable database. We believe that such a trusted database would be a reasonable component to build many larger application services. Our evaluation shows that SQLite can be fully executed inside an SGX enclave via WebAssembly and existing system interface, with similar average performance overheads. We estimate that the performance penalties measured are largely compensated by the additional security guarantees and its full compatibility with standard WebAssembly. An indepth analysis of our results indicates that performance can be greatly improved by modifying some of the underlying libraries. We describe and implement one such modification in the paper, showing up to 4.1 × speedup. Twine is open-source, available at GitHub along with instructions to reproduce our experiments. Jämes Ménétrey, Marcelo Pasin, Pascal Felber, Valerio Schiavoni |
ICDE | 4 |
| 2021 | Montsalvat: Intel SGX shielding for GraalVM native imagesabstractThe popularity of the Java programming language has led to its wide adoption in cloud computing infrastructures. However, Java applications running in untrusted clouds are vulnerable to various forms of privileged attacks. The emergence of trusted execution environments (TEEs) such as Intel SGX mitigates this problem. TEEs protect code and data in secure enclaves inaccessible to untrusted software, including the kernel and hypervisors. To efficiently use TEEs, developers must manually partition their applications into trusted and untrusted parts, in order to reduce the size of the trusted computing base (TCB) and minimise the risks of security vulnerabilities. However, partitioning applications poses two important challenges: (i) ensuring efficient object communication between the partitioned components, and (ii) ensuring the consistency of garbage collection between the parts, especially with memory-managed languages such as Java. We present Montsalvat, a tool which provides a practical and intuitive annotation-based partitioning approach for Java applications destined for secure enclaves. Montsalvat provides an RMI-like mechanism to ensure inter-object communication, as well as consistent garbage collection across the partitioned components. We implement Montsalvat with GraalVM native-image, a tool for compiling Java applications ahead-of-time into standalone native executables that do not require a JVM at runtime. Our extensive evaluation with micro- and macro-benchmarks shows our partitioning approach to boost performance in real-world applications up to 6.6x (PalDB) and 2.2x (GraphChi) as compared to solutions that naively include the entire applications in the enclave. Peterson Yuhala, Jämes Ménétrey, Pascal Felber, Valerio Schiavoni, Alain Tchana, Gaël Thomas 0001, Hugo Guiroux, Jean-Pierre Lozi |
Middleware | 4 |
| 2021 | Scrooge Attack: Undervolting ARM Processors for Profit: Practical experience reportabstractLatest ARM processors are approaching the computational power of x86 architectures while consuming much less energy. Consequently, supply follows demand with Amazon EC2, Equinix Metal and Microsoft Azure offering ARM-based instances, while Oracle Cloud Infrastructure is about to add such support. We expect this trend to continue, with an increasing number of cloud providers offering ARM-based cloud instances. ARM processors are more energy-efficient leading to substantial electricity savings for cloud providers. However, a malicious cloud provider could intentionally reduce the CPU voltage to further lower its costs. Running applications malfunction when the undervolting goes below critical thresholds. By avoiding critical voltage regions, a cloud provider can run undervolted instances in a stealthy manner. This practical experience report describes a novel attack scenario: an attack launched by the cloud provider against its users to aggressively reduce the processor voltage for saving energy to the last penny. We call it the Scrooge Attack and show how it could be executed using ARM-based computing instances. We mimic ARM-based cloud instances by deploying our own ARM-based devices using different generations of Raspberry Pi. Using realistic and synthetic workloads, we demonstrate to which degree of aggressiveness the attack is relevant. The attack is unnoticeable by our detection method up to an offset of −50 mV. We show that the attack may even remain completely stealthy for certain workloads. Finally, we propose a set of client-based detection methods that can identify undervolted instances. We support experimental reproducibility and provide instructions to reproduce our results. Christian Göttel, Konstantinos Parasyris, Osman S. Unsal, Pascal Felber, Marcelo Pasin, Valerio Schiavoni |
SRDS | 6 |
| 2021 | MinervaFS: A User-Space File System for Generalised Deduplication: (Practical experience report)abstractDeduplication exploits the presence of similar data chunks to reduce storage overhead. Generalised deduplication (GD) uses transformation functions to split data into a basis (common to millions of chunks) and a deviation with respect to the basis. Doing so, it avoids computing additional hashes, comparing or differentiating against previously stored chunks. Minervafs is the first FUSE-based file system for GD. We implement and evaluate it using several real-world datasets, e.g., satellite images and virtual machine images, comparing against classical deduplication approaches (ZFS, SDFS), delta compression (xdelta) or compression (Gzip). Compared to ZFS, Minervafs achieves up to 63.53% (average of 27.38%) saving in storage usage and a speedup of 16% in read-heavy workloads. For VM images, MINERVAFS's data compression is on par with Gzip, while outperforming ZFS by severalfold. In contrast to ZFS’ growing RAM costs when more data is stored, MinervaFS’ RAM usage is independent from the amount of data stored, making it well suited to handle growing storage demands. Lars Nielsen, Dorian Burihabwa, Valerio Schiavoni, Pascal Felber, Daniel Enrique Lucani |
SRDS | 3 |
| 2020 | ZipLine: in-network compression at line speedabstractNetwork appliances continue to offer novel opportunities to offload processing from computing nodes directly into the data plane. One popular concern of network operators and their customers is to move data increasingly faster. A common technique to increase data throughput is to compress it before its transmission. However, this requires compression of the data---a time and energy demanding preprocessing phase---and decompression upon reception---a similarly resource consuming operation. Moreover, if multiple nodes transfer similar data chunks across the network hop (e.g., a given pair of switches), each node effectively wastes resources by executing similar steps. This paper proposes ZipLine, an approach to design and implement (de)compression at line speed leveraging the Tofino hardware platform which is programmable using the P416 language. We report on lessons learned while building the system and show throughput, latency and compression measurements on synthetic and real-world traces, showcasing the benefits and trade-offs of our design. Sébastien Vaucher, Niloofar Yazdani, Pascal Felber, Daniel Enrique Lucani, Valerio Schiavoni |
CoNEXT | 5 |
| 2020 | LEGaTO: Low-Energy, Secure, and Resilient Toolset for Heterogeneous ComputingabstractThe LEGaTO project leverages task-based programming models to provide a software ecosystem for Made in-Europe heterogeneous hardware composed of CPUs, GPUs, FPGAs and dataflow engines. The aim is to attain one order of magnitude energy savings from the edge to the converged cloud/HPC, balanced with the security and resilience challenges. LEGaTO is an ongoing three-year EU H2020 project started in December 2017. Behzad Salami 0001, Konstantinos Parasyris, Adrián Cristal, Osman S. Unsal, Xavier Martorell, Raúl de la Cruz, Leonardo Arturo Bautista-Gomez, Daniel A. Jiménez, Carlos Álvarez 0001, Seyed Saber Nabavi Larimi, Sergi Madonar, Miquel Pericàs, Pedro Trancoso, Mustafa Abdul Jabbar, Jing Chen 0038, Pirah Noor Soomro, Madhavan Manivannan, Micha vor dem Berge, Stefan Krupop, Frank Klawonn, Al Mekhlafi, Sigrun May, Tobias Becker, Georgi Gaydadjiev, Hans Salomonsson, Devdatt P. Dubhashi, Oron Port, Yoav Etsion, Do Le Quoc, Christof Fetzer, Martin Kaiser, Nils Kucza, Jens Hagemeyer, René Griessl, Lennart Tigges, Kevin Mika, A. Hüffmeier, Marcelo Pasin, Valerio Schiavoni, Isabelly Rocha, Christian Göttel, Pascal Felber |
DATE | 40 |
| 2020 | Trust Management as a Service: Enabling Trusted Execution in the Face of Byzantine StakeholdersabstractTrust is arguably the most important challenge for critical services both deployed as well as accessed remotely over the network. These systems are exposed to a wide diversity of threats, ranging from bugs to exploits, active attacks, rogue operators, or simply careless administrators. To protect such applications, one needs to guarantee that they are properly configured and securely provisioned with the "secrets" (e.g., encryption keys) necessary to preserve not only the confidentiality, integrity and freshness of their data but also their code. Furthermore, these secrets should not be kept under the control of a single stakeholder—which might be compromised and would represent a single point of failure—and they must be protected across software versions in the sense that attackers cannot get access to them via malicious updates. Traditional approaches for solving these challenges often use ad hoc techniques and ultimately rely on a hardware security module (HSM) as root of trust. We propose a more powerful and generic approach to trust management that instead relies on trusted execution environments (TEEs) and a set of stakeholders as root of trust. Our system, PALÆMON, can operate as a managed service deployed in an untrusted environment, i.e., one can delegate its operations to an untrusted cloud provider with the guarantee that data will remain confidential despite not trusting any individual human (even with root access) nor system software. PALÆMON addresses in a secure, efficient and cost-effective way five main challenges faced when developing trusted networked applications and services. Our evaluation on a range of benchmarks and real applications shows that PALÆMON performs efficiently and can protect secrets of services without any change to their source code. Franz Gregor, Wojciech Ozga, Sébastien Vaucher, Rafael Pires 0001, Do Le Quoc, Sergei Arnautov, André Martin, Valerio Schiavoni, Pascal Felber, Christof Fetzer |
DSN | 8 |
| 2020 | Kollaps: decentralized and dynamic topology emulationabstractThe performance and behavior of large-scale distributed applications is highly influenced by network properties such as latency, bandwidth, packet loss, and jitter. For instance, an engineer might need to answer questions such as: What is the impact of an increase in network latency in application response time? How does moving a cluster between geographical regions affect application throughput? What is the impact of network dynamics on application stability? Currently, answering these questions in a systematic and reproducible way is very hard due to the variability and lack of control over the underlying network. Unfortunately, state-of-the-art network emulation or testbed environments do not scale beyond a single machine or small cluster (i.e., MiniNet), are focused exclusively on the control-plane (i.e., CrystalNet) or lack support for network dynamics (i.e., EmuLab). Paulo Gouveia, Carlos Segarra, Luca Liechti, Shady Issa, Valerio Schiavoni, Miguel Matos |
EuroSys | 6 |
| 2020 | TEEMon: A continuous performance monitoring framework for TEEsabstractTrusted Execution Environments (TEEs), such as Intel Software Guard eXtensions (SGX), are considered as a promising approach to resolve security challenges in clouds. TEEs protect the confidentiality and integrity of application code and data even against privileged attackers with root and physical access by providing an isolated secure memory area, i.e., enclaves. The security guarantees are provided by the CPU, thus even if system software is compromised, the attacker can never access the enclave's content. While this approach ensures strong security guarantees for applications, it also introduces a considerable runtime overhead in part by the limited availability of protected memory (enclave page cache). Currently, only a limited number of performance measurement tools for TEE-based applications exist and none offer performance monitoring and analysis during runtime. Robert Krahn, Donald Dragoti, Franz Gregor, Do Le Quoc, Valerio Schiavoni, Pascal Felber, Clenimar Souza, Andrey Brito, Christof Fetzer |
Middleware | 5 |
| 2020 | PipeTune: Pipeline Parallelism of Hyper and System Parameters Tuning for Deep Learning ClustersabstractDNN learning jobs are common in today's clusters due to the advances in AI driven services such as machine translation and image recognition. The most critical phase of these jobs for model performance and learning cost is the tuning of hyperparameters. Existing approaches make use of techniques such as early stopping criteria to reduce the tuning impact on learning cost. However, these strategies do not consider the impact that certain hyperparameters and systems parameters have on training time. This paper presents PipeTune, a framework for DNN learning jobs that addresses the trade-offs between these two types of parameters. PipeTune takes advantage of the high parallelism and recurring characteristics of such jobs to minimize the learning cost via a pipelined simultaneous tuning of both hyper and system parameters. Our experimental evaluation using three different types of workloads indicates that PipeTune achieves up to 22.6% reduction and 1.7× speed up on tuning and training time, respectively. PipeTune not only improves performance but also lowers energy consumption up to 29%. Isabelly Rocha, Nathaniel Morris, Lydia Y. Chen, Pascal Felber, Robert Birke, Valerio Schiavoni |
Middleware | 6 |
| 2020 | TZ4Fabric: Executing Smart Contracts with ARM TrustZone : (Practical Experience Report)abstractBlockchain technology promises to revolutionize manufacturing industries. For example, several supply chain use cases may benefit from transparent asset tracking and automated processes using smart contracts. Several real-world deployments exist where the transparency aspect of a blockchain is both an advantage and a disadvantage at the same time. The exposure of assets and business interaction represent critical risks. However, there are typically no confidentiality guarantees to protect the smart contract logic as well as the processed data. Trusted execution environments (TEE) are an emerging technology available in both edge or mobile-grade processors (e.g., ARM TrustZone) and server-grade processors (e.g., Intel SGX). TEEs shield both code and data from malicious attackers. This practical experience report presents TZ4FABRIC, an extension of Hyperledger Fabric to leverage ARM TrustZone for the secure execution of smart contracts. Our design minimizes the trusted computing base executed by avoiding the execution of a whole Hyperledger Fabric node inside the TEE, which continues to run in untrusted environment. Instead, we restrict it to the execution of only the smart contract. The TZ4FABRIC prototype exploits the opensource OP-TEE framework, as it supports deployments on cheap low-end devices (e.g., Raspberry Pis). Our experimental results highlight the performance trade-off due to the additional security guarantees provided by ARM TrustZone. TZ4FABRIC will be released as open source. Christina Müller, Marcus Brandenburger, Christian Cachin, Pascal Felber, Christian Göttel, Valerio Schiavoni |
SRDS | 6 |
| 2020 | MQT-TZ: Hardening IoT Brokers Using ARM TrustZone : (Practical Experience Report)abstractThe publish-subscribe paradigm is an efficient communication scheme with strong decoupling between the nodes, that is especially fit for large-scale deployments. It adapts natively to very dynamic settings and it is used in a diversity of real-world scenarios, including finance, smart cities, medical environments, or IoT sensors. Several of the mentioned application scenarios require increasingly stringent security guarantees due to the sensitive nature of the exchanged messages as well as the privacy demands of the clients/stakeholders/receivers. MQTT is a lightweight topic-based publish-subscribe protocol popular in edge and IoT settings, a de-facto standard widely adopted nowadays by the industry and researchers. However, MQTT brokers must process data in clear, hence exposing a large attack surface. This paper presents MQT-TZ, a secure MQTT broker leveraging ARM TRUSTZONE, a trusted execution environment (TEE) commonly found even on inexpensive devices largely available on the market (such as Raspberry Pi units). We define a mutual TLS-based handshake and a two-layer encryption for end-to-end security using the TEE as a trusted proxy. The experimental evaluation of our fully implemented prototype with micro-, macro-benchmarks, as well as with real-world industrial workloads from a MedTech use-case, highlights several tradeoffs using TRUSTZONE TEE. We report several lessons learned while building and evaluating our system. We release MQT-TZ as open-source. Carlos Segarra, Ricard Delgado-Gonzalo, Valerio Schiavoni |
SRDS | 3 |
| 2019 | On the Performance of ARM TrustZone - (Practical Experience Report)
Julien Amacher, Valerio Schiavoni |
DAIS | 2 |
| 2019 | Developing Secure Services for IoT with OP-TEE: A First Look at Performance and Usability
Christian Göttel, Pascal Felber, Valerio Schiavoni |
DAIS | 3 |
| 2019 | Using Trusted Execution Environments for Secure Stream Processing of Medical Data - (Case Study Paper)
Carlos Segarra, Ricard Delgado-Gonzalo, Mathieu Lemay, Pierre-Louis Aublin, Peter R. Pietzuch, Valerio Schiavoni |
DAIS | 6 |
| 2019 | Differential Approximation and Sprinting for Multi-Priority Big Data EnginesabstractToday's big data clusters based on the MapReduce paradigm are capable of executing analysis jobs with multiple priorities, providing differential latency guarantees. Traces from production systems show that the latency advantage of high-priority jobs comes at the cost of severe latency degradation of low-priority jobs as well as daunting resource waste caused by repetitive eviction and re-execution of low-priority jobs. We advocate a new resource management design that exploits the idea of differential approximation and sprinting. The unique combination of approximation and sprinting avoids the eviction of low-priority jobs and its consequent latency degradation and resource waste. To this end, we designed, implemented and evaluated DiAS, an extension of the Spark processing engine to support deflate jobs by dropping tasks and to sprint jobs. Our experiments on scenarios with two and three priority classes indicate that DiAS achieves up to 90% and 60% latency reduction for low- and high-priority jobs, respectively. DiAS not only eliminates resource waste but also (surprisingly) lowers energy consumption up to 30% at only a marginal accuracy loss for low-priority jobs. Robert Birke, Isabelly Rocha, Juan F. Pérez, Valerio Schiavoni, Pascal Felber, Lydia Y. Chen |
Middleware | 4 |
| 2019 | Heats: Heterogeneity-and Energy-Aware Task-Based SchedulingabstractCloud providers usually offer diverse types of hardware for their users. Customers exploit this option to deploy cloud instances featuring GPUs, FPGAs, architectures other than x86 (e.g., ARM, IBM Power8), or featuring certain specific extensions (e.g., Intel SGX). We consider in this work the instances used by customers to deploy containers, nowadays the de facto standard for micro-services, or to execute computing tasks. In doing so, the underlying container orchestrator (e.g., Kubernetes) should be designed so as to take into account and exploit this hardware diversity. In addition, besides the feature range provided by different machines, there is an often overlooked diversity in the energy requirements introduced by hardware heterogeneity, which is simply ignored by default container orchestrator's placement strategies. We introduce Heats, a new task-oriented and energy-aware orchestrator for containerized applications targeting heterogeneous clusters. Heats allows customers to trade performance vs. energy requirements. Our system first learns the performance and energy features of the physical hosts. Then, it monitors the execution of tasks on the hosts and opportunistically migrates them onto different cluster nodes to match the customer-required deployment trade-offs. Our Heats prototype is implemented within Google's Kubernetes. The evaluation with synthetic traces in our cluster indicate that our approach can yield considerable energy savings (up to 8.5%) and only marginally affect the overall runtime of deployed tasks (by at most 7%). Heats is released as open-source. Isabelly Rocha, Christian Göttel, Pascal Felber, Marcelo Pasin, Romain Rouvoy, Valerio Schiavoni |
PDP | 6 |
| 2019 | Blockchain-Based Metadata Protection for Archival SystemsabstractLong-term archival storage systems must protect data from powerful attackers that might try to corrupt or censor (part of) the documents. They must also protect the corresponding metadata information, which is essential to maintain and rebuild the stored data. In this practical experience report, we present metablock, a metadata protection system leveraging the Ethereum distributed ledger. We combine metablock with an existing secure long-term data archival system to provide a scalable design that allows external auditing, data validation and efficient data repair. We reflect on our experiences in using a blockchain for metadata protection, with the goal of providing valuable insights and lessons for developers of such secure systems, by highlighting the potential and limitations of the approach. Our prototype is available at https://github.com/ArnaudLhutereau/mb. Arnaud L'Hutereau, Dorian Burihabwa, Pascal Felber, Hugues Mercier, Valerio Schiavoni |
SRDS | 5 |
| 2019 | THUNDERSTORM: A Tool to Evaluate Dynamic Network Topologies on Distributed SystemsabstractNetwork dynamics, such as sudden changes in latency or available bandwidth, have a significant impact on the performance of distributed systems. While such dynamics are common, especially in WAN deployments, existing tools lack the capabilities to systematically evaluate the impact of such changes in real systems. We present THUNDERSTORM, a tool to evaluate the impact of dynamic network topologies on the performance of large-scale distributed systems. THUNDERSTORM is a fully functional tool that integrates with Kubernetes and can be used to evaluate off-the-shelf applications. THUNDERSTORM defines an easy-to-use language to describe arbitrarily complex network topologies and dynamic events used to enrich the default container composition descriptors. Our evaluation, using micro-and macro-benchmarks, as well as off-the-shelf unmodified systems (e.g., Apache Cassandra, MariaDB) shows that THUNDERSTORM is easy to use, accurate in reproducing dynamic behaviours and that it can help researchers uncover unexpected behaviours otherwise very costly to reproduce in real deployments typically captured only during malfunctioning periods. Luca Liechti, Paulo Gouveia, Peter G. Kropf, Miguel Matos, Valerio Schiavoni |
SRDS | 6 |
| 2019 | ABEONA: An Architecture for Energy-Aware Task Migrations from the Edge to the CloudabstractThis paper presents our preliminary results with ABEONA, an edge-to-cloud architecture that allows migrating tasks from low-energy, resource-constrained devices on the edge up to the cloud. Our preliminary results on artificial and real-world datasets show that it is possible to execute workloads in a more efficient manner energy-wise by scaling horizontally at the edge, without negatively affecting the execution runtime. Isabelly Rocha, Gabriel Vinha, Andrey Brito, Pascal Felber, Marcelo Pasin, Valerio Schiavoni |
SRDS | 6 |
| 2019 | iperfTZ: Understanding Network Bottlenecks for TrustZone-Based Trusted Applications
Christian Göttel, Pascal Felber, Valerio Schiavoni |
SSS | 3 |
| 2018 | LEGaTO: towards energy-efficient, secure, fault-tolerant toolset for heterogeneous computingabstractLEGaTO is a three-year EU H2020 project which started in December 2017. The LEGaTO project will leverage task-based programming models to provide a software ecosystem for Made-in-Europe heterogeneous hardware composed of CPUs, GPUs, FPGAs and dataflow engines. The aim is to attain one order of magnitude energy savings from the edge to the converged cloud/HPC. Adrián Cristal, Osman S. Unsal, Xavier Martorell, Raúl de la Cruz, Leonardo Arturo Bautista-Gomez, Daniel Jiménez-González, Carlos Álvarez 0001, Behzad Salami 0001, Sergi Madonar, Miquel Pericàs, Pedro Trancoso, Micha vor dem Berge, Gunnar Billung-Meyer, Stefan Krupop, Wolfgang Christmann, Frank Klawonn, Amani Mihklafi, Tobias Becker, Georgi Gaydadjiev, Hans Salomonsson, Devdatt P. Dubhashi, Oron Port, Yoav Etsion, Vesna Nowack, Christof Fetzer, Jens Hagemeyer, Thorsten Jungeblut, Nils Kucza, Martin Kaiser, Mario Porrmann, Marcelo Pasin, Valerio Schiavoni, Isabelly Rocha, Christian Göttel, Pascal Felber |
CF | 33 |
| 2018 | SGX-FS: Hardening a File System in User-Space with Intel SGXabstractFile systems have long benefited from hardware acceleration to improve their performance. In order to leverage such hardware capabilities, file systems rely on direct and trusted support from the underlying operating system. However, this assumes that the OS and the associated kernel drivers, which access the accelerators, are trustworthy. The recent introduction of the Intel software guard extensions (SGX) instruction set allows application developers to lift part of these assumptions, in conjunction with the widespread availability of these new extensions in mass-market CPUs. With SGX, programmers can design secure applications under a stronger adversarial model, such as a compromised OS or kernel module. Code executes inside enclaves and is protected from privileged processes, including the OS itself. This paper presents SGX-FS, a new user-space file system that leverages SGX data sealing capabilities for secure in-memory and persistent storage. It combines the FUSE framework with SGX to securely protect user data. In particular, SGX-FS efficiently encrypts and decrypts the application data within the enclaves. We fully implement an open-source SGX-FS prototype and evaluate its performance by means of a representative set of nano-and micro-benchmarks. Dorian Burihabwa, Pascal Felber, Hugues Mercier, Valerio Schiavoni |
CloudCom | 4 |
| 2018 | Boosting Transactional Memory with Stricter Serializability
Pierre Sutra, Patrick Marlier, Valerio Schiavoni, François Trahay |
COORDINATION | 3 |
| 2018 | RECAST: Random Entanglement for Censorship-Resistant Archival STorageabstractUsers entrust an increasing amount of data to online cloud systems for archival purposes. Existing storage systems designed to preserve user data unaltered for decades do not, however, provide strong security guarantees - at least at a reasonable cost. This paper introduces RECAST, an anti-censorship data archival system based on random data entanglement. Documents are mixed together using an entanglement scheme that exploits erasure codes for secure and tamper-proof long-term archival. Data is intertwined in such a way that it becomes virtually impossible to delete a specific document that has been stored long enough in the system, without also erasing a substantial fraction of the whole archive, which requires a very powerful adversary and openly exposes the attack. We validate RECAST entanglement approach via simulations and we present and evaluate a full-fledged prototype deployed in a local cluster. In one of our settings, we show that RECAST, configured with the same storage overhead as triple replication, can withstand 10% of storage node failures without any data loss. Furthermore, we estimate that the effort required from a powerful censor to delete a specific target document is two orders of magnitude larger than for triple replication. Roberta Barbi, Dorian Burihabwa, Pascal Felber, Hugues Mercier, Valerio Schiavoni |
DSN | 5 |
| 2018 | EndBox: Scalable Middlebox Functions Using Client-Side Trusted ExecutionabstractMany organisations enhance the performance, security, and functionality of their managed networks by deploying middleboxes centrally as part of their core network. While this simplifies maintenance, it also increases cost because middlebox hardware must scale with the number of clients. A promising alternative is to outsource middlebox functions to the clients themselves, thus leveraging their CPU resources. Such an approach, however, raises security challenges for critical middlebox functions such as firewalls and intrusion detection systems. We describe EndBox, a system that securely executes middlebox functions on client machines at the network edge. Its design combines a virtual private network (VPN) with middlebox functions that are hardware-protected by a trusted execution environment (TEE), as offered by Intel's Software Guard Extensions (SGX). By maintaining VPN connection endpoints inside SGX enclaves, EndBox ensures that all client traffic, including encrypted communication, is processed by the middlebox. Despite its decentralised model, EndBox's middlebox functions remain maintainable: they are centrally controlled and can be updated efficiently. We demonstrate EndBox with two scenarios involving (i) a large company; and (ii) an Internet service provider that both need to protect their network and connected clients. We evaluate EndBox by comparing it to centralised deployments of common middlebox functions, such as load balancing, intrusion detection, firewalling, and DDoS prevention. We show that EndBox achieves up to 3.8x higher throughput and scales linearly with the number of clients. David Goltzsche, Signe Rüsch, Manuel Nieke, Sébastien Vaucher, Nico Weichbrodt, Valerio Schiavoni, Pierre-Louis Aublin, Paolo Costa, Christof Fetzer, Pascal Felber, Peter R. Pietzuch, Rüdiger Kapitza |
DSN | 6 |
| 2018 | CYCLOSA: Decentralizing Private Web Search through SGX-Based Browser ExtensionsabstractBy regularly querying Web search engines, users (unconsciously) disclose large amounts of their personal data as part of their search queries, among which some might reveal sensitive information (e.g. health issues, sexual, political or religious preferences). Several solutions exist to allow users querying search engines while improving privacy protection. However, these solutions suffer from a number of limitations: some are subject to user re-identification attacks, while others lack scalability or are unable to provide accurate results. This paper presents CYCLOSA, a secure, scalable and accurate private Web search solution. CYCLOSA improves security by relying on trusted execution environments (TEEs) as provided by Intel SGX. Further, CYCLOSA proposes a novel adaptive privacy protection solution that reduces the risk of user re-identification. CYCLOSA sends fake queries to the search engine and dynamically adapts their count according to the sensitivity of the user query. In addition, CYCLOSA meets scalability as it is fully decentralized, spreading the load for distributing fake queries among other nodes. Finally, CYCLOSA achieves accuracy of Web search as it handles the real query and the fake queries separately, in contrast to other existing solutions that mix fake and real query results. Rafael Pires 0001, David Goltzsche, Sonia Ben Mokhtar, Sara Bouchenak, Antoine Boutet, Pascal Felber, Rüdiger Kapitza, Marcelo Pasin, Valerio Schiavoni |
ICDCS | 9 |
| 2018 | SGX-Aware Container Orchestration for Heterogeneous ClustersabstractContainers are becoming the de facto standard to package and deploy applications and micro-services in the cloud. Several cloud providers (e.g., Amazon, Google, Microsoft) begin to offer native support on their infrastructure by integrating container orchestration tools within their cloud offering. At the same time, the security guarantees that containers offer to applications remain questionable. Customers still need to trust their cloud provider with respect to data and code integrity. The recent introduction by Intel of Software Guard Extensions (SGX) into the mass market offers an alternative to developers, who can now execute their code in a hardware-secured environment without trusting the cloud provider. This paper provides insights regarding the support of SGX inside Kubernetes, an industry-standard container orchestrator. We present our contributions across the whole stack supporting execution of SGX-enabled containers. We provide details regarding the architecture of the scheduler and its monitoring framework, the underlying operating system support and the required kernel driver extensions. We evaluate our complete implementation on a private cluster using the real-world Google Borg traces. Our experiments highlight the performance trade-offs that will be encountered when deploying SGX-enabled micro-services in the cloud. Sébastien Vaucher, Rafael Pires 0001, Pascal Felber, Marcelo Pasin, Valerio Schiavoni, Christof Fetzer |
ICDCS | 5 |
| 2018 | PubSub-SGX: Exploiting Trusted Execution Environments for Privacy-Preserving Publish/Subscribe SystemsabstractThis paper presents PUBSUB-SGX, a content-based publish-subscribe system that exploits trusted execution environments (TEEs), such as Intel SGX, to guarantee confidentiality and integrity of data as well as anonymity and privacy of publishers and subscribers. We describe the technical details of our Python implementation, as well as the required system support introduced to deploy our system in a container-based runtime. Our evaluation results show that our approach is sound, while at the same time highlighting the performance and scalability trade-offs. In particular, by supporting just-in-time compilation inside of TEEs, Python programs inside of TEEs are in general faster than when executed natively using standard CPython. Sergei Arnautov, Andrey Brito, Pascal Felber, Christof Fetzer, Franz Gregor, Robert Krahn, Wojciech Ozga, André Martin, Valerio Schiavoni, Marcus Tenorio, Nikolaus Thummel |
SRDS | 9 |
| 2018 | Security, Performance and Energy Trade-Offs of Hardware-Assisted Memory Protection MechanismsabstractThe deployment of large-scale distributed systems, e.g., publish-subscribe platforms, that operate over sensitive data using the infrastructure of public cloud providers, is nowadays heavily hindered by the surging lack of trust toward the cloud operators. Although purely software-based solutions exist to protect the confidentiality of data and the processing itself, such as homomorphic encryption schemes, their performance is far from being practical under real-world workloads. The performance trade-offs of two novel hardware-assisted memory protection mechanisms, namely AMD SEV and Intel SGX - currently available on the market to tackle this problem, are ADD described in this practical experience. Specifically, we implement and evaluate a publish/subscribe use-case and evaluate the impact of the memory protection mechanisms and the resulting performance. This paper reports on the experience gained while building this system, in particular when having to cope with the technical limitations imposed by SEV and SGX. Several tradeoffs that provide valuable insights in terms of latency, throughput, processing time and energy requirements are exhibited by means of micro-and macro-benchmarks. Christian Göttel, Rafael Pires 0001, Isabelly Rocha, Sébastien Vaucher, Pascal Felber, Marcelo Pasin, Valerio Schiavoni |
SRDS | 7 |
| 2018 | Security, Performance and Energy Implications of Hardware-Assisted Memory Protection Mechanisms on Event-Based Streaming SystemsabstractMajor cloud providers such as Amazon [1], Google [2] and Microsoft [3] provide nowadays some form of infrastructure as a service (IaaS) which allows deploying services in the form of virtual machines [4], containers [5] or bare-metal [6] instances. Although software-based solutions like homomorphic encryption exit, privacy concerns [7] greatly hinder the deployment of such services over public clouds. It is particularly difficult for homomorphic encryption to match performance requirements of modern workloads [8]. Evaluating simple operations on basic data types with HElib [9], a homomorphic encryption library, against their unencrypted counter part reveals, that homomorphic encryption is still impractical under realistic workloads. Christian Göttel, Rafael Pires 0001, Isabelly Rocha, Sébastien Vaucher, Pascal Felber, Marcelo Pasin, Valerio Schiavoni |
SRDS | 7 |
| 2017 | Block Placement Strategies for Fault-Resilient Distributed Tuple Spaces: An Experimental Study - (Practical Experience Report)
Roberta Barbi, Vitaly Buravlev, Claudio Antares Mezzina, Valerio Schiavoni |
DAIS | 4 |
| 2017 | SecureCloud: Secure big data processing in untrusted cloudsabstractWe present the SecureCloud EU Horizon 2020 project, whose goal is to enable new big data applications that use sensitive data in the cloud without compromising data security and privacy. For this, SecureCloud designs and develops a layered architecture that allows for (i) the secure creation and deployment of secure micro-services; (ii) the secure integration of individual micro-services to full-fledged big data applications; and (iii) the secure execution of these applications within untrusted cloud environments. To provide security guarantees, SecureCloud leverages novel security mechanisms present in recent commodity CPUs, in particular, Intel's Software Guard Extensions (SGX). SecureCloud applies this architecture to big data applications in the context of smart grids. We describe the SecureCloud approach, initial results, and considered use cases. Florian Kelbert, Franz Gregor, Rafael Pires 0001, Stefan Köpsell, Marcelo Pasin, Aurelien Havet, Valerio Schiavoni, Pascal Felber, Christof Fetzer, Peter R. Pietzuch |
DATE | 7 |
| 2017 | GENPACK: A Generational Scheduler for Cloud Data CentersabstractCloud data centers largely rely on virtualization to provision resources and host services across their infrastructure. The scheduling problem has been widely studied and is well understood when the resource requirements and the expected lifetime of services are known beforehand. In contrast, when workloads are not known in advance, effective scheduling of services, and more generally system containers, becomes much more complex. In this paper, we propose GENPACK, a framework for system containers scheduling in cloud data centers that leverages principles from generational garbage collection (GC). It combines runtime monitoring of system containers to learn their requirements and properties, and a scheduler that manages different generations of servers. The population of these generations may vary over time depending on the global load, hence they are subject to being shut down when idle to save energy. We implemented GENPACK and tested it in a dedicated data center, showing that it can be up to 23% more energy-efficient that SWARM's built-in scheduling policies on a real-world trace. Aurelien Havet, Valerio Schiavoni, Pascal Felber, Maxime Colmant, Romain Rouvoy, Christof Fetzer |
IC2E | 2 |
| 2017 | Introducing SECURESTREAMS: Scalable Middleware for Reactive and Secure Data Stream ProcessingabstractWe introduce SECURESTREAMS, a middleware framework for secure stream processing. Its design builds on Intel's Secure Guard Extensions (SGX) to guarantee the privacy and the integrity of the data being processed. Our initial experimental results of SECURESTREAMS are promising: the framework is easy to use, and delivers high throughput, enabling developers to implement complex processing pipelines in a few lines of scripting code. Aurelien Havet, Valerio Schiavoni, Pascal Felber, Romain Rouvoy |
IC2E | 2 |
| 2017 | X-search: revisiting private web search using intel SGXabstractThe exploitation of user search queries by search engines is at the heart of their economic model. As consequence, offering private Web search functionalities is essential to the users who care about their privacy. Nowadays, there exists no satisfactory approach to enable users to access search engines in a privacy-preserving way. Existing solutions are either too costly due to the heavy use of cryptographic mechanisms (e.g., private information retrieval protocols), subject to attacks (e.g., Tor, TrackMeNot, GooPIR) or rely on weak adversarial models (e.g., PEAS). This paper introduces X-Search, a novel private Web search mechanism building on the disruptive Software Guard Extensions (SGX) proposed by Intel. We compare X-Search to its closest competitors, Tor and PEAS, using a dataset of real web search queries. Our evaluation shows that: (1) X-Search offers stronger privacy guarantees than its competitors as it operates under a stronger adversarial model; (2) it better resists state-of-the-art re-identification attacks; and (3) from the performance perspective, X-Search outperforms its competitors both in terms of latency and throughput by orders of magnitude. Sonia Ben Mokhtar, Antoine Boutet, Pascal Felber, Marcelo Pasin, Rafael Pires 0001, Valerio Schiavoni |
Middleware | 6 |
| 2017 | SafeFS: a modular architecture for secure user-space file systems: one FUSE to rule them allabstractThe exponential growth of data produced, the ever faster and ubiquitous connectivity, and the collaborative processing tools lead to a clear shift of data stores from local servers to the cloud. This migration occurring across different application domains and types of users---individual or corporate---raises two immediate challenges. First, out-sourcing data introduces security risks, hence protection mechanisms must be put in place to provide guarantees such as privacy, confidentiality and integrity. Second, there is no "one-size-fits-all" solution that would provide the right level of safety or performance for all applications and users, and it is therefore necessary to provide mechanisms that can be tailored to the various deployment scenarios. Rogerio Pontes, Dorian Burihabwa, Francisco Maia 0001, João Paulo 0001, Valerio Schiavoni, Pascal Felber, Hugues Mercier, Rui Oliveira 0001 |
SYSTOR | 5 |
| 2016 | A Performance Evaluation of Erasure Coding Libraries for Cloud-Based Data Stores - (Practical Experience Report)abstractErasure codes have been widely used over the last decade to implement reliable data stores. They offer interesting trade-offs between efficiency, reliability, and storage overhead. Indeed, a distributed data store holding encoded data blocks can tolerate the failure of multiple nodes while requiring only a fraction of the space necessary for plain replication, albeit at an increased encoding and decoding cost. There exists nowadays a number of libraries implementing several variations of erasure codes, which notably differ in terms of complexity and implementation-specific optimizations. Seven years ago, Plank et al. [ 14 ] have conducted a comprehensive performance evaluation of open-source erasure coding libraries available at the time to compare their raw performance and measure the impact of different parameter configurations. In the present experimental study, we take a fresh perspective at the state of the art of erasure coding libraries. Not only do we cover a wider set of libraries running on modern hardware, but we also consider their efficiency when used in realistic settings for cloud-based storage, namely when deployed across several nodes in a data centre. Our measurements therefore account for the end-to-end costs of data accesses over several distributed nodes, including the encoding and decoding costs, and shed light on the performance one can expect from the various libraries when deployed in a real system. Our results reveal important differences in the efficiency of the different libraries, notably due to the type of coding algorithm and the use of hardware-specific optimizations. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Dorian Burihabwa, Pascal Felber, Hugues Mercier, Valerio Schiavoni |
DAIS | 4 |
| 2016 | Evaluating the Cost and Robustness of Self-organizing Distributed Hash TablesabstractSelf-organizing construction principles are a natural fit for large-scale distributed system in unpredictable deployment environments. These principles allow a system to systematically converge to a global state by means of simple, uncoordinated actions by individual peers. Indexing services based on the distributed hash table (DHT) abstraction have been established as a solid foundation for large-scale distributed applications. For most DHTs, the creation and maintenance of the overlay structure relies on the exploration and update of an already stabilized structure. We evaluate in this paper the practical interest of self-organizing principles, and in particular gossip-based overlay construction protocols, to bootstrap and maintain various DHT implementations. Based on the seminal work on T-Chord, a self-organizing version of Chord using the T-Man overlay construction service, we contribute three additional self-organizing DHTs: T-Pastry, T-Kademlia and T-Kelips. We conduct an experimental evaluation of the cost and performance of each of these designs using a prototype implementation. Our conclusion is that, while providing equivalent performance in a stabilized system, self-organizing DHTs are able to sustain and recover from higher level of churn than their explicitly-created counterparts, and should therefore be considered as a method of choice for deploying robust indexing layers in adverse environments. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Sveta Krasikova, Raziel Carvajal-Gomez, Heverson B. Ribeiro, Etienne Rivière, Valerio Schiavoni |
DAIS | 5 |
| 2016 | On the Cost of Safe Storage for Public Clouds: An Experimental EvaluationabstractCloud-based storage services such as Dropbox, Google Drive and OneDrive are increasingly popular for storing enterprise data, and they have already become the de facto choice for cloud-based backup of hundreds of millions of regular users. Drawn by the wide range of services they provide, no upfront costs and 24/7 availability across all personal devices, customers are well-aware of the benefits that these solutions can bring. However, most users tend to forget—or worse ignore—some of the main drawbacks of such cloud-based services, namely in terms of privacy. Data entrusted to these providers can be leaked by hackers, disclosed upon request from a governmental agency's subpoena, or even accessed directly by the storage providers (e.g., for commercial benefits). While there exist solutions to prevent or alleviate these problems, they typically require direct intervention from the clients, like encrypting their data before storing it, and reduce the benefits provided such as easily sharing data between users. This practical experience report studies a wide range of security mechanisms that can be used atop standard cloud-based storage services. We present the details of our evaluation testbed and discuss the design choices that have driven its implementation. We evaluate several state-of-the-art techniques with varying security guarantees responding to user-assigned security and privacy criteria. Our results reveal the various trade-offs of the different techniques by means of representative workloads on top of industry-grade storage services. Dorian Burihabwa, Rogerio Pontes, Pascal Felber, Francisco Maia 0001, Hugues Mercier, Rui Oliveira 0001, João Paulo 0001, Valerio Schiavoni |
SRDS | 8 |
| 2016 | GlobalFS: A Strongly Consistent Multi-site File SystemabstractThis paper introduces GlobalFS, a POSIX-compliant geographically distributed file system. GlobalFS builds on two fundamental building blocks, an atomic multicast group communication abstraction and multiple instances of a single-site data store. We define four execution modes and show how all file system operations can be implemented with these modes while ensuring strong consistency and tolerating failures. We describe the GlobalFS prototype in detail and report on an extensive performance assessment. We have deployed GlobalFS across all EC2 regions and show that the system scales geographically, providing performance comparable to other state-of-the-art distributed file systems for local commands and allowing for strongly consistent operations over the whole system. The code of GlobalFS is available as open source. Leandro Pacheco de Sousa, Raluca Halalai, Valerio Schiavoni, Fernando Pedone, Etienne Rivière, Pascal Felber |
SRDS | 3 |
| 2015 | UniCrawl: A Practical Geographically Distributed Web CrawlerabstractAs the wealth of information available on the web keeps growing, being able to harvest massive amounts of data has become a major challenge. Web crawlers are the core components to retrieve such vast collections of publicly available data. The key limiting factor of any crawler architecture is however its large infrastructure cost. To reduce this cost, and in particular the high upfront investments, we present in this paper a geo-distributed crawler solution, UniCrawl. UniCrawl orchestrates several geographically distributed sites. Each site operates an independent crawler and relies on well-established techniques for fetching and parsing the content of the web. UniCrawl splits the crawled domain space across the sites and federates their storage and computing resources, while minimizing thee inter-site communication cost. To assess our design choices, we evaluate UniCrawl in a controlled environment using the ClueWeb12 dataset, and in the wild when deployed over several remote locations. We conducted several experiments over 3 sites spread across Germany. When compared to a centralized architecture with a crawler simply stretched over several locations, UniCrawl shows a performance improvement of 93.6% in terms of network bandwidth consumption, and a speedup factor of 1.75. Do Le Quoc, Christof Fetzer, Pascal Felber, Etienne Rivière, Valerio Schiavoni, Pierre Sutra |
CLOUD | 5 |
| 2014 | Autonomous Multi-dimensional Slicing for Large-Scale Distributed Systems
Mathieu Pasquet, Francisco Maia 0001, Etienne Rivière, Valerio Schiavoni |
DAIS | 4 |
| 2014 | LAYSTREAM: Composing standard gossip protocols for live video streamingabstractGossip-based live streaming is a popular topic, as attested by the vast literature on the subject. Despite the particular merits of each proposal, all need to implement and deal with common challenges such as membership management, topology construction and video packets dissemination. Well-principled gossip-based protocols have been proposed in the literature for each of these aspects. Our goal is to assess the feasibility of building a live streaming system, LAYSTREAM, as a composition of these existing protocols, to deploy the resulting system on real testbeds, and report on lessons learned in the process. Unlike previous evaluations conducted by simulations and considering each protocol independently, we use real deployments. We evaluate protocols both independently and as a layered composition, and unearth specific problems and challenges associated with deployment and composition. We discuss and present solutions for these, such as a novel topology construction mechanism able to cope with the specificities of a large-scale and delay-sensitive environment, but also with requirements from the upper layer. Our implementation and data are openly available to support experimental reproducibility. Miguel Matos, Valerio Schiavoni, Etienne Rivière, Pascal Felber, Rui Oliveira 0001 |
P2P | 2 |
| 2014 | On the Support of Versioning in Distributed Key-Value StoresabstractThe ability to access and query data stored in multiple versions is an important asset for many applications, such as Web graph analysis, collaborative editing platforms, data forensics, or correlation mining. The storage and retrieval of versioned data requires a specific API and support from the storage layer. The choice of the data structures used to maintain versioned data has a fundamental impact on the performance of insertions and queries. The appropriate data structure also depends on the nature of the versioned data and the nature of the access patterns. In this paper we study the design and implementation space for providing versioning support on top of a distributed key-value store (KVS). We define an API for versioned data access supporting multiple writers and show that a plain KVS does not offer the necessary synchronization power for implementing this API. We leverage the support for listeners at the KVS level and propose a general construction for implementing arbitrary types of data structures for storing and querying versioned data. We explore the design space of versioned data storage ranging from a flat data structure to a distributed sharded index. The resulting system, ALEPH, is implemented on top of an industrial-grade open-source KVS, Infinispan. Our evaluation, based on real-world Wikipedia access logs, studies the performance of each versioning mechanisms in terms of load balancing, latency and storage overhead in the context of different access scenarios. Pascal Felber, Marcelo Pasin, Etienne Rivière, Valerio Schiavoni, Pierre Sutra, Fábio Coelho 0001, Rui Oliveira 0001, Miguel Matos, Ricardo Vilaça |
SRDS | 4 |
| 2013 | SplayNet: Distributed User-Space Topology Emulation
Valerio Schiavoni, Etienne Rivière, Pascal Felber |
Middleware | 1 |
| 2013 | Lightweight, efficient, robust epidemic dissemination
Miguel Matos, Valerio Schiavoni, Pascal Felber, Rui Oliveira 0001, Etienne Rivière |
J. Parallel Distributed Comput. | 2 |
| 2013 | CoFeed: privacy-preserving Web search recommendation based on collaborative aggregation of interest feedbackabstractSUMMARY Search engines essentially rely on the structure of the graph of hyperlinks. Although accurate for the main trend, this is not effective when some query is ambiguous. Leveraging semantic information by the mean of interest matching allows proposing complementary results that are tailored to the user's expectations. This paper proposes a collaborative search companion system, CoFeed, that collects user search queries and that considers feedback to build user‐centric and document‐centric profiling information. Over time, the system constructs ranked collections of elements that maintain the required information diversity and enhance the user search experience by presenting additional results tailored to the user's interest space. This collaborative search companion requires a supporting architecture adapted to large user populations generating high request loads. To that end, it integrates mechanisms for ensuring scalability and load balancing of the service under varying loads and user interest distributions. Moreover, collecting the recommendation data poses the problem of users’ privacy, and the bias one peer can induce to the system by sending fake recommendations. To that end, CoFeed ensures both publisher anonymity and rate limitation. With the former, the origin of the data is never known by the server that processes it, even if several servers collude to spy on some user. The latter, combined with decoupled authentication, allows to minimize the influence of cheating peers sending fake recommendations. Experiments with a deployed prototype highlight the efficiency of the system by analyzing improvement in search relevance, computational cost, scalability and load balancing. Copyright © 2011 John Wiley & Sons, Ltd. Pascal Felber, Peter G. Kropf, Lorenzo Leonini, Toan Luu, Martin Rajman, Etienne Rivière, Valerio Schiavoni, José Valerio |
Softw. Pract. Exp. | 7 |
| 2012 | BRISA: Combining Efficiency and Reliability in Epidemic Data DisseminationabstractThere is an increasing demand for efficient and robust systems able to cope with today's global needs for intensive data dissemination, e.g., media content or news feeds. Unfortunately, traditional approaches tend to focus on one end of the efficiency/robustness design spectrum, by either leveraging rigid structures such as trees to achieve efficient distribution, or using loosely-coupled epidemic protocols to obtain robustness. In this paper we present BRISA, a hybrid approach combining the robustness of epidemic-based dissemination with the efficiency of tree-based structured approaches. This is achieved by having dissemination structures such as trees implicitly emerge from an underlying epidemic substrate by a judicious selection of links. These links are chosen with local knowledge only and in such a way that the completeness of data dissemination is not compromised, i.e., the resulting structure covers all nodes. Failures are treated as an integral part of the system as the dissemination structures can be promptly compensated and repaired thanks to the underlying epidemic substrate. Besides presenting the protocol design, we conduct an extensive evaluation in a real environment, analyzing the effectiveness of the structure creation mechanism and its robustness under faults and churn. Results confirm BRISA as an efficient and robust approach to data dissemination in the large scale. Miguel Matos, Valerio Schiavoni, Pascal Felber, Rui Oliveira 0001, Etienne Rivière |
IPDPS | 2 |
| 2012 | A component-based middleware platform for reconfigurable service-oriented architecturesabstractSUMMARY ThetextitService Component Architecture (SCA) is a technology‐independent standard for developing distributed Service‐oriented Architectures (SOA). The SCA standard promotes the use of components and architecture descriptors, and mostly covers the lifecycle steps of implementation and deployment. Unfortunately, SCA does not address the governance of SCA applications and provides no support for the maintenance of deployed components. This article covers this issue and introduces the F RA SCA TI platform, a run‐time support for SCA with dynamic reconfiguration capabilities and run‐time management features. This article presents the internal component‐based architecture of the F RA SCA TI platform, and highlights its key features. The component‐based design of the F RA SCA TI platform introduces many degrees of flexibility and configurability in the platform itself and it can host the SOA applications. This article reports on micro‐benchmarks highlighting that run‐time manageability in the F RA SCA TI platform does not decrease its performance when compared with the de facto reference SCA implementation: Apache T USCANY . Finally, a smart home scenario illustrates the extension capabilities and the various reconfigurations of the F RA SCA TI platform. Copyright © 2011 John Wiley & Sons, Ltd. Lionel Seinturier, Philippe Merle, Romain Rouvoy, Daniel Romero 0002, Valerio Schiavoni, Jean-Bernard Stefani |
Softw. Pract. Exp. | 5 |
| 2011 | WHISPER: Middleware for Confidential Communication in Large-Scale NetworksabstractA wide range of distributed applications requires some form of confidential communication between groups of users. In particular, the messages exchanged between the users and the identity of group members should not be visible to external observers. Classical approaches to confidential group communication rely upon centralized servers, which limit scalability and represent single points of failure. In this paper, we present WHISPER, a fully decentralized middleware that supports confidential communications within groups of nodes in large-scale systems. It builds upon a peer sampling service that takes into account network limitations such as NAT and firewalls. WHISPER implements confidentiality in two ways: it protects the content of messages exchanged between the members of a group, and it keeps the group memberships secret to external observers. Using multi-hops paths allows these guarantees to hold even if attackers can observe the link between two nodes, or be used as content relays for NAT bypassing. Evaluation in real-world settings indicates that the price of confidentiality remains reasonable in terms of network load and processing costs. Valerio Schiavoni, Etienne Rivière, Pascal Felber |
ICDCS | 1 |
| 2011 | Exploiting Node Connection Regularity for DHT ReplicationabstractDistributed Hash-Tables (DHTs) provide an efficient way to store objects in large-scale peer-to-peer systems. To guarantee that objects are reliably stored, DHTs rely on replication. Several replication strategies have been proposed in the last years. The most efficient ones use predictions about the availability of nodes to reduce the number of object migrations that need to be performed: objects are preferably stored on highly available nodes. This paper proposes an alternative replication strategy. Rather than exploiting highly available nodes, we propose to leverage nodes that exhibit regularity in their connection pattern. Roughly speaking, the strategy consists in replicating each object on a set of nodes that is built in such a way that, with high probability, at any time, there are always at least k nodes in the set that are available. We evaluate this replication strategy using traces of two real-world systems: eDonkey and Skype. The evaluation shows that our regularity-based replication strategy induces a systematically lower network usage than existing state of the art replication strategies. Alessio Pace, Vivien Quéma, Valerio Schiavoni |
SRDS | 3 |
| 2010 | SPADS: Publisher Anonymization for DHT StorageabstractMany distributed applications, such as collaborative Web mapping, collaborative feedback and ranking, or bug reporting systems, rely on the aggregation of privacy-sensitive information gathered from human users. This information is typically aggregated at servers and later used as the basis for some collaborative service. Expecting that clients trust that the user-centric information will not be used for malevolent purposes is not realistic in a fully distributed setting where nodes are not under the control of a single administrative domain. Moreover, most of the time the origin of the data is of small importance when computing the aggregation onto which these services are based. Trust problems can be evinced by ensuring that the identity of the user is dropped before the data can actually be used, a process called publisher anonymization. Such a property shall be guaranteed even if a set of servers is colluding to spy on some user. This also requires that malevolent users cannot harm the service by sending any number of items without being traceable due to publisher anonymization. Rate limitation and decoupled authentication are the two mechanisms that ensure that these cheating users have a limited impact on the system. This paper presents SPADS, a system that interfaces to any DHT and supports the three objectives of publisher anonymization, rate limitation and decoupled authentication. The evaluation of a deployed prototype on a cluster assesses its performance and small footprint. Pascal Felber, Martin Rajman, Etienne Rivière, Valerio Schiavoni, José Valerio |
Peer-to-Peer Computing | 4 |
| 2009 | NAT-resilient Gossip Peer SamplingabstractGossip peer sampling protocols now represent a solid basis to build and maintain peer to peer (p2p) overlay networks. They provide peers with a random sample of the network and maintain connectivity in highly dynamic settings. They rely on the assumption that, at any time, each peer is able to communicate with any other peer. Yet, this ignores the fact that there is a significant proportion of peers that now sit behind NAT devices, preventing direct communication without specific mechanisms. In this paper, we propose a NAT-resilient gossip peer sampling protocol called Nylon, that accounts for the presence of NATs. Nylon is fully decentralized and spreads evenly among peers the extra load caused by the presence of NATs. Nylon ensures that a peer can always communicate with any peer in its sample. This is achieved through a simple, yet efficient mechanism, establishing a path of relays between peers. Our results show that the randomness of the generated samples is preserved, and that the connectivity is not impacted even in the presence of high churn and a high ratio of peers sitting behind NAT devices. Anne-Marie Kermarrec, Alessio Pace, Vivien Quéma, Valerio Schiavoni |
ICDCS | 4 |