EDBT 2026 Demo / reviewers in the wild / expert
Peter R. Pietzuch
dblp:45/4887 · also Peter Robert Pietzuch
· DBLP profile ↗
84ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-6963-5640ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 7 since 2021Databases, data management, data science and information retrieval · 20 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 5 since 2021Computer networks · 12 · 2 since 2021Security and privacy · 4Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHARM: Chiplet Heterogeneity-Aware Runtime Mapping SystemabstractPublisher Copyright: © 2026 Copyright held by the owner/author(s) Alessandro Fogli, Bo Zhao 0019, Peter R. Pietzuch, Jana Giceva |
EuroSys | 3 |
| 2025 | GRANNY: Granular Management of Compute-Intensive Applications in the Cloud
Carlos Segarra, Simon Shillaker, Eleftheria Mappoura, Rodrigo Bruno, Lluís Vilanova, Peter R. Pietzuch |
NSDI | 7 |
| 2025 | Tempo: Compiled Dynamic Deep Learning with Symbolic Dependence GraphsabstractDeep learning (DL) algorithms are often defined in terms of temporal relationships: a tensor at one timestep may depend on tensors from earlier or later timesteps. Such dynamic dependencies (and corresponding dynamic tensor shapes) are difficult to express and optimize: while eager DL systems support such dynamism, they cannot apply compiler-based optimizations; graph-based systems require static tensor shapes, which forces users to pad tensors or break-up programs into multiple static graphs. Pedro F. Silvestre, Peter R. Pietzuch |
SOSP | 2 |
| 2025 | Front Matter
Themis Palpanas, Peter R. Pietzuch, Nesime Tatbul, Peter Triantafillou |
Proc. VLDB Endow. | 2 |
| 2024 | Is It Time To Put Cold Starts In The Deep Freeze?abstractCold-start times have been the "end-all, be-all" metric for research in serverless cloud computing over the past decade. Reducing the impact of cold starts matters, because they can be the biggest contributor to a serverless function's end-to-end execution time. Recent studies from cloud providers, however, indicate that, in practice, a majority of serverless functions are triggered by non-interactive workloads. To substantiate this, we study the types of serverless functions used in 35 publications and find that over 80% of functions are not semantically latency sensitive. If a function is non-interactive and latency insensitive, is end-to-end execution time the right metric to optimize in serverless? What if cold starts do not matter that much, after all? Carlos Segarra, Ivan Durev, Peter R. Pietzuch |
SoCC | 3 |
| 2024 | Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor CollectionsabstractDeep learning (DL) jobs use multi-dimensional parallelism, i.e., combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experience changes to their GPU allocation: (i) resource elasticity during training adds or removes GPUs; (ii) hardware maintenance may require redeployment on different GPUs; and (iii) GPU failures force jobs to run with fewer devices. Current DL frameworks tie jobs to a set of GPUs and thus lack support for these scenarios. In particular, they cannot change the multi-dimensional parallelism of an already-running job in an efficient and model-independent way. Marcel Wagenländer, Bo Zhao 0019, Luo Mai, Peter R. Pietzuch |
SOSP | 5 |
| 2024 | OLAP on Modern Chiplet-Based ProcessorsabstractChiplet-based CPUs, which combine multiple independent dies on a single package, allow hardware to scale to higher CPU core counts at the cost of more memory heterogeneity and performance variability. This introduces challenges when existing query engines are deployed on chiplet-based CPUs, as current designs make assumptions about uniform memory access, cache locality and consistent core performance, e.g., leading to ineffective CPU utilization. In this paper, we analyse the performance impact when query engines ignore chiplet-specific properties. We demonstrate that a naïve deployment can result in a significant degradation of query processing efficiency, exhibiting non-linear scaling even within a single CPU socket domain. Based on comprehensive experiments, we explore approaches to deploy query engines on chiplet-based CPUs with improved performance: we show that distributing processing tasks according to a chiplet-aware strategy achieves higher resource utilization and scalability, yielding an up to 7× speedup compared to hardware-oblivious approaches. Alessandro Fogli, Bo Zhao 0019, Peter R. Pietzuch, Maximilian Bandle, Jana Giceva |
Proc. VLDB Endow. | 3 |
| 2023 | ORC: Increasing Cloud Memory Density via Object Reuse with Capabilities
Vasily A. Sartakov, Lluís Vilanova, Munir Geden, David M. Eyers, Takahiro Shinagawa, Peter R. Pietzuch |
OSDI | 6 |
| 2023 | Translation Pass-Through for Near-Native Paging Performance in VMs
Shai Bergman, Mark Silberstein, Takahiro Shinagawa, Peter R. Pietzuch, Lluís Vilanova |
USENIX ATC | 4 |
| 2023 | MSRL: Distributed Reinforcement Learning with Dataflow Fragments
Huanzhou Zhu, Bo Zhao 0019, Yaodong Yang 0001, Peter R. Pietzuch |
USENIX ATC | 8 |
| 2022 | IA-CCF: Individual Accountability for Permissioned Ledgers
Alex Shamis, Peter R. Pietzuch, Burcu Canakci, Miguel Castro 0001, Cédric Fournet, Edward Ashton, Amaury Chamayou, Sylvan Clebsch, Antoine Delignat-Lavaud, Matthew Kerner, Julien Maffre, Olga Vrousgou, Christoph M. Wintersteiger, Manuel Costa, Mark Russinovich |
NSDI | 2 |
| 2022 | CAP-VMs: Capability-Based Isolation and Sharing in the Cloud
Vasily A. Sartakov, Lluís Vilanova, David M. Eyers, Takahiro Shinagawa, Peter R. Pietzuch |
OSDI | 5 |
| 2021 | CubicleOS: a library OS with software componentisation for practical isolationabstractLibrary OSs have been proposed to deploy applications isolated inside containers, VMs, or trusted execution environments. They often follow a highly modular design in which third-party components are combined to offer the OS functionality needed by an application, and they are customised at compilation and deployment time to fit application requirements. Yet their monolithic design lacks isolation across components: when applications and OS components contain security-sensitive data (e.g., cryptographic keys or user data), the lack of isolation renders library OSs open to security breaches via malicious or vulnerable third-party components. Vasily A. Sartakov, Lluís Vilanova, Peter R. Pietzuch |
ASPLOS | 3 |
| 2021 | rkt-io: a direct I/O stack for shielded executionabstractThe shielding of applications using trusted execution environments (TEEs) can provide strong security guarantees in untrusted cloud environments. When executing I/O operations, today's shielded execution frameworks, however, exhibit performance and security limitations: they assign resources to the I/O path inefficiently, perform redundant data copies, use untrusted host I/O stacks with security risks and performance overheads. This prevents TEEs from running modern I/O-intensive applications that require high-performance networking and storage. Jörg Thalheim, Harshavardhan Unnibhavi, Christian Priebe, Pramod Bhatotia, Peter R. Pietzuch |
EuroSys | 5 |
| 2021 | Spons & Shields: practical isolation for trusted executionabstractTrusted execution environments (TEEs) promise a cost-effective, “lift-and-shift” solution for deploying security-sensitive applications in untrusted clouds. For this, they must support rich, multi-component applications, but a large trusted computing base (TCB) inside the TEE risks that attackers can compromise application security. Fine-grained compartmentalisation can increase security through defense-in-depth, but current solutions either run all software components unprotected in the same TEE, lack efficient shared memory support, or isolate application processes using separate TEEs, impacting performance and compatibility. Vasily A. Sartakov, Dan O'Keeffe, David M. Eyers, Lluís Vilanova, Peter R. Pietzuch |
VEE | 5 |
| 2021 | Scabbard: Single-Node Fault-Tolerant Stream ProcessingabstractSingle-node multi-core stream processing engines (SPEs) can process hundreds of millions of tuples per second. Yet making them fault-tolerant with exactly-once semantics while retaining this performance is an open challenge: due to the limited I/O bandwidth of a single-node, it becomes infeasible to persist all stream data and operator state during execution. Instead, single-node SPEs rely on upstream distributed systems, such as Apache Kafka, to recover stream data after failure, necessitating complex cluster-based deployments. This lack of built-in fault-tolerance features has hindered the adoption of single-node SPEs. We describe Scabbard, the first single-node SPE that supports exactly-once fault-tolerance semantics despite limited local I/O bandwidth. Scabbard achieves this by integrating persistence operations with the query workload. Within the operator graph, Scabbard determines when to persist streams based on the selectivity of operators: by persisting streams after operators that discard data, it can substantially reduce the required I/O bandwidth. As part of the operator graph, Scabbard supports parallel persistence operations and uses markers to decide when to discard persisted data. The persisted data volume is further reduced using workload-specific compression: Scabbard monitors stream statistics and dynamically generates computationally efficient compression operators. Our experiments show that Scabbard can execute stream queries that process over 200 million tuples per second while recovering from failures with sub-second latencies. Georgios Theodorakis, Fotios Kounelis, Peter R. Pietzuch, Holger Pirk |
Proc. VLDB Endow. | 3 |
| 2020 | SlideSide: A fast Incremental Stream Processing Algorithm for Multiple QueriesabstractAggregate window computations lie at the core of online analyt-ics in both academic and industrial applications. To efficientlycompute sliding windows, the state-of-the-art algorithms utilizeincremental processing that avoids the recomputation of windowresults from scratch. In this paper, we propose a novel algorithm,calledSlideSide, that extendsTwoStacksfor multiple concur-rent aggregate queries over the same data stream. Our approachuses different yet similar processing schemes for invertible andnon-invertible functions and exhibits up to 2×better through-put compared to the state-of-the-art incremental techniques in amulti-query environment. Georgios Theodorakis, Peter R. Pietzuch, Holger Pirk |
EDBT | 2 |
| 2020 | KungFu: Making Training in Distributed Machine Learning Adaptive
Luo Mai, Marcel Wagenländer, Konstantinos Fertakis, Andrei-Octavian Brabete, Peter R. Pietzuch |
OSDI | 6 |
| 2020 | LightSaber: Efficient Window Aggregation on Multi-core ProcessorsabstractWindow aggregation queries are a core part of streaming applications. To support window aggregation efficiently, stream processing engines face a trade-off between exploiting parallelism (at the instruction/multi-core levels) and incremental computation (across overlapping windows and queries). Existing engines implement ad-hoc aggregation and parallelization strategies. As a result, they only achieve high performance for specific queries depending on the window definition and the type of aggregation function. We describe a general model for the design space of window aggregation strategies. Based on this, we introduce LightSaber, a new stream processing engine that balances parallelism and incremental processing when executing window aggregation queries on multi-core CPUs. Its design generalizes existing approaches: (i) for parallel processing, LightSaber constructs a parallel aggregation tree (PAT) that exploits the parallelism of modern processors. The PAT divides window aggregation into intermediate steps that enable the efficient use of both instruction-level (i.e., SIMD) and task-level (i.e., multi-core) parallelism; and (ii) to generate efficient incremental code from the PAT, LightSaber uses a generalized aggregation graph (GAG), which encodes the low-level data dependencies required to produce aggregates over the stream. A GAG thus generalizes state-of-the-art approaches for incremental window aggregation and supports work-sharing between overlapping windows. LightSaber achieves up to an order of magnitude higher throughput compared to existing systems-on a 16-core server, it processes 470 million records/s with 132 ?s average latency. Georgios Theodorakis, Alexandros Koliousis, Peter R. Pietzuch, Holger Pirk |
SIGMOD Conference | 3 |
| 2020 | Faasm: Lightweight Isolation for Efficient Stateful Serverless Computing
Simon Shillaker, Peter R. Pietzuch |
USENIX ATC | 2 |
| 2019 | Thriving in the No Man's Land between Compilers and Databases
Holger Pirk, Jana Giceva, Peter R. Pietzuch |
CIDR | 3 |
| 2019 | Neptune: Scheduling Suspendable Tasks for Unified Stream/Batch ApplicationsabstractDistributed dataflow systems allow users to express a wide range of computations, including batch, streaming, and machine learning. A recent trend is to unify different computation types as part of a single stream/batch application that combines latency-sensitive ("stream") and latency-tolerant ("batch") jobs. This sharing of state and logic across jobs simplifies application development. Examples include machine learning applications that perform batch training and low-latency inference, and data analytics applications that include batch data transformations and low-latency querying. Existing execution engines, however, were not designed for unified stream/batch applications. As we show, they fail to schedule and execute them efficiently while respecting their diverse requirements. Panagiotis Garefalakis, Konstantinos Karanasos, Peter R. Pietzuch |
SoCC | 3 |
| 2019 | Using Trusted Execution Environments for Secure Stream Processing of Medical Data - (Case Study Paper)
Carlos Segarra, Ricard Delgado-Gonzalo, Mathieu Lemay, Pierre-Louis Aublin, Peter R. Pietzuch, Valerio Schiavoni |
DAIS | 5 |
| 2019 | Teechain: a secure payment network with asynchronous blockchain accessabstractBlockchains such as Bitcoin and Ethereum execute payment transactions securely, but their performance is limited by the need for global consensus. Payment networks overcome this limitation through off-chain transactions. Instead of writing to the blockchain for each transaction, they only settle the final payment balances with the underlying blockchain. When executing off-chain transactions in current payment networks, parties must access the blockchain within bounded time to detect misbehaving parties that deviate from the protocol. This opens a window for attacks in which a malicious party can steal funds by deliberately delaying other parties' blockchain access and prevents parties from using payment networks when disconnected from the blockchain. Joshua Lind, Oded Naor, Ittay Eyal, Florian Kelbert, Emin Gün Sirer, Peter R. Pietzuch |
SOSP | 6 |
| 2019 | Crossbow: Scaling Deep Learning with Small Batch Sizes on Multi-GPU ServersabstractDeep learning models are trained on servers with many GPUs, and training must scale with the number of GPUs. Systems such as TensorFlow and Caffe2 train models with parallel synchronous stochastic gradient descent: they process a batch of training data at a time, partitioned across GPUs, and average the resulting partial gradients to obtain an updated global model. To fully utilise all GPUs, systems must increase the batch size, which hinders statistical efficiency. Users tune hyper-parameters such as the learning rate to compensate for this, which is complex and model-specific. We describe Crossbow, a new single-server multi-GPU system for training deep learning models that enables users to freely choose their preferred batch size---however small---while scaling to multiple GPUs. Crossbow uses many parallel model replicas and avoids reduced statistical efficiency through a new synchronous training method. We introduce SMA, a synchronous variant of model averaging in which replicas independently explore the solution space with gradient descent, but adjust their search synchronously based on the trajectory of a globally-consistent average model. Crossbow achieves high hardware efficiency with small batch sizes by potentially training multiple model replicas per GPU, automatically tuning the number of replicas to maximise throughput. our experiments show that Crossbow improves the training time of deep learning models on an 8-GPU server by 1.3--4X compared to TensorFlow. Alexandros Koliousis, Pijika Watcharapichat, Matthias Weidlich 0001, Luo Mai, Paolo Costa, Peter R. Pietzuch |
Proc. VLDB Endow. | 6 |
| 2019 | Faces in the Clouds: Long-Duration, Multi-User, Cloud-Assisted Video ConferencingabstractMulti-user video conferencing is a ubiquitous technology. Increasingly end-hosts in a conference are assisted by cloud-based servers that improve the quality of experience for end users. This paper evaluates the impact of strategies for placement of such servers on user experience and deployment cost. We consider scenarios based upon the Amazon EC2 infrastructure as well as future scenarios in which cloud instances can be located at a larger number of possible sites across the planet. We compare a number of possible strategies for choosing which cloud locations should host services and how traffic should route through them. Our study is driven by real data to create demand scenarios with realistic geographical user distributions and diurnal behaviour. We conclude that on the EC2 infrastructure a well chosen static selection of servers performs well but as more cloud locations are available a dynamic choice of servers becomes important. Richard G. Clegg, Raul Landa, David Griffin 0001, Miguel Rio, Peter Hughes, Ian Kegel, Tim Stevens, Peter R. Pietzuch, Doug Williams |
IEEE Trans. Cloud Comput. | 8 |
| 2018 | EndBox: Scalable Middlebox Functions Using Client-Side Trusted ExecutionabstractMany organisations enhance the performance, security, and functionality of their managed networks by deploying middleboxes centrally as part of their core network. While this simplifies maintenance, it also increases cost because middlebox hardware must scale with the number of clients. A promising alternative is to outsource middlebox functions to the clients themselves, thus leveraging their CPU resources. Such an approach, however, raises security challenges for critical middlebox functions such as firewalls and intrusion detection systems. We describe EndBox, a system that securely executes middlebox functions on client machines at the network edge. Its design combines a virtual private network (VPN) with middlebox functions that are hardware-protected by a trusted execution environment (TEE), as offered by Intel's Software Guard Extensions (SGX). By maintaining VPN connection endpoints inside SGX enclaves, EndBox ensures that all client traffic, including encrypted communication, is processed by the middlebox. Despite its decentralised model, EndBox's middlebox functions remain maintainable: they are centrally controlled and can be updated efficiently. We demonstrate EndBox with two scenarios involving (i) a large company; and (ii) an Internet service provider that both need to protect their network and connected clients. We evaluate EndBox by comparing it to centralised deployments of common middlebox functions, such as load balancing, intrusion detection, firewalling, and DDoS prevention. We show that EndBox achieves up to 3.8x higher throughput and scales linearly with the number of clients. David Goltzsche, Signe Rüsch, Manuel Nieke, Sébastien Vaucher, Nico Weichbrodt, Valerio Schiavoni, Pierre-Louis Aublin, Paolo Costa, Christof Fetzer, Pascal Felber, Peter R. Pietzuch, Rüdiger Kapitza |
DSN | 11 |
| 2018 | LibSEAL: revealing service integrity violations using trusted executionabstractUsers of online services such as messaging, code hosting and collaborative document editing expect the services to uphold the integrity of their data. Despite providers' best efforts, data corruption still occurs, but at present service integrity violations are excluded from SLAs. For providers to include such violations as part of SLAs, the competing requirements of clients and providers must be satisfied. Clients need the ability to independently identify and prove service integrity violations to claim compensation. At the same time, providers must be able to refute spurious claims. Pierre-Louis Aublin, Florian Kelbert, Dan O'Keeffe, Divya Muthukumaran, Christian Priebe, Joshua Lind, Robert Krahn, Christof Fetzer, David M. Eyers, Peter R. Pietzuch |
EuroSys | 10 |
| 2018 | Medea: scheduling of long running applications in shared production clustersabstractThe rise in popularity of machine learning, streaming, and latency-sensitive online applications in shared production clusters has raised new challenges for cluster schedulers. To optimize their performance and resilience, these applications require precise control of their placements, by means of complex constraints, e.g., to collocate or separate their long-running containers across groups of nodes. In the presence of these applications, the cluster scheduler must attain global optimization objectives, such as maximizing the number of deployed applications or minimizing the violated constraints and the resource fragmentation, but without affecting the scheduling latency of short-running containers. Panagiotis Garefalakis, Konstantinos Karanasos, Peter R. Pietzuch, Arun Suresh, Sriram Rao |
EuroSys | 3 |
| 2018 | Meta-Dataflows: Efficient Exploratory Dataflow JobsabstractDistributed dataflow systems such as Apache Spark and Apache Flink are used to derive new insights from large datasets. While they efficiently execute concrete data processing workflows, expressed as dataflow graphs, they lack generic support for exploratory workflows : if a user is uncertain about the correct processing pipeline, e.g. in terms of data cleaning strategy or choice of model parameters, they must repeatedly submit modified jobs to the system. This, however, misses out on optimisation opportunities for exploratory workflows, both in terms of scheduling and memory allocation. Raul Castro Fernandez, William Culhane, Pijika Watcharapichat, Matthias Weidlich 0001, Victoria Lopez Morales, Peter R. Pietzuch |
SIGMOD Conference | 6 |
| 2018 | Teechain: Reducing Storage Costs on the Blockchain With Offline Payment ChannelsabstractNo abstract available. Joshua Lind, Oded Naor, Ittay Eyal, Florian Kelbert, Peter R. Pietzuch, Emin Gün Sirer |
SYSTOR | 5 |
| 2018 | Frontier: Resilient Edge Processing for the Internet of ThingsabstractIn an edge deployment model, Internet-of-Things (IoT) applications, e.g. for building automation or video surveillance, must process data locally on IoT devices without relying on permanent connectivity to a cloud backend. The ability to harness the combined resources of multiple IoT devices for computation is influenced by the quality of wireless network connectivity. An open challenge is how practical edge-based IoT applications can be realised that are robust to changes in network bandwidth between IoT devices, due to interference and intermittent connectivity. We present Frontier , a distributed and resilient edge processing platform for IoT devices. The key idea is to express data-intensive IoT applications as continuous data-parallel streaming queries and to improve query throughput in an unreliable wireless network by exploiting network path diversity : a query includes operator replicas at different IoT nodes, which increases possible network paths for data. Frontier dynamically routes stream data to operator replicas based on network path conditions. Nodes probe path throughput and use backpressure stream routing to decide on transmission rates, while exploiting multiple operator replicas for data-parallelism. If a node loses network connectivity, a transient disconnection recovery mechanism reprocesses the lost data. Our experimental evaluation of Frontier shows that network path diversity improves throughput by 1.3×−2.8×for different IoT applications, while being resilient to intermittent network connectivity. Dan O'Keeffe, Theodoros Salonidis, Peter R. Pietzuch |
Proc. VLDB Endow. | 3 |
| 2017 | SecureCloud: Secure big data processing in untrusted cloudsabstractWe present the SecureCloud EU Horizon 2020 project, whose goal is to enable new big data applications that use sensitive data in the cloud without compromising data security and privacy. For this, SecureCloud designs and develops a layered architecture that allows for (i) the secure creation and deployment of secure micro-services; (ii) the secure integration of individual micro-services to full-fledged big data applications; and (iii) the secure execution of these applications within untrusted cloud environments. To provide security guarantees, SecureCloud leverages novel security mechanisms present in recent commodity CPUs, in particular, Intel's Software Guard Extensions (SGX). SecureCloud applies this architecture to big data applications in the context of smart grids. We describe the SecureCloud approach, initial results, and considered use cases. Florian Kelbert, Franz Gregor, Rafael Pires 0001, Stefan Köpsell, Marcelo Pasin, Aurelien Havet, Valerio Schiavoni, Pascal Felber, Christof Fetzer, Peter R. Pietzuch |
DATE | 10 |
| 2017 | SwiftAnalytics: Optimizing Object Storage for Big Data AnalyticsabstractDue to their scalability and low cost, object-based storage systems are an attractive storage solution and widely deployed. To gain valuable insight from the data residing in object storage but avoid expensive copying to a distributed filesystem (e.g. HDFS), it would be natural to directly use them as a storage backend for data-parallel analytics frameworks such as Spark or MapReduce. Unfortunately, executing data-parallel frameworks on object storage exhibits severe performance problems, reducing average job completion times by up to 6.5×. We identify the two most severe performance problems when running data-parallel frameworks on the OpenStack Swift object storage system in comparison to the HDFS distributed filesystem: (i) the fixed mapping of object names to storage nodes prevents local writes and adds delay when objects are renamed, (ii) the coarser granularity of objects compared to blocks reduces data locality during reads. We propose the SwiftAnalytics object storage system to address them: (i) it uses locality-aware writes to control an object's location and eliminate unnecessary I/O related to renames during job completion, speeding up analytics jobs by up to 5.1×, (ii) it transparently chunks objects into smaller sized parts to improve data-locality, leading to up to 3.4× faster reads. Lukas Rupprecht, Bill Owen, Peter R. Pietzuch, Dean Hildebrand |
IC2E | 4 |
| 2017 | Glamdring: Automatic Application Partitioning for Intel SGX
Joshua Lind, Christian Priebe, Divya Muthukumaran, Dan O'Keeffe, Pierre-Louis Aublin, Florian Kelbert, Tobias Reiher, David Goltzsche, David M. Eyers, Rüdiger Kapitza, Christof Fetzer, Peter R. Pietzuch |
USENIX ATC | 12 |
| 2017 | Emu: Rapid Prototyping of Networking Services
Nik Sultana, Salvator Galea, David Greaves, Marcin Wójcik, Jonny Shipton, Richard G. Clegg, Luo Mai, Pietro Bressana, Robert Soulé, Richard Mortier, Paolo Costa, Peter R. Pietzuch, Jon Crowcroft, Andrew W. Moore 0002, Noa Zilberman |
USENIX ATC | 12 |
| 2017 | SquirrelJoin: Network-Aware Distributed Join Processing with Lazy PartitioningabstractTo execute distributed joins in parallel on compute clusters, systems partition and exchange data records between workers. With large datasets, workers spend a considerable amount of time transferring data over the network. When compute clusters are shared among multiple applications, workers must compete for network bandwidth with other applications. These variances in the available network bandwidth lead to network skew , which causes straggling workers to prolong the join completion time. We describe SquirrelJoin , a distributed join processing technique that uses lazy partitioning to adapt to transient network skew in clusters. Workers maintain in-memory lazy partitions to withhold a subset of records, i.e. not sending them immediately to other workers for processing. Lazy partitions are then assigned dynamically to other workers based on network conditions: each worker takes periodic throughput measurements to estimate its completion time, and lazy partitions are allocated as to minimise the join completion time. We implement SquirrelJoin as part of the Apache Flink distributed dataflow framework and show that, under transient network contention in a shared compute cluster, SquirrelJoin speeds up join completion times by up to 2.9× with only a small, fixed overhead. Lukas Rupprecht, William Culhane, Peter R. Pietzuch |
Proc. VLDB Endow. | 3 |
| 2016 | Towards enabling hyper-responsive mobile apps through network edge assistanceabstractPoor Internet performance currently undermines the efficiency of hyper-responsive mobile apps such as augmented reality clients and online games, which require low-latency access to real-time backend services. While edge-assisted execution, i.e. moving entire services to the edge of an access network, helps eliminate part of the communication overhead involved, this does not scale to the number of users that share an edge infrastructure. This is due to a mismatch between the scarce availability of resources in access networks and the aggregate demand for computational power from client applications. Instead, this paper proposes a hybrid edge-assisted deployment model in which only part of a service executes on LTE edge servers. We provide insights about the conditions that must hold for such a model to be effective by investigating in simulation different deployment and application scenarios. In particular, we show that using LTE edge servers with modest capabilities, performance can improve significantly as long as at most 50% of client requests are processed at the edge. Moreover, we argue that edge servers should be installed at the core of a mobile network, rather than the mobile base station: the difference in performance is negligible, whereas the latter choice entails high deployment costs. Finally, we verify that, for the proposed model, the impact of user mobility on TCP performance is low. Miguel Baguena, George Samaras, Andreas Pamboris, Mihail L. Sichitiu, Peter R. Pietzuch, Pietro Manzoni |
CCNC | 5 |
| 2016 | Ako: Decentralised Deep Learning with Partial Gradient ExchangeabstractDistributed systems for the training of deep neural networks (DNNs) with large amounts of data have vastly improved the accuracy of machine learning models for image and speech recognition. DNN systems scale to large cluster deployments by having worker nodes train many model replicas in parallel; to ensure model convergence, parameter servers periodically synchronise the replicas. This raises the challenge of how to split resources between workers and parameter servers so that the cluster CPU and network resources are fully utilised without introducing bottlenecks. In practice, this requires manual tuning for each model configuration or hardware type. Pijika Watcharapichat, Victoria Lopez Morales, Raul Castro Fernandez, Peter R. Pietzuch |
SoCC | 4 |
| 2016 | AsyncShock: Exploiting Synchronisation Bugs in Intel SGX Enclaves
Nico Weichbrodt, Anil Kurmus, Peter R. Pietzuch, Rüdiger Kapitza |
ESORICS (1) | 3 |
| 2016 | Java2SDG: Stateful big data processing for the massesabstractBig data processing is no longer restricted to specially-trained engineers. Instead, domain experts, data scientists and data users all want to benefit from applying data mining and machine learning algorithms at scale. A considerable obstacle towards this “democratisation of big data” are programming models: current scalable big data processing platforms such as Spark, Naiad and Flink require users to learn custom functional or declarative programming models, which differ fundamentally from popular languages such as Java, Matlab, Python or C++. An open challenge is how to provide a big data programming model for users that are not familiar with functional programming, while maintaining performance, scalability and fault tolerance. We describe JAVA2SDG, a compiler that translates annotated Java programs to stateful dataflow graphs (SDGs) that can execute on a compute cluster in a data-parallel and fault-tolerant fashion. Compared to existing distributed dataflow models, a distinguishing feature of SDGs is that their computational tasks can access distributed mutable state, thus allowing SDGs to capture the semantics of stateful Java programs. As part of the demonstration, we provide examples of machine learning programs in Java, including collaborative filtering and logistic regression, and we explain how they are translated to SDGs and executed on a large set of machines. Raul Castro Fernandez, Panagiotis Garefalakis, Peter R. Pietzuch |
ICDE | 3 |
| 2016 | SecureKeeper: Confidential ZooKeeper using Intel SGX
Stefan Brenner, Colin Wulf, David Goltzsche, Nico Weichbrodt, Matthias Lorenz, Christof Fetzer, Peter R. Pietzuch, Rüdiger Kapitza |
Middleware | 7 |
| 2016 | BrowserFlow: Imprecise Data Flow Tracking to Prevent Accidental Data Disclosure
Ioannis Papagiannis, Pijika Watcharapichat, Divya Muthukumaran, Peter R. Pietzuch |
Middleware | 4 |
| 2016 | SCONE: Secure Linux Containers with Intel SGX
Sergei Arnautov, Bohdan Trach, Franz Gregor, Thomas Knauth, André Martin, Christian Priebe, Joshua Lind, Divya Muthukumaran, Dan O'Keeffe, Mark Stillwell, David Goltzsche, David M. Eyers, Rüdiger Kapitza, Peter R. Pietzuch, Christof Fetzer |
OSDI | 14 |
| 2016 | THEMIS: Fairness in Federated Stream Processing under OverloadabstractFederated stream processing systems, which utilise nodes from multiple independent domains, can be found increasingly in multi-provider cloud deployments, internet-of-things systems, collaborative sensing applications and large-scale grid systems. To pool resources from several sites and take advantage of local processing, submitted queries are split into query fragments, which are executed collaboratively by different sites. When supporting many concurrent users, however, queries may exhaust available processing resources, thus requiring constant load shedding. Given that individual sites have autonomy over how they allocate query fragments on their nodes, it is an open challenge how to ensure global fairness on processing quality experienced by queries in a federated scenario. Evangelia Kalyvianaki, Marco Fiscato, Theodoros Salonidis, Peter R. Pietzuch |
SIGMOD Conference | 4 |
| 2016 | SABER: Window-Based Hybrid Stream Processing for Heterogeneous ArchitecturesabstractModern servers have become heterogeneous, often combining multi-core CPUs with many-core GPGPUs. Such heterogeneous architectures have the potential to improve the performance of data-intensive stream processing applications, but they are not supported by current relational stream processing engines. For an engine to exploit a heterogeneous architecture, it must execute streaming SQL queries with sufficient data-parallelism to fully utilise all available heterogeneous processors, and decide how to use each in the most effective way. It must do this while respecting the semantics of streaming SQL queries, in particular with regard to window handling. Alexandros Koliousis, Matthias Weidlich 0001, Raul Castro Fernandez, Alexander L. Wolf, Paolo Costa, Peter R. Pietzuch |
SIGMOD Conference | 6 |
| 2016 | AT-GIS: Highly Parallel Spatial Query Processing with Associative TransducersabstractUsers in many domains, including urban planning, transportation, and environmental science want to execute analytical queries over continuously updated spatial datasets. Current solutions for large-scale spatial query processing either rely on extensions to RDBMS, which entails expensive loading and indexing phases when the data changes, or distributed map/reduce frameworks, running on resource-hungry compute clusters. Both solutions struggle with the sequential bottleneck of parsing complex, hierarchical spatial data formats, which frequently dominates query execution time. Our goal is to fully exploit the parallelism offered by modern multi-core CPUs for parsing and query execution, thus providing the performance of a cluster with the resources of a single machine. We describe AT-GIS, a highly-parallel spatial query processing system that scales linearly to a large number of CPU cores. AT-GIS integrates the parsing and querying of spatial data using a new computational abstraction called associative transducers (ATs). ATs can form a single data-parallel pipeline for computation without requiring the spatial input data to be split into logically independent blocks. Using ATs, AT-GIS can execute, in parallel, spatial query operators on the raw input data in multiple formats, without any pre-processing. On a single 64-core machine, AT-GIT provides 3x the performance of an 8-node Hadoop cluster with 192 cores for containment queries, and 10x for aggregation queries. Peter Ogden, David B. Thomas, Peter R. Pietzuch |
SIGMOD Conference | 3 |
| 2016 | FLICK: Developing and Running Application-Specific Network Services
Abdul Alim, Richard G. Clegg, Luo Mai, Lukas Rupprecht, Eric Seckler, Paolo Costa, Peter R. Pietzuch, Alexander L. Wolf, Nik Sultana, Jon Crowcroft, Anil Madhavapeddy, Andrew W. Moore 0002, Richard Mortier, Masoud Koleini, Luis Oviedo, Matteo Migliavacca, Derek McAuley |
USENIX ATC | 7 |
| 2016 | C-RAM: Breaking Mobile Device Memory Barriers Using the CloudabstractMobile applications are constrained by the available memory of mobile devices. We present C-RAM, a system that uses cloud-based memory to extend the memory of mobile devices. Itsplits application stateand its associated computation between a mobile device and a cloud node to allow applications to consume more memory, while minimizing the performance impact. C-RAM thus enables developers to realize new applications or port legacy desktop applications with a large memory footprint to mobile platforms without explicitly designing them to account for memory limitations. To handle network failures with partitioned application state, C-RAM uses a new snapshot-based fault tolerance mechanism in which changes to remote memory objects are periodically backed up to the device. After failure, or when network usage exceeds a given limit, the device rolls back execution to continue from the last snapshot. C-RAM supports local execution with an application state that exceeds the available device memory through a user-level virtual memory: objects are loaded on-demand from snapshots in flash memory. Our C-RAM prototype supports Objective-C applications on the unmodified iOS platform. With C-RAM, applications can consume 10$\times$more memory than the device capacity, with a negligible impact on application performance. In some cases, C-RAM even achieves a significant speed-up in execution time (up to 9.7$\times$). Andreas Pamboris, Peter R. Pietzuch |
IEEE Trans. Mob. Comput. | 2 |
| 2015 | FlowWatcher: Defending against Data Disclosure Vulnerabilities in Web ApplicationsabstractBugs in the authorisation logic of web applications can expose the data of one user to another. Such data disclosure vulnerabilities are common---they can be caused by a single omitted access control check in the application. We make the observation that, while the implementation of the authorisation logic is complex and therefore error-prone, most web applications only use simple access control models, in which each piece of data is accessible by a user or a group of users. This makes it possible to validate the correct operation of the authorisation logic externally, based on the observed data in HTTP traffic to and from an application. Divya Muthukumaran, Dan O'Keeffe, Christian Priebe, David M. Eyers, Brian Shand, Peter R. Pietzuch |
CCS | 6 |
| 2015 | Liquid: Unifying Nearline and Offline Big Data Integration
Raul Castro Fernandez, Peter R. Pietzuch, Jay Kreps, Neha Narkhede, Jun Rao, Joel Koshy, Dong Lin, Chris Riccomini, Guozhang Wang |
CIDR | 2 |
| 2015 | CloudScope: Diagnosing and Managing Performance Interference in Multi-tenant CloudsabstractVirtual machine consolidation is attractive in cloud computing platforms for several reasons including reduced infrastructure costs, lower energy consumption and ease of management. However, the interference between co-resident workloads caused by virtualization can violate the service level objectives (SLOs) that the cloud platform guarantees. Existing solutions to minimize interference between virtual machines (VMs) are mostly based on comprehensive micro-benchmarks or online training which makes them computationally intensive. In this paper, we present CloudScope, a system for diagnosing interference for multi-tenant cloud systems in a lightweight way. CloudScope employs a discrete-time Markov Chain model for the online prediction of performance interference of co-resident VMs. It uses the results to optimally (re)assign VMs to physical machines and to optimize the hypervisor configuration, e.g. the CPU share it can use, for different workloads. We have implemented CloudScope on top of the Xen hypervisor and conducted experiments using a set of CPU, disk, and network intensive workloads and a real system (MapReduce). Our results show that CloudScope interference prediction achieves an average error of 9%. The interference-aware scheduler improves VM performance by up to 10% compared to the default scheduler. In addition, the hypervisor reconfiguration can improve network throughput by up to 30%. Xi Chen 0015, Lukas Rupprecht, Rasha Osman, Peter R. Pietzuch, Felipe Franciosi, William J. Knottenbelt |
MASCOTS | 4 |
| 2015 | Better Performance in LTE Networks with Edge Assistance: The World of Warcraft CaseabstractTo improve the performance of Massively Multiplayer Online Games (MMOGs) in mobile networks, we explore the potential benefits of an edge-assisted deployment model: part of the MMOG backend service executes closer to the end user at the edge of the LTE network. We investigate the impact on game late Miguel Baguena, Andreas Pamboris, Peter R. Pietzuch, Mihail L. Sichitiu, Pietro Manzoni |
MobiQuitous | 3 |
| 2015 | Demo: : NOMAD: An Edge Cloud Platform for Hyper-Responsive Mobile AppsabstractNo abstract available. Andreas Pamboris, Miguel Baguena, Alexander L. Wolf, Pietro Manzoni, Peter R. Pietzuch |
MobiSys | 5 |
| 2014 | NetAgg: Using Middleboxes for Application-specific On-path Aggregation in Data CentresabstractData centre applications for batch processing (e.g. map/reduce frameworks) and online services (e.g. search engines) scale by distributing data and computation across many servers. They typically follow a partition/aggregation pattern: tasks are first partitioned across servers that process data locally, and then those partial results are aggregated. This data aggregation step, however, shifts the performance bottleneck to the network, which typically struggles to support many-to-few, high-bandwidth traffic between servers. Luo Mai, Lukas Rupprecht, Abdul Alim, Paolo Costa, Matteo Migliavacca, Peter R. Pietzuch, Alexander L. Wolf |
CoNEXT | 6 |
| 2014 | Outsourcing multi-version key-value stores with verifiable data freshnessabstractIn the age of big data, key-value data updated by intensive write streams is increasingly common, e.g., in social event streams. To serve such data in a cost-effective manner, a popular new paradigm is to outsource it to the cloud and store it in a scalable key-value store while serving a large user base. Due to the limited trust in third-party cloud infrastructures, data owners have to sign the data stream so that the data users can verify the authenticity of query results from the cloud. In this paper, we address the problem of verifiable freshness for multi-version key-value data. We propose a memory-resident digest structure that utilizes limited memory effectively and can have efficient verification performance. The proposed structure is named IncBM-Tree because it can INCrementally build a Bloom filter-embedded Merkle Tree. We have demonstrated the superior performance of verification under small memory footprints for signing, which is typical in an outsourcing scenario where data owners and users have limited resources. Yuzhe Tang, Ling Liu 0001, Ting Wang 0006, Xin Hu 0001, Reiner Sailer, Peter R. Pietzuch |
ICDE | 6 |
| 2014 | Managing Expectations: Runtime Negotiation of Information Quality Requirements in Event-Based Systems
Sebastian Frischbier, Peter R. Pietzuch, Alejandro P. Buchmann |
ICSOC | 2 |
| 2014 | Making State Explicit for Imperative Big Data Processing
Raul Castro Fernandez, Matteo Migliavacca, Evangelia Kalyvianaki, Peter R. Pietzuch |
USENIX ATC | 4 |
| 2014 | Information Flow Control for Secure Cloud ComputingabstractSecurity concerns are widely seen as an obstacle to the adoption of cloud computing solutions. Information Flow Control (IFC) is a well understood Mandatory Access Control methodology. The earliest IFC models targeted security in a centralised environment, but decentralised forms of IFC have been designed and implemented, often within academic research projects. As a result, there is potential for decentralised IFC to achieve better cloud security than is available today. In this paper we describe the properties of cloud computing-Platform-as-a-Service clouds in particular-and review a range of IFC models and implementations to identify opportunities for using IFC within a cloud computing context. Since IFC security is linked to the data that it protects, both tenants and providers of cloud services can agree on security policy, in a manner that does not require them to understand and rely on the particulars of the cloud software stack in order to effect enforcement. Jean Bacon, David M. Eyers, Thomas Pasquier, Jatinder Singh, Ioannis Papagiannis, Peter R. Pietzuch |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2014 | SymbexNet: Testing Network Protocol Implementations with Symbolic Execution and Rule-Based SpecificationsabstractImplementations of network protocols, such as DNS, DHCP and Zeroconf, are prone to flaws, security vulnerabilities and interoperability issues caused by developer mistakes and ambiguous requirements in protocol specifications. Detecting such problems is not easy because (i) many bugs manifest themselves only after prolonged operation; (ii) reasoning about semantic errors requires a machine-readable specification; and (iii) the state space of complex protocol implementations is large. This article presents a novel approach that combines symbolic execution and rule-based specifications to detect various types of flaws in network protocol implementations. The core idea behind our approach is to (1) automatically generate high-coverage test input packets for a network protocol implementation using single- and multi-packet exchange symbolic execution (targeting stateless and stateful protocols, respectively) and then (2) use these packets to detect potential violations of manual rules derived from the protocol specification, and check the interoperability of different implementations of the same network protocol. We present a system based on these techniques, SymbexNet, and evaluate it on multiple implementations of two network protocols: Zeroconf, a service discovery protocol, and DHCP, a network configuration protocol. SymbexNet is able to discover non-trivial bugs as well as interoperability problems, most of which have been confirmed by the developers. Jaeseung Song, Cristian Cadar, Peter R. Pietzuch |
IEEE Trans. Software Eng. | 3 |
| 2013 | Supporting application-specific in-network processing in data centresabstractNo abstract available. Luo Mai, Lukas Rupprecht, Paolo Costa, Matteo Migliavacca, Peter R. Pietzuch, Alexander L. Wolf |
SIGCOMM | 5 |
| 2013 | Integrating scale out and fault tolerance in stream processing using operator state managementabstractAs users of "big data" applications expect fresh results, we witness a new breed of stream processing systems (SPS) that are designed to scale to large numbers of cloud-hosted machines. Such systems face new challenges: (i) to benefit from the "pay-as-you-go" model of cloud computing, they must scale out on demand, acquiring additional virtual machines (VMs) and parallelising operators when the workload increases; (ii) failures are common with deployments on hundreds of VMs-systems must be fault-tolerant with fast recovery times, yet low per-machine overheads. An open question is how to achieve these two goals when stream queries include stateful operators, which must be scaled out and recovered without affecting query results. Raul Castro Fernandez, Matteo Migliavacca, Evangelia Kalyvianaki, Peter R. Pietzuch |
SIGMOD Conference | 4 |
| 2013 | Scalable XML Query Processing using Parallel Pushdown TransducersabstractIn online social networking, network monitoring and financial applications, there is a need to query high rate streams of XML data, but methods for executing individual XPath queries on streaming XML data have not kept pace with multicore CPUs. For data-parallel processing, a single XML stream is typically split into well-formed fragments, which are then processed independently. Such an approach, however, introduces a sequential bottleneck and suffers from low cache locality, limiting its scalability across CPU cores. We describe a data-parallel approach for the processing of streaming XPath queries based on pushdown transducers. Our approach permits XML data to be split into arbitrarilysized chunks, with each chunk processed by a parallel automaton instance. Since chunks may be malformed, our automata consider all possible starting states for XML elements and build mappings from starting to finishing states. These mappings can be constructed independently for each chunk by different CPU cores. For streaming queries from the XPathMark benchmark, we show a processing throughput of 2.5 GB/s, with near linear scaling up to 64 CPU cores. Peter Ogden, David B. Thomas, Peter R. Pietzuch |
Proc. VLDB Endow. | 3 |
| 2012 | Chams: Churn-aware overlay construction for media streaming
Mouna Allani, Benoît Garbinato, Peter R. Pietzuch |
Peer-to-Peer Netw. Appl. | 3 |
| 2011 | IO Tetris: Deep Storage Consolidation for the Cloud via Fine-Grained Workload AnalysisabstractIntelligent workload consolidation in storage systems leads to better Return On Investment (ROI), in terms of more efficient use of data center resources, better Quality of Service (QoS), and lower power consumption. This is particularly significant yet challenging in a cloud environment, in which a large set of different workloads multiplex on a shared, heterogeneous infrastructure. However, the increasing availability of fine grained workload logging facilities allows better insights to be gained from workload profiles. As a consequence, consolidation can be done more deeply, according to a detailed understanding of how well given workloads mix. We describe IO Tetris, which takes a first look at fine-grained consolidation in large-scale storage systems by leveraging temporal patterns found in real-world I/O traces gathered from enterprise storage environments. The core functionality of IO Tetris consists of two stages. A grouping stage performs hierarchical grouping of storage workloads to find complementary groupings that consolidate well together over time and conflicting ones that do not. After that, a migration stage examines the discovered groupings to determine how to maximize resource utilization efficiency while minimizing migration costs. Experiments based on customer I/O traces from a high-end enterprise class IBM storage controller show that a non-trivial number of IO Tetris groupings exist in real-world storage workloads, and that these groupings can be leveraged to achieve better storage consolidation in a cloud setting. Ramani Routray, David M. Eyers, David D. Chambliss, Prasenjit Sarkar, Douglas Willcocks, Peter R. Pietzuch |
IEEE CLOUD | 7 |
| 2011 | CusComNet: A customisable network for reconfigurable heterogeneous clustersabstractComputer clusters equipped with reconfigurable accelerators have shown promise in high performance computing. This paper explores novel ways of customising data communication between accelerator nodes, which is often a bottleneck when scaling up the cluster size. Based on the direct connection of high speed serial links between advanced reconfigurable devices, we develop and evaluate CusComNet, a scalable, flexible and efficient communication framework. The CusComNet framework is built around customisable, packet-based communication and supports three main types of customisation: packet protocol customisation, system-level customisation, and prioritised communication customisation. A performance model for estimating CusComNet's communication latency is proposed and demonstrated. Our framework is applied to a 16-node cluster, each node of which contains an FPGA accelerator which can be connected directly to other FPGA accelerators. The proposed framework can be used to improve the scalability of a reconfigurable cluster by involving more nodes in a single application. Performance measurements show high efficiency data throughput for both large and small data volumes, as well as low communication overhead. Stewart Denholm, Kuen Hung Tsoi, Peter R. Pietzuch, Wayne Luk |
ASAP | 3 |
| 2011 | Rule-Based Verification of Network Protocol Implementations Using Symbolic ExecutionabstractThe secure and correct implementation of network protocols for resource discovery, device configuration and network management is complex and error-prone. Protocol specifications contain ambiguities, leading to implementation flaws and security vulnerabilities in network daemons. Such problems are hard to detect because they are often triggered by complex sequences of packets that occur only after prolonged operation. The goal of this work is to find semantic bugs in network daemons. Our approach is to replay a set of input packets that result in high source code coverage of the daemon and observe potential violations of rules derived from the protocol specification. We describe SYMNV, a practical verification tool that first symbolically executes a network daemon to generate high coverage input packets and then checks a set of rules constraining permitted input and output packets. We have applied SYMNV to three different implementations of the Zeroconf protocol and show that it is able to discover non-trivial bugs. Jaeseung Song, Tiejun Ma, Cristian Cadar, Peter R. Pietzuch |
ICCCN | 4 |
| 2011 | SQPR: Stream query planning with reuseabstractWhen users submit new queries to a distributed stream processing system (DSPS), a query planner must allocate physical resources, such as CPU cores, memory and network bandwidth, from a set of hosts to queries. Allocation decisions must provide the correct mix of resources required by queries, while achieving an efficient overall allocation to scale in the number of admitted queries. By exploiting overlap between queries and reusing partial results, a query planner can conserve resources but has to carry out more complex planning decisions. In this paper, we describe SQPR, a query planner that targets DSPSs in data centre environments with heterogeneous resources. SQPR models query admission, allocation and reuse as a single constrained optimisation problem and solves an approximate version to achieve scalability. It prevents individual resources from becoming bottlenecks by re-planning past allocation decisions and supports different allocation objectives. As our experimental evaluation in comparison with a state-of-the-art planner shows SQPR makes efficient resource allocation decisions, even with a high utilisation of resources, with acceptable overheads. Evangelia Kalyvianaki, Wolfram Wiesemann, Quang Hieu Vu, Daniel Kuhn 0001, Peter R. Pietzuch |
ICDE | 5 |
| 2011 | SafeWeb: A Middleware for Securing Ruby-Based Web Applications
Petr Hosek 0001, Matteo Migliavacca, Ioannis Papagiannis, David M. Eyers, David Evans 0002, Brian Shand, Jean Bacon, Peter R. Pietzuch |
Middleware | 8 |
| 2011 | Femtocell Coverage Optimisation Using Statistical Verification
Tiejun Ma, Peter R. Pietzuch |
Networking (1) | 2 |
| 2011 | On the Feasibility of Bandwidth Detouring
Thom Haddow, Sing Wang Ho, Jonathan Ledlie, Cristian Lumezanu, Moez Draief, Peter R. Pietzuch |
PAM | 6 |
| 2011 | Configuring large-scale storage using a middleware with machine learningabstractSUMMARY The proliferation of cloud services and other forms of service‐oriented computing continues to accelerate. Alongside this development is an ever‐increasing need for storage within the data centres that host these services. Management applications used by cloud providers to configure their infrastructure should ideally operate in terms of high‐level policy goals, and not burden administrators with the details presented by particular instances of storage systems. One common technology used by cloud providers is the Storage Area Network (SAN). Support for seamless scalability is engineered into SAN devices. However, SAN infrastructure has a very large parameter space: their optimal deployment is a difficult challenge, and subsequent management in cloud storage continues to be difficult. parindent = 10pt In this article, we discuss our work in SAN configuration middleware, which aims to provide users of large‐scale storage infrastructure such as cloud providers with tools to assist them in their management and evolution of heterogeneous SAN environments. We propose a middleware rather than a stand‐alone tool so that the middleware can be a proxy for interacting with, and informing, a central repository of SAN configurations. Storage system users can have their SAN configurations validated against a knowledge base of best practices that are contained within the central repository. Desensitized information is exported from local management applications to the repository, and the local middleware can subscribe to updates that proactively notify storage users should particular configurations be updated to be considered as sub‐optimal, or unsafe. Copyright © 2011 John Wiley & Sons, Ltd. David M. Eyers, Ramani Routray, Douglas Willcocks, Peter R. Pietzuch |
Concurr. Comput. Pract. Exp. | 5 |
| 2010 | Exploring algorithmic trading in reconfigurable hardwareabstractThis paper describes an algorithmic trading engine based on reconfigurable hardware, derived from a software implementation. Our approach exploits parallelism and reconfigurability of field-programmable gate array (FPGA) technology. FPGAs offer many benefits over software solutions, including a reduction in latency, while increasing overall throughput and computational density. All of which are important attributes to a successful algorithmic trading engine. Experiments show that the peak performance of our hardware architecture for algorithmic trading is 133 times faster than the corresponding software implementation. Six implementations can operate simultaneously on a Xilinx Vertex 5 xc5vlx30 FPGA on average, maximising performance and available resource usage. Stephen Wray, Wayne Luk, Peter R. Pietzuch |
ASAP | 3 |
| 2010 | Run-Time Reconfiguration for a Reconfigurable Algorithmic Trading EngineabstractIn this paper we present an analysis of using run-time reconfiguration of reconfigurable hardware to modify trading algorithms during use. This provides flexibility in algorithm design, enabling the implementation to be reactive to changes in market conditions, increasing in performance. We study what can be achieved to reduce performance loss in algorithms while reconfiguration takes place, such as buffering information during this time. Our results show our average partial reconfiguration time is 0.002091 seconds, using historic highest market data rates would result in about 5,000 messages being missed or require buffering. This is the worst case scenario, normally the system would only require a fraction of messages. The reconfiguration time is acceptable if it is under the required limit by the user to prevent business performance suffering. Stephen Wray, Wayne Luk, Peter R. Pietzuch |
FPL | 3 |
| 2010 | Enforcing End-to-End Application Security in the Cloud - (Big Ideas Paper)
Jean Bacon, David Evans 0002, David M. Eyers, Matteo Migliavacca, Peter R. Pietzuch, Brian Shand |
Middleware | 5 |
| 2010 | Distributed Middleware Enforcement of Event Flow Security Policy
Matteo Migliavacca, Ioannis Papagiannis, David M. Eyers, Brian Shand, Jean Bacon, Peter R. Pietzuch |
Middleware | 6 |
| 2010 | DEFCON: High-Performance Event Processing with Information Security
Matteo Migliavacca, Ioannis Papagiannis, David M. Eyers, Brian Shand, Jean Bacon, Peter R. Pietzuch |
USENIX ATC | 6 |
| 2009 | Introduction
Dejan Kostic, Guillaume Pierre, Flavio Paiva Junqueira, Peter R. Pietzuch |
Euro-Par | 4 |
| 2008 | Distributed content delivery using load-aware network coordinatesabstractTo scale to millions of Internet users with good performance, content delivery networks (CDNs) must balance requests between content servers while assigning clients to nearby servers. In this paper, we describe a new CDN design that associates synthetic load-aware coordinates with clients and content servers and uses them to direct content requests to cached content. This approach helps achieve good performance when request workloads and resource availability in the CDN are dynamic. A deployment and evaluation of our system on PlanetLab demonstrates how it achieves low request times with high cache hit ratios when compared to other CDN approaches. Nicholas Ball, Peter R. Pietzuch |
CoNEXT | 2 |
| 2007 | Cobra: Content-based Filtering and Aggregation of Blogs and RSS Feeds
Ian Rose, Rohan Murty, Peter R. Pietzuch, Jonathan Ledlie, Mema Roussopoulos, Matt Welsh |
NSDI | 3 |
| 2006 | Stable and Accurate Network CoordinatesabstractNetwork coordinates provide a scalable way to estimate latencies among large numbers of hosts. While there are several algorithms for producing coordinates, none account for the fact that nodes observe a stream of distinct observations that may vary by as much as three orders-ofmagnitude. With such variable data, coordinate systems are prone to high error and instability in live deployments. In addition, dynamics such as triangle violations can lead to coordinate oscillations, producing further instability and making it difficult for applications to know when their coordinates have truly changed. Because simulation results demonstrate that network coordinates are capable of providing low cost and sufficiently accurate answers to common queries, it is vital that we develop the ability to obtain similar results in practice. We propose two filters which combined to improve network coordinate accuracy by 54% and coordinate stability by 96% when run on a real, largescale network. Jonathan Ledlie, Peter R. Pietzuch, Margo I. Seltzer |
ICDCS | 2 |
| 2006 | Network-Aware Operator Placement for Stream-Processing SystemsabstractTo use their pool of resources efficiently, distributed stream-processing systems push query operators to nodes within the network. Currently, these operators, ranging from simple filters to custom business logic, are placed manually at intermediate nodes along the transmission path to meet application-specific performance goals. Determining placement locations is challenging because network and node conditions change over time and because streams may interact with each other, opening venues for reuse and repositioning of operators. This paper describes a stream-based overlay network (SBON), a layer between a stream-processing system and the physical network that manages operator placement for stream-processing systems. Our design is based on a cost space, an abstract representation of the network and on-going streams, which permits decentralized, large-scale multi-query optimization decisions. We present an evaluation of the SBON approach through simulation, experiments on PlanetLab, and an integration with Borealis, an existing stream-processing engine. Our results show that an SBON consistently improves network utilization, provides low stream latency, and enables dynamic optimization at low engineering cost. Peter R. Pietzuch, Jonathan Ledlie, Jeffrey Shneidman, Mema Roussopoulos, Matt Welsh, Margo I. Seltzer |
ICDE | 1 |
| 2003 | Congestion Control in a Reliable Scalable Message-Oriented Middleware
Peter R. Pietzuch, Sumeer Bhola |
Middleware | 1 |
| 2003 | A Framework for Event Composition in Distributed Systems
Peter R. Pietzuch, Brian Shand, Jean Bacon |
Middleware | 1 |