VLDB 2026 Research / reviewers in the wild / expert
Antony I. T. Rowstron
dblp:r/AITRowstron · also Antony Ian Taylor Rowstron
· DBLP profile ↗
72ranked-venue papers
11as first author
9since 2021 · last 2025
0009-0009-5936-6895ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 1 first-author · 4 since 2021Computer networks · 21 · 2 since 2021Software engineering, systems software and programming languages · 15 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Good things come in small packages: Should we build AI clusters with Lite-GPUs?abstractTo match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages. However, there is growing uncertainty about the scalability of individual GPUs and thus AI clusters, as state-of-the-art GPUs are already displaying packaging, yield, and cooling limitations. We propose to rethink the design and scaling of AI clusters through efficiently-connected large clusters of Lite-GPUs, GPUs with single, small dies and a fraction of the capabilities of larger GPUs. We think recent advances in co-packaged optics can enable distributing AI workloads onto many Lite-GPUs through high bandwidth and efficient communication. In this paper, we present the key benefits of Lite-GPUs on manufacturing cost, blast radius, yield, and power efficiency; and discuss systems opportunities and challenges around resource, workload, memory, and network management. Burcu Canakci, Xingbo Wu, Nathanael Cheriere, Paolo Costa, Sergey Legtchenko, Dushyanth Narayanan, Antony I. T. Rowstron |
HotOS | 8 |
| 2025 | Storage Class Memory is Dead, All Hail Managed-Retention Memory: Rethinking Memory for the AI EraabstractAI clusters today are one of the major uses of High Bandwidth Memory (HBM). However, HBM is suboptimal for AI workloads for several reasons. Analysis shows HBM is overprovisioned on write performance, but underprovisioned on density and read bandwidth, and also has significant energy per bit overheads. It is also expensive, with lower yield than DRAM due to manufacturing complexity. We propose a new memory class: Managed-Retention Memory (MRM), which is more optimized to store key data structures for AI inference workloads. We believe that MRM may finally provide a path to viability for technologies that were originally proposed to support Storage Class Memory (SCM). These technologies traditionally offered long-term persistence (10+ years) but provided poor IO performance and/or endurance. MRM makes different trade-offs, and by understanding the workload IO patterns, MRM foregoes long-term data retention and write performance for better potential performance on the metrics important for these workloads. Sergey Legtchenko, Ioan A. Stefanovici, Richard Black, Antony I. T. Rowstron, Paolo Costa, Burcu Canakci, Dushyanth Narayanan, Xingbo Wu |
HotOS | 4 |
| 2025 | Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and PlanningabstractRobotic manipulation in cluttered environments presents significant challenges, particularly when the clutter includes thin, deformable objects like cables, which complicate perception and decision-making processes. In the context of datacenters, the automation of networking tasks often involves the manipulation of optical transceivers within densely packed cable configurations. Such environments are characterized by an abundance of delicate, overlapping, and intersecting cables, leading to frequent occlusions. This paper introduces an innovative system designed for the manipulation of optical transceivers in environments cluttered by cables. Our integrated approach combines advanced 3D scene understanding with a heuristic-based pushing policy to effectively manipulate optical transceivers amidst clutter. The system's perception component utilizes image segmentation and 3D reconstruction to accurately model the transceivers and surrounding cables. Meanwhile, the planning aspect employs a search algorithm with task-specific heuristics, to navigate the gripper, displace obstructing cables, and safely achieve a precise pre-grasp position in front of the target transceiver. We have conducted extensive evaluations of our methodology in both simulated and real-world settings, demonstrating its high success rates, robustness, and proficiency in addressing the unique challenges posed by cable-occluded environments within datacenters. Iason Sarantopoulos, Bohong Weng, Sicheng Xu, Jiaolong Yang, Xin Tong 0001, Fabian Otto, David Sweeney, Andromachi Chatzieleftheriou, Antony I. T. Rowstron |
ICRA | 11 |
| 2025 | Mosaic: Breaking the Optics versus Copper Trade-off with a Wide-and-Slow Architecture and MicroLEDsabstractLink technologies in today's data center networks impose a fundamental trade-off between reach, power, and reliability. Copper links are power-efficient and reliable but have very limited reach (< 2 m). Optical links offer longer reach but at the expense of high power consumption and lower reliability. As network speeds increase, this trade-off becomes more pronounced, constraining future scalability. Kaoutar Benyahya, Ariel Gomez Diaz, Vassily Lyutsarev, Marianna Pantouvaki, Kai Shi 0003, Shawn Yohanes Siew, Hitesh Ballani, Thomas Burridge, Daniel Cletheroe, Thomas Karagiannis, Antony I. T. Rowstron, Mengyang Yang, Paolo Costa |
SIGCOMM | 13 |
| 2025 | Project Silica: Towards Sustainable Cloud Archival Storage in GlassabstractSustainable and cost-effective long-term storage remains an unsolved problem. The most widely used storage technologies today are magnetic (hard disk drives and tape). They use media that degrades over time and has a limited lifetime, which leads to inefficient, wasteful, and costly solutions for long-lived data. This article presents Silica: the first cloud storage system for archival data underpinned by quartz glass, an extremely resilient media that allows data to be left in situ indefinitely. The hardware and software of Silica have been co-designed and co-optimized from the media up to the service level with sustainability as a primary objective. The design follows a cloud-first, data-driven methodology underpinned by principles derived from analyzing the archival workload of a large public cloud service. Silica can support a wide range of archival storage workloads and ushers in a new era of sustainable, cost-effective storage. Patrick Anderson 0001, Erika Blancada Aranas, Youssef Assaf, Raphael Behrendt, Richard Black, Marco Caballero, Pashmina Cameron, Burcu Canakci, Andromachi Chatzieleftheriou, Rebekah Storan Clarke, James Clegg, Daniel Cletheroe, Bridgette Cooper, Thales De Carvalho, Tim Deegan, Austin Donnelly, Rokas Drevinskas, Alexander L. Gaunt, Christos Gkantsidis, Ariel Gomez Diaz, István Haller, Freddie Hong, Teodora Ilieva, Shashidhar Joshi, Russell Joyce, Mint Kunkel, David Lara Alabazares, Sergey Legtchenko, Fanglin Linda Liu, Bruno Magalhães, Alana Marzoev, Marvin McNett, Jayashree Mohan, Michael Myrah, Sebastian Nowozin, Aaron Ogus, Hiske Overweg, Antony I. T. Rowstron, Maneesh Sah, Masaaki Sakakura, Peter Scholtz, Nina Schreiner, Omer Sella, Ioan A. Stefanovici, David Sweeney, Benn C. Thomsen, Govert Verkes, Phil Wainman, Jonathan Westcott, Luke Weston, Charles Whittaker, Pablo Wilke Berenguer, Hugh Williams, Stefan Winzeck |
ACM Trans. Storage | 39 |
| 2025 | Holographic Storage for the Cloud: advances and challengesabstractHolographic Storage is an old idea that has always promised high density and fast random access, but has never been commercially competitive with Hard Disk Drives (HDDs) and Solid State Devices (SSDs). In Project HSD at Microsoft Research we asked the question: “Does holographic storage finally make sense for cloud storage?” This article describes our journey toward answering this question. We achieved 1.8× higher density than the previous state-of-the-art, using commodity components available today and leveraging machine learning to compensate for the noise and distortions introduced by commodity components. This uncovered two new challenges which are the focus of this article: achieving high end-to-end energy efficiency without sacrificing capacity, and spatial multiplexing without mechanical movement. Improving end-to-end energy efficiency requires joint optimization across low-level media parameters and higher-level system parameters that govern background maintenance operations such as read refresh and garbage collection. We developed new physics models of the media; analytic and simulation models of the media access and background media maintenance; and workload-driven optimization to find optimal parameter combinations. These techniques resulted in a 14× improvement over the previous approach for typical workloads without sacrificing capacity. We also designed the first scalable and mechanical movement free spatial multiplexing system for holographic storage. Despite these advances, we conclude that currently, holographic storage is still far from the combination of density, capacity scaling, and energy efficiency needed to compete with the incumbent technologies. We need fundamental advances in the physical media that improve energy efficiency by another 1–2 orders of magnitude without reducing data density. Further advances in optics are also required to achieve spatial multiplexing that is simultaneously scalable, low-loss, and high-density. Nathanael Cheriere, Jiaqi Chu, Grace Brennan, Pashmina Cameron, Pedro Da Costa, Jannes Gladrow, Guilherme Ilunga, Douglas J. Kelly, Joowon Lim, Giorgio Maltese, Tony Mason, Greg O'Shea, Soujanya Ponnapalli, Michael Rudow, Alan Sanders, Theano Stavrinos, Xingbo Wu, Mengyang Yang, Dushyanth Narayanan, Benn C. Thomsen, Antony I. T. Rowstron |
ACM Trans. Storage | 22 |
| 2024 | Self-maintaining [networked] systems: The rise of datacenter robotics!abstractThe vision of self-maintaining systems is to make cloud hardware automatically servicing and repairing using robotics. We define a self-maintaining system as one where software can control robotics that can automatically perform hardware maintenance tasks and repair operations. This reduces failure service windows and lowers the risk of repairs causing further cascading failures and outages. Self-maintaining systems are not purely reactive to failures, but also do proactive maintenance before failures occur which reduces future hardware failures. Operating an entire datacenter as a self-maintaining system is many years away, and we present four stages of automation, analogous to levels used for autonomous vehicles, required to reach the full vision for datacenters. Freddie Hong, Iason Sarantopoulos, Elliott Hogg, Hugh Williams, David Sweeney, Andromachi Chatzieleftheriou, Antony I. T. Rowstron |
HotNets | 9 |
| 2024 | RASCAL: A Scalable, High-redundancy Robot for Automated Storage and Retrieval SystemsabstractAutomated storage and retrieval systems (ASRS) are a key component of the modern storage industry, and are used in a wide range of applications, carrying anything from lightweight tape cartridges to entire pallets of goods. Many of these systems are under pressure to maximise the use of space by growing in height and density, but this can create challenges for the the robots that service them. In this context, we present RASCAL, a novel ASRS robot for small payload items in structured environments, with a focus on system-level scalability and redundancy. We describe the design objectives of RASCAL and how they address some of the limitations of existing robotic systems in this area, such as scalability and redundancy. We then demonstrate the viability of our design with a proof-of-concept implementation of a data centre storage media robot, and show through a series of experiments that its design, speed, accuracy, and energy efficiency are appropriate for this application. Richard Black, Marco Caballero, Andromachi Chatzieleftheriou, Tim Deegan, Philip Heard, Freddie Hong, Russell Joyce, Sergey Legtchenko, Antony I. T. Rowstron, David Sweeney, Hugh Williams |
ICRA | 9 |
| 2023 | Project Silica: Towards Sustainable Cloud Archival Storage in GlassabstractSustainable and cost-effective long-term storage remains an unsolved problem. The most widely used storage technologies today are magnetic (hard disk drives and tape). They use media that degrades over time and has a limited lifetime, which leads to inefficient, wasteful, and costly solutions for long-lived data. This paper presents Silica: the first cloud storage system for archival data underpinned by quartz glass, an extremely resilient media that allows data to be left in situ indefinitely. The hardware and software of Silica have been co-designed and co-optimized from the media up to the service level with sustainability as a primary objective. The design follows a cloud-first, data-driven methodology underpinned by principles derived from analyzing the archival workload of a large public cloud service. Silica can support a wide range of archival storage workloads and ushers in a new era of sustainable, cost-effective storage. Patrick Anderson 0001, Erika Blancada Aranas, Youssef Assaf, Raphael Behrendt, Richard Black, Marco Caballero, Pashmina Cameron, Burcu Canakci, Thales De Carvalho, Andromachi Chatzieleftheriou, Rebekah Storan Clarke, James Clegg, Daniel Cletheroe, Bridgette Cooper, Tim Deegan, Austin Donnelly, Rokas Drevinskas, Alexander L. Gaunt, Christos Gkantsidis, Ariel Gomez Diaz, István Haller, Freddie Hong, Teodora Ilieva, Shashidhar Joshi, Russell Joyce, Mint Kunkel, David Lara Alabazares, Sergey Legtchenko, Fanglin Linda Liu, Bruno Magalhães, Alana Marzoev, Marvin McNett, Jayashree Mohan, Michael Myrah, Sebastian Nowozin, Aaron Ogus, Hiske Overweg, Antony I. T. Rowstron, Maneesh Sah, Masaaki Sakakura, Peter Scholtz, Nina Schreiner, Omer Sella, Ioan A. Stefanovici, David Sweeney, Benn C. Thomsen, Govert Verkes, Phil Wainman, Jonathan Westcott, Luke Weston, Charles Whittaker, Pablo Wilke Berenguer, Hugh Williams, Stefan Winzeck |
SOSP | 39 |
| 2020 | Could cloud storage be disrupted in the next decade?
Andromachi Chatzieleftheriou, Ioan A. Stefanovici, Dushyanth Narayanan, Benn C. Thomsen, Antony I. T. Rowstron |
HotStorage | 5 |
| 2018 | Glass: A New Media for a New Era?
Patrick Anderson 0001, Richard Black, Ausra Cerkauskaite, Andromachi Chatzieleftheriou, James Clegg, Chris Dainty, Raluca Diaconu, Rokas Drevinskas, Austin Donnelly, Alexander L. Gaunt, Andreas Georgiou, Ariel Gomez Diaz, Peter G. Kazansky, David Lara Alabazares, Sergey Legtchenko, Sebastian Nowozin, Aaron Ogus, Douglas Phillips, Antony I. T. Rowstron, Masaaki Sakakura, Ioan A. Stefanovici, Benn C. Thomsen, Hugh Williams, Mengyang Yang |
HotStorage | 19 |
| 2018 | Larry: Practical Network Reconfigurability in the Data Center
Andromachi Chatzieleftheriou, Sergey Legtchenko, Hugh Williams, Antony I. T. Rowstron |
NSDI | 4 |
| 2017 | Understanding Rack-Scale Disaggregated Storage
Sergey Legtchenko, Hugh Williams, Kaveh Razavi, Austin Donnelly, Richard Black, Andrew Douglas, Nathanael Cheriere, Daniel Fryer, Kai Mast, Angela Demke Brown, Ana Klimovic, Andy Slowey, Antony I. T. Rowstron |
HotStorage | 13 |
| 2016 | Flamingo: Enabling Evolvable HDD-based Near-Line Storage
Sergey Legtchenko, Antony I. T. Rowstron, Austin Donnelly, Richard Black |
FAST | 3 |
| 2016 | Feeding the Pelican: Using Archival Hard Drives for Cold Storage Racks
Richard Black, Austin Donnelly, David Harper, Aaron Ogus, Antony I. T. Rowstron |
HotStorage | 5 |
| 2016 | XFabric: A Reconfigurable In-Rack Network for Rack-Scale Computers
Sergey Legtchenko, Nicholas Chen, Daniel Cletheroe, Antony I. T. Rowstron, Hugh Williams, Xiaohan Zhao |
NSDI | 4 |
| 2015 | Software-defined caching: managing caches in multi-tenant data centersabstractIn data centers, caches work both to provide low IO latencies and to reduce the load on the back-end network and storage. But they are not designed for multi-tenancy; system-level caches today cannot be configured to match tenant or provider objectives. Exacerbating the problem is the increasing number of un-coordinated caches on the IO data plane. The lack of global visibility on the control plane to coordinate this distributed set of caches leads to inefficiencies, increasing cloud provider cost. Ioan A. Stefanovici, Eno Thereska, Greg O'Shea, Bianca Schroeder, Hitesh Ballani, Thomas Karagiannis, Antony I. T. Rowstron, Tom Talpey |
SoCC | 7 |
| 2014 | Towards Paravirtualized Network File Systems
Raja Appuswamy, Sergey Legtchenko, Antony I. T. Rowstron |
HotStorage | 3 |
| 2014 | Pelican: A Building Block for Exascale Cold Data Storage
Shobana Balakrishnan, Richard Black, Austin Donnelly, Paul England, Adam Glass, David Harper, Sergey Legtchenko, Aaron Ogus, Eric Peterson, Antony I. T. Rowstron |
OSDI | 10 |
| 2014 | Decentralized task-aware scheduling for data center networksabstractMany data center applications perform rich and complex tasks (e.g., executing a search query or generating a user's news-feed). From a network perspective, these tasks typically comprise multiple flows, which traverse different parts of the network at potentially different times. Most network resource allocation schemes, however, treat all these flows in isolation -- rather than as part of a task -- and therefore only optimize flow-level metrics. Fahad R. Dogar, Thomas Karagiannis, Hitesh Ballani, Antony I. T. Rowstron |
SIGCOMM | 4 |
| 2013 | Scale-up vs scale-out for Hadoop: time to rethink?abstractIn the last decade we have seen a huge deployment of cheap clusters to run data analytics workloads. The conventional wisdom in industry and academia is that scaling out using a cluster of commodity machines is better for these workloads than scaling up by adding more resources to a single server. Popular analytics infrastructures such as Hadoop are aimed at such a cluster scale-out environment. Raja Appuswamy, Christos Gkantsidis, Dushyanth Narayanan, Orion Hodson, Antony I. T. Rowstron |
SoCC | 5 |
| 2013 | CamCubeOS: a key-based network stack for 3D torus cluster topologies
Paolo Costa, Austin Donnelly, Greg O'Shea, Antony I. T. Rowstron |
HPDC | 4 |
| 2013 | Rhea: Automatic Filtering for Unstructured Cloud Storage
Christos Gkantsidis, Dimitrios Vytiniotis, Orion Hodson, Dushyanth Narayanan, Florin Dinu, Antony I. T. Rowstron |
NSDI | 6 |
| 2013 | IOFlow: a software-defined storage architectureabstractIn data centers, the IO path to storage is long and complex. It comprises many layers or "stages" with opaque interfaces between them. This makes it hard to enforce end-to-end policies that dictate a storage IO flow's performance (e.g., guarantee a tenant's IO bandwidth) and routing (e.g., route an untrusted VM's traffic through a sanitization middlebox). These policies require IO differentiation along the flow path and global visibility at the control plane. We design IOFlow, an architecture that uses a logically centralized control plane to enable high-level flow policies. IOFlow adds a queuing abstraction at data-plane stages and exposes this to the controller. The controller can then translate policies into queuing rules at individual stages. It can also choose among multiple stages for policy enforcement. Eno Thereska, Hitesh Ballani, Greg O'Shea, Thomas Karagiannis, Antony I. T. Rowstron, Tom Talpey, Richard Black, Timothy Zhu |
SOSP | 5 |
| 2013 | Measurement-Based Design of Roadside Content Delivery SystemsabstractWith today's ubiquity of thin computing devices, mobile users are accustomed to having rich location-aware information at their fingertips, such as restaurant menus, shopping mall maps, movie showtimes, and trailers. However, delivering rich content is challenging, particularly for highly mobile users in vehicles. Technologies such as cellular-3G provide limited bandwidth at significant costs. In contrast, providers can cheaply and easily deploy a small number of WiFi infostations that quickly deliver large content to vehicles passing by for future offline browsing. While several projects have proposed systems for disseminating content via roadside infostations, most use simplified models and simulations to guide their design for scalability. Many suspect that scalability with increasing vehicle density is the major challenge for infostations, but few if any have studied the performance of these systems via real measurements. Intuitively, per-vehicle throughput for unicast infostations degrades with the number of vehicles near the infostation, while broadcast infostations are unreliable, and lack rate adaptation. In this work, we collect over 200 h of detailed highway measurements with a fleet of WiFi-enabled vehicles. We use analysis of these results to explore the design space of WiFi infostations, in order to determine whether unicast or broadcast should be used to build high-throughput infostations that scale with device density. Our measurement results demonstrate the limitations of both approaches. Our insights lead to Starfish, a high-bandwidth and scalable infostation system that incorporates device-to-device data scavenging, where nearby vehicles share data received from the infostation. Data scavenging increases dissemination throughput by a factor of 2-6, allowing both broadcast and unicast throughput to scale with device density. Vinod Kone, Haitao Zheng 0001, Antony I. T. Rowstron, Greg O'Shea, Ben Y. Zhao |
IEEE Trans. Mob. Comput. | 3 |
| 2012 | Bridging the tenant-provider gap in cloud servicesabstractThe disconnect between the resource-centric interface exposed by today's cloud providers and tenant goals hurts both entities. Tenants are encumbered by having to translate their performance and cost goals into the corresponding resource requirements, while providers suffer revenue loss due to un-informed resource selection by tenants. Instead, we argue for a "job-centric" cloud whereby tenants only specify high-level goals regarding their jobs and applications. To illustrate our ideas, we present Bazaar, a cloud framework offering a job-centric interface for data analytics applications. Virajith Jalaparti, Hitesh Ballani, Paolo Costa, Thomas Karagiannis, Antony I. T. Rowstron |
SoCC | 5 |
| 2012 | Camdoop: Exploiting In-network Aggregation for Big Data Applications
Paolo Costa, Austin Donnelly, Antony I. T. Rowstron, Greg O'Shea |
NSDI | 3 |
| 2011 | The price is right: towards location-independent costs in datacentersabstractThe performance and cost for tenants in today's datacenters depends on the location of their virtual machines within the datacenter. However, a tenant's location is a knob for the provider and is of no interest to the tenant. Hence, this paper argues for location independent tenant costs in datacenters. We show how a change in today's IaaS offerings, coupled with a simple pricing scheme, can achieve this. We discuss how such a pricing model can be implemented and show that the consequent increase in system throughput can lead to a win-win situation- tenant costs are location independent and lower while provider revenue increases too. Hitesh Ballani, Paolo Costa, Thomas Karagiannis, Antony I. T. Rowstron |
HotNets | 4 |
| 2011 | Towards predictable datacenter networksabstractThe shared nature of the network in today's multi-tenant datacenters implies that network performance for tenants can vary significantly. This applies to both production datacenters and cloud environments. Network performance variability hurts application performance which makes tenant costs unpredictable and causes provider revenue loss. Motivated by these factors, this paper makes the case for extending the tenant-provider interface to explicitly account for the network. We argue this can be achieved by providing tenants with a virtual network connecting their compute instances. To this effect, the key contribution of this paper is the design of virtual network abstractions that capture the trade-off between the performance guarantees offered to tenants, their costs and the provider revenue. Hitesh Ballani, Paolo Costa, Thomas Karagiannis, Antony I. T. Rowstron |
SIGCOMM | 4 |
| 2011 | Better never than late: meeting deadlines in datacenter networksabstractThe soft real-time nature of large scale web applications in today's datacenters, combined with their distributed workflow, leads to deadlines being associated with the datacenter application traffic. A network flow is useful, and contributes to application throughput and operator revenue if, and only if, it completes within its deadline. Today's transport pro- tocols (TCP included), given their Internet origins, are agnostic to such flow deadlines. Instead, they strive to share network resources fairly. We show that this can hurt application performance. Christo Wilson, Hitesh Ballani, Thomas Karagiannis, Antony I. T. Rowstron |
SIGCOMM | 4 |
| 2011 | The Impact of Infostation Density on Vehicular Data DisseminationabstractVehicle-to-Vehicle and Vehicle-to-Roadside communications are going to become an indispensable part of the modern day automotive experience. For people on the move, vehicular networks can provide critical network connectivity and access to real-time information. Infostations play a vital role in these networks by acting as gateways to the Internet and by extending network connectivity. In this context, an important question is “What is the minimum number of infostations that need to be deployed in an area in order to support vehicular applications?” Optimizing infostation density is vital to understanding and reducing the cost of deployment and management. In this paper, we examine the required infostation density in a highway scenario using different data dissemination models. We start from a simple analysis that captures the required density under idealized assumptions. These models are validated by an event-driven simulator. We then run detailed QualNet simulations on both controlled and realistic vehicular traces to observe the information density trends in practical environments, and consequently propose techniques to improve dissemination performance and reduce the required infostation density. Vinod Kone, Haitao Zheng 0001, Antony I. T. Rowstron, Ben Y. Zhao |
Mob. Networks Appl. | 3 |
| 2010 | Hermes: clustering users in large-scale e-mail servicesabstractHermes is an optimization engine for large-scale enterprise e-mail services. Such services could be hosted by a virtualized e-mail service provider, or by dedicated enterprise data centers. In both cases we observe that the pattern of e-mails between employees of an enterprise forms an implicit social graph. Hermes tracks this implicit social graph, periodically identifies clusters of strongly connected users within the graph, and co-locates such users on the same server. Co-locating the users reduces storage requirements: senders and recipients who reside on the same server can share a single copy of an e-mail. Co-location also reduces inter-server bandwidth usage. Thomas Karagiannis, Christos Gkantsidis, Dushyanth Narayanan, Antony I. T. Rowstron |
SoCC | 4 |
| 2010 | Symbiotic routing in future data centersabstractBuilding distributed applications that run in data centers is hard. The CamCube project explores the design of a shipping container sized data center with the goal of building an easier platform on which to build these applications. CamCube replaces the traditional switch-based network with a 3D torus topology, with each server directly connected to six other servers. As in other proposals, e.g. DCell and BCube, multi-hop routing in CamCube requires servers to participate in packet forwarding. To date, as in existing data centers, these approaches have all provided a single routing protocol for the applications. Hussam Abu-Libdeh, Paolo Costa, Antony I. T. Rowstron, Greg O'Shea, Austin Donnelly |
SIGCOMM | 3 |
| 2009 | Feasibility of content dissemination between devices in moving vehiclesabstractWe investigate the feasibility of content distribution between devices mounted in moving vehicles using commodity WiFi. We assume that each device stores content in a set of files, and that each file has a version number. When two devices come into wireless range, they attempt to synchronize the latest versions of any files they have in common. This is challenging because connections are often short-lived and have variable link quality. Prior work demonstrates that current protocols perform badly under these conditions. To motivate this work, we use the example of Personal Navigation Devices (PNDs), or SatNavs, where the content to be exchanged includes maps and points-of-interest files. Thomas Zahn, Greg O'Shea, Antony I. T. Rowstron |
CoNEXT | 3 |
| 2009 | What Is in a Namespace?
Antony I. T. Rowstron |
Euro-Par | 1 |
| 2009 | Migrating server storage to SSDs: analysis of tradeoffsabstractRecently, flash-based solid-state drives (SSDs) have become standard options for laptop and desktop storage, but their impact on enterprise server storage has not been studied. Provisioning server storage is challenging. It requires optimizing for the performance, capacity, power and reliability needs of the expected workload, all while minimizing financial costs. In this paper we analyze a number of workload traces from servers in both large and small data centers, to decide whether and how SSDs should be used to support each. We analyze both complete replacement of disks by SSDs, as well as use of SSDs as an intermediate tier between disks and DRAM. We describe an automated tool that, given device models and a block-level trace of a workload, determines the least-cost storage configuration that will support the workload's performance, capacity, and fault-tolerance requirements. We found that replacing disks by SSDs is not a costeffective option for any of our workloads, due to the low capacity per dollar of SSDs. Depending on the workload, the capacity per dollar of SSDs needs to increase by a factor of 3-3000 for an SSD-based solution to break even with a diskbased solution. Thus, without a large increase in SSD capacity per dollar, only the smallest volumes, such as system boot volumes, can be cost-effectively migrated to SSDs. The benefit of using SSDs as an intermediate caching tier is also limited: fewer than 10% of our workloads can reduce provisioning costs by using an SSD tier at today's capacity per dollar, and fewer than 20% can do so at any SSD capacity per dollar. Although SSDs are much more energy-efficient than enterprise disks, the energy savings are outweighed by the hardware costs, and comparable energy savings are achievable with low-power SATA disks. Dushyanth Narayanan, Eno Thereska, Austin Donnelly, Sameh Elnikety, Antony I. T. Rowstron |
EuroSys | 5 |
| 2008 | Write Off-Loading: Practical Power Management for Enterprise Storage
Dushyanth Narayanan, Austin Donnelly, Antony I. T. Rowstron |
FAST | 3 |
| 2008 | Everest: Scaling Down Peak Loads Through I/O Off-Loading
Dushyanth Narayanan, Austin Donnelly, Eno Thereska, Sameh Elnikety, Antony I. T. Rowstron |
OSDI | 5 |
| 2008 | Network exception handlers: host-network control in enterprise networksabstractEnterprise network architecture and management have followed the Internet's design principles despite different requirements and characteristics: enterprise hosts are administered by a single authority, which intrinsically assigns different values to traffic from different business applications. Thomas Karagiannis, Richard Mortier, Antony I. T. Rowstron |
SIGCOMM | 3 |
| 2008 | Vigilante: End-to-end containment of Internet worm epidemicsabstractWorm containment must be automatic because worms can spread too fast for humans to respond. Recent work proposed network-level techniques to automate worm containment; these techniques have limitations because there is no information about the vulnerabilities exploited by worms at the network level. We propose Vigilante, a new end-to-end architecture to contain worms automatically that addresses these limitations. In Vigilante, hosts detect worms by instrumenting vulnerable programs to analyze infection attempts. We introduce dynamic data-flow analysis : a broad-coverage host-based algorithm that can detect unknown worms by tracking the flow of data from network messages and disallowing unsafe uses of this data. We also show how to integrate other host-based detection mechanisms into the Vigilante architecture. Upon detection, hosts generate self-certifying alerts (SCAs), a new type of security alert that can be inexpensively verified by any vulnerable host. Using SCAs, hosts can cooperate to contain an outbreak, without having to trust each other. Vigilante broadcasts SCAs over an overlay network that propagates alerts rapidly and resiliently. Hosts receiving an SCA protect themselves by generating filters with vulnerability condition slicing : an algorithm that performs dynamic analysis of the vulnerable program to identify control-flow conditions that lead to successful attacks. These filters block the worm attack and all its polymorphic mutations that follow the execution path identified by the SCA. Our results show that Vigilante can contain fast-spreading worms that exploit unknown vulnerabilities, and that Vigilante's filters introduce a negligible performance overhead. Vigilante does not require any changes to hardware, compilers, operating systems, or the source code of vulnerable programs; therefore, it can be used to protect current software binaries. Manuel Costa, Jon Crowcroft, Miguel Castro 0001, Antony I. T. Rowstron, Lidong Zhou, Paul Barham 0001 |
ACM Trans. Comput. Syst. | 4 |
| 2008 | Write off-loading: Practical power management for enterprise storageabstractIn enterprise data centers power usage is a problem impacting server density and the total cost of ownership. Storage uses a significant fraction of the power budget and there are no widely deployed power-saving solutions for enterprise storage systems. The traditional view is that enterprise workloads make spinning disks down ineffective because idle periods are too short. We analyzed block-level traces from 36 volumes in an enterprise data center for one week and concluded that significant idle periods exist, and that they can be further increased by modifying the read/write patterns using write off-loading . Write off-loading allows write requests on spun-down disks to be temporarily redirected to persistent storage elsewhere in the data center. The key challenge is doing this transparently and efficiently at the block level, without sacrificing consistency or failure resilience. We describe our write off-loading design and implementation that achieves these goals. We evaluate it by replaying portions of our traces on a rack-based testbed. Results show that just spinning disks down when idle saves 28--36% of energy, and write off-loading further increases the savings to 45--60%. Dushyanth Narayanan, Austin Donnelly, Antony I. T. Rowstron |
ACM Trans. Storage | 3 |
| 2008 | Delay aware querying with Seaweed
Dushyanth Narayanan, Austin Donnelly, Richard Mortier, Antony I. T. Rowstron |
VLDB J. | 4 |
| 2006 | Network coding with traffic engineeringabstractIn network coding, a router in the network mixes information from different flows. In the seminal work by Ahlswede et al [1], network coding is established as a technique to potentially increase the network capacity. Miguel Castro 0001, Jon Crowcroft, Greg O'Shea, Antony I. T. Rowstron |
CoNEXT | 5 |
| 2006 | POS: A Practical Order Statistics Service forWireless Sensor NetworksabstractWe present the design and implementation of POS, an in-network service that computes accurate order statistics energy-efficiently. POS returns a stream of periodic samples from any order statistic. It initially computes the value of the order statistic and then periodically runs a validation protocol to determine whether the value is still valid. If not, it uses an optimized binary search to determine the new value and then resumes periodic validation. POS uses in-network aggregation and transmission suppression to reduce communication complexity. Results from both experiments on a mote testbed and simulations show that POS can compute order statistics accurately while consuming less energy than the best techniques to compute averages in common cases. Landon P. Cox, Miguel Castro 0001, Antony I. T. Rowstron |
ICDCS | 3 |
| 2006 | Virtual ring routing: network routing inspired by DHTsabstractThis paper presents Virtual Ring Routing (VRR), a new network routing protocol that occupies a unique point in the design space. VRR is inspired by overlay routing algorithms in Distributed Hash Tables (DHTs) but it does not rely on an underlying network routing protocol. It is implemented directly on top of the link layer. VRR provides both raditional point-to-point network routing and DHT routing to the node responsible for a hash table key.VRR can be used with any link layer technology but this paper describes a design and several implementations of VRR that are tuned for wireless networks. We evaluate the performance of VRR using simulations and measurements from a sensor network and an 802.11a testbed. The experimental results show that VRR provides robust performance across a wide range of environments and workloads. It performs comparably to, or better than, the best wireless routing protocol in each experiment. VRR performs well because of its unique features: it does not require network flooding or trans-lation between fixed identifiers and location-dependent addresses. Matthew Caesar 0001, Miguel Castro 0001, Ed Nightingale, Greg O'Shea, Antony I. T. Rowstron |
SIGCOMM | 5 |
| 2006 | Delay Aware Querying with Seaweed
Dushyanth Narayanan, Austin Donnelly, Richard Mortier, Antony I. T. Rowstron |
VLDB | 4 |
| 2005 | Topic 15 - Peer-to-Peer and Web Computing
Anne-Marie Kermarrec, Márk Jelasity, Antony I. T. Rowstron, Henrique João L. Domingos |
Euro-Par | 3 |
| 2005 | Debunking Some Myths About Structured and Unstructured Overlays
Miguel Castro 0001, Manuel Costa, Antony I. T. Rowstron |
NSDI | 3 |
| 2005 | Cashmere: Resilient Anonymous Routing
Li Zhuang, Ben Y. Zhao, Antony I. T. Rowstron |
NSDI | 4 |
| 2005 | Vigilante: end-to-end containment of internet wormsabstractWorm containment must be automatic because worms can spread too fast for humans to respond. Recent work has proposed network-level techniques to automate worm containment; these techniques have limitations because there is no information about the vulnerabilities exploited by worms at the network level. We propose Vigilante, a new end-to-end approach to contain worms automatically that addresses these limitations. Vigilante relies on collaborative worm detection at end hosts, but does not require hosts to trust each other. Hosts run instrumented software to detect worms and broadcast self-certifying alerts (SCAs) upon worm detection. SCAs are proofs of vulnerability that can be inexpensively verified by any vulnerable host. When hosts receive an SCA, they generate filters that block infection by analysing the SCA-guided execution of the vulnerable software. We show that Vigilante can automatically contain fast-spreading worms that exploit unknown vulnerabilities without blocking innocuous traffic. Manuel Costa, Jon Crowcroft, Miguel Castro 0001, Antony I. T. Rowstron, Lidong Zhou, Paul Barham 0001 |
SOSP | 4 |
| 2004 | Performance and Dependability of Structured Peer-to-Peer OverlaysabstractStructured peer-to-peer (P2P) overlay networks provide a useful substrate for building distributed applications. They map object keys to overlay nodes and offer a primitive to send a message to the node responsible for a key. They can implement, for example, distributed hash tables and multicast trees. However, there are concerns about the performance and dependability of these overlays in realistic environments. Several studies have shown that current P2P environments have high churn rates: nodes join and leave the overlay continuously. This paper presents techniques that continuously detect faults and repair the overlay to achieve high dependability and good performance in realistic environments. The techniques are evaluated using large-scale network simulation experiments with fault injection guided by real traces of node arrivals and departures. The results show that previous concerns are unfounded; our techniques can achieve dependable routing in realistic environments with an average delay stretch below two and a maintenance overhead of less than half a message per second per node. Miguel Castro 0001, Manuel Costa, Antony I. T. Rowstron |
DSN | 3 |
| 2004 | PIC: Practical Internet Coordinates for Distance EstimationabstractWe introduce PIC, a practical coordinate-based mechanism to estimate Internet network distance (i.e., round-trip delay or network hops). Network distance estimation is important in many applications; for example, network-aware overlay construction and server selection. There are several proposals for distance estimation in the Internet but they all suffer from problems that limit their benefit. Most rely on a small set of infrastructure nodes that are a single point of failure and limit scalability. Others use sets of peers to compute coordinates but these coordinates can be arbitrarily wrong if one of these peers is malicious. While it may be reasonable to secure a small set of infrastructure nodes, it is unreasonable to secure all peers. PIC addresses these problems: it does not rely on infrastructure nodes and it can compute accurate coordinates even when some peers are malicious. We present PIC's design, experimental evaluation, and an application to network-aware overlay construction and maintenance. Manuel Costa, Miguel Castro 0001, Antony I. T. Rowstron, Peter B. Key |
ICDCS | 3 |
| 2003 | An Evaluation of Scalable Application-Level Multicast Built Using Peer-To-Peer OverlaysabstractStructured peer-to-peer overlay networks such as CAN, Chord, Pastry, and Tapestry can be used to implement Internet-scale application-level multicast. There are two general approaches to accomplishing this: tree building and flooding. This paper evaluates these two approaches using two different types of structured overlay: 1) overlays which use a form of generalized hypercube routing, e.g., Chord, Pastry and Tapestry, and 2) overlays which use a numerical distance metric to route through a Cartesian hyperspace, e.g., CAN. Pastry and CAN are chosen as the representatives of each type of overlay. To the best of our knowledge, this paper reports the first head-to-head comparison of CAN-style versus Pastry-style overlay networks, using multicast communication workloads running on an identical simulation infrastructure. The two approaches to multicast are independent of overlay network choice, and we provide a comparison of flooding versus tree-based multicast on both overlays. Results show that the tree-based approach consistently outperforms the flooding approach. Finally, for tree-based multicast, we show that Pastry provides better performance than CAN. Miguel Castro 0001, Michael B. Jones, Anne-Marie Kermarrec, Antony I. T. Rowstron, Marvin Theimer, Helen J. Wang, Alec Wolman |
INFOCOM | 4 |
| 2003 | SplitStream: high-bandwidth multicast in cooperative environmentsabstractIn tree-based multicast systems, a relatively small number of interior nodes carry the load of forwarding multicast messages. This works well when the interior nodes are highly-available, dedicated infrastructure routers but it poses a problem for application-level multicast in peer-to-peer systems. SplitStream addresses this problem by striping the content across a forest of interior-node-disjoint multicast trees that distributes the forwarding load among all participating peers. For example, it is possible to construct efficient SplitStream forests in which each peer contributes only as much forwarding bandwidth as it receives. Furthermore, with appropriate content encodings, SplitStream is highly robust to failures because a node failure causes the loss of a single stripe on average. We present the design and implementation of SplitStream and show experimental results obtained on an Internet testbed and via large-scale network simulation. The results show that SplitStream distributes the forwarding load among all peers and can accommodate peers with different bandwidth capacities while imposing low overhead for forest construction and maintenance. Miguel Castro 0001, Peter Druschel, Anne-Marie Kermarrec, Animesh Nandi, Antony I. T. Rowstron, Atul Singh |
SOSP | 5 |
| 2003 | Using mobile code to provide fault tolerance in tuple space based coordination languages
Antony I. T. Rowstron |
Sci. Comput. Program. | 1 |
| 2002 | State- and Event-Based Reactive Programming in Shared Dataspaces
Nadia Busi, Antony I. T. Rowstron, Gianluigi Zavattaro |
COORDINATION | 2 |
| 2002 | Secure Routing for Structured Peer-to-Peer Overlay Networks
Miguel Castro 0001, Peter Druschel, Ayalvadi J. Ganesh, Antony I. T. Rowstron, Dan S. Wallach |
OSDI | 4 |
| 2002 | Squirrel: a decentralized peer-to-peer web cacheabstractThis paper presents a decentralized, peer-to-peer web cache called Squirrel. The key idea is to enable web browsers on desktop machines to share their local caches, to form an efficient and scalable web cache, without the need for dedicated hardware and the associated administrative cost. We propose and evaluate decentralized web caching algorithms for Squirrel, and discover that it exhibits performance comparable to a centralized web cache in terms of hit ratio, bandwidth usage and latency. It also achieves the benefits of decentralization, such as being scalable, self-organizing and resilient to node failures, while imposing low overhead on the participating nodes. Sitaram Iyer, Antony I. T. Rowstron, Peter Druschel |
PODC | 2 |
| 2002 | Scribe: a large-scale and decentralized application-level multicast infrastructureabstractThis paper presents Scribe, a scalable application-level multicast infrastructure. Scribe supports large numbers of groups, with a potentially large number of members per group. Scribe is built on top of Pastry, a generic peer-to-peer object location and routing substrate overlayed on the Internet, and leverages Pastry's reliability, self-organization, and locality properties. Pastry is used to create and manage groups and to build efficient multicast trees for the dissemination of messages to each group. Scribe provides best-effort reliability guarantees, and we outline how an application can extend Scribe to provide stronger reliability. Simulation results, based on a realistic network topology model, show that Scribe scales across a wide range of groups and group sizes. Also, it balances the load on the nodes while achieving acceptable delay and link stress when compared with Internet protocol multicast. Miguel Castro 0001, Peter Druschel, Anne-Marie Kermarrec, Antony I. T. Rowstron |
IEEE J. Sel. Areas Commun. | 4 |
| 2001 | PAST: A large-scale, persistent peer-to-peer storage utilityabstractThis paper sketches the design of PAST, a large-scale, Internet-based, global storage utility that provides scalability, high availability, persistence and security. PAST is a peer-to-peer Internet application and is entirely self-organizing. PAST nodes serve as access points for clients, participate in the routing of client requests, and contribute storage to the system. Nodes are not trusted, they may join the system at any time and may silently leave the system without warning. Yet, the system is able to provide strong assurances, efficient storage access, load balancing and scalability. Among the most interesting aspects of PAST's design are (1) the Pastry location and routing scheme, which reliably and efficiently routes client requests among the PAST nodes, has good network locality properties and automatically resolves node failures and node additions; (2) the use of randomization to ensure diversity in the set of nodes that store a file's replicas and to provide load balancing; and (3) the optional use of smartcards, which are held by each PAST user and issued by a third party called a broker The smartcards support a quota system that balances supply and demand of storage in the system. Peter Druschel, Antony I. T. Rowstron |
HotOS | 2 |
| 2001 | Probabilistic Modelling of Replica DivergenceabstractIt is common in distributed systems to replicate data. In many cases this data evolves in a consistent fashion, and this evolution can be modelled. A probabilistic model of the evolution allows us to estimate the divergence of the replicas and can be used by the application to alter its behaviour, for example to control synchronisation times, to determine the propagation of writes, and to convey to the user information about how much the data may have evolved. In this paper, we describe how the evolution of the data may be modelled and outline how the probabilistic model may be utilised in various applications, concentrating on a news database example. Antony I. T. Rowstron, Neil D. Lawrence, Christopher M. Bishop |
HotOS | 1 |
| 2001 | Pastry: Scalable, Decentralized Object Location, and Routing for Large-Scale Peer-to-Peer Systems
Antony I. T. Rowstron, Peter Druschel |
Middleware | 1 |
| 2001 | Optimising Synchronisation Times for Mobile DevicesabstractWith the increasing number of users of mobile computing devices (e.g. personal digital assistants) and the advent of third generation mobile phones, wireless communications are becoming increasingly important. Many applications rely on the device maintaining a replica of a data-structure which is stored on a server, for exam(cid:173) ple news databases, calendars and e-mail. ill this paper we explore the question of the optimal strategy for synchronising such replicas. We utilise probabilistic models to represent how the data-structures evolve and to model user behaviour. We then formulate objective functions which can be minimised with respect to the synchronisa(cid:173) tion timings. We demonstrate, using two real world data-sets, that a user can obtain more up-to-date information using our approach. Neil D. Lawrence, Antony I. T. Rowstron, Christopher M. Bishop, M. J. Taylor |
NIPS | 2 |
| 2001 | The IceCube approach to the reconciliation of divergent replicasabstractWe describe a novel approach to log-based reconciliation called IceCube. It is general and is parameterised by application and object semantics. IceCube considers more flexible orderings and is designed to ease the burden of reconciliation on the application programmers. IceCube captures the static and dynamic reconciliation constraints between all pairs of actions, proposes schedules that satisfy the static constraints, and validates them against the dynamic constraints. Anne-Marie Kermarrec, Antony I. T. Rowstron, Marc Shapiro 0001, Peter Druschel |
PODC | 2 |
| 2001 | Storage Management and Caching in PAST, A Large-scale, Persistent Peer-to-peer Storage UtilityabstractThis paper presents and evaluates the storage management and caching in PAST, a large-scale peer-to-peer persistent storage utility. PAST is based on a self-organizing, Internet-based overlay network of storage nodes that cooperatively route file queries, store multiple replicas of files, and cache additional copies of popular files.In the PAST system, storage nodes and files are each assigned uniformly distributed identifiers, and replicas of a file are stored at nodes whose identifier matches most closely the file's identifier. This statistical assignment of files to storage nodes approximately balances the number of files stored on each node. However, non-uniform storage node capacities and file sizes require more explicit storage load balancing to permit graceful behavior under high global storage utilization; likewise, non-uniform popularity of files requires caching to minimize fetch distance and to balance the query load.We present and evaluate PAST, with an emphasis on its storage management and caching system. Extensive trace-driven experiments show that the system minimizes fetch distance, that it balances the query load for popular files, and that it displays graceful degradation of performance as the global storage utilization increases beyond 95%. Antony I. T. Rowstron, Peter Druschel |
SOSP | 1 |
| 2000 | Proving the Correctness of Optimising Destructive and Non-destructive Reads over Tuple Spaces
Rocco De Nicola, Rosario Pugliese, Antony I. T. Rowstron |
COORDINATION | 3 |
| 1999 | Mobile Co-ordination: Providing Fault Tolerance in Tuple Space Based Co-ordination Languages
Antony I. T. Rowstron |
COORDINATION | 1 |
| 1998 | The Cambridge University Robot Football Team Description
Antony I. T. Rowstron, Bem Bradshaw, Dave Cosby, Tim Edmonds, Steve Hodges 0001, Andy Hopper, Steve Lloyd, Stuart Wray |
RoboCup | 1 |
| 1998 | Solving the Linda Multiple rd Problem Using the Copy-Collect Primitive
Antony I. T. Rowstron, Alan Wood |
Sci. Comput. Program. | 1 |
| 1998 | WCL: A Co-ordination Language for Geographically Distributed Agents
Antony I. T. Rowstron |
World Wide Web | 1 |
| 1997 | Using Asynchronous Tuple-Space Access Primitives (BONITA Primitives) for Process Co-ordination
Antony I. T. Rowstron |
COORDINATION | 1 |
| 1996 | Solving the LINDA Multiple rd Problem
Antony I. T. Rowstron, Alan Wood |
COORDINATION | 1 |