Hakim Weatherspoon

dblp:w/HakimWeatherspoon · DBLP profile ↗
← Back
53ranked-venue papers
3as first author
12since 2021 · last 2024
0000-0002-6361-7687ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 2 first-author · 4 since 2021Computer networks · 19 · 4 since 2021Software engineering, systems software and programming languages · 4Databases, data management, data science and information retrieval · 4 · 1 first-authorSecurity and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Seam Work and Simulacra of Societal Impact in Networking Research: A Critical Technical Practice Approach
abstract
This paper explores how conceptions of societal impact are produced and performed during academic computer science research, by leveraging critical technical practice while building a digital agriculture networking platform. Our findings reveal how everyday practices of envisioning and building infrastructure require working across disciplinary and institutional seams, leading us as computer scientists to continuously reconceptualize the intended societal impact. By self-reflectively analyzing how we accrue resources for projects, produce research systems, write about them, and maintain alignments with stakeholders, we demonstrate that this seam work produces shifting simulacra of societal impact around which the system’s success is narrated. HCI researchers frequently suggest that technical systems’ impact could be improved by motivating computer scientists to consider impact in system-building. Our findings show that institutional and disciplinary structures significantly shape how computer scientists can enact societal impact in their work. This work suggests opportunities for structural interventions to shape the impact of computing systems.
Gloire Rubambiza, Phoebe Sengers, Hakim Weatherspoon, Jen Liu
CHI3
2024 Towards Swap-Free, Continuous Ballooning for Fast, Cloud-Based Virtual Machine Migrations
abstract
We have a production need to reduce the time for customers to live migrate their application virtual machine (VM) in the cloud. A single customer of ours migrates their nested, cloud-based, user virtual machines tens of thousands of times a month.
Kevin Alarcón Negy, Tycho Nightingale, Hakim Weatherspoon, Zhiming Shen
SoCC3
2024 Semi-Oblivious Reconfigurable Datacenter Networks
abstract
Reconfigurable datacenter networks use fast optical circuit switches to provide high bandwidths at low cost, therefore emerging as a compelling alternative to packet switching. These switches offer micro- and nano-second reconfiguration, and reacting to demand at this time scale is infeasible. Proposed designs have therefore largely been oblivious, supporting arbitrary traffic patterns. However, this imposes a fundamental latency-throughput tradeoff that significantly limits the benefits of these switches.
Nitika Saran, Daniel Amir, Tegan Wilson, Robert D. Kleinberg, Vishal Shrivastav, Hakim Weatherspoon
HotNets6
2024 Shale: A Practical, Scalable Oblivious Reconfigurable Network
abstract
Circuit-switched technologies have long been proposed for handling high-throughput traffic in datacenter networks, but recent developments in nanosecond-scale reconfiguration have created the enticing possibility of handling low-latency traffic as well. The novel Oblivious Reconfigurable Network (ORN) design paradigm promises to deliver on this possibility. Prior work in ORN designs achieved latencies that scale linearly with system size, making them unsuitable for large-scale deployments. Recent theoretical work showed that ORNs can achieve far better latency scaling, proposing theoretical ORN designs that are Pareto optimal in latency and throughput.
Daniel Amir, Nitika Saran, Tegan Wilson, Robert D. Kleinberg, Vishal Shrivastav, Hakim Weatherspoon
SIGCOMM6
2024 A Decentralized SDN Architecture for the WAN
abstract
Motivated by our experiences operating a global WAN, we argue that SDN's reliance on infrastructure external to the data plane has substantially complicated the challenge of maintaining high availability. We propose a new decentralized SDN (dSDN) architecture in which SDN control logic instead runs within routers, eliminating the control plane's reliance on external infrastructure and restoring fate-sharing between control and data planes. We present dSDN as a simpler approach to realizing the benefits of SDN in the WAN. Despite its much simpler design, we show that dSDN is practical from an implementation viewpoint, and outperforms centralized SDN in terms of routing convergence and SLO impact.
Alexander Krentsel, Nitika Saran, Bikash Koley, Subhasree Mandal, Ashok Narayanan, Sylvia Ratnasamy, Ali Al-Shabibi, Anees Shaikh, Rob Shakir, Ankit Singla, Hakim Weatherspoon
SIGCOMM11
2024 Breaking the VLB Barrier for Oblivious Reconfigurable Networks
abstract
In a landmark 1981 paper, Valiant and Brebner gave birth to the study of oblivious routing and, simultaneously, introduced its most powerful and ubiquitous method: Valiant load balancing (VLB). By routing messages through a randomly sampled intermediate node, VLB lengthens routing paths by a factor of two but gains the crucial property of obliviousness: it balances load in a completely decentralized manner, with no global knowledge of the communication pattern. Forty years later, with datacenters handling workloads whose communication pattern varies too rapidly to allow centralized coordination, oblivious routing is as relevant as ever, and VLB continues to take center stage as a widely used — and in some settings, provably optimal — way to balance load in the network obliviously to the traffic demands. However, the ability of the network to rapidly reconfigure its interconnection topology gives rise to new possibilities.
Tegan Wilson, Daniel Amir, Nitika Saran, Robert D. Kleinberg, Vishal Shrivastav, Hakim Weatherspoon
STOC6
2023 Poster: Scalability and Congestion Control in Oblivious Reconfigurable Networks
abstract
Traditional datacenter networks have been designed primarily using packet switches. However, due to the end of Moore's Law and Denard Scaling, packet switches face increasing difficulty in scaling to meet network demands without consuming unnecessarily large amounts of power, both within high-density racks[14] and throughout the datacenter[1]. As a result, many emerging network designs have intentionally avoided using packet switches [5, 7, 9, 10, 12, 15, 16]. Circuit switches present an exciting alternative to packet switches due to their reduced power consumption[1, 14], and potential to scale to arbitrary bandwidth (in the case of optical switches). While slow reconfiguration times have historically made circuit switches unable to support low-latency traffic, recent circuit switch design have emerged that are capable of nanosecond-scale reconfiguration times, including both electrical [11] and optical [3, 4, 6] switches. Unfortunately, conventional, dynamically-reconfiguring circuit-switched network designs have inherent latencies both for computing which circuits to deploy and for coordinating switches and nodes, limiting the benefits of this new capability.
Daniel Amir, Tegan Wilson, Vishal Shrivastav, Hakim Weatherspoon, Robert D. Kleinberg
SIGCOMM4
2023 Comosum: An Extensible, Reconfigurable, and Fault-Tolerant IoT Platform for Digital Agriculture
Gloire Rubambiza, Shiang-Wan Chin, Mueed Rehman, Sachille Atapattu, José F. Martínez, Hakim Weatherspoon
USENIX ATC6
2023 Introduction to the Special Section on USENIX OSDI 2022
abstract
No abstract available.
Marcos K. Aguilera, Hakim Weatherspoon
ACM Trans. Storage2
2022 Seamless Visions, Seamful Realities: Anticipating Rural Infrastructural Fragility in Early Design of Digital Agriculture
abstract
Rural infrastructure is known to be more prone to breakdown than urban infrastructure. This paper explores how the fragility of rural infrastructure is reproduced through the process of engineering design. Building on values in design, we examine how eventual use is anticipated by engineering researchers building on emerging infrastructure for digital agriculture (DA). Our approach combines critically reflective technical systems-building with interviews with other practitioners to understand and address moments early in the design process where the eventual effects of DA systems may be being built-in. Our findings contrast researchers’ visions of seamless farming technologies with the seamful realities of their work to produce them. We trace how, when anticipating future use, the seams that researchers themselves experience disappear, other seams are hidden from view by institutional support, and seams end users may face are too distant to be in sight. We develop suggestions for the design of these technologies grounded in a more artful management of seamfulness and seamlessness during the process of design and development.
Gloire Rubambiza, Phoebe Sengers, Hakim Weatherspoon
CHI3
2022 Optimal oblivious reconfigurable networks
abstract
Oblivious routing has a long history in both the theory and practice of networking. In this work we initiate the formal study of oblivious routing in the context of reconfigurable networks, a new architecture that has recently come to the fore in datacenter networking. These networks allow a rapidly changing bounded-degree pattern of interconnections between nodes, but the network topology and the selection of routing paths must both be oblivious to the traffic demand matrix. Our focus is on the trade-off between maximizing throughput and minimizing latency in these networks. For every constant throughput rate, we characterize (up to a constant factor) the minimum latency achievable by an oblivious reconfigurable network design that satisfies the given throughput guarantee. The trade-off between these two objectives turns out to be surprisingly subtle: the curve depicting it has an unexpected scalloped shape reflecting the fact that load-balancing becomes more difficult when the average length of routing paths is not an integer because equalizing all the path lengths is not possible. The proof of our lower bound uses LP duality to verify that Valiant load balancing is the most efficient oblivious routing scheme when used in combination with an optimally-designed reconfigurable network topology. The proof of our upper bound uses an algebraic construction in which the network nodes are identified with vectors over a finite field, the network topology is described by either the elementary basis or a sequence of Vandermonde matrices, and routing paths are constructed by selecting columns of these matrices to yield the appropriate mixture of path lengths within the shortest possible time interval.
Daniel Amir, Tegan Wilson, Vishal Shrivastav, Hakim Weatherspoon, Robert D. Kleinberg, Rachit Agarwal 0001
STOC4
2021 CacheInspector: Reverse Engineering Cache Resources in Public Clouds
abstract
Infrastructure-as-a-Service cloud providers sell virtual machines that are only specified in terms of number of CPU cores, amount of memory, and I/O throughput. Performance-critical aspects such as cache sizes and memory latency are missing or reported in ways that make them hard to compare across cloud providers. It is difficult for users to adapt their application’s behavior to the available resources. In this work, we aim to increase the visibility that cloud users have into shared resources on public clouds. Specifically, we present CacheInspector , a lightweight runtime that determines the performance and allocated capacity of shared caches on multi-tenant public clouds. We validate CacheInspector ’s accuracy in a controlled environment, and use it to study the characteristics and variability of cache resources in the cloud, across time, instances, availability regions, and cloud providers. We show that CacheInspector ’s output allows cloud users to tailor their application’s behavior, including their output quality, to avoid suboptimal performance when resources are scarce.
Weijia Song, Christina Delimitrou, Zhiming Shen, Robbert van Renesse, Hakim Weatherspoon, Lotfi Benmohamed, Frederic J. de Vaulx, Charif Mahmoudi
ACM Trans. Archit. Code Optim.5
2020 P4xos: Consensus as a Network Service
abstract
In this paper, we explore how a programmable forwarding plane offered by a new breed of network switches might naturally accelerate consensus protocols, specifically focusing on Paxos. The performance of consensus protocols has long been a concern. By implementing Paxos in the forwarding plane, we are able to significantly increase throughput and reduce latency. Our P4-based implementation running on an ASIC in isolation can process over 2.5 billion consensus messages per second, a four orders of magnitude improvement in throughput over a widely-used software implementation. This effectively removes consensus as a bottleneck for distributed applications in data centers. Beyond sheer performance, our approach offers several other important benefits: it readily lends itself to formal verification; it does not rely on any additional network hardware; and as a full Paxos implementation, it makes only very weak assumptions about the network.
Huynh Tu Dang, Pietro Bressana, Han Wang 0009, Ki Suh Lee, Noa Zilberman, Hakim Weatherspoon, Marco Canini, Fernando Pedone, Robert Soulé
IEEE/ACM Trans. Netw.6
2020 Introduction to the Special Issue on USENIX FAST 2019
abstract
No abstract available.
Arif Merchant, Hakim Weatherspoon
ACM Trans. Storage2
2019 X-Containers: Breaking Down Barriers to Improve Performance and Isolation of Cloud-Native Containers
abstract
"Cloud-native" container platforms, such as Kubernetes, have become an integral part of production cloud environments. One of the principles in designing cloud-native applications is called Single Concern Principle, which suggests that each container should handle a single responsibility well. In this paper, we propose X-Containers as a new security paradigm for isolating single-concerned cloud-native containers. Each container is run with a Library OS (LibOS) that supports multi-processing for concurrency and compatibility. A minimal exokernel ensures strong isolation with small kernel attack surface. We show an implementation of the X-Containers architecture that leverages Xen paravirtualization (PV) to turn Linux kernel into a LibOS. Doing so results in a highly efficient LibOS platform that does not require hardware-assisted virtualization, improves inter-container isolation, and supports binary compatibility and multi-processing. By eliminating some security barriers such as seccomp and Meltdown patch, X-Containers have up to 27X higher raw system call throughput compared to Docker containers, while also achieving competitive or superior performance on various benchmarks compared to recent container platforms such as Google's gVisor and Intel's Clear Containers.
Zhiming Shen, Gur-Eyal Sela, Eugene Bagdasarian, Christina Delimitrou, Robbert van Renesse, Hakim Weatherspoon
ASPLOS7
2019 Shoal: A Network Architecture for Disaggregated Racks
Vishal Shrivastav, Asaf Valadarsky, Hitesh Ballani, Paolo Costa, Ki Suh Lee, Han Wang 0009, Rachit Agarwal 0001, Hakim Weatherspoon
NSDI8
2019 Globally Synchronized Time via Datacenter Networks
abstract
Synchronized time is critical to distributed systems and network applications in a datacenter network. Unfortunately, many clock synchronization protocols in datacenter networks such as NTP and PTP are fundamentally limited by the characteristics of packet-switched networks. In particular, network jitter, packet buffering and scheduling in switches, and network stack overheads add non-deterministic variances to the round trip time, which must be accurately measured to synchronize clocks precisely. We present the Datacenter Time Protocol (DTP), a clock synchronization protocol that does not use packets at all, but is able to achieve nanosecond precision. In essence, the DTP uses the physical layer of network devices to implement a decentralized clock synchronization protocol. By doing so, the DTP eliminates most non-deterministic elements in clock synchronization protocols and has virtually zero protocol overhead since it does not add load at layer-2 or higher at all. It does require replacing network devices, which can be done incrementally and with very small amount of hardware resource consumption. We demonstrate that the precision provided by DTP in hardware is bounded by 4TD where D is the longest distance between any two nodes in a network in terms of number of hops and T is the period of the fastest clock. The precision can be further improved by combining DTP with frequency synchronization. By contrast, the precision of the state-of-the-art protocol (PTP) is not bounded: The precision is hundreds of nanoseconds in an idle network and can decrease to hundreds of microseconds in a heavily congested network.
Vishal Shrivastav, Ki Suh Lee, Han Wang 0009, Hakim Weatherspoon
IEEE/ACM Trans. Netw.4
2018 Untethered: Deployable Blockchains for IoT Environments
abstract
No abstract available.
Kolbeinn Karlsson, Danny Adams, Gloire Rubambiza, Zangyueyang Xian, Robbert van Renesse, Hakim Weatherspoon, Stephen B. Wicker
SoCC6
2018 Vegvisir: A Partition-Tolerant Blockchain for the Internet-of-Things
abstract
While the intersection of blockchains and the Internet of Things (IoT) have received considerable research interest lately, Nakamoto-style blockchains possess a number of qualities that make them poorly suited for many IoT scenarios. Specifically, they require high network connectivity and are power-intensive. This is a drawback in IoT environments where battery-constrained nodes form an unreliable ad hoc network such as in digital agriculture. In this paper we present Vegvisir, a partition-tolerant blockchain for use in power-constrained IoT environments with limited network connectivity. It is a permissioned, directed acyclic graph (DAG)-structured blockchain that can be used to create a shared, tamperproof data repository that keeps track of data provenance. We discuss the use cases, architecture, and challenges of such a blockchain.
Kolbeinn Karlsson, Weitao Jiang, Stephen B. Wicker, Danny Adams, Edwin Ma, Robbert van Renesse, Hakim Weatherspoon
ICDCS7
2017 Towards an emergency edge supercloud
abstract
The "cloud paradigm" can provide a wealth of sophisticated emergency communication services that are gamechangers in emergency response, but its current implementation is not suitable to the challenging environments in which these responses often take place. The networking infrastructure may be all but unavailable, and access to centralized datacenters may be impossible.
Kolbeinn Karlsson, Zhiming Shen, Weijia Song, Hakim Weatherspoon, Robbert van Renesse, Stephen B. Wicker
SoCC4
2017 Supercloud: A Library Cloud for Exploiting Cloud Diversity
abstract
Infrastructure-as-a-Service (IaaS) cloud providers hide available interfaces for virtual machine (VM) placement and migration, CPU capping, memory ballooning, page sharing, and I/O throttling, limiting the ways in which applications can optimally configure resources or respond to dynamically shifting workloads. Given these interfaces, applications could migrate VMs in response to diurnal workloads or changing prices, adjust resources in response to load changes, and so on. This article proposes a new abstraction that we call a Library Cloud and that allows users to customize the diverse available cloud resources to best serve their applications. We built a prototype of a Library Cloud that we call the Supercloud . The Supercloud encapsulates applications in a virtual cloud under users’ full control and can incorporate one or more availability zones within a cloud provider or across different providers. The Supercloud provides virtual machine, storage, and networking complete with a full set of management operations, allowing applications to optimize performance. In this article, we demonstrate various innovations enabled by the Library Cloud.
Zhiming Shen, Qin Jia, Gur-Eyal Sela, Weijia Song, Hakim Weatherspoon, Robbert van Renesse
ACM Trans. Comput. Syst.5
2017 Isotope: ACID Transactions for Block Storage
abstract
Existing storage stacks are top heavy and expect little from block storage. As a result, new high-level storage abstractions—and new designs for existing abstractions—are difficult to realize, requiring developers to implement from scratch complex functionality such as failure atomicity and fine-grained concurrency control. In this article, we argue that pushing transactional isolation into the block store (in addition to atomicity and durability) is both viable and broadly useful, resulting in simpler high-level storage systems that provide strong semantics without sacrificing performance. We present Isotope, a new block store that supports ACID transactions over block reads and writes. Internally, Isotope uses a new multiversion concurrency control protocol that exploits fine-grained, subblock parallelism in workloads and offers both strict serializability and snapshot isolation guarantees. We implemented several high-level storage systems over Isotope, including two key-value stores that implement the LevelDB API over a hash table and B-tree, respectively, and a POSIX file system. We show that Isotope’s block-level transactions enable systems that are simple (100s of lines of code), robust (i.e., providing ACID guarantees), and fast (e.g., 415MB/s for random file writes). We also show that these systems can be composed using Isotope, providing applications with transactions across different high-level constructs such as files, directories, and key-value pairs.
Ji-Yong Shin, Mahesh Balakrishnan 0001, Tudor Marian, Hakim Weatherspoon
ACM Trans. Storage4
2016 Follow the Sun through the Clouds: Application Migration for Geographically Shifting Workloads
abstract
Global cloud services have to respond to workloads that shift geographically as a function of time-of-day or in response to special events. While many such services have support for adding nodes in one region and removing nodes in another, we demonstrate that such mechanisms can lead to significant performance degradation. Yet other services do not support application-level migration at all. Live VM migration between availability zones or even across cloud providers would be ideal, but cloud providers do not support this flexible mechanism.
Zhiming Shen, Qin Jia, Gur-Eyal Sela, Ben Rainero, Weijia Song, Robbert van Renesse, Hakim Weatherspoon
SoCC7
2016 Towards Weakly Consistent Local Storage Systems
abstract
Heterogeneity is a fact of life for modern storage servers. For example, a server may spread terabytes of data across many different storage media, ranging from magnetic disks, DRAM, NAND-based solid state drives (SSDs), as well as hybrid drives that package various combinations of these technologies. It follows that access latencies to data can vary hugely depending on which media the data resides on. At the same time, modern storage systems naturally retain older versions of data due to the prevalence of log-structured designs and caches in software and hardware layers. In a sense, a contemporary storage system is very similar to a small-scale distributed system, opening the door to consistency/performance trade-offs. In this paper, we propose a class of local storage systems called StaleStores that support relaxed consistency, returning stale data for better performance. We describe several examples of StaleStores, and show via emulations that serving stale data can improve access latency by between 35% and 20X. We describe a particular StaleStore called Yogurt, a weakly consistent local block storage system. Depending on the application's consistency requirements (e.g. bounded staleness, mono-tonic reads, read-my-writes, etc.), Yogurt queries the access costs for different versions of data within tolerable staleness bounds and returns the fastest version. We show that a distributed key-value store running on top of Yogurt obtains a 6X speed-up for access latency by trading off consistency and performance within individual storage servers.
Ji-Yong Shin, Mahesh Balakrishnan 0001, Tudor Marian, Jakub Szefer, Hakim Weatherspoon
SoCC5
2016 Isotope: Transactional Isolation for Block Storage
Ji-Yong Shin, Mahesh Balakrishnan 0001, Tudor Marian, Hakim Weatherspoon
FAST4
2016 Globally Synchronized Time via Datacenter Networks
abstract
In this paper, we present Datacenter Time Protocol (DTP), a clock synchronization protocol that does not use packets at all, but is able to achieve nanosecond precision. In essence, DTP uses the physical layer of network devices to implement a decentralized clock synchronization protocol. By doing so, DTP eliminates most non-deterministic elements in clock synchronization protocols. Further, DTP uses control messages in the physical layer for communicating hundreds of thousands of protocol messages without interfering with higher layer packets. Thus, DTP has virtually zero overhead since it does not add load at layers 2 or higher layers. It does require replacing network devices, which can be done incrementally. We demonstrate that the precision provided by DTP is bounded by 25.6 nanoseconds for directly connected nodes, and in general, is bounded by 4TD where D is the longest distance between any two servers in a network in terms of number of hops and T is the period of the fastest clock (≈ 6.4ns). Moreover, in software, a DTP daemon can access the DTP clock with usually better than 4T (≈ 25.6ns) precision. As a result, the end-to-end precision can be better than 4T D + 8T nanoseconds. By contrast, the precision of the state of the art protocol is not bounded: The precision is hundreds of nanoseconds when a network is idle and can decrease to hundreds of microseconds when a network is heavily congested.
Ki Suh Lee, Han Wang 0009, Vishal Shrivastav, Hakim Weatherspoon
SIGCOMM4
2014 Timing is Everything: Accurate, Minimum Overhead, Available Bandwidth Estimation in High-speed Wired Networks
abstract
Active end-to-end available bandwidth estimation is intrusive, expensive, inaccurate, and does not work well with bursty cross traffic or on high capacity links. Yet, it is important for designing high performant networked systems, improving network protocols, building distributed systems, and improving application performance. In this paper, we present minProbe which addresses unsolved issues that have plagued available bandwidth estimation. As a middlebox, minProbe measures and estimates available bandwidth with high-fidelity, minimal-cost, and in userspace; thus, enabling cheaper (virtually no overhead) and more accurate available bandwidth estimation. MinProbe performs accurately on high capacity networks up to 10 Gbps and with bursty cross traffic. We evaluated the performance and accuracy of minProbe over a wide-area network, the National Lambda Rail (NLR), and within our own network testbed. Results indicate that minProbe can estimate available bandwidth with error typically no more than 0.4 Gbps in a 10 Gbps network.
Han Wang 0009, Ki Suh Lee, Erluo Li, Chiunlin Lim, Ao Tang, Hakim Weatherspoon
Internet Measurement Conference6
2014 PHY Covert Channels: Can you see the Idles?
Ki Suh Lee, Han Wang 0009, Hakim Weatherspoon
NSDI3
2013 SuperCloud: economical cloud service on multiple vendors
abstract
Today, Infrastructure-as-a-Service (IaaS) cloud providers such as Amazon's Elastic Compute Engine (EC2), Google's Compute Engine, and Microsoft's Azure offer elastic and isolated compute resources via virtualization and users often choose one of these providers based on price, locality, performance, and features. Typically, a user will choose the same provider for computation and storage to minimize latency and networking costs. Unfortunately, it can be difficult to switch providers once one is selected due to vendor lock-in [2].
Qin Jia, Robbert van Renesse, Hakim Weatherspoon
SoCC3
2013 Gecko: contention-oblivious disk arrays for cloud storage
Ji-Yong Shin, Mahesh Balakrishnan 0001, Tudor Marian, Hakim Weatherspoon
FAST4
2013 SoNIC: Precise Realtime Software Access and Control of Wired Networks
Ki Suh Lee, Han Wang 0009, Hakim Weatherspoon
NSDI3
2013 Integrated Approach to Data Center Power Management
abstract
Energy accounts for a significant fraction of the operational costs of a data center, and data center operators are increasingly interested in moving toward low-power designs. Two distinct approaches have emerged toward achieving this end: the power-proportional approach focuses on reducing disk and server power consumption, while the green data center approach focuses on reducing power consumed by support-infrastructure like cooling equipment, power distribution units, and power backup equipment. We propose an integrated approach, which combines the benefits of both. Our solution enforces power-proportionality at the granularity of a rack or even an entire containerized data center; thus, we power down not only idle IT equipment, but also their associated support-infrastructure. We show that it is practical today to design data centers to power down idle racks or containers-and in fact, current online service trends strongly enable this model. Finally, we show that our approach combines the energy savings of power-proportional and green data center approaches, while performance remains unaffected.
Lakshmi Ganesh, Hakim Weatherspoon, Tudor Marian, Kenneth P. Birman
IEEE Trans. Computers2
2013 On the Feasibility of Completely Wirelesss Datacenters
abstract
Conventional datacenters, based on wired networks, entail high wiring costs, suffer from performance bottlenecks, and have low resilience to network failures. In this paper, we investigate a radically new methodology for building wire-free datacenters based on emerging 60-GHz radio frequency (RF) technology. We propose a novel rack design and a resulting network topology inspired by Cayley graphs that provide a dense interconnect. Our exploration of the resulting design space shows that wireless datacenters built with this methodology can potentially attain higher aggregate bandwidth, lower latency, and substantially higher fault tolerance than a conventional wired datacenter while improving ease of construction and maintenance.
Ji-Yong Shin, Emin Gün Sirer, Hakim Weatherspoon, Darko Kirovski
IEEE/ACM Trans. Netw.3
2012 NetBump: user-extensible active queue management with bumps on the wire
abstract
Engineering large-scale data center applications built from thousands of commodity nodes requires both an underlying network that supports a wide variety of traffic demands, and low latency at microsecond timescales. Many ideas for adding innovative functionality to networks, especially active queue management strategies, require either modifying packets or performing alternative queuing to packets in-flight on the data plane. However, configuring packet queuing, marking, and dropping is challenging, since buffering in commercial switches and routers is not programmable.
Mohammad Al-Fares, Rishi Kapoor, George Porter, Sambit Das, Hakim Weatherspoon, Balaji Prabhakar, Amin Vahdat
ANCS5
2012 NetSlices: scalable multi-core packet processing in user-space
abstract
Modern commodity operating systems do not provide developers with user-space abstractions for building high-speed packet processing applications. The conventional raw socket is inefficient and unable to take advantage of the emerging hardware, like multi-core processors and multi-queue network adapters. In this paper we present the NetSlice operating system abstraction. Unlike the conventional raw socket, NetSlice tightly couples the hardware and software packet processing resources, and provides the application with control over these resources. To reduce shared resource contention, NetSlice performs domain specific, coarse-grained, spatial partitioning of CPU cores, memory, and NICs. Moreover, it provides a streamlined communication channel between NICs and user-space. Although backward compatible with the conventional socket API, the NetSlice API also provides batched (multi-) send / receive operations to amortize the cost of protection domain crossings. We show that complex user-space packet processors---like a protocol accelerator and an IPsec gateway---built from commodity components can scale linearly with the number of cores and operate at 10Gbps network line speeds.
Tudor Marian, Ki Suh Lee, Hakim Weatherspoon
ANCS3
2012 On the feasibility of completely wireless datacenters
abstract
Conventional datacenters, based on wired networks, entail high wiring costs, suffer from performance bottlenecks, and have low resilience to network failures. In this paper, we investigate a radically new methodology for building wire-free datacenters based on emerging 60GHz RF technology. We propose a novel rack design and a resulting network topology inspired by Cayley graphs that provide a dense interconnect. Our exploration of the resulting design space shows that wireless datacenters built with this methodology can potentially attain higher aggregate bandwidth, lower latency, and substantially higher fault tolerance than a conventional wired datacenter while improving ease of construction and maintenance.
Ji-Yong Shin, Emin Gün Sirer, Hakim Weatherspoon, Darko Kirovski
ANCS3
2012 The Xen-Blanket: virtualize once, run everywhere
abstract
Current Infrastructure as a Service (IaaS) clouds operate in isolation from each other. Slight variations in the virtual machine (VM) abstractions or underlying hypervisor services prevent unified access and control across clouds. While standardization efforts aim to address these issues, they will take years to be agreed upon and adopted, if ever. Instead of standardization, which is by definition provider-centric, we advocate a user-centric approach that gives users an unprecedented level of control over the virtualization layer. We introduce the Xen-Blanket, a thin, immediately deployable virtualization layer that can homogenize today's diverse cloud infrastructures. We have deployed the Xen-Blanket across Amazon's EC2, an enterprise cloud, and a private setup at Cornell University. We show that a user-centric approach to homogenize clouds can achieve similar performance to a paravirtualized environment while enabling previously impossible tasks like cross-provider live migration. The Xen-Blanket also allows users to exploit resource management opportunities like oversubscription, and ultimately can reduce costs for users.
Dan Williams 0001, Hani Jamjoom, Hakim Weatherspoon
EuroSys3
2012 Gecko: A Contention-Oblivious Design for Cloud Storage
Ji-Yong Shin, Mahesh Balakrishnan 0001, Lakshmi Ganesh, Tudor Marian, Hakim Weatherspoon
HotStorage5
2012 Fmeter: Extracting Indexable Low-Level System Signatures by Counting Kernel Function Calls
Tudor Marian, Hakim Weatherspoon, Ki Suh Lee, Abhishek Sagar
Middleware2
2011 Beyond Power Proportionality: Designing Power-Lean Cloud Storage
abstract
We present a power-lean storage system, where racks of servers, or even entire data center shipping containers, can be powered down to save energy. We show that racks and containers are more than the sum of their servers, and demonstrate the feasibility of designing a storage system that powers them up and down on demand further, we show that such a system would save an order of magnitude more energy than current disk-based power-proportional storage systems. Our simulation results using file system traces from the Internet Archive show over 44% energy savings, a 5x improvement over disk-based power management systems, without performance impact. We explore the tradeoffs in choosing the right unit to power off/on, and present an automated framework to compute the optimal power management unit for different scenarios.
Lakshmi Ganesh, Hakim Weatherspoon, Kenneth P. Birman
NCA2
2011 Overdriver: handling memory overload in an oversubscribed cloud
abstract
With the intense competition between cloud providers, oversubscription is increasingly important to maintain profitability. Oversubscribing physical resources is not without consequences: it increases the likelihood of overload. Memory overload is particularly damaging. Contrary to traditional views, we analyze current data center logs and realistic Web workloads to show that overload is largely transient: up to 88.1% of overloads last for less than 2 minutes. Regarding overload as a continuum that includes both transient and sustained overloads of various durations points us to consider mitigation approaches also as a continuum, complete with tradeoffs with respect to application performance and data center overhead. In particular, heavyweight techniques, like VM migration, are better suited to sustained overloads, whereas lightweight approaches, like network memory, are better suited to transient overloads. We present Overdriver, a system that adaptively takes advantage of these tradeoffs, mitigating all overloads within 8% of well-provisioned performance. Furthermore, under reasonable oversubscription ratios, where transient overload constitutes the vast majority of overloads, Overdriver requires 15% of the excess space and generates a factor of four less network traffic than a migration-only approach.
Dan Williams 0001, Hani Jamjoom, Yew-Huey Liu, Hakim Weatherspoon
VEE4
2011 Maelstrom: transparent error correction for communication between data centers
abstract
The global network of data centers is emerging as an important distributed systems paradigm-commodity clusters running high-performance applications, connected by high-speed “lambda” networks across hundreds of milliseconds of network latency. Packet loss on long-haul networks can cripple applications and protocols: A loss rate as low as 0.1% is sufficient to reduce TCP/IP throughput by an order of magnitude on a 1-Gb/s link with 50-ms one-way latency. Maelstrom is an edge appliance that masks packet loss transparently and quickly from intercluster protocols, aggregating traffic for high-speed encoding and using a new forward error correction scheme to handle bursty loss.
Mahesh Balakrishnan 0001, Tudor Marian, Kenneth P. Birman, Hakim Weatherspoon, Lakshmi Ganesh
IEEE/ACM Trans. Netw.4
2010 RACS: a case for cloud storage diversity
abstract
The increasing popularity of cloud storage is leading organizations to consider moving data out of their own data centers and into the cloud. However, success for cloud storage providers can present a significant risk to customers; namely, it becomes very expensive to switch storage providers. In this paper, we make a case for applying RAID-like techniques used by disks and file systems, but at the cloud storage level. We argue that striping user data across multiple providers can allow customers to avoid vendor lock-in, reduce the cost of switching providers, and better tolerate provider outages or failures. We introduce RACS, a proxy that transparently spreads the storage load over many providers. We evaluate a prototype of our system and estimate the costs incurred and benefits reaped. Finally, we use trace-driven simulations to demonstrate how RACS can reduce the cost of switching storage vendors for a large organization such as the Internet Archive by seven-fold or more by varying erasure-coding parameters.
Hussam Abu-Libdeh, Lonnie Princehouse, Hakim Weatherspoon
SoCC3
2010 Empirical characterization of uncongested optical lambda networks and 10GbE commodity endpoints
abstract
High-bandwidth, semi-private optical lambda networks carry growing volumes of data on behalf of large data centers, both in cloud computing environments and for scientific, financial, defense, and other enterprises. This paper undertakes a careful examination of the end-to-end characteristics of an uncongested lambda network running at high speeds over long distances, identifying scenarios associated with loss, latency variations, and degraded throughput at attached end-hosts. We use identical fast commodity source and destination platforms, hence expect the destination to receive more or less what we send. We observe otherwise: degraded performance is common and easily provoked. In particular, the receiver loses packets even when the sender employs relatively low data rates. Data rates of future optical network components are projected to outpace clock speeds of commodity end-host processors, hence more and more end-to-end applications will confront the same issue we encounter. Our work thus poses a new challenge for those hoping to achieve dependable performance in higher-end networked settings.
Tudor Marian, Daniel A. Freedman, Kenneth P. Birman, Hakim Weatherspoon
DSN4
2010 Exact temporal characterization of 10 Gbps optical wide-area network
abstract
We design and implement a novel class of highly precise network instrumentation and apply this tool to perform the first exact packet-timing measurements of a wide-area network ever undertaken, capturing 10 Gigabit Ethernet packets in flight on optical fiber. Through principled design, we improve timing precision by two to six orders of magnitude over existing techniques. Our observations contest several common assumptions about behavior of wide-area networks and the relationship between their input and output traffic flows. Further, we identify and characterize emergent packet chains as a mechanism to explain previously observed anomalous packet loss on receiver endpoints of such networks.
Daniel A. Freedman, Tudor Marian, Jennifer H. Lee, Kenneth P. Birman, Hakim Weatherspoon, Chris Xu
Internet Measurement Conference5
2009 Smoke and Mirrors: Reflecting Files at a Geographically Remote Location Without Loss of Performance
Hakim Weatherspoon, Lakshmi Ganesh, Tudor Marian, Mahesh Balakrishnan 0001, Kenneth P. Birman
FAST1
2008 Maelstrom: Transparent Error Correction for Lambda Networks
Mahesh Balakrishnan 0001, Tudor Marian, Kenneth P. Birman, Hakim Weatherspoon, Einar Vollset
NSDI4
2007 Antiquity: exploiting a secure log for wide-area distributed storage
abstract
Antiquity is a wide-area distributed storage system designed to provide a simple storage service for applications like file systems and back-up. The design assumes that all servers eventually fail and attempts to maintain data despite those failures. Antiquity uses a secure log to maintain data integrity, replicates each log on multiple servers for durability, and uses dynamic Byzantine fault-tolerant quorum protocols to ensure consistency among replicas. We present Antiquity's design and an experimental evaluation with global and local testbeds. Antiquity has been running for over two months on 400+ PlanetLab servers storing nearly 20,000 logs totaling more than 84 GB of data. Despite constant server churn, all logs remain durable.
Hakim Weatherspoon, Patrick R. Eaton, Byung-Gon Chun, John Kubiatowicz
EuroSys1
2007 Optimizing Power Consumption in Large Scale Storage Systems
Lakshmi Ganesh, Hakim Weatherspoon, Mahesh Balakrishnan 0001, Kenneth P. Birman
HotOS2
2006 Efficient Replica Maintenance for Distributed Storage Systems
Byung-Gon Chun, Frank Dabek, Andreas Haeberlen, Emil Sit, Hakim Weatherspoon, M. Frans Kaashoek, John Kubiatowicz, Robert Morris 0005
NSDI5
2003 Pond: The OceanStore Prototype
Sean C. Rhea, Patrick R. Eaton, Dennis Geels, Hakim Weatherspoon, Ben Y. Zhao, John Kubiatowicz
FAST4
2002 Introspective Failure Analysis: Avoiding Correlated Failures in Peer-to-Peer Systems
abstract
Failure independence is an important assumption for many fault tolerance techniques. Unfortunately, real systems exhibit correlated failures. In this paper, we present a framework for online discovery of groups of server nodes that are maximally independent in their failure characteristics. We discuss the framework in detail and provide a preliminary evaluation.
Hakim Weatherspoon, Tal Moscovitz, John Kubiatowicz
SRDS1
2000 OceanStore: An Architecture for Global-Scale Persistent Storage
abstract
OceanStore is a utility infrastructure designed to span the globe and provide continuous access to persistent information. Since this infrastructure is comprised of untrusted servers, data is protected through redundancy and cryptographic techniques. To improve performance, data is allowed to be cached anywhere, anytime. Additionally, monitoring of usage patterns allows adaptation to regional outages and denial of service attacks; monitoring also enhances performance through pro-active movement of data. A prototype implementation is currently under development.
John Kubiatowicz, David Bindel, Yan Chen 0004, Steven E. Czerwinski, Patrick R. Eaton, Dennis Geels, Ramakrishna Gummadi, Sean C. Rhea, Hakim Weatherspoon, Westley Weimer, Chris Wells, Ben Y. Zhao
ASPLOS9