VLDB 2026 Research / reviewers in the wild / expert
Amin Vahdat
dblp:v/AminVahdat
· DBLP profile ↗
165ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-4866-1698ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 84 · 16 since 2021Systems, architecture and hardware · 45 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 31 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorSecurity and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSIG: Congestion Signaling for Datacenter TransportsabstractOptimizing burst-heavy datacenter workloads necessitates finegrained network control and visibility. We introduce CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header. The architecture captures μsgranularity switch metrics, such as available bandwidth, and signals them to end-hosts using in-band, line-rate operations. We propose Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production. Beyond transport-level performance, CSIG enables flow-aware observability by embedding μs-scale metrics into every packet, allowing individual application transfers to pinpoint their bottleneck location, such as the topology tier limiting their performance. CSIG thus transforms network telemetry from post-hoc correlation into a real time, context-aware capability. We demonstrate CSIG's broad deployability by validating it across five generations of commodity switch hardware (up to 102.4 Tbps), four NIC generations, and five transport stacks. Our design proves that a streamlined Layer 2 approach, focusing exclusively on the principal path bottleneck, provides transport-agnostic gains without requiring forklift hardware upgrades. Abhiram Ravi, Nandita Dukkipati, Weiwu Pang, Neal Cardwell, Brad Karp, Mohammad Jafar Akhbarizadeh, Weida Huang, Konstantinos Prasopoulos, Kok-Kiong Yap, Amin Vahdat |
SIGCOMM | 11 |
| 2025 | Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski, Liqun Cheng, Amin Vahdat |
NSDI | 14 |
| 2025 | Firefly: Scalable, Ultra-Accurate Clock Synchronization for DatacentersabstractCloud-based financial exchanges require sub-10ns device-to-device clock synchronization accuracy while adhering to Coordinated Universal Time (UTC). Existing clock sync techniques struggle to meet this demand at scale and are vulnerable to clock drift, jitter, and path asymmetries. Firefly, a software-driven datacenter clock sync system, scalably, cost-effectively, and reliably achieves very high clock sync accuracy. It employs a distributed consensus algorithm on a random overlay graph to rapidly converge to a common time while applying gradual adjustments to device hardware clocks. To realize consistent sync-to-UTC (external sync) across devices while maintaining a stable device-to-device internal sync, Firefly uses a novel technique, layered synchronization, that decouples internal and external syncs. In a 248-machine Clos network, Firefly achieves sub-10ns device-to-device and ≤1μs device-to-UTC sync, and is resilient to time server failure and unstable clocks. Pooria Namyar, Nandita Dukkipati, KK Yap, Junzhi Gong, Peixuan Gao, Devdeep Ray, Gautam Kumar 0001, Ramesh Govindan, Amin Vahdat |
SIGCOMM | 13 |
| 2025 | Falcon: A Reliable, Low Latency Hardware TransportabstractHardware transports such as RoCE deliver high performance with minimal host CPU, but are best suited to special-purpose deployments that limit their use, e.g., backend networks or Ethernet with Priority Flow Control (PFC). We introduce Falcon, the first hardware transport that supports multiple Upper Layer Protocols (ULPs) and heterogeneous application workloads in general-purpose Ethernet datacenter environments (with losses and without special switch support). Key design elements include: delay-based congestion control with multipath load balancing; a layered design with a simple request-response transaction interface for multi-ULP support; hardware-based retransmissions and error-handling for scalability; and a programmable engine for flexibility. The first Falcon hardware implementation delivers a peak performance of 200 Gbps, 120 Mops/sec, with near-optimal operation completion times that are up to 8× lower than CX-7 RoCE under network congestion, and up to 65% higher goodput under lossy conditions. Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan M. G. Wassel, Naveen Kr. Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar 0003, Behnam Montazeri, Neelesh Bansod, Sarin Thomas, Inho Cho, Hyojeong Lee Seibert, Baijun Wu, Rui Yang 0034, Qianwen Yin, Srinivas Vaduvatha, Weihuang Wang, Masoud Moshref, David Wetherall, Amin Vahdat |
SIGCOMM | 25 |
| 2024 | Societal infrastructure in the age of Artificial General IntelligenceabstractToday, we are at an inflection point in computing where emerging Generative AI services are placing unprecedented demand for compute while the existing architectural patterns for improving efficiency have stalled. In this talk, we will discuss the likely needs of the next generation of computing infrastructure and use recent examples at Google from networks to accelerators to servers to illustrate the challenges and opportunities ahead. Taken together, we chart a course where computing must be increasingly specialized and co-optimized with algorithms and software, all while fundamentally focusing on security and sustainability. Amin Vahdat |
ASPLOS (3) | 1 |
| 2024 | Harmony: A Congestion-free Datacenter Architecture
Saksham Agarwal, Qizhe Cai, Rachit Agarwal 0001, David B. Shmoys, Amin Vahdat |
NSDI | 5 |
| 2024 | CAPA: An Architecture For Operating Cluster Networks With High Availability
Bingzhe Liu, Colin Scott, Mukarram Tariq, Andrew D. Ferguson, Phillipa Gill, Richard Alimi, Omid Alipourfard, Deepak Arulkannan, Virginia Beauregard, Patrick Conner, Brighten Godfrey, Xander Lin, Joon Ong, Mayur Patel, Amr Sabaa, Alex Smirnov, Manish Verma, Prerepa V. Viswanadham, Amin Vahdat |
NSDI | 20 |
| 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScaleabstractFast DRAM increasingly dominates infrastructure spend in large scale computing environments and this trend will likely worsen without an architectural shift. The cost of deployed memory can be reduced by replacing part of the conventional DRAM with lower cost albeit slower memory media, thus creating a tiered memory system where both tiers are directly addressable and cached. But, this poses numerous challenges in a highly multi-tenant warehouse-scale computing setting. The diversity and scale of its applications motivates an application-transparent solution in the general case, adaptable to specific workload demands. Padmapriya Duraisamy, Scott Hare, Ravi Rajwar, David E. Culler, Zhiyi Xu, Jianing Fan, Chris Kennelly, Bill McCloskey, Danijela Mijailovic, Brian Morris, Chiranjit Mukherjee, Jingliang Ren, Greg Thelen, Carlos Villavieja, Parthasarathy Ranganathan, Amin Vahdat |
ASPLOS (3) | 18 |
| 2023 | Exciting Directions for ML Models and the Implications for Computing HardwareabstractIn recent years, ML has completely changed our expectations of what is possible with computers Jeffrey Dean, Amin Vahdat |
HCS | 2 |
| 2023 | Towards Modern Development of Cloud ApplicationsabstractWhen writing a distributed application, conventional wisdom says to split your application into separate services that can be rolled out independently. This approach is well-intentioned, but a microservices-based architecture like this often backfires, introducing challenges that counteract the benefits the architecture tries to achieve. Fundamentally, this is because microservices conflate logical boundaries (how code is written) with physical boundaries (how code is deployed). In this paper, we propose a different programming methodology that decouples the two in order to solve these challenges. With our approach, developers write their applications as logical monoliths, offload the decisions of how to distribute and run applications to an automated runtime, and deploy applications atomically. Our prototype implementation reduces application latency by up to 15× and reduces cost by up to 9× compared to the status quo. Sanjay Ghemawat, Robert Grandl, Srdjan Petrovic, Michael Whittaker, Parveen Patel, Ivan Posva, Amin Vahdat |
HotOS | 7 |
| 2023 | Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning SystemsabstractWe describe our experience developing what we believe to be the world's first large-scale production deployments of lightwave fabrics used for both datacenter networking and machine-learning (ML) applications. Using optical circuit switches (OCSes) and optical transceivers developed in-house, we employ hardware and software codesign to integrate the fabrics into our network and computing infrastructure. Key to our design is a high degree of multiplexing enabled by new kinds of wavelength-division-multiplexing (WDM) and optical circulators that support high-bandwidth bidirectional traffic on a single strand of optical fiber. The development of the requisite OCS and optical transceiver technologies leads to a synchronous lightwave fabric that is reconfigurable, low latency, rate agnostic, and highly available. These fabrics have provided substantial benefits for long-lived traffic patterns in our datacenter networks and predictable traffic patterns in tightly-coupled machine learning clusters. We report results for a large-scale ML superpod with 4096 tensor processing unit (TPU) V4 chips that has more than one ExaFLOP of computing power. For this use case, the deployment of a lightwave fabric provides up to 3× better system availability and model-dependent performance improvements of up to 3.3× compared to a static fabric, despite constituting less than 6% of the total system cost. Ryohei Urata, Kevin Yasumura, Roy Bannon, Jill Berger, Pedram Dashti, Norman P. Jouppi, Cedric F. Lam, Sheng Li 0007, Erji Mao, Daniel Nelson, George Papen, Muhammad Mukarram Bin Tariq, Amin Vahdat |
SIGCOMM | 15 |
| 2023 | Fathom: Understanding Datacenter Application Network PerformanceabstractWe describe our experience with Fathom, a system for identifying the network performance bottlenecks of any service running in the Google fleet. Fathom passively samples RPCs, the principal unit of work for services. It segments the overall latency into host and network components with kernel and RPC stack instrumentation. It records these detailed latency metrics, along with detailed transport connection state, for every sampled RPC. This lets us determine if the completion is constrained by the client, network or server. To scale while enabling analysis, we also aggregate samples into distributions that retain multi-dimensional breakdowns. This provides us with a macroscopic view of individual services. Fathom runs globally in our datacenters for all production traffic, where it monitors billions of TCP connections 24x7. For five years Fathom has been our primary tool for troubleshooting service network issues and assessing network infrastructure changes. We present case studies to show how it has helped us improve our production services. Mubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh, Yousuk Seung, Neal Cardwell, Willem de Bruijn, Van Jacobson, Jasleen Kaur 0001, David Wetherall, Amin Vahdat |
SIGCOMM | 11 |
| 2023 | Improving Network Availability with Protective ReRouteabstractWe present PRR (Protective ReRoute), a transport technique for shortening user-visible outages that complements routing repair. It can be added to any transport to provide benefits in multipath networks. PRR responds to flow connectivity failure signals, e.g., retransmission timeouts, by changing the FlowLabel on packets of the flow, which causes switches and hosts to choose a different network path that may avoid the outage. To enable it, we shifted our IPv6 network architecture to use the FlowLabel, so that hosts can change the paths of their flows without application involvement. PRR is deployed fleetwide at Google for TCP and Pony Express, where it has been protecting all production traffic for several years. It is also available to our Cloud customers. We find it highly effective for real outages. In a measurement study on our network backbones, adding PRR reduced the cumulative region-pair outage time for RPC traffic by 63--84%. This is the equivalent of adding 0.4--0.8 "nines" of availability. David Wetherall, Abdul Kabbani, Van Jacobson, Jim Winget, Yuchung Cheng, Charles B. Morrey III, Uma Parthavi Moravapalle, Phillipa Gill, Steven Knight, Amin Vahdat |
SIGCOMM | 10 |
| 2023 | Change Management in Physical Network Lifecycle Automation
Mohammad Al-Fares, Virginia Beauregard, Kevin Grant, Angus Griffith, Jahangir Hasan, Quan Leng, Alexander Lin, Zhuotao Liu, Ahmed Mansy, Bill Martinusen, Nikil Mehta, Jeffrey C. Mogul, Andrew Narver, Anshul Nigham, Melanie Obenberger, Kurt Steinkraus, Edward Thiele, Amin Vahdat |
USENIX ATC | 22 |
| 2022 | Understanding host interconnect congestionabstractWe present evidence and characterization of host congestion in production clusters: adoption of high-bandwidth access links leading to emergence of bottlenecks within the host interconnect (NIC-to-CPU data path). We demonstrate that contention on existing IO memory management units and/or the memory subsystem can significantly reduce the available NIC-to-CPU bandwidth, resulting in hundreds of microseconds of queueing delays and eventual packet drops at hosts (even when running a state-of-the-art congestion control protocol that accounts for CPU-induced host congestion). We also discuss implications of host interconnect congestion to design of future host architecture, network stacks and network protocols. Saksham Agarwal, Rachit Agarwal 0001, Behnam Montazeri, Masoud Moshref, Khaled Elmeleegy, Luigi Rizzo, Marc de Kruijf, Gautam Kumar 0001, Sylvia Ratnasamy, David E. Culler, Amin Vahdat |
HotNets | 11 |
| 2022 | Aquila: A unified, low-latency fabric for datacenter networks
Dan Gibson, Hema Hariharan, Eric Lance, Moray McLaren, Behnam Montazeri, Hassan M. G. Wassel, Zhehua Wu, Sunghwan Yoo, Raghuraman Balasubramanian, Prashant Chandra, Michael Cutforth, Peter Cuy, David Decotigny, Rakesh Gautam, Alex Iriza, Milo M. K. Martin, Rick Roy, Zuowei Shen, Monica Wong-Chan, Joe Zbiciak, Amin Vahdat |
NSDI | 25 |
| 2022 | Carbink: Fault-Tolerant Far Memory
Yang Zhou 0008, Hassan M. G. Wassel, Sihang Liu 0001, James W. Mickens, Minlan Yu, Chris Kennelly, David E. Culler, Henry M. Levy, Amin Vahdat |
OSDI | 11 |
| 2022 | Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingabstractWe present a decade of evolution and production experience with Jupiter datacenter network fabrics. In this period Jupiter has delivered 5x higher speed and capacity, 30% reduction in capex, 41% reduction in power, incremental deployment and technology refresh all while serving live production traffic. A key enabler for these improvements is evolving Jupiter from a Clos to a direct-connect topology among the machine aggregation blocks. Critical architectural changes for this include: A datacenter interconnection layer employing Micro-Electro-Mechanical Systems (MEMS) based Optical Circuit Switches (OCSes) to enable dynamic topology reconfiguration, centralized Software-Defined Networking (SDN) control for traffic engineering, and automated network operations for incremental capacity delivery and topology engineering. We show that the combination of traffic and topology engineering on direct-connect fabrics achieves similar throughput as Clos fabrics for our production traffic patterns. We also optimize for path lengths: 60% of the traffic takes direct path from source to destination aggregation blocks, while the remaining transits one additional block, achieving an average block-level path length of 1.4 in our fleet today. OCS also achieves 3x faster fabric reconfiguration compared to pre-evolution Clos fabrics that used a patch panel based interconnect. Leonid B. Poutievski, Omid Mashayekhi, Joon Ong, Muhammad Mukarram Bin Tariq, Rui Wang 0025, Virginia Beauregard, Patrick Conner, Steve D. Gribble, Rishi Kapoor, Stephen Kratzer, Nanfang Li, Karthik Nagaraj, Jason Ornstein, Samir Sawhney, Ryohei Urata, Lorenzo Vicisano, Kevin Yasumura, Shidong Zhang, Junlan Zhou, Amin Vahdat |
SIGCOMM | 23 |
| 2022 | Aequitas: admission control for performance-critical RPCs in datacentersabstractWith the increasing popularity of disaggregated storage and microservice architectures, high fan-out and fan-in Remote Procedure Calls (RPCs) now generate most of the traffic in modern datacenters. While the network plays a crucial role in RPC performance, traditional traffic classification categories cannot sufficiently capture their importance due to wide variations in RPC characteristics. As a result, meeting service-level objectives (SLOs), especially for performance-critical (PC) RPCs, remains challenging. Yiwen Zhang 0008, Gautam Kumar 0001, Nandita Dukkipati, Xian Wu 0001, Priyaranjan Jha, Mosharaf Chowdhury, Amin Vahdat |
SIGCOMM | 7 |
| 2022 | Hashing Design in Modern Networks: Challenges and Mitigation Techniques
Yunhong Xu, Keqiang He, Rui Wang 0025, Minlan Yu, Nick G. Duffield, Hassan M. G. Wassel, Shidong Zhang, Leonid B. Poutievski, Junlan Zhou, Amin Vahdat |
USENIX ATC | 10 |
| 2021 | Cores that don't countabstractWe are accustomed to thinking of computers as fail-stop, especially the cores that execute instructions, and most system software implicitly relies on that assumption. During most of the VLSI era, processors that passed manufacturing tests and were operated within specifications have insulated us from this fiction. As fabrication pushes towards smaller feature sizes and more elaborate computational structures, and as increasingly specialized instruction-silicon pairings are introduced to improve performance, we have observed ephemeral computational errors that were not detected during manufacturing tests. These defects cannot always be mitigated by techniques such as microcode updates, and may be correlated to specific components within the processor, allowing small code changes to effect large shifts in reliability. Worse, these failures are often "silent" - the only symptom is an erroneous computation. Peter Hochschild, Jeffrey C. Mogul, Rama Govindaraju, Parthasarathy Ranganathan, David E. Culler, Amin Vahdat |
HotOS | 7 |
| 2021 | Orion: Google's Software-Defined Networking Control Plane
Andrew D. Ferguson, Steve D. Gribble, Chi-Yao Hong, Chip Killian, Waqar Mohsin, Henrik Mühe, Joon Ong, Leonid B. Poutievski, Lorenzo Vicisano, Richard Alimi, Shawn Shuoshuo Chen, Mike Conley, Subhasree Mandal, Karthik Nagaraj, Kondapa Naidu Bollineni, Amr Sabaa, Shidong Zhang, Amin Vahdat |
NSDI | 20 |
| 2021 | SiP-ML: high-bandwidth optical network interconnects for machine learning trainingabstractThis paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3--9.1x. Mehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, Eiman Ebrahimi |
SIGCOMM | 7 |
| 2021 | CliqueMap: productionizing an RMA-based distributed caching systemabstractDistributed in-memory caching is a key component of modern Internet services. Such caches are often accessed via remote procedure call (RPC), as RPC frameworks provide rich support for productionization, including protocol versioning, memory efficiency, auto-scaling, and hitless upgrades. However, full-featured RPC limits performance and scalability as it incurs high latencies and CPU overheads. Remote Memory Access (RMA) offers a promising alternative, but meeting productionization requirements can be a significant challenge with RMA-based systems due to limited programmability and narrow RMA primitives. Arjun Singhvi, Aditya Akella, Maggie Anderson, Rob Cauble, Harshad Deshmukh, Dan Gibson, Milo M. K. Martin, Amanda Strominger, Thomas F. Wenisch, Amin Vahdat |
SIGCOMM | 10 |
| 2020 | Sundial: Fault-tolerant Clock Synchronization for Datacenters
Gautam Kumar 0001, Hema Hariharan, Hassan M. G. Wassel, Peter Hochschild, Dave Platt, Simon L. Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, Amin Vahdat |
OSDI | 11 |
| 2020 | Annulus: A Dual Congestion Control Loop for Datacenter and WAN Traffic AggregatesabstractCloud services are deployed in datacenters connected though high-bandwidth Wide Area Networks (WANs). We find that WAN traffic negatively impacts the performance of datacenter traffic, increasing tail latency by 2.5x, despite its small bandwidth demand. This behavior is caused by the long round-trip time (RTT) for WAN traffic, combined with limited buffering in datacenter switches. The long WAN RTT forces datacenter traffic to take the full burden of reacting to congestion. Furthermore, datacenter traffic changes on a faster time-scale than the WAN RTT, making it difficult for WAN congestion control to estimate available bandwidth accurately. Ahmed Saeed 0001, Prateesh Goyal, Milad Sharif, Mostafa H. Ammar, Ellen Zegura, Keon Jang, Mohammad Alizadeh, Abdul Kabbani, Amin Vahdat |
SIGCOMM | 11 |
| 2020 | Swift: Delay is Simple and Effective for Congestion Control in the DatacenterabstractWe report on experiences with Swift congestion control in Google datacenters. Swift targets an end-to-end delay by using AIMD control, with pacing under extreme congestion. With accurate RTT measurement and care in reasoning about delay targets, we find this design is a foundation for excellent performance when network distances are well-known. Importantly, its simplicity helps us to meet operational challenges. Delay is easy to decompose into fabric and host components to separate concerns, and effortless to deploy and maintain as a congestion signal while the datacenter evolves. In large-scale testbed experiments, Swift delivers a tail latency of <50μs for short RPCs, with near-zero packet drops, while sustaining ~100Gbps throughput per server. This is a tail of <3x the minimal latency at a load close to 100%. In production use in many different clusters, Swift achieves consistently low tail completion times for short RPCs, while providing high throughput for long RPCs. It has loss rates that are at least 10x lower than a DCTCP protocol, and handles O(10k) incasts that sharply degrade with DCTCP. Gautam Kumar 0001, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel, Xian Wu 0001, Behnam Montazeri, Yaogong Wang, Kevin Springborn, Christopher Alfeld, Michael Ryan, David Wetherall, Amin Vahdat |
SIGCOMM | 12 |
| 2020 | 1RMA: Re-envisioning Remote Memory Access for Multi-tenant DatacentersabstractRemote Direct Memory Access (RDMA) plays a key role in supporting performance-hungry datacenter applications. However, existing RDMA technologies are ill-suited to multi-tenant datacenters, where applications run at massive scales, tenants require isolation and security, and the workload mix changes over time. Our experiences seeking to operationalize RDMA at scale indicate that these ills are rooted in standard RDMA's basic design attributes: connectionorientedness and complex policies baked into hardware. Arjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch, Monica Wong-Chan, Sean Clark 0003, Milo M. K. Martin, Moray McLaren, Prashant Chandra, Rob Cauble, Hassan M. G. Wassel, Behnam Montazeri, Simon L. Sabato, Joel Scherpelz, Amin Vahdat |
SIGCOMM | 15 |
| 2019 | Using P4 on Fixed-Pipeline and Programmable Stratum SwitchesabstractStratum is an open source network operating system (NOS)that provides a common implementation of P4Runtime and OpenConfig interfaces for white box switches. This demonstration will show an SDN leaf-spine fabric of Stratum-enabled white box switches managed by the ONOS SDN controller. The switching chips (ASICs)and platforms will come from different vendors, but they will share a common P4-defined pipeline and set of OpenConfig models. Brian O'Connor, Tomek Madejski, Jim Wanderer, Amin Vahdat, Yi Tseng, Maximilian Pudelko, Carmelo Cascone, Abhilash Endurthi, Alireza Ghaffarkhah, Devjit Gopalpur, Tom Everman |
ANCS | 4 |
| 2019 | Eiffel: Efficient and Flexible Software Packet Scheduling
Ahmed Saeed 0001, Yimeng Zhao, Nandita Dukkipati, Ellen Zegura, Mostafa H. Ammar, Khaled A. Harras, Amin Vahdat |
NSDI | 7 |
| 2019 | SIMON: A Simple and Scalable Method for Sensing, Inference and Measurement in Data Center Networks
Yilong Geng, Zi Yin, Ashish Naik, Balaji Prabhakar, Mendel Rosenblum, Amin Vahdat |
NSDI | 7 |
| 2019 | Minimal Rewiring: Efficient Live Expansion for Clos Data Center Networks
Shizhen Zhao, Rui Wang 0025, Junlan Zhou, Joon Ong, Jeffrey C. Mogul, Amin Vahdat |
NSDI | 6 |
| 2019 | PicNIC: predictable virtualized NICabstractNetwork virtualization stacks are the linchpins of public clouds. A key goal is to provide performance isolation so that workloads on one Virtual Machine (VM) do not adversely impact the network experience of another VM. Using data from a major public cloud provider, we systematically characterize how performance isolation can break in current virtualization stacks and find a fundamental tradeoff between isolation and resource multiplexing for efficiency. In order to provide predictable performance, we propose a new system called PicNIC that shares resources efficiently in the common case while rapidly reacting to ensure isolation. PicNIC builds on three constructs to quickly detect isolation breakdown and to enforce it when necessary: CPU-fair weighted fair queues at receivers, receiver-driven congestion control for backpressure, and sender-side admission control with shaping. Based on an extensive evaluation, we show that this combination ensures isolation for VMs at sub-millisecond timescales with negligible overhead. Praveen Kumar 0003, Nandita Dukkipati, Nathan Lewis, Yaogong Wang, Chonggang Li, Valas Valancius, Jake Adriaens, Steve D. Gribble, Nate Foster, Amin Vahdat |
SIGCOMM | 11 |
| 2019 | Risk based planning of network changes in evolving data centersabstractData center networks evolve as they serve customer traffic. When applying network changes, operators risk impacting customer traffic because the network operates at reduced capacity and is more vulnerable to failures and traffic variations. The impact on customer traffic ultimately translates to operator cost (e.g., refunds to customers). However, planning a network change while minimizing the risks is challenging as we need to adapt to a variety of traffic dynamics and cost functions while scaling to large networks and large changes. Today, operators often use plans that maximize the residual capacity (MRC), which often incurs a high cost under different traffic dynamics. Instead, we propose Janus, which searches the large planning space by leveraging the high degree of symmetry in data center networks. Our evaluation on large Clos networks and Facebook traffic traces shows that Janus generates plans in real-time only needing 33~71% of the cost of MRC planners while adapting to a variety of settings. Omid Alipourfard, Jérémie Koenig, Christopher Harshaw, Amin Vahdat, Minlan Yu |
SOSP | 5 |
| 2019 | Snap: a microkernel approach to host networkingabstractThis paper presents our design and experience with a microkernel-inspired approach to host networking called Snap. Snap is a userspace networking system that supports Google's rapidly evolving needs with flexible modules that implement a range of network functions, including edge packet switching, virtualization for our cloud platform, traffic shaping policy enforcement, and a high-performance reliable messaging and RDMA-like service. Snap has been running in production for over three years, supporting the extensible communication needs of several large and critical systems. Michael Marty, Marc de Kruijf, Jacob Adriaens, Christopher Alfeld, Sean Bauer, Carlo Contavalli, Michael Dalton, Nandita Dukkipati, William C. Evans, Steve D. Gribble, Nicholas Kidd, Roman Kononov, Gautam Kumar 0001, Carl Mauer, Emily Musick, Lena E. Olson, Erik Rubow, Michael Ryan, Kevin Springborn, Valas Valancius, Amin Vahdat |
SOSP | 23 |
| 2018 | Andromeda: Performance, Isolation, and Velocity at Scale in Cloud Network Virtualization
Michael Dalton, David Schultz, Jacob Adriaens, Ahsan Arefin, Anshuman Gupta, Brian Fahs, Dima Rubinstein, Enrique Cauich Zermeno, Erik Rubow, James Alexander Docauer, Jesse Alpert, Jing Ai, Jon Olson, Kevin DeCabooter, Marc de Kruijf, Nan Hua, Nathan Lewis, Nikhil Kasinadhuni, Riccardo Crepaldi, Srinivas Krishnan, Subbaiah Venkata, Yossi Richter, Uday Naik, Amin Vahdat |
NSDI | 24 |
| 2018 | Exploiting a Natural Network Effect for Scalable, Fine-grained Clock Synchronization
Yilong Geng, Zi Yin, Ashish Naik, Balaji Prabhakar, Mendel Rosenblum, Amin Vahdat |
NSDI | 7 |
| 2018 | Sincronia: near-optimal network design for coflowsabstractWe present Sincronia, a near-optimal network design for coflows that can be implemented on top on any transport layer (for flows) that supports priority scheduling. Sincronia achieves this using a key technical result --- we show that given a "right" ordering of coflows, any per-flow rate allocation mechanism achieves average coflow completion time within 4X of the optimal as long as (co)flows are prioritized with respect to the ordering. Saksham Agarwal, Shijin Rajakrishnan, Akshay Narayan 0001, Rachit Agarwal 0001, David B. Shmoys, Amin Vahdat |
SIGCOMM | 6 |
| 2018 | B4 and after: managing hierarchy, partitioning, and asymmetry for availability and scale in google's software-defined WANabstractPrivate WANs are increasingly important to the operation of enterprises, telecoms, and cloud providers. For example, B4, Google's private software-defined WAN, is larger and growing faster than our connectivity to the public Internet. In this paper, we present the five-year evolution of B4. We describe the techniques we employed to incrementally move from offering best-effort content-copy services to carrier-grade availability, while concurrently scaling B4 to accommodate 100x more traffic. Our key challenge is balancing the tension introduced by hierarchy required for scalability, the partitioning required for availability, and the capacity asymmetry inherent to the construction and operation of any large-scale network. We discuss our approach to managing this tension: i) we design a custom hierarchical network topology for both horizontal and vertical software scaling, ii) we manage inherent capacity asymmetry in hierarchical topologies using a novel traffic engineering algorithm without packet encapsulation, and iii) we re-architect switch forwarding rules via two-stage matching/hashing to deal with asymmetric network failures at scale. Chi-Yao Hong, Subhasree Mandal, Mohammad Al-Fares, Richard Alimi, Kondapa Naidu Bollineni, Chandan Bhagat, Sourabh Jain, Jay Kaimal, Shiyu Liang, Kirill Mendelev, Steve Padgett, Faro Rabe, Saikat Ray, Malveeka Tewari, Matt Tierney, Monika Zahn, Jonathan Zolla, Joon Ong, Amin Vahdat |
SIGCOMM | 20 |
| 2017 | Carousel: Scalable Traffic Shaping at End HostsabstractTraffic shaping, including pacing and rate limiting, is fundamental to the correct and efficient operation of both datacenter and wide area networks. Sample use cases include policy-based bandwidth allocation to flow aggregates, rate-based congestion control algorithms, and packet pacing to avoid bursty transmissions that can overwhelm router buffers. Driven by the need to scale to millions of flows and to apply complex policies, traffic shaping is moving from network switches into the end hosts, typically implemented in software in the kernel networking stack. Ahmed Saeed 0001, Nandita Dukkipati, Vytautas Valancius, Vinh The Lam, Carlo Contavalli, Amin Vahdat |
SIGCOMM | 6 |
| 2017 | Taking the Edge off with Espresso: Scale, Reliability and Programmability for Global Internet PeeringabstractWe present the design of Espresso, Google's SDN-based Internet peering edge routing infrastructure. This architecture grew out of a need to exponentially scale the Internet edge cost-effectively and to enable application-aware routing at Internet-peering scale. Espresso utilizes commodity switches and host-based routing/packet processing to implement a novel fine-grained traffic engineering capability. Overall, Espresso provides Google a scalable peering edge that is programmable, reliable, and integrated with global traffic systems. Espresso also greatly accelerated deployment of new networking features at our peering edge. Espresso has been in production for two years and serves over 22% of Google's total traffic to the Internet. Kok-Kiong Yap, Murtaza Motiwala, Jeremy Rahe, Steve Padgett, Matthew J. Holliman, Gary Baldus, Marcus Hines, Taeeun Kim, Ashok Narayanan, Victor Lin, Colin Rice, Brian Rogan, Bert Tanaka, Manish Verma, Puneet Sood, Muhammad Mukarram Bin Tariq, Matt Tierney, Dzevad Trumic, Vytautas Valancius, Calvin Ying, Mahesh Kallahalla, Bikash Koley, Amin Vahdat |
SIGCOMM | 25 |
| 2017 | Local Recovery for High Availability in Strongly Consistent Cloud ServicesabstractEmerging cloud-based network services must deliver both good performance and high availability. Achieving both of these goals requires content replication across multiple sites. Many cloud-based services either require or would benefit from the semantics and simplicity of strong consistency. However, replication techniques for strong consistency can severely limit the availability of replicated services when recovering large data objects over wide-area links. To address this problem, we present the design and implementation of ZORFU, a hierarchical system architecture for replication across data centers. The primary contribution of ZORFU is a local recovery technique that significantly increases availability of replicated strongly consistent services. Local recovery achieves this by reducing the recovery time by an order of magnitude, while imposing only a negligible latency overhead. Experimental results show that ZORFU can recover a 100 MB object in 4 ms. James W. Anderson, Hein Meling, Alexander Rasmussen, Amin Vahdat, Keith Marzullo |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2016 | Evolve or Die: High-Availability Design Principles Drawn from Googles Network InfrastructureabstractMaintaining the highest levels of availability for content providers is challenging in the face of scale, network evolution and complexity. Little, however, is known about failures large content providers are susceptible to, and what mechanisms they employ to ensure high availability. From a detailed analysis of over 100 high-impact failure events in a global-scale content provider encompassing several data centers and two WANs, we quantify several dimensions of availability failures. We find that failures are evenly distributed across different network types and planes, but that a large number of failures happen when a management operation is in progress within the network. We discuss some of these failures in detail, and also describe our design principles for high availability motivated by these failures, including using defense in depth, maintaining consistency across planes, failing open on large failures, carefully preventing and avoiding failures, and assessing root cause quickly. Our findings suggest that, as networks become more complicated, failures lurk everywhere, and, counter-intuitively, continuous incremental evolution of the network can, when applied together with our design principles, result in a more robust network. Ramesh Govindan, Ina Minei, Mahesh Kallahalla, Bikash Koley, Amin Vahdat |
SIGCOMM | 5 |
| 2016 | Trumpet: Timely and Precise Triggers in Data CentersabstractAs data centers grow larger and strive to provide tight performance and availability SLAs, their monitoring infrastructure must move from passive systems that provide aggregated inputs to human operators, to active systems that enable programmed control. In this paper, we propose Trumpet, an event monitoring system that leverages CPU resources and end-host programmability, to monitor every packet and report events at millisecond timescales. Trumpet users can express many *network-wide events*, and the system efficiently detects these events using *triggers* at end-hosts. Using careful design, Trumpet can evaluate triggers by inspecting every packet at full line rate even on future generations of NICs, scale to thousands of triggers per end-host while bounding packet processing delay to a few microseconds, and report events to a controller within 10 milliseconds, even in the presence of attacks. We demonstrate these properties using an implementation of Trumpet, and also show that it allows operators to describe new network events such as detecting correlated bursts and loss, identifying the root cause of transient congestion, and detecting short-term anomalies at the scale of a data center tenant. Masoud Moshref, Minlan Yu, Ramesh Govindan, Amin Vahdat |
SIGCOMM | 4 |
| 2015 | Achieving cost-efficient, data-intensive computing in the cloudabstractCloud computing providers have recently begun to offer high-performance virtualized flash storage and virtualized network I/O capabilities, which have the potential to increase application performance. Since users pay for only the resources they use, these new resources have the potential to lower overall cost. Yet achieving low cost requires choosing the right mixture of resources, which is only possible if their performance and scaling behavior is known. Michael Conley, Amin Vahdat, George Porter |
SoCC | 2 |
| 2015 | FastLane: making short flows shorter with agile drop notificationabstractThe drive towards richer and more interactive web content places increasingly stringent requirements on datacenter network performance. Applications running atop these networks typically partition an incoming query into multiple subqueries, and generate the final result by aggregating the responses for these subqueries. As a result, a large fraction --- as high as 80% --- of the network flows in such workloads are short and latency-sensitive. The speed with which existing networks respond to packet drops limits their ability to meet high-percentile flow completion time SLOs. Indirect notifications indicating packet drops (e.g., duplicates in an end-to-end acknowledgement sequence) are an important limitation to the agility of response to packet drops. David Zats, Anand Padmanabha Iyer, Ganesh Ananthanarayanan, Rachit Agarwal 0001, Randy H. Katz, Ion Stoica, Amin Vahdat |
SoCC | 7 |
| 2015 | SCREAM: sketch resource allocation for software-defined measurementabstractSoftware-defined networks can enable a variety of concurrent, dynamically instantiated, measurement tasks, that provide fine-grain visibility into network traffic. Recently, there have been many proposals for using sketches for network measurement. However, sketches in hardware switches use constrained resources such as SRAM memory, and the accuracy of measurement tasks is a function of the resources devoted to them on each switch. This paper presents SCREAM, a system for allocating resources to sketch-based measurement tasks that ensures a user-specified minimum accuracy. SCREAM estimates the instantaneous accuracy of tasks so as to dynamically adapt the allocated resources for each task. Thus, by finding the right amount of resources for each task on each switch and correctly merging sketches at the controller, SCREAM can multiplex resources among network-wide measurement tasks. Simulations with three measurement tasks (heavy hitter, hierarchical heavy hitter, and super source/destination detection) show that SCREAM can support more measurement tasks with higher accuracy than existing approaches. Masoud Moshref, Minlan Yu, Ramesh Govindan, Amin Vahdat |
CoNEXT | 4 |
| 2015 | BwE: Flexible, Hierarchical Bandwidth Allocation for WAN Distributed ComputingabstractWAN bandwidth remains a constrained resource that is economically infeasible to substantially overprovision. Hence, it is important to allocate capacity according to service priority and based on the incremental value of additional allocation. For example, it may be the highest priority for one service to receive 10Gb/s of bandwidth but upon reaching such an allocation, incremental priority may drop sharply favoring allocation to other services. Motivated by the observation that individual flows with fixed priority may not be the ideal basis for bandwidth allocation, we present the design and implementation of Bandwidth Enforcer (BwE), a global, hierarchical bandwidth allocation infrastructure. BwE supports: i) service-level bandwidth allocation following prioritized bandwidth functions where a service can represent an arbitrary collection of flows, ii)independent allocation and delegation policies according to user-defined hierarchy, all accounting for a global view of bandwidth and failure conditions, iii) multi-path forwarding common in traffic-engineered networks, and iv) a central administrative point to override (perhaps faulty) policy during exceptional conditions. BwE has delivered more service efficient bandwidth utilization and simpler management in production for multiple years. Sushant Jain, Uday Naik, Anand Raghuraman, Nikhil Kasinadhuni, Enrique Cauich Zermeno, C. Stephen Gunn, Jing Ai, Björn Carlin, Mihai Amarandei-Stavila, Mathieu Robin, Aspi Siganporia, Stephen Stuart, Amin Vahdat |
SIGCOMM | 14 |
| 2015 | TIMELY: RTT-based Congestion Control for the DatacenterabstractDatacenter transports aim to deliver low latency messaging together with high throughput. We show that simple packet delay, measured as round-trip times at hosts, is an effective congestion signal without the need for switch feedback. First, we show that advances in NIC hardware have made RTT measurement possible with microsecond accuracy, and that these RTTs are sufficient to estimate switch queueing. Then we describe how TIMELY can adjust transmission rates using RTT gradients to keep packet latency low while delivering high bandwidth. We implement our design in host software running over NICs with OS-bypass capabilities. We show using experiments with up to hundreds of machines on a Clos network topology that it provides excellent performance: turning on TIMELY for OS-bypass messaging over a fabric with PFC lowers 99 percentile tail latency by 9X while maintaining near line-rate throughput. Our system also outperforms DCTCP running in an optimized kernel, reducing tail latency by $13$X. To the best of our knowledge, TIMELY is the first delay-based congestion control protocol for use in the datacenter, and it achieves its results despite having an order of magnitude fewer RTT signals (due to NIC offload) than earlier delay-based schemes such as Vegas. Radhika Mittal, Vinh The Lam, Nandita Dukkipati, Emily R. Blem, Hassan M. G. Wassel, Manya Ghobadi, Amin Vahdat, Yaogong Wang, David Wetherall, David Zats |
SIGCOMM | 7 |
| 2015 | Condor: Better Topologies Through Declarative DesignabstractThe design space for large, multipath datacenter networks is large and complex, and no one design fits all purposes. Network architects must trade off many criteria to design cost-effective, reliable, and maintainable networks, and typically cannot explore much of the design space. We present Condor, our approach to enabling a rapid, efficient design cycle. Condor allows architects to express their requirements as constraints via a Topology Description Language (TDL), rather than having to directly specify network structures. Condor then uses constraint-based synthesis to rapidly generate candidate topologies, which can be analyzed against multiple criteria. We show that TDL supports concise descriptions of topologies such as fat-trees, BCube, and DCell; that we can generate known and novel variants of fat-trees with simple changes to a TDL file; and that we can synthesize large topologies in tens of seconds. We also show that Condor supports the daunting task of designing multi-phase network expansions that can be carried out on live networks. Brandon Schlinker, Radhika Niranjan Mysore, Jeffrey C. Mogul, Amin Vahdat, Minlan Yu, Ethan Katz-Bassett, Michael Rubin |
SIGCOMM | 5 |
| 2015 | Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google's Datacenter NetworkabstractWe present our approach for overcoming the cost, operational complexity, and limited scale endemic to datacenter networks a decade ago. Three themes unify the five generations of datacenter networks detailed in this paper. First, multi-stage Clos topologies built from commodity switch silicon can support cost-effective deployment of building-scale networks. Second, much of the general, but complex, decentralized network routing and management protocols supporting arbitrary deployment scenarios were overkill for single-operator, pre-planned datacenter networks. We built a centralized control mechanism based on a global configuration pushed to all datacenter switches. Third, modular hardware design coupled with simple, robust software allowed our design to also support inter-cluster and wide-area networks. Our datacenter networks run at dozens of sites across the planet, scaling in capacity by 100x over ten years to more than 1Pbps of bisection bandwidth. Joon Ong, Glen Anderson, Ashby Armistead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, Amin Vahdat |
SIGCOMM | 18 |
| 2014 | WCMP: weighted cost multipathing for improved fairness in data centersabstractData Center topologies employ multiple paths among servers to deliver scalable, cost-effective network capacity. The simplest and the most widely deployed approach for load balancing among these paths, Equal Cost Multipath (ECMP), hashes flows among the shortest paths toward a destination. ECMP leverages uniform hashing of balanced flow sizes to achieve fairness and good load balancing in data centers. However, we show that ECMP further assumes a balanced, regular, and fault-free topology, which are invalid assumptions in practice that can lead to substantial performance degradation and, worse, variation in flow bandwidths even for same size flows. Junlan Zhou, Malveeka Tewari, Abdul Kabbani, Leonid B. Poutievski, Amin Vahdat |
EuroSys | 7 |
| 2014 | Cutting the cord: a robust wireless facilities network for data centersabstractToday's network control and management traffic are limited by their reliance on existing data networks. Fate sharing in this context is highly undesirable, since control traffic has very different availability and traffic delivery requirements. In this paper, we explore the feasibility of building a dedicated wireless facilities network for data centers. We propose Angora, a low-latency facilities network using low-cost, 60GHz beamforming radios that provides robust paths decoupled from the wired network, and flexibility to adapt to workloads and network dynamics. We describe our solutions to address challenges in link coordination, link interference and network failures. Our testbed measurements and simulation results show that Angora enables large number of low-latency control paths to run concurrently, while providing low latency end-to-end message delivery with high tolerance for radio and rack failures. Yibo Zhu 0001, Zengbin Zhang, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 5 |
| 2014 | SENIC: Scalable NIC for End-Host Rate Limiting
Sivasankar Radhakrishnan, Yilong Geng, Vimalkumar Jeyakumar, Abdul Kabbani, George Porter, Amin Vahdat |
NSDI | 6 |
| 2014 | Libra: Divide and Conquer to Verify Forwarding Tables in Huge Networks
Hongyi Zeng, Shidong Zhang, Vimalkumar Jeyakumar, Mickey Ju, Junda Liu, Nick McKeown, Amin Vahdat |
NSDI | 8 |
| 2014 | DREAM: dynamic resource allocation for software-defined measurementabstractSoftware-defined networks can enable a variety of concurrent, dynamically instantiated, measurement tasks, that provide fine-grain visibility into network traffic. Recently, there have been many proposals to configure TCAM counters in hardware switches to monitor traffic. However, the TCAM memory at switches is fundamentally limited and the accuracy of the measurement tasks is a function of the resources devoted to them on each switch. This paper describes an adaptive measurement framework, called DREAM, that dynamically adjusts the resources devoted to each measurement task, while ensuring a user-specified level of accuracy. Since the trade-off between resource usage and accuracy can depend upon the type of tasks, their parameters, and traffic characteristics, DREAM does not assume an a priori characterization of this trade-off, but instead dynamically searches for a resource allocation that is sufficient to achieve a desired level of accuracy. A prototype implementation and simulations with three network-wide measurement tasks (heavy hitter, hierarchical heavy hitter and change detection) and diverse traffic show that DREAM can support more concurrent tasks with higher accuracy than several other alternatives. Masoud Moshref, Minlan Yu, Ramesh Govindan, Amin Vahdat |
SIGCOMM | 4 |
| 2014 | Gestalt: Fast, Unified Fault Localization for Networked Systems
Radhika Niranjan Mysore, Ratul Mahajan, Amin Vahdat, George Varghese |
USENIX ATC | 3 |
| 2013 | Dahu: Commodity switches for direct connect data center networksabstractSolving “Big Data” problems requires bridging massive quantities of compute, memory, and storage, which requires a very high bandwidth network. Recently proposed direct connect networks like HyperX [1] and Flattened Butterfly [20] offer large capacity through paths of varying lengths between servers, and are highly cost effective for common data center workloads. However data center deployments are constrained to multi-rooted tree topologies like Fat-tree [2] and VL2 [16] due to shortest path routing and the limitations of commodity data center switch silicon. In this work we present Dahu1, simple enhancements to commodity Ethernet switches to support direct connect networks in data centers. Dahu avoids congestion hot-spots by dynamically spreading traffic uniformly across links, and forwarding traffic over non-minimal paths where possible. By performing load balancing primarily using local information, Dahu can act more quickly than centralized approaches, and responds to failure gracefully. Our evaluation shows that Dahu delivers up to 500% improvement in throughput over ECMP in large scale HyperX networks with over 130,000 servers, and up to 50% higher throughput in an 8,192 server Fat-tree network. Sivasankar Radhakrishnan, Malveeka Tewari, Rishi Kapoor, George Porter, Amin Vahdat |
ANCS | 5 |
| 2013 | FasTrak: enabling express lanes in multi-tenant data centersabstractThe shared nature of multi-tenant cloud networks requires providing tenant isolation and quality of service, which in turn requires enforcing thousands of network-level rules, policies, and traffic rate limits. Enforcing these rules in virtual machine hypervisors imposes significant computational overhead, as well as increased latency. In FasTrak, we seek to exploit temporal locality in flows and flow sizes to offload a subset of network virtualization functionality from the hypervisor into switch hardware freeing up the hypervisor. FasTrak manages the required hardware and hypervisor rules as a unified set, moving rules back and forth to minimize the overhead of network virtualization, and focusing on flows (or flow aggregates) that are either most latency sensitive or exhibit the highest packets-per-second rates. Radhika Niranjan Mysore, George Porter, Amin Vahdat |
CoNEXT | 3 |
| 2013 | Aspen trees: balancing data center fault tolerance, scalability and costabstractFault recovery is a key issue in modern data centers. In a fat tree topology, a single link failure can disconnect a set of end hosts from the rest of the network until updated routing information is disseminated to every switch in the topology. The time for re-convergence can be substantial, leaving hosts disconnected for long periods of time and significantly reducing the overall availability of the data center. Moreover, the message overhead of sending updated routing information to the entire topology may be unacceptable at scale. We present techniques to modify hierarchical data center topologies to enable switches to react to failures locally, thus reducing both the convergence time and control overhead of failure recovery. We find that for a given network size, decreasing a topology's convergence time results in a proportional decrease to its scalability (e.g. the number of hosts supported). On the other hand, reducing convergence time without affecting scalability necessitates the introduction of additional switches and links. We explore the tradeoffs between fault tolerance, scalability and network size, and propose a range of modified multi-rooted tree topologies that provide significantly reduced convergence time while retaining most of the traditional fat tree's desirable properties. Meg Walraed-Sullivan, Amin Vahdat, Keith Marzullo |
CoNEXT | 2 |
| 2013 | Peer-to-Peer Keyword Search: A Retrospective
Patrick Reynolds, Amin Vahdat |
Middleware | 2 |
| 2013 | B4: experience with a globally-deployed software defined wanabstractWe present the design, implementation, and evaluation of B4, a private WAN connecting Google's data centers across the planet. B4 has a number of unique characteristics: i) massive bandwidth requirements deployed to a modest number of sites, ii) elastic traffic demand that seeks to maximize average bandwidth, and iii) full control over the edge servers and network, which enables rate limiting and demand measurement at the edge. Sushant Jain, Subhasree Mandal, Joon Ong, Leonid B. Poutievski, Subbaiah Venkata, Jim Wanderer, Junlan Zhou, Jonathan Zolla, Urs Hölzle, Stephen Stuart, Amin Vahdat |
SIGCOMM | 14 |
| 2013 | Integrating microsecond circuit switching into the data centerabstractRecent proposals have employed optical circuit switching (OCS) to reduce the cost of data center networks. However, the relatively slow switching times (10--100 ms) assumed by these approaches, and the accompanying latencies of their control planes, has limited its use to only the largest data center networks with highly aggregated and constrained workloads. As faster switch technologies become available, designing a control plane capable of supporting them becomes a key challenge. George Porter, Richard D. Strong, Nathan Farrington, Alex Forencich, Pang-Chen Sun, Tajana Rosing, Yeshaiahu Fainman, George Papen, Amin Vahdat |
SIGCOMM | 9 |
| 2013 | TritonSort: A Balanced and Energy-Efficient Large-Scale Sorting SystemabstractWe present TritonSort, a highly efficient, scalable sorting system. It is designed to process large datasets, and has been evaluated against as much as 100TB of input data spread across 832 disks in 52 nodes at a rate of 0.938TB/min. When evaluated against the annual Indy GraySort sorting benchmark, TritonSort is 66% better in absolute performance and has over six times the per-node throughput of the previous record holder. When evaluated against the 100TB Indy JouleSort benchmark, TritonSort sorted 9703 records/Joule. In this article, we describe the hardware and software architecture necessary to operate TritonSort at this level of efficiency. Through careful management of system resources to ensure cross-resource balance, we are able to sort data at approximately 80% of the disks’ aggregate sequential write speed. We believe the work holds a number of lessons for balanced system design and for scale-out architectures in general. While many interesting systems are able to scale linearly with additional servers, per-server performance can lag behind per-server capacity by more than an order of magnitude. Bridging the gap between high scalability and high performance would enable either significantly less expensive systems that are able to do the same work or provide the ability to address significantly larger problem sets with the same infrastructure. Alexander Rasmussen, George Porter, Michael Conley, Harsha V. Madhyastha, Radhika Niranjan Mysore, Alexander Pucher, Amin Vahdat |
ACM Trans. Comput. Syst. | 7 |
| 2012 | NetBump: user-extensible active queue management with bumps on the wireabstractEngineering large-scale data center applications built from thousands of commodity nodes requires both an underlying network that supports a wide variety of traffic demands, and low latency at microsecond timescales. Many ideas for adding innovative functionality to networks, especially active queue management strategies, require either modifying packets or performing alternative queuing to packets in-flight on the data plane. However, configuring packet queuing, marking, and dropping is challenging, since buffering in commercial switches and routers is not programmable. Mohammad Al-Fares, Rishi Kapoor, George Porter, Sambit Das, Hakim Weatherspoon, Balaji Prabhakar, Amin Vahdat |
ANCS | 7 |
| 2012 | xOMB: extensible open middleboxes with commodity serversabstractThis paper presents the design and implementation of an incrementally scalable architecture for middleboxes based on commodity servers and operating systems. xOMB, the eXtensible Open MiddleBox, employs general programmable network processing pipelines, with user-defined C++ modules responsible for parsing, transforming, and forwarding network flows. We implement three processing pipelines in xOMB, demonstrating good performance for load balancing, protocol acceleration, and application integration. In particular, our xOMB load balancing switch is able to match or outperform a commercial programmable switch and popular open-source reverse proxy while still providing a more flexible programming model. James W. Anderson, Ryan Braud, Rishi Kapoor, George Porter, Amin Vahdat |
ANCS | 5 |
| 2012 | Chronos: predictable low latency for data center applicationsabstractIn data center applications, predictability in service time and controlled latency, especially tail latency, are essential for building performant applications. This is especially true for applications or services built by accessing data across thousands of servers to generate a user response. Current practice has been to run such services at low utilization to rein in latency outliers, which decreases efficiency and limits the number of service invocations developers can issue while still meeting tight latency budgets. Rishi Kapoor, George Porter, Malveeka Tewari, Geoffrey M. Voelker, Amin Vahdat |
SoCC | 5 |
| 2012 | Themis: an I/O-efficient MapReduceabstract"Big Data" computing increasingly utilizes the MapReduce programming model for scalable processing of large data collections. Many MapReduce jobs are I/O-bound, and so minimizing the number of I/O operations is critical to improving their performance. In this work, we present Themis, a MapReduce implementation that reads and writes data records to disk exactly twice, which is the minimum amount possible for data sets that cannot fit in memory. Alexander Rasmussen, Vinh The Lam, Michael Conley, George Porter, Rishi Kapoor, Amin Vahdat |
SoCC | 6 |
| 2012 | Practical TDMA for datacenter ethernetabstractCloud computing is placing increasingly stringent demands on datacenter networks. Applications like MapReduce and Hadoop demand high bisection bandwidth to support their all-to-all shuffle communication phases. Conversely, Web services often rely on deep chains of relatively lightweight RPCs. While HPC vendors market niche hardware solutions, current approaches to providing high-bandwidth and low-latency communication in the datacenter exhibit significant inefficiencies on commodity Ethernet hardware. Bhanu Chandra Vattikonda, George Porter, Amin Vahdat, Alex C. Snoeren |
EuroSys | 3 |
| 2012 | scc: cluster storage provisioning informed by application characteristics and SLAs
Harsha V. Madhyastha, John McCullough, George Porter, Rishi Kapoor, Stefan Savage, Alex C. Snoeren, Amin Vahdat |
FAST | 7 |
| 2012 | Hunting mice with microsecond circuit switchesabstractRecently, there have been proposals for constructing hybrid data center networks combining electronic packet switching with either wireless or optical circuit switching, which are ideally suited for supporting bulk traffic. Previous work has relied on a technique called hotspot scheduling, in which the traffic matrix is measured, hotspots identified, and circuits established to automatically offload traffic from the packet-switched network. While this hybrid approach does reduce CAPEX and OPEX, it still relies on having a well-provisioned packet-switched network to carry the remaining traffic. In this paper, we describe a generalization of hotspot scheduling, called traffic matrix scheduling, where most or even all bulk traffic is routed over circuits. In other words, we don't just hunt elephants, we also hunt mice. Traffic matrix scheduling rapidly time-shares circuits across many destinations at microsecond time scales. The traffic matrix scheduling algorithm can route arbitrary traffic patterns and runs in polynomial time. We briefly describe a working implementation of traffic matrix scheduling using a custom-built data center optical circuit switch with a 2.8 microsecond switching time. Nathan Farrington, George Porter, Yeshaiahu Fainman, George Papen, Amin Vahdat |
HotNets | 5 |
| 2012 | Less Is More: Trading a Little Bandwidth for Ultra-Low Latency in the Data Center
Mohammad Alizadeh, Abdul Kabbani, Tom Edsall, Balaji Prabhakar, Amin Vahdat, Masato Yasuda |
NSDI | 5 |
| 2012 | A demonstration of ultra-low-latency data center optical circuit switchingabstractWe designed and constructed a 24x24-port optical circuit switch (OCS) prototype with a programming time of 68.5 μs, a switching time of 2.8 μs, and a receiver electronics initialization time of 8.7 μs [1]. We demonstrate the operation of this prototype switch in a data center testbed under various workloads. Nathan Farrington, George Porter, Pang-Chen Sun, Alex Forencich, Joseph E. Ford, Yeshaiahu Fainman, George Papen, Amin Vahdat |
SIGCOMM | 8 |
| 2012 | Mirror mirror on the ceiling: flexible wireless links for data centersabstractModern data centers are massive, and support a range of distributed applications across potentially hundreds of server racks. As their utilization and bandwidth needs continue to grow, traditional methods of augmenting bandwidth have proven complex and costly in time and resources. Recent measurements show that data center traffic is often limited by congestion loss caused by short traffic bursts. Thus an attractive alternative to adding physical bandwidth is to augment wired links with wireless links in the 60 GHz band. Zengbin Zhang, Yibo Zhu 0001, Saipriya Kumar, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 6 |
| 2012 | Symbiosis in scale out networking and data managementabstractThis talk highlights the symbiotic relationship between data management and networking through a study of two seemingly independent trends in the traditionally separate communities: large-scale data processing and software defined networking. First, data processing at scale increasingly runs across hundreds or thousands of servers. We show that balancing network performance with computation and storage is a prerequisite to both efficient and scalable data processing. We illustrate the need for scale out networking in support of data management through a case study of TritonSort, currently the record holder for several sorting benchmarks, including GraySort and JouleSort. Our TritonSort experience shows that disk-bound workloads require 10 Gb/s provisioned bandwidth to keep up with modern processors while emerging flash workloads require 40 Gb/s fabrics at scale. Amin Vahdat |
SIGMOD Conference | 1 |
| 2011 | Switching the optical divide: fundamental challenges for hybrid electrical/optical datacenter networksabstractRecent proposals to build hybrid electrical (packet-switched) and optical (circuit switched) data center interconnects promise to reduce the cost, complexity, and energy requirements of very large data center networks. Supporting realistic traffic patterns, however, exposes a number of unexpected and difficult challenges to actually deploying these systems "in the wild." In this paper, we explore several of these challenges, uncovered during a year of experience using hybrid interconnects. We discuss both the problems that must be addressed to make these interconnects truly useful, and the implications of these challenges on what solutions are likely to be ultimately feasible. Hamid Hajabdolali Bazzaz, Malveeka Tewari, George Porter, T. S. Eugene Ng, David G. Andersen, Michael Kaminsky, Michael A. Kozuch, Amin Vahdat |
SoCC | 9 |
| 2011 | ALIAS: scalable, decentralized label assignment for data centersabstractModern data centers can consist of hundreds of thousands of servers and millions of virtualized end hosts. Managing address assignment while simultaneously enabling scalable communication is a challenge in such an environment. We present ALIAS, an addressing and communication protocol that automates topology discovery and address assignment for the hierarchical topologies that underlie many data center network fabrics. Addresses assigned by ALIAS interoperate with a variety of scalable communication techniques. ALIAS is fully decentralized, scales to large network sizes, and dynamically recovers from arbitrary failures, without requiring modifications to hosts or to commodity switch hardware. We demonstrate through simulation that ALIAS quickly and correctly configures networks that support up to hundreds of thousands of hosts, even in the face of failures and erroneous cabling, and we show that ALIAS is a practical solution for auto-configuration with our NetFPGA testbed implementation. Meg Walraed-Sullivan, Radhika Niranjan Mysore, Malveeka Tewari, Ying Zhang 0022, Keith Marzullo, Amin Vahdat |
SoCC | 6 |
| 2011 | TritonSort: A Balanced Large-Scale Sorting System
Alexander Rasmussen, George Porter, Michael Conley, Harsha V. Madhyastha, Radhika Niranjan Mysore, Alexander Pucher, Amin Vahdat |
NSDI | 7 |
| 2011 | Brief Announcement: A Randomized Algorithm for Label Assignment in Dynamic Networks
Meg Walraed-Sullivan, Radhika Niranjan Mysore, Keith Marzullo, Amin Vahdat |
DISC | 4 |
| 2011 | DieCast: Testing Distributed Systems with an Accurate Scale ModelabstractLarge-scale network services can consist of tens of thousands of machines running thousands of unique software configurations spread across hundreds of physical networks. Testing such services for complex performance problems and configuration errors remains a difficult problem. Existing testing techniques, such as simulation or running smaller instances of a service, have limitations in predicting overall service behavior at such scales. Testing large services should ideally be done at the same scale and configuration as the target deployment, which can be technically and economically infeasible. We present DieCast , an approach to scaling network services in which we multiplex all of the nodes in a given service configuration as virtual machines across a much smaller number of physical machines in a test harness. We show how to accurately scale CPU, network, and disk to provide the illusion that each VM matches a machine in the original service in terms of both available computing resources and communication behavior. We present the architecture and evaluation of a system we built to support such experimentation and discuss its limitations. We show that for a variety of services---including a commercial high-performance cluster-based file system---and resource utilization levels, DieCast matches the behavior of the original service while using a fraction of the physical resources. Diwaker Gupta, Kashi Venkatesh Vishwanath, Marvin McNett, Amin Vahdat, Ken Yocum, Alex C. Snoeren, Geoffrey M. Voelker |
ACM Trans. Comput. Syst. | 4 |
| 2011 | Distributed application configuration, management, and visualization with plushabstractSupport for distributed application management in large-scale networked environments remains in its early stages. Although a number of solutions exist for subtasks of application deployment, monitoring, and maintenance in distributed environments, few tools provide a unified framework for application management. Many of the existing tools address the management needs of a single type of application or service that runs in a specific environment, and these tools are not adaptable enough to be used for other applications or platforms. To this end, we present the design and implementation of Plush, a fully configurable application management infrastructure designed to meet the general requirements of several different classes of distributed applications. Plush allows developers to specifically define the flow of control needed by their computations using application building blocks. Through an extensible resource management interface, Plush supports execution in a variety of environments, including both live deployment platforms and emulated clusters. Plush also uses relaxed synchronization primitives for improving fault tolerance and liveness in failure-prone environments. To gain an understanding of how Plush manages different classes of distributed applications, we take a closer look at specific applications and evaluate how Plush provides support for each. Jeannie R. Albrecht, Christopher Tuttle, Ryan Braud, Darren Dao, Nikolay Topilski, Alex C. Snoeren, Amin Vahdat |
ACM Trans. Internet Techn. | 7 |
| 2010 | NetEx: efficient and cost-effective internet bulk content deliveryabstractThe Internet is witnessing explosive growth in traffic due to bulk content transfers, such as multimedia and software downloads, and online sharing of personal, commercial, and scientific data. Yet bulk data transfers remain very expensive and inefficient. As a result, huge amounts of digital data continue to be delivered outside of the Internet using hard drives, optical media or tapes. Meanwhile, large reserves of spare bandwidth lie unutilized in today's networks, where links are overprovisioned for peak load. We designed NetEx, a bulk transfer system that opportunistically exploits the excess capacities of network links to deliver bulk content cheaply and efficiently. Our results based on data from both a commercial tier-1 ISP and the Abilene network suggest that NetEx can considerably increase the capacity of the network, and at the same time it can provide good average performance to bulk transfers. Massimiliano Marcon, Nuno Santos 0001, Krishna P. Gummadi, Nikolaos Laoutaris, Pablo Rodriguez 0001, Amin Vahdat |
ANCS | 6 |
| 2010 | Chimpp: a click-based programming and simulation environment for reconfigurable networking hardwareabstractReconfigurable network hardware makes it easier to experiment with and prototype high-speed networking systems. However, these devices are still relatively hard to program; for example, requiring users to develop in Verilog or VHDL. Further, these devices are commonly designed to work with software on a host computer, requiring the co-development of these hardware and software components. Erik Rubow, Rick McGeer, Jeffrey C. Mogul, Amin Vahdat |
ANCS | 4 |
| 2010 | Greedy Forwarding in Dynamic Scale-Free Networks Embedded in Hyperbolic Metric SpacesabstractWe show that complex (scale-free) network topologies naturally emerge from hyperbolic metric spaces. Hyperbolic geometry facilitates maximally efficient greedy forwarding in these networks. Greedy forwarding is topology-oblivious. Nevertheless, greedy packets find their destinations with 100% probability following almost optimal shortest paths. This remarkable efficiency sustains even in highly dynamic networks. Our findings suggest that forwarding information through complex networks, such as the Internet, is possible without the overhead of existing routing protocols, and may also find practical applications in overlay networks for tasks such as application-level routing, information sharing, and data distribution. Fragkiskos Papadopoulos, Dmitri V. Krioukov, Marián Boguñá, Amin Vahdat |
INFOCOM | 4 |
| 2010 | Operator and radio resource sharing in multi-carrier environmentsabstractToday's mobile networks prevent users from freely accessing all available networks. Instead, seamless network composition could present a win-win situation for both users and operators. Users can gain better quality of service with more resources to choose from, while each individual operator can provision lesser bandwidth since resources can be shared during times of peak demand. In this paper, we analyze the benefits of operator cooperation using real trace data of cellular data access. We leverage the difference in burstiness at small timescales across network providers to shed the peak usage of one operator on to another. Our results show that even when an operator provisions network capacity below the peak load, cooperation with other network providers can help maintain quality of service for most sessions. In addition, we investigate the performance delivered by various kinds of cellular data cards. Our results confirm that WiFi 802.11b/g consistently delivers superior performance compared to 3G. It will take the next generation 4G technologies such as LTE to deliver end-user performance comparable to widely-deployed 802.11 networks. Pongsakorn Teeraparpwong, Per Johansson, Harsha V. Madhyastha, Amin Vahdat |
NOMS | 4 |
| 2010 | Hedera: Dynamic Flow Scheduling for Data Center Networks
Mohammad Al-Fares, Sivasankar Radhakrishnan, Barath Raghavan, Nelson Huang, Amin Vahdat |
NSDI | 5 |
| 2010 | Helios: a hybrid electrical/optical switch architecture for modular data centersabstractThe basic building block of ever larger data centers has shifted from a rack to a modular container with hundreds or even thousands of servers. Delivering scalable bandwidth among such containers is a challenge. A number of recent efforts promise full bisection bandwidth between all servers, though with significant cost, complexity, and power consumption. We present Helios, a hybrid electrical/optical switch architecture that can deliver significant reductions in the number of switching elements, cabling, cost, and power consumption relative to recently proposed data center network architectures. We explore architectural trade offs and challenges associated with realizing these benefits through the evaluation of a fully functional Helios prototype. Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, Amin Vahdat |
SIGCOMM | 8 |
| 2010 | Neon: system support for derived data managementabstractModern organizations face increasingly complex information management requirements. A combination of commercial needs, legal liability and regulatory imperatives has created a patchwork of mandated policies. Among these, personally identifying customer records must be carefully access-controlled, sensitive files must be encrypted on mobile computers to guard against physical theft, and intellectual property must be protected from both exposure and "poisoning." However, enforcing such policies can be quite difficult in practice since users routinely share data over networks and derive new files from these inputs--incidentally laundering any policy restrictions. In this paper, we describe a virtual machine monitor system called Neon that transparently labels derived data using byte-level "tints" and tracks these labels end to end across commodity applications, operating systems and networks. Our goal with Neon is to explore the viability and utility of transparent information flow tracking within conventional networked systems when used in the manner in which they were intended. We demonstrate that this mechanism allows the enforcement of a variety of data management policies, including data-dependent confinement, mandatory I/O encryption, and intellectual property management. Qing Zhang 0012, John McCullough, Justin Ma, Nabil Schear, Michael Vrable, Amin Vahdat, Alex C. Snoeren, Geoffrey M. Voelker, Stefan Savage |
VEE | 6 |
| 2009 | Live Debugging of Distributed Systems
Darren Dao, Jeannie R. Albrecht, Chip Killian, Amin Vahdat |
CC | 4 |
| 2009 | A Platform for Content-based Partial Replication
Venugopalan Ramasubramanian, Thomas L. Rodeheffer, Douglas B. Terry, Meg Walraed-Sullivan, Ted Wobber, Catherine C. Marshall, Amin Vahdat |
NSDI | 7 |
| 2009 | Application Management and Visualization with PlushabstractDeploying, running, and maintaining applications running on a distributed set of resources is a challenging task. Software developers often spend a significant amount of time dealing with the complexities associated with software configuration and management in these environments. Distributed application management systems are designed to automate the process, and to ultimately help developers cope with the common problems that arise during the design, implementation, and evaluation of distributed systems. In this talk, we highlight the key features of Plush, an application management system for PlanetLab and ModelNet, and describe how Plush simplifies peer-to-peer system visualization and evaluation. Jeannie R. Albrecht, Ryan Braud, Alex C. Snoeren, Amin Vahdat |
Peer-to-Peer Computing | 4 |
| 2009 | Building Distributed Systems Using MaceabstractMace, MaceMC andMacePCwork together to make it easier to build correct, high performance distributed systems implementations. Mace developers find that it now takes them a fraction of the time previously needed to go from design to implementation of a new distributed system. Together with ModelNet and Plush, the whole toolkit is among the best in the world for implementing, testing, evaluating, and deploying distributed and peer-to-peer systems. Mace represents six years of development work and has been publicly available for five years. In addition to the Mace research contributions, the Mace distribution also represents many of the best-quality publicly-available implementations of the included services, and by itself represents a practical contribution that users worldwide recognize and utilize. Chip Killian, James W. Anderson, Ryan Braud, Ranjit Jhala, Amin Vahdat |
Peer-to-Peer Computing | 5 |
| 2009 | ModelNet: Towards a DataCenter Emulation EnvironmentabstractModelNet is a network emulator designed for repeatable, large-scale experimentation with real networked systems. This talk introduces the ideas behind ModelNet that have made it a successful experimental platform. Beyond these core concepts, the talk highlights the latest additions to our methodology to test the next generation of network protocols and applications. Many of these developments address the datacenter compute environment: high-capacity networks, sophisticated infrastructure software (storage and virtualization), and complex network load. While these efforts significantly extend ModelNet's capabilities, there remain a number of open challenges, including incorporating new performance objectives (energy) and multicore architectures. Kashi Venkatesh Vishwanath, Amin Vahdat, Ken Yocum, Diwaker Gupta |
Peer-to-Peer Computing | 2 |
| 2009 | Evaluating the impact of inaccurate information in utility-based schedulingabstractProponents of utility-based scheduling policies have shown the potential for a 100--1400% increase in value-delivered to users when used in lieu of traditional approaches such as FCFS, backfill or priority queues. However, perhaps due to concerns about their potential fragility, these policies are rarely used in practice. We present an evaluation of a utility-based scheduling policy based upon real workload data from both an auction-based resource infrastructure, and a supercomputing cluster. We model potential sources of imperfect operating conditions for a utility-based policy: user uncertainty and wealth inequity. Through simulation, we find that while the value delivered by a utility-based policy can degrade to half that of traditional approaches in the worst case, the policy we study provides 20--100% improvement under realistic operating conditions. We conclude that future efforts in designing utility-based allocation mechanisms and policies must explicitly consider the fidelity of elicited job value information from users. Alvin AuYoung, Amin Vahdat, Alex C. Snoeren |
SC | 2 |
| 2009 | PortLand: a scalable fault-tolerant layer 2 data center network fabricabstractThis paper considers the requirements for a scalable, easily manageable, fault-tolerant, and efficient data center network fabric. Trends in multi-core processors, end-host virtualization, and commodities of scale are pointing to future single-site data centers with millions of virtual end points. Existing layer 2 and layer 3 network protocols face some combination of limitations in such a setting: lack of scalability, difficult management, inflexible communication, or limited support for virtual machine migration. To some extent, these limitations may be inherent for Ethernet/IP style protocols when trying to support arbitrary topologies. We observe that data center networks are often managed as a single logical network fabric with a known baseline topology and growth model. We leverage this observation in the design and implementation of PortLand, a scalable, fault tolerant layer 2 routing and forwarding protocol for data center environments. Through our implementation and evaluation, we show that PortLand holds promise for supporting a ``plug-and-play" large-scale, data center network. Radhika Niranjan Mysore, Andreas Pamboris, Nathan Farrington, Nelson Huang, Pardis Miri, Sivasankar Radhakrishnan, Vikram Subramanya, Amin Vahdat |
SIGCOMM | 8 |
| 2009 | Swing: realistic and responsive network traffic generation
Kashi Venkatesh Vishwanath, Amin Vahdat |
IEEE/ACM Trans. Netw. | 2 |
| 2008 | DieCast: Testing Distributed Systems with an Accurate Scale Model
Diwaker Gupta, Kashi Venkatesh Vishwanath, Amin Vahdat |
NSDI | 3 |
| 2008 | Difference Engine: Harnessing Memory Redundancy in Virtual Machines
Diwaker Gupta, Michael Vrable, Stefan Savage, Alex C. Snoeren, George Varghese, Geoffrey M. Voelker, Amin Vahdat |
OSDI | 8 |
| 2008 | A scalable, commodity data center network architectureabstractToday's data centers may contain tens of thousands of computers with significant aggregate bandwidth requirements. The network architecture typically consists of a tree of routing and switching elements with progressively more specialized and expensive equipment moving up the network hierarchy. Unfortunately, even when deploying the highest-end IP switches/routers, resulting topologies may only support 50% of the aggregate bandwidth available at the edge of the network, while still incurring tremendous cost. Non-uniform bandwidth among data center nodes complicates application design and limits overall system performance. Mohammad Al-Fares, Alexander Loukissas, Amin Vahdat |
SIGCOMM | 3 |
| 2008 | Evaluating Distributed Systems: Does Background Traffic Matter?
Kashi Venkatesh Vishwanath, Amin Vahdat |
USENIX ATC | 2 |
| 2008 | High-bandwidth data dissemination for large-scale distributed systemsabstractThis article focuses on the multireceiver data dissemination problem. Initially, IP multicast formed the basis for efficiently supporting such distribution. More recently, overlay networks have emerged to support point-to-multipoint communication. Both techniques focus on constructing trees rooted at the source to distribute content among all interested receivers. We argue, however, that trees have two fundamental limitations for data dissemination. First, since all data comes from a single parent, participants must often continuously probe in search of a parent with an acceptable level of bandwidth. Second, due to packet losses and failures, available bandwidth is monotonically decreasing down the tree. To address these limitations, we present Bullet, a data dissemination mesh that takes advantage of the computational and storage capabilities of end hosts to create a distribution structure where a node receives data in parallel from multiple peers. For the mesh to deliver improved bandwidth and reliability, we need to solve several key problems: (i) disseminating disjoint data over the mesh, (ii) locating missing content, (iii) finding who to peer with (peering strategy), (iv) retrieving data at the right rate from all peers (flow control), and (v) recovering from failures and adapting to dynamically changing network conditions. Additionally, the system should be self-adjusting and should have few user-adjustable parameter settings. We describe our approach to addressing all of these problems in a working, deployed system across the Internet. Bullet outperforms state-of-the-art systems, including BitTorrent, by 25-70% and exhibits strong performance and reliability in a range of deployment settings. In addition, we find that, relative to tree-based solutions, Bullet reduces the need to perform expensive bandwidth probing. Dejan Kostic, Alex C. Snoeren, Amin Vahdat, Ryan Braud, Chip Killian, James W. Anderson, Jeannie R. Albrecht, Adolfo Rodriguez, Erik Vandekieft |
ACM Trans. Comput. Syst. | 3 |
| 2008 | Design and implementation trade-offs for wide-area resource discoveryabstractWe describe the design and implementation of SWORD, a scalable resource discovery service for wide-area distributed systems. In contrast to previous systems, SWORD allows users to describe desired resources as a topology of interconnected groups with required intragroup, intergroup, and per-node characteristics, along with the utility that the application derives from specified ranges of metric values. This design gives users the flexibility to find geographically distributed resources for applications that are sensitive to both node and network characteristics, and allows the system to rank acceptable configurations based on their quality for that application. Rather than evaluating a single implementation of SWORD, we explore a variety of architectural designs that deliver the required functionality in a scalable and highly available manner. We discuss the trade-offs of using a centralized architecture as compared to a fully decentralized design to perform wide-area resource discovery. To summarize our results, we found that a centralized architecture based on 4-node server cluster sites at network-peering facilities outperforms a decentralized DHT-based resource discovery infrastructure with respect to query latency for all but the smallest number of sites. However, although a centralized architecture shows significant promise in stable environments, we find that our decentralized implementation has acceptable performance and also benefits from the DHT's self-healing properties in more volatile environments. We evaluate the advantages and disadvantages of centralized and distributed resource discovery architectures on 1000 hosts in emulation and on approximately 200 PlanetLab nodes spread across the Internet. Jeannie R. Albrecht, David Oppenheimer, Amin Vahdat, David A. Patterson 0001 |
ACM Trans. Internet Techn. | 3 |
| 2007 | A Performance Analysis of Indirect RoutingabstractIndirect routing involves sending messages between Internet end nodes through a specified intermediate node to effect a different end-to-end route than the default "direct" route. Indirect routing offers the potential for significant improvements in throughput performance by exploiting the Internet's richness in throughput diversity. We present a performance analysis of indirect routing, showing that it can result in a 33-49% increase in average throughput performance while incurring low overhead. Joshua M. Opos, Sriram Ramabhadran, Andrew Terry, Joseph Pasquale, Alex C. Snoeren, Amin Vahdat |
IPDPS | 6 |
| 2007 | Remote Control: Distributed Application Configuration, Management, and Visualization with Plush
Jeannie R. Albrecht, Ryan Braud, Darren Dao, Nikolay Topilski, Christopher Tuttle, Alex C. Snoeren, Amin Vahdat |
LISA | 7 |
| 2007 | Usher: An Extensible Framework for Managing Clusters of Virtual Machines
Marvin McNett, Diwaker Gupta, Amin Vahdat, Geoffrey M. Voelker |
LISA | 3 |
| 2007 | Life, Death, and the Critical Transition: Finding Liveness Bugs in Systems Code (Awarded Best Paper)
Chip Killian, James W. Anderson, Ranjit Jhala, Amin Vahdat |
NSDI | 4 |
| 2007 | Mace: language support for building distributed systemsabstractBuilding distributed systems is particularly difficult because of the asynchronous, heterogeneous, and failure-prone environment where these systemsmust run. Tools for building distributed systems must strike a compromise between reducing programmer effort and increasing system efficiency. We present Mace, a C++ language extension and source-to-source compiler that translates a concise but expressive distributed system specification into a C++ implementation. Mace overcomes the limitations of low-level languages by providing a unified framework for networking and event handling, and the limitations of high-level languages by allowing programmers to write program components in a controlled and structured manner in C++. By imposing structure and restrictions on how applications can be written, Mace supports debugging at a higher level, including support for efficient model checking and causal-path debugging. Because Mace programs compile to C++, programmers can use existing C++ tools, including optimizers, profilers, and debuggers to analyze their systems. Chip Killian, James W. Anderson, Ryan Braud, Ranjit Jhala, Amin Vahdat |
PLDI | 5 |
| 2007 | Orbis: rescaling degree correlations to generate annotated internet topologiesabstractResearchers involved in designing network services and protocols rely on results from simulation and emulation environments to understand their application performance and scalability. To better understand the behavior of these applications and predict their performance when deployed on the actual Internet, the generated topologies must closely match real network characteristics, not just in terms of graph structure (node interconnectivity) but also with respect to various node and link annotations. Relevant annotations include link latencies, AS membership and whether a router is a peering or internal router. Finally, it should be possible to rescale a given topology to a variety of sizes while still maintaining its essential characteristics. Priya Mahadevan, Calvin Hubble, Dmitri V. Krioukov, Bradley Huffaker, Amin Vahdat |
SIGCOMM | 5 |
| 2007 | Modeling and generating realistic streaming media server workloads
Wenting Tang, Yun Fu 0003, Ludmila Cherkasova, Amin Vahdat |
Comput. Networks | 4 |
| 2006 | Distributed Network Querying with Bounded Approximate Caching
Badrish Chandramouli, Jun Yang 0001, Amin Vahdat |
DASFAA | 3 |
| 2006 | Glavlit: Preventing Exfiltration at Wire Speed
Nabil Schear, Carmelo Kintana, Qing Zhang 0012, Amin Vahdat |
HotNets | 4 |
| 2006 | Enforcing Performance Isolation Across Virtual Machines in Xen
Diwaker Gupta, Ludmila Cherkasova, Rob Gardner, Amin Vahdat |
Middleware | 4 |
| 2006 | To Infinity and Beyond: Time-Warped Network Emulation
Diwaker Gupta, Ken Yocum, Marvin McNett, Alex C. Snoeren, Amin Vahdat, Geoffrey M. Voelker |
NSDI | 5 |
| 2006 | Pip: Detecting the Unexpected in Distributed Systems
Patrick Reynolds, Chip Killian, Janet L. Wiener, Jeffrey C. Mogul, Mehul A. Shah, Amin Vahdat |
NSDI | 6 |
| 2006 | Systematic topology analysis and generation using degree correlationsabstractResearchers have proposed a variety of metrics to measure important graph properties, for instance, in social, biological, and computer networks. Values for a particular graph metric may capture a graph's resilience to failure or its routing efficiency. Knowledge of appropriate metric values may influence the engineering of future topologies, repair strategies in the face of failure, and understanding of fundamental properties of existing networks. Unfortunately, there are typically no algorithms to generate graphs matching one or more proposed metrics and there is little understanding of the relationships among individual metrics or their applicability to different settings. We present a new, systematic approach for analyzing network topologies. We first introduce the dK-series of probability distributions specifying all degree correlations within d-sized subgraphs of a given graph G. Increasing values of d capture progressively more properties of G at the cost of more complex representation of the probability distribution. Using this series, we can quantitatively measure the distance between two graphs and construct random graphs that accurately reproduce virtually all metrics proposed in the literature. The nature of the dK-series implies that it will also capture any future metrics that may be proposed. Using our approach, we construct graphs for d=0, 1, 2, 3 and demonstrate that these graphs reproduce, with increasing accuracy, important properties of measured and modeled Internet topologies. We find that the d=2 case is sufficient for most practical purposes, while d=3 essentially reconstructs the Internet AS-and router-level topologies exactly. We hope that a systematic method to analyze and synthesize topologies offers a significant improvement to the set of tools available to network topology and protocol researchers. Priya Mahadevan, Dmitri V. Krioukov, Kevin R. Fall, Amin Vahdat |
SIGCOMM | 4 |
| 2006 | Realistic and responsive network traffic generationabstractThis paper presents Swing, a closed-loop, network-responsive traffic generator that accurately captures the packet interactions of a range of applications using a simple structural model. Starting from observed traffic at a single point in the network, Swing automatically extracts distributions for user, application, and network behavior. It then generates live traffic corresponding to the underlying models in a network emulation environment running commodity network protocol stacks. We find that the generated traces are statistically similar to the original traces. Further, to the best of our knowledge, we are the first to reproduce burstiness in traffic across a range of timescales using a model applicable to a variety of network settings. An initial sensitivity analysis reveals the importance of capturing and recreating user, application, and network characteristics to accurately reproduce such burstiness. Finally, we explore Swing's ability to vary user characteristics, application properties, and wide-area network conditions to project traffic characteristics into alternate scenarios. Kashi Venkatesh Vishwanath, Amin Vahdat |
SIGCOMM | 2 |
| 2006 | Loose Synchronization for Large-Scale Networked Systems
Jeannie R. Albrecht, Christopher Tuttle, Alex C. Snoeren, Amin Vahdat |
USENIX ATC, General Track | 4 |
| 2006 | Service Placement in a Shared Wide-Area Platform
David Oppenheimer, Brent N. Chun, David A. Patterson 0001, Alex C. Snoeren, Amin Vahdat |
USENIX ATC, General Track | 5 |
| 2006 | WAP5: black-box performance debugging for wide-area systemsabstractWide-area distributed applications are challenging to debug, optimize, and maintain. We present Wide-Area Project 5 (WAP5), which aims to make these tasks easier by exposing the causal structure of communication within an application and by exposing delays that imply bottlenecks. These bottlenecks might not otherwise be obvious, with or without the application's source code. Previous research projects have presented algorithms to reconstruct application structure and the corresponding timing information from black-box message traces of local-area systems. In this paper we present (1) a new algorithm for reconstructing application structure in both local- and wide-area distributed systems, (2) an infrastructure for gathering application traces in PlanetLab, and (3) our experiences tracing and analyzing three systems: CoDeeN and Coral, two content-distribution networks in PlanetLab; and Slurpee, an enterprise-scale incident-monitoring system. Patrick Reynolds, Janet L. Wiener, Jeffrey C. Mogul, Marcos K. Aguilera, Amin Vahdat |
WWW | 5 |
| 2006 | A unified benchmarking and model-based framework for building QoS-aware streaming media services
Ludmila Cherkasova, Wenting Tang, Amin Vahdat |
Multim. Syst. | 3 |
| 2006 | The costs and limits of availability for replicated servicesabstractAs raw system performance continues to improve at exponential rates, the utility of many services is increasingly limited by availability rather than performance. A key approach to improving availability involves replicating the service across multiple, wide-area sites. However, replication introduces well-known trade-offs between service consistency and availability. Thus, this article explores the benefits of dynamically trading consistency for availability using a continuous consistency model . In this model, applications specify a maximum deviation from strong consistency on a per-replica basis. In this article, we: i) evaluate the availability of a prototype replication system running across the Internet as a function of consistency level, consistency protocol, and failure characteristics, ii) demonstrate that simple optimizations to existing consistency protocols result in significant availability improvements (more than an order of magnitude in some scenarios), iii) use our experience with these optimizations to prove tight upper bound on the availability of services, and iv) show that maximizing availability typically entails remaining as close to strong consistency as possible during times of good connectivity, resulting in a communication versus availability trade-off. Amin Vahdat |
ACM Trans. Comput. Syst. | 2 |
| 2005 | Why Markets Could (But Don't Currently) Solve Resource Allocation Problems in Systems
Jeffrey Shneidman, Chaki Ng, David C. Parkes, Alvin AuYoung, Alex C. Snoeren, Amin Vahdat, Brent N. Chun |
HotOS | 6 |
| 2005 | Distributed application management using PlushabstractRecent computing trends have shown an increase in the demand for large-scale, distributed, federated computing environments. Two of the more popular environments that have emerged are the grid and PlanetLab. At a high level, these systems are similar in many ways; both are comprised of a set of heterogeneous interconnected machines that allows secure resource sharing for a variety of different users and applications. However, at a lower level, the systems are very distinct in the sense that they were designed to solve different types of problems, and therefore have fundamental differences that make it difficult to develop and deploy applications on both platforms. As a result, application designers and researchers create software that runs on either the grid or PlanetLab, but not both. We propose to solve this problem by describing a common abstraction for both PlanetLab and grid applications. Further, we present Plush - a tool that implements the distributed application abstraction by providing a pluggable and extensible infrastructure allowing users to customize their environment for running experiments on both PlanetLab and the grid. Jeannie R. Albrecht, Christopher Tuttle, Alex C. Snoeren, Amin Vahdat |
HPDC | 4 |
| 2005 | Design and implementation tradeoffs for wide-area resource discoveryabstractThis paper describes the design and implementation of SWORD, a scalable resource discovery service for wide-area distributed systems. In contrast to previous systems, SWORD allows users to describe desired resources as a topology of interconnected groups with required intragroup, intergroup, and per-node characteristics, along with the utility that the application derives from various ranges of values of those characteristics. This design gives users the flexibility to find geographically distributed resources for applications that are sensitive to both node and network characteristics, and allows the system to rank acceptable configurations based on their quality for that application. We explore a variety of architectures to deliver SWORD's functionality in a scalable and highly-available manner. A 1000-node ModelNet evaluation using a workload of measurements collected from PlanetLab shows that an architecture based on 4-node server cluster sites at network peering facilities outperforms a decentralized DHT-based resource discovery infrastructure for all but the smallest number of sites. While such a centralized architecture shows significant promise, we find that our decentralized implementation, both in emulation and running continuously on over 200 PlanetLab nodes, performs well while benefiting from the DHT's self-healing properties. David Oppenheimer, Jeannie R. Albrecht, David A. Patterson 0001, Amin Vahdat |
HPDC | 4 |
| 2005 | Designing incentives for peer-to-peer routingabstractIn a peer-to-peer network, nodes are typically required to route packets for each other. This leads to a problem of "free-loaders", nodes that use the network but refuse to route other nodes' packets. In this paper we study ways of designing incentives to discourage free-loading. We model the interactions between nodes as a "random matching game", and describe a simple reputation system that provides incentives for good behavior. Under certain assumptions, we obtain a stable subgame-perfect equilibrium. We use simulations to investigate the robustness of this scheme in the presence of noise and malicious nodes, and we examine some of the design trade-offs. We also evaluate some possible adversarial strategies, and discuss how our results might apply to real peer-to-peer systems. Alberto Blanc, Yi-Kai Liu 0001, Amin Vahdat |
INFOCOM | 3 |
| 2005 | Brief announcement: the overlay network content distribution problemabstractMany overlay multicast protocols have been designed and deployed across the Internet to support content distribution. To our knowledge, however, none have provided a rigorous analysis of the problem or the effectiveness of their proposed solutions. We define the Overlay Network Content Distribution (OCD) problem to allow such analyses. Chip Killian, Michael Vrable, Alex C. Snoeren, Amin Vahdat, Joseph Pasquale |
PODC | 4 |
| 2005 | To infinity and beyond: time warped network emulationabstractThis work explores the viability and benefits of time dilation - providing the illusion to an operating system and its applications that time is passing at a rate different from real time. For example, we may wish to convince a system that for every 10 seconds of wall clock time, only one second of time passes in the host's dilated time frame. This enables external stimuli to appear to take place at higher rates than would be physically possible. For example, a host dilated by a factor of 10 receiving data from a network interface at a real rate of 1-Gbps believes it is receiving data at 10-Gbps. Diwaker Gupta, Ken Yocum, Marvin McNett, Alex C. Snoeren, Amin Vahdat, Geoffrey M. Voelker |
SOSP | 5 |
| 2005 | Service placement in shared wide-area platformsabstractFederated geographically-distributed computing platforms such as PlanetLab [1] and the Grid [2, 3] have recently become popular for evaluating and deploying network services and scientific computations. As the size, reach, and user population of such infrastructures grow, resource discovery and resource selection become increasingly important. Although a number of resource discovery and allocation services have been built, there is little data on the utilization of the distributed computing platforms they target. Yet the design and efficacy of such services depends on the characteristics of the target platform. David Oppenheimer, Brent N. Chun, David A. Patterson 0001, Alex C. Snoeren, Amin Vahdat |
SOSP | 5 |
| 2005 | Experiences with Pip: finding unexpected behavior in distributed systemsabstractBugs in complex distributed systems are often hard to find. Many bugs reflect discrepancies between a system's behavior and the programmer's assumptions about that behavior. Differences may be in correctness, in performance characteristics, or both. Our debugging framework, Pip, compares actual behavior with expected behavior and visualizes both. Pip consists of two tools to help reconcile assumptions and actual behavior: an automatic expectations checker and an interactive behavior-explorer GUI. Patrick Reynolds, Janet L. Wiener, Jeffrey C. Mogul, Mehul A. Shah, Chip Killian, Amin Vahdat |
SOSP | 6 |
| 2005 | Maintaining High-Bandwidth Under Dynamic Network Conditions
Dejan Kostic, Ryan Braud, Chip Killian, Erik Vandekieft, James W. Anderson, Alex C. Snoeren, Amin Vahdat |
USENIX ATC, General Track | 7 |
| 2005 | Consistent and automatic replica regenerationabstractReducing management costs and improving the availability of large-scale distributed systems require automatic replica regeneration , that is, creating new replicas in response to replica failures. A major challenge to regeneration is maintaining consistency when the replica group changes. Doing so is particularly difficult across the wide area where failure detection is complicated by network congestion and node overload.In this context, this article presents Om, the first read/write peer-to-peer, wide-area storage system that achieves high availability and manageability through online automatic regeneration while still preserving consistency guarantees. We achieve these properties through the following techniques. First, by utilizing the limited view divergence property in today's Internet and by adopting the witness model , Om is able to regenerate from any single replica, rather than requiring a majority quorum, at the cost of a small (10 −6 in our experiments) probability of violating consistency during each regeneration. As a result, Om can deliver high availability with a small number of replicas, while traditional designs would significantly increase the number of replicas. Next, we distinguish failure-free reconfigurations from failure-induced ones, enabling common reconfigurations to proceed with a single round of communication. Finally, we use a lease graph among the replicas and a two-phase write protocol to optimize for reads, so that reads in Om can be processed by any single replica. Experiments on PlanetLab show that consistent regeneration in Om completes in approximately 20 seconds. Amin Vahdat |
ACM Trans. Storage | 2 |
| 2004 | Scalability in Adaptive Multi-Metric OverlaysabstractIncreasing application requirements have placed heavy emphasis on building overlay networks to efficiently deliver data to multiple receivers. A key performance challenge is simultaneously achieving adaptivity to changing network conditions and scalability to large numbers of users. In addition, most current algorithms focus on a single performance metric, such as delay or bandwidth, particular to individual application requirements. We introduce a two-fold approach for creating robust, high-performance overlays called adaptive multimetric overlays (AMMO). First, AMMO uses an adaptive, highly-parallel, and metric-independent protocol, TreeMaint, to build and maintain overlay trees. Second, AMMO provides a mechanism for comparing overlay edges along specified application performance goals to guide TreeMaint transformations. We have used AMMO to implement and evaluate a single-metric (bandwidth-optimized) tree similar to Overcast and a two-metric (delay-constrained, cost-optimized) overlay. Adolfo Rodriguez, Dejan Kostic, Amin Vahdat |
ICDCS | 3 |
| 2004 | MACEDON: Methodology for Automatically Creating, Evaluating, and Designing Overlay Networks
Adolfo Rodriguez, Chip Killian, Sooraj Bhat, Dejan Kostic, Amin Vahdat |
NSDI | 5 |
| 2004 | Consistent and Automatic Replica Regeneration
Amin Vahdat |
NSDI | 2 |
| 2003 | Efficient Peer-to-Peer Keyword Searching
Patrick Reynolds, Amin Vahdat |
Middleware | 2 |
| 2003 | MediSyn: a synthetic streaming media service workload generatorabstractCurrently, Internet hosting centers and content distribution networks leverage statistical multiplexing to meet the performance requirements of a number of competing hosted network services. Developing efficient resource allocation mechanisms for such services requires an understanding of both the short-term and long-term behavior of client access patterns to these competing services. At the same time, streaming media services are becoming increasingly popular, presenting new challenges for designers of shared hosting services. These new challenges result from fundamentally new characteristics of streaming media relative to traditional web objects, principally different client access patterns and significantly larger computational and bandwidth overhead associated with a streaming request. To understand the characteristics of these new workloads we use two long-term traces of streaming media services to develop MediSyn, a publicly available streaming media workload generator. In summary, this paper makes the following contributions: i) we model the long-term behavior of network services capturing the process of file introduction and changing file popularity, ii) we present a novel generalized Zipf-like distribution that captures recently-observed popularity of both web objects and streaming media not captured by existing Zipf-like distributions, and iii) we capture a number of characteristics unique to streaming media services, including file duration, encoding bit rate, session duration and non-stationary popularity of media accesses. Wenting Tang, Yun Fu 0003, Ludmila Cherkasova, Amin Vahdat |
NOSSDAV | 4 |
| 2003 | SHARP: an architecture for secure resource peeringabstractThis paper presents Sharp, a framework for secure distributed resource management in an Internet-scale computing infrastructure. The cornerstone of Sharp is a construct to represent cryptographically protected resource claims ---promises or rights to control resources for designated time intervals---together with secure mechanisms to subdivide and delegate claims across a network of resource managers. These mechanisms enable flexible resource peering : sites may trade their resources with peering partners or contribute them to a federation according to local policies. A separation of claims into tickets and leases allows coordinated resource management across the system while preserving site autonomy and local control over resources. Sharp also introduces mechanisms for controlled, accountable oversubscription of resource claims as a fundamental tool for dependable, efficient resource management. We present experimental results from a Sharp prototype for PlanetLab, and illustrate its use with a decentralized barter economy for global PlanetLab resources. The results demonstrate the power and practicality of the architecture, and the effectiveness of oversubscription for protecting resource availability in the presence of failures. Yun Fu 0003, Jeffrey S. Chase, Brent N. Chun, Stephen Schwab, Amin Vahdat |
SOSP | 5 |
| 2003 | Bullet: high bandwidth data dissemination using an overlay meshabstractIn recent years, overlay networks have become an effective alternative to IP multicast for efficient point to multipoint communication across the Internet. Typically, nodes self-organize with the goal of forming an efficient overlay tree, one that meets performance targets without placing undue burden on the underlying network. In this paper, we target high-bandwidth data distribution from a single source to a large number of receivers. Applications include large-file transfers and real-time multimedia streaming. For these applications, we argue that an overlay mesh, rather than a tree, can deliver fundamentally higher bandwidth and reliability relative to typical tree structures. This paper presents Bullet, a scalable and distributed algorithm that enables nodes spread across the Internet to self-organize into a high bandwidth overlay mesh. We construct Bullet around the insight that data should be distributed in a disjoint manner to strategic points in the network. Individual Bullet receivers are then responsible for locating and retrieving the data from multiple points in parallel.Key contributions of this work include: i) an algorithm that sends data to different points in the overlay such that any data object is equally likely to appear at any node, ii) a scalable and decentralized algorithm that allows nodes to locate and recover missing data items, and iii) a complete implementation and evaluation of Bullet running across the Internet and in a large-scale emulation environment reveals up to a factor two bandwidth improvements under a variety of circumstances. In addition, we find that, relative to tree-based solutions, Bullet reduces the need to perform expensive bandwidth probing. In a tree, it is critical that a node's parent delivers a high rate of application data to each child. In Bullet however, nodes simultaneously receive data from multiple sources in parallel, making it less important to locate any single source capable of sustaining a high transmission rate. Dejan Kostic, Adolfo Rodriguez, Jeannie R. Albrecht, Amin Vahdat |
SOSP | 4 |
| 2003 | Currentcy: A Unifying Abstraction for Expressing Energy Management Policies
Heng Zeng, Carla Schlatter Ellis, Alvin R. Lebeck, Amin Vahdat |
USENIX ATC, General Track | 4 |
| 2003 | Measuring and characterizing end-to-end Internet service performanceabstractFundamental to the design of reliable, high-performance network services is an understanding of the performance characteristics of the service as perceived by the client population as a whole. Understanding and measuring such end-to-end service performance is a challenging task. Current techniques include periodic sampling of service characteristics from strategic locations in the network and instrumenting Web pages with code that reports client-perceived latency back to a performance server. Limitations to these approaches include potentially nonrepresentative access patterns in the first case and determining the location of a performance bottleneck in the second.This paper presents EtE monitor, a novel approach to measuring Web site performance. Our system passively collects packet traces from a server site to determine service performance characteristics. We introduce a two-pass heuristic and a statistical filtering mechanism to accurately reconstruct different client page accesses and to measure performance characteristics integrated across all client accesses. Relative to existing approaches, EtE monitor offers the following benefits: i) a latency breakdown between the network and server overhead of retrieving a Web page, ii) longitudinal information for all client accesses, not just the subset probed by a third party, iii) characteristics of accesses that are aborted by clients, iv) an understanding of the performance breakdown of accesses to dynamic, multitiered services, and v) quantification of the benefits of network and browser caches on server performance. Our initial implementation and performance analysis across three different commercial Web sites confirm the utility of our approach. Ludmila Cherkasova, Yun Fu 0003, Wenting Tang, Amin Vahdat |
ACM Trans. Internet Techn. | 4 |
| 2002 | ECOSystem: managing energy as a first class operating system resourceabstractEnergy consumption has recently been widely recognized as a major challenge of computer systems design. This paper explores how to support energy as a first-class operating system resource. Energy, because of its global system nature, presents challenges beyond those of conventional resource management. To meet these challenges we propose the Currentcy Model that unifies energy accounting over diverse hardware components and enables fair allocation of available energy among applications. Our particular goal is to extend battery lifetime by limiting the average discharge rate and to share this limited resource among competing task according to user preferences. To demonstrate how our framework supports explicit control over the battery resource we implemented ECOSystem, a modified Linux, that incorporates our currentcy model. Experimental results show that ECOSystem accurately accounts for the energy consumed by asynchronous device operation, can achieve a target battery lifetime, and proportionally shares the limited energy resource among competing tasks. Heng Zeng, Carla Schlatter Ellis, Alvin R. Lebeck, Amin Vahdat |
ASPLOS | 4 |
| 2002 | Scalability and Accuracy in a Large-Scale Network Emulator
Amin Vahdat, Ken Yocum, Kevin Walsh 0001, Priya Mahadevan, Dejan Kostic, Jeffrey S. Chase |
OSDI | 1 |
| 2002 | Minimal replication cost for availabilityabstractToday, the utility of many replicated Internet services is limited by availability rather than raw performance. To better understand the effects of replica placement on availability, we propose the problem of minimal replication cost for availability. Let replication cost be the cost associated with replica deployment, dynamic replica creation and teardown at n candidate locations. Given client access patterns, replica failure patterns, network partition patterns, a required consistency level and a target level of availability, the minimal replication cost is the lower bound on a system's replication cost. Solving this problem also answers the dual question of optimal availability given a constraint on replication cost.We design the first algorithm we are aware of to solve the problem, through reduction to integer linear programming and enumeration of pruned serialization orders. Using practical faultloads and workloads, we demonstrate that the exponential complexity of our algorithm is tractable for practical problems with hundreds of candidate locations. The lower bound computed by our algorithm is tight, but the tightness can be sacrificed by a proposed optimization for large problems. We also show that with low replica creation and teardown costs, the bound is close to tight in practical problems even with the optimization. Amin Vahdat |
PODC | 2 |
| 2002 | Application-specific Network Management for Energy-Aware Streaming of Popular Multimedia Formats
Surendar Chandra, Amin Vahdat |
USENIX ATC, General Track | 2 |
| 2002 | EtE: Passive End-to-End Internet Service Performance Monitoring
Yun Fu 0003, Amin Vahdat, Ludmila Cherkasova, Wenting Tang |
USENIX ATC, General Track | 2 |
| 2002 | The Trickle-Down Effect: Web Caching and Server Request Distribution
Ronald P. Doyle, Jeffrey S. Chase, Syam Gadde, Amin Vahdat |
Comput. Commun. | 4 |
| 2002 | Interposed request routing for scalable network storageabstractThis paper explores interposed request routing in Slice, a new storage system architecture for high-speed networks incorporating network-attached block storage. Slice interposes a request switching filter---called a μproxy---along each client's network path to the storage service (e.g., in a network adapter or switch). The μproxy intercepts request traffic and distributes it across a server ensemble. We propose request routing schemes for I/O and file service traffic, and explore their effect on service structure. The Slice prototype uses a packet filter μproxy to virtualize the standard Network File System (NFS) protocol, presenting to NFS clients a unified shared file volume with scalable bandwidth and capacity. Experimental results from the industry-standard SPECsfs97 workload demonstrate that the architecture enables construction of powerful network-attached storage services by aggregating cost-effective components on a switched Gigabit Ethernet LAN. Darrell C. Anderson, Jeffrey S. Chase, Amin Vahdat |
ACM Trans. Comput. Syst. | 3 |
| 2002 | Design and evaluation of a conit-based continuous consistency model for replicated servicesabstractThe tradeoffs between consistency, performance, and availability are well understood. Traditionally, however, designers of replicated systems have been forced to choose from either strong consistency guarantees or none at all. This paper explores the semantic space between traditional strong and optimistic consistency models for replicated services. We argue that an important class of applications can tolerate relaxed consistency, but benefit from bounding the maximum rate of inconsistent access in an application-specific manner. Thus, we develop a conit-based continuous consistency model to capture the consistency spectrum using three application-independent metrics, numerical error , order error , and staleness . We then present the design and implementation of TACT, a middleware layer that enforces arbitrary consistency bounds among replicas using these metrics. We argue that the TACT consistency model can simultaneously achieve the often conflicting goals of generality and practicality by describing how a broad range of applications can express their consistency semantics using TACT and by demonstrating that application-independent algorithms can efficiently enforce target consistency levels. Finally, we show that three replicated applications running across the Internet demonstrate significant semantic and performance benefits from using our framework. Amin Vahdat |
ACM Trans. Comput. Syst. | 2 |
| 2001 | Anypoint Communication ProtocolabstractSummary form only given. We are developing the Anypoint Communication Protocol (ACP). ACP clients establish connections to abstract services, represented at the network edge by Anypoint intermediaries. The intermediary is an intelligent network switch that acts as an extension of the service; it encapsulates a service-specific policy for distributing requests among servers in the active set for each service. The switch routes incoming requests on each ACP connection to any active server at the discretion of the service routing policy, hence the name "Anypoint". The Anypoint abstraction and ACP protocol enable virtualization using intermediaries for a general class of wide-area network services based on request/response communication over persistent transport connections. Potential applications include scalable IP-based network storage protocols and next-generation Web services. Ken Yocum, Jeffrey S. Chase, Amin Vahdat |
HotOS | 3 |
| 2001 | Combining Generality and Practicality in a Conit-Based Continuous Consistency Model for Wide-Area ReplicationabstractReplication is a key approach to scaling wide-area applications. However, the overhead associated with large-scale replication quickly becomes prohibitive across wide-area networks. One effective approach to addressing this limitation is to allow applications to dynamically trade reduced consistency for increased performance and availability. Although extensive study has been performed on relaxed consistency models in traditional replicated databases, none of the models can simultaneously achieve the following two typically conflicting requirements imposed by wide-area applications: generality (capturing application-specific consistency semantics) and practicality (enabling efficient application-independent consistency protocols to be designed and providing natural ways to express application semantics). We propose a conit-based continuous consistency model designed to simultaneously achieve generality and practicality. Our conit theory provides generality, where application-specific consistency requirements are exported as conits. Practicality is achieved by using a simple, spanning set of metrics for conit consistency and by using a per-write weight specification. We demonstrate the generality of our model through representative wide-area applications and by showing that a number of existing models can be expressed as instances of our model. Our efficient, application-independent consistency protocols and prototype implementation verify its practicality. Amin Vahdat |
ICDCS | 2 |
| 2001 | A chat room assignment for teaching network securityabstractThis paper describes a chat room application suitable for teaching basic network programming and security protocols. A client/server design illustrates the structure of current scalable network services while a multicast version demonstrates the need for efficient simultaneous distribution of network content to multiple receivers (e.g., as required by video broadcasts). The system also includes implementations of two security protocols, one similar to Kerberos and another based on public key encryption. W. Garrett Mitchener, Amin Vahdat |
SIGCSE | 2 |
| 2001 | Managing Energy and Server Resources in Hosting CentresabstractInternet hosting centers serve multiple service sites from a common hardware base. This paper presents the design and implementation of an architecture for resource management in a hosting center operating system, with an emphasis on energy as a driving resource management issue for large server clusters. The goals are to provision server resources for co-hosted services in a way that automatically adapts to offered load, improve the energy efficiency of server clusters by dynamically resizing the active server set, and respond to power supply disruptions or thermal events by degrading service in accordance with negotiated Service Level Agreements (SLAs).Our system is based on an economic approach to managing shared server resources, in which services "bid" for resources as a function of delivered performance. The system continuously monitors load and plans resource allotments by estimating the value of their effects on service performance. A greedy resource allocation algorithm adjusts resource prices to balance supply and demand, allocating resources to their most efficient use. A reconfigurable server switching infrastructure directs request traffic to the servers assigned to each service. Experimental results from a prototype confirm that the system adapts to offered load and resource availability, and can reduce server energy usage by 29% or more for a typical Web workload. Jeffrey S. Chase, Darrell C. Anderson, Prachi N. Thakar, Amin Vahdat, Ronald P. Doyle |
SOSP | 4 |
| 2001 | The Costs and Limits of Availability for Replicated ServicesabstractAs raw system and network performance continues to improve at exponential rates, the utility of many services is increasingly limited by availability rather than performance. A key approach to improving availability involves replicating the service across multiple, wide-area sites. However, replication introduces well-known tradeoffs between service consistency and availability. Thus, this paper explores the benefits of dynamically trading consistency for availability using a continuous consistency model. In this model, applications specify a maximum deviation from strong consistency on a per-replica basis. In this paper, we: i) evaluate availability of a prototype replication system running across the Internet as a function of consistency level, consistency protocol, and failure characteristics, ii) demonstrate that simple optimizations to existing consistency protocols result in significant availability improvements (more than an order of magnitude in some scenarios), iii) use our experience with these optimizations to prove tight upper bounds on the availability of services, and iv) show that maximizing availability typically entails remaining as close to strong consistency as possible during times of good connectivity, resulting in a communication versus availability trade-off. Amin Vahdat |
SOSP | 2 |
| 2000 | Differentiated Multimedia Web Services Using Quality Aware TranscodingabstractThe ability of a Web service to provide low-latency access to its contents is constrained by available network bandwidth. It is important for the service to manage available bandwidth wisely. While providing differentiated quality of service (QoS) is typically enforced through network mechanisms, in this paper we introduce a robust mechanism for managing network resources at the application level. We use transcoding to allow Web servers to customize the size of objects constituting a Web page, and hence the bandwidth consumed by that page, by dynamically varying the size of multimedia objects on a per-client basis. We leverage earlier work on characterizing quality versus size tradeoffs in transcoding JPEG images to dynamically determine the quality and size of the object to transmit. We evaluate the performance benefits of incorporating this information in a series of bandwidth management policies. We develop metrics to measure the performance of our system. We use realistic workloads and access scenarios to drive our system. The principal contribution of this work is the demonstration that it is possible to use informed transcoding techniques to provide differentiated service and to dynamically allocate available bandwidth among different client classes, while delivering a high degree of information content (quality factor) for all clients. Surendar Chandra, Carla Schlatter Ellis, Amin Vahdat |
INFOCOM | 3 |
| 2000 | Interposed Request Routing for Scalable Network Storage
Darrell C. Anderson, Jeffrey S. Chase, Amin Vahdat |
OSDI | 3 |
| 2000 | Design and Evaluation of a Continuous Consistency Model for Replicated Services
Amin Vahdat |
OSDI | 2 |
| 2000 | Efficient Numerical Error Bounding for Replicated Network Services
Amin Vahdat |
VLDB | 2 |
| 2000 | Application-level differentiated multimedia Web services using quality aware transcodingabstractThe ability of a Web service to provide low-latency access to its content is constrained by available network bandwidth. While providing differentiated quality of service (QoS) is typically enforced through network mechanisms, in this paper we introduce a robust mechanism for managing network resources using application-specific characteristics of Web services. We use transcoding to allow Web servers to customize the size of objects constituting a Web page, and hence the bandwidth consumed by that page, by dynamically varying the size of multimedia objects on a per-client basis. We leverage our earlier work on characterizing quality versus size tradeoffs in transcoding JPEG images to supply more information for determining the quality and size of the object to transmit. We evaluate the performance benefits of incorporating this information in a series of bandwidth management policies using realistic workloads and access scenarios to drive our system. The principal contribution of this paper is the demonstration that it is possible to use informed transcoding techniques to provide differentiated service and to dynamically allocate available bandwidth among different client classes, while delivering good quality of information content for all clients. We also show that it is possible to customize multimedia objects to the highly variable network conditions experienced by mobile clients in order to provide acceptable quality and latency depending on the networks used in accessing the service. We show that policies that aggressively transcode the larger images can produce images with quality factor values that closely follow the untranscoded base case while still saving as much as 150 kB. A transcoding policy that has knowledge of the characteristics of the link to the client can avoid as many as 40% of (unnecessary) transcodings. Surendar Chandra, Carla Schlatter Ellis, Amin Vahdat |
IEEE J. Sel. Areas Commun. | 3 |
| 1998 | WebOS: Operating System Services for Wide Area ApplicationsabstractDemonstrates the power of providing a common set of operating system services to wide-area applications, including mechanisms for naming, persistent storage, remote process execution, resource management, authentication and security. On a single machine, application developers can rely on the local operating system to provide these abstractions. In the wide area, however, application developers are forced to build these abstractions themselves or to do without. This ad-hoc approach often results in individual programmers implementing non-optimal solutions, wasting both programmer effort and system resources. To address these problems, we are building a system, WebOS, that provides the basic operating systems services needed to build applications that are geographically distributed, highly available, incrementally scalable and dynamically reconfigurable. Experience with a number of applications developed under WebOS indicates that it simplifies system development and improves resource utilization. In particular, we use WebOS to implement Rent-A-Server to provide dynamic replication of overloaded Web services across the wide area in response to client demands. Amin Vahdat, Thomas E. Anderson, Michael Dahlin, Eshwar Belani, David E. Culler, Paul Eastham, Chad Yoshikawa |
HPDC | 1 |
| 1998 | Transparent Result Caching
Amin Vahdat, Thomas E. Anderson |
USENIX ATC | 1 |
| 1998 | The CRISIS Wide Area Security Architecture
Eshwar Belani, Amin Vahdat, Thomas E. Anderson, Michael Dahlin |
USENIX Security Symposium | 2 |
| 1998 | GLUnix: A Global Layer Unix for a Network of WorkstationsabstractRecent improvements in network and workstation performance have made workstation clusters an attractive architecture for diverse workloads, including interactive sequential and parallel applications. Although viable hardware solutions are available today, the largest challenge in making such a cluster usable lies in the system software. This paper describes the design and implementation of GLUnix, operating system middleware for a cluster of workstations. GLUnix was designed to provide transparent remote execution, support for interactive parallel and sequential jobs, load ballancing, and backward compatibility for existing application binaries. GLUnix was constructed to be easily portable to a number of platforms. GLUnix has been in daily use for over two and a half years and is currently running on a 100-node cluster of Sun UltraSPARCs. This paper relates our experiences with designing, building, and operating GLUnix. We discuss three important design tradeoffs faced by any cluster system, and present the reasons for our choices. Each of these design decisions is then re-evaluated in light of both our experience and recent technological advancements. We then describe the user-level, centralized, event-driven architecture of GLUnix and highlight a number of aspects of the implementation. Performance and scalability measurements of the system indicate that a centralized, user-level design can scale gracefully to significant cluster sizes, incurring only an additional 220 μs of overhead per node for remote execution. The discussion focuses on the successes and failures we encountered while building and maintaining the system, including a characterization of the limitations of a user-level implementation and various features that were added to satisfy the user community. © 1998 John Wiley & Sons, Ltd. David Petrou, Steven H. Rodrigues, Amin Vahdat, Thomas E. Anderson |
Softw. Pract. Exp. | 3 |
| 1997 | Effects of Communication Latency, Overhead, and Bandwidth in a Cluster ArchitectureabstractThis work provides a systematic study of the impact of communication performance on parallel applications in a high performance network of workstations. We develop an experimental system in which the communication latency, overhead, and bandwidth can be independently varied to observe the effects on a wide range of applications. Our results indicate that current efforts to improve cluster communication performance to that of tightly integrated parallel machines results in significantly improved application performance. We show that applications demonstrate strong sensitivity to overhead, slowing down by a factor of 60 on 32 processors when overhead is increased from 3 to 103 µs. Applications in this study are also sensitive to per-message bandwidth, but are surprisingly tolerant of increased latency and lower per-byte bandwidth. Finally, most applications demonstrate a highly linear dependence to both overhead and per-message bandwidth, indicating that further improvements in communication performance will continue to improve application performance. Richard P. Martin, Amin Vahdat, David E. Culler, Thomas E. Anderson |
ISCA | 2 |
| 1995 | The Interaction of Parallel and Sequential Workloads on a Network of WorkstationsabstractThis paper examines the plausibility of using a network of workstations (NOW) for a mixture of parallel and sequential jobs. Through simulations, our study examines issues that arise when combining these two workloads on a single platform. Starting from a dedicated NOW just for parallel programs, we incrementally relax uniprogramming restrictions until we have a multi-programmed, multi-user NOW for both interactive sequential users and parallel programs. We show that a number of issues associated with the distributed NOW environment (e.g., daemon activity, coscheduling skew) can have a small but noticeable effect on parallel program performance. We also find that efficient migration to idle workstations is necessary to maintain acceptable parallel application performance. Furthermore, we present a methodology for deriving an optimal delay time for recruiting idle machines for use by parallel programs; this recruitment threshold was just 3 minutes for the research cluster we measured. Finally, we quantify the effects of the additional parallel load upon interactive users by keeping track of the potential number of user delays in our simulations. When we limit the maximum number of delays per user, we can still maintain acceptable parallel program performance. In summary, we find that for our workloads a 2:1 rule applies: a NOW cluster of approximately 60 machines can sustain a 32-node parallel workload in addition to the sequential load placed upon it by interactive users. Remzi H. Arpaci-Dusseau, Andrea C. Arpaci-Dusseau, Amin Vahdat, Lok T. Liu, Thomas E. Anderson, David A. Patterson 0001 |
SIGMETRICS | 3 |
| 1993 | Tools for the Development of Application-Specific Virtual Memory ManagementabstractThe operating system's virtual memory management policy is increasingly important to application performance because gains in processing speed are far outstripping improvements in disk latency. Indeed, large applications can gain large performance benefits from using a virtual memory policy tuned to their specific memory access patterns rather than a general policy provided by the operating system. As a result, a number of schemes have been proposed to allow for application-specific extensions to virtual memory management. These schemes have the potential to improve performance; however, to realize this performance gain, application developers must implement their own virtual memory module, a non-trivial programming task. Current operating systems and programming tools are inadequate for developing application-specific policies. Our work combines (i) an extensible user-level virtual memory system based on a metaobject protocol with (ii) an innovative graphical performance monitor to make the task of implementing a new application-specific page replacement policy considerably simpler. The techniques presented for opening up operating system virtual memory policy to user control are general; they could be used to build application-specific implementations of other operating system policies Keith Krueger, David Loftesness, Amin Vahdat, Thomas E. Anderson |
OOPSLA | 3 |